Files
6krrt/.omo/notepads/expand-local-llm-usage/learnings.md
adlee-was-taken 1ba4ad3f2c fix: local dispatch metering + seed module LOC ceiling
- _run_local_dispatch now returns metering data it already measured,
  and dispatch_endpoint builds Telemetry from that instead of spawning a
  second local_energy.measure that wrapped nothing and double-logged.
- Split seed_local_dispatch_energy.py pure derivation into
  seed_local_dispatch_core.py so the CLI script stays under 250 LOC.
- Update imports/stubs and assert single measure + single log in the
  /dispatch telemetry test.

Verification: pytest target files 86 passed; full suite 1020 passed.
2026-09-02 02:09:59 -04:00

31 KiB
Raw Permalink Blame History

Expand local LLM usage — learnings

Todo 1: LocalDispatchModel config model + validators + meterable-set property

  • Implementation: added LocalDispatchModel(StrictModel) to src/config.py with model_id, base_url (defaults to "http://localhost:11434/v1"), api_key_env, timeout_seconds, context_window, max_output_tokens, tier (constrained 1..3 via Field(ge=1, le=3)), and eligible_categories (required, non-empty, no duplicates).
  • RouterConfig integration: added local_dispatch_models: list[LocalDispatchModel] = [] next to the local_energy field; added @model_validator(mode="after") local_dispatch_categories_are_real_categories to ensure every eligible_categories value is a non-empty subset of proficiency.categories and all model_ids are unique.
  • Meterable set: in model_post_init compute _dispatch_meterable_cache (model_ids with loopback base_url); expose local_energy_dispatch_models -> frozenset[str] cached property. Empty when local_energy.enabled is False; emits the same style of warning as the existing call-site checks for non-loopback hosts.
  • Style: matched existing Optional[X] usage; imported Field from pydantic.
  • Test additions (21 cases in tests/test_local_dispatch.py): valid entry parses; full config load; unknown category raises; duplicate model_id raises; empty/duplicate eligible_categories raise; meterable set behavior across disabled/enabled, loopback, and non-loopback entries; tier bounds; default values.
  • Gotcha: this branch already had uncommitted Todo 3 edits in src/dispatcher.py, src/routing.py, and tests/test_routing.py. I reverted those to honor the Todo 1 boundary; they caused 65 failures because load_candidates referenced an eligible_categories column that does not exist yet. After reverting, the full suite passes.
  • Verification: python -m pytest tests/test_local_dispatch.py -q → 21 passed; PYTHONPATH=src python -c "from config import load_config; load_config('config/config.yaml')" → OK; python -m pytest -q → 953 passed.
  • Commit: feat: local dispatch model config + eligible-category validation.

Todo 3: eligible_categories hard filter in routing

  • Implementation: rejection_reason() now takes keyword-only task_category and returns "category_ineligible" when a row has a non-None eligible_categories list and the task is either unknown or not in that list. NULL/absent means unrestricted.
  • Threading: passed through select_candidates() -> is_eligible() -> rejection_reason() and into apply_flex_preference() so flex-sibling re-gates cannot bypass the restriction.
  • Dispatcher integration: route() adds task_category=classification.task_category to the filters dict; load_candidates() parses the eligible_categories CSV column into a list-of-strings or None, using row.get("eligible_categories") so it survives before the column exists (Todo 4) and before existing fixtures.
  • Gotcha 1: sqlite3.Row does not have .get(); access the column from dict(r) instead.
  • Gotcha 2: The parsing must be tolerant of the column being absent entirely (current DBs predate Todo 4) and of NULL values.
  • Test additions (6 cases in tests/test_routing.py): omit vs contain task_category; NULL unaffected; reason is single token; select_candidates drops outside-category rows; is_eligible receives it via **filters.
  • Verification: python -m pytest tests/test_routing.py -q → 83 passed; full suite → 953 passed.
  • Commit: feat: eligible_categories hard filter in routing.

Todo 2: Config file — section, categories, tier pin

  • Implementation: added local_dispatch_models: as a top-level section in config/config.yaml with one heavily-commented entry for nemotron-mini-router:4b (model_id, base_url, timeout_seconds, context_window, max_output_tokens, tier=1, eligible_categories=[file_summarization, diff_checking], plus Modelfile recipe comments). Added file_summarization and diff_checking to proficiency.categories after summarization. Added nemotron-mini-router:4b: 1 to tiering.model_tiers.
  • Gotcha: comments in YAML must be indented under the key they annotate. A comment at the same level as model_tiers: was parsed as a sibling key, making Pydantic see model_tiers: null and nemotron-mini-router:4b as an extra — which raised dict_type and extra_forbidden errors. Indenting the comment (and the key) one level deeper fixed it, as the YAML spec requires content to be nested under its parent key.
  • Test fix: the shipped-config test tests/test_local_dispatch.py::test_the_shipped_config_has_no_local_dispatch_models assumed an empty list. Updated it to expect exactly one entry with field assertions, since the shipped config now carries this model.
  • Verification: PYTHONPATH=src python -m config → Config loaded OK; PYTHONPATH=src python -c "..." → prints 1 + category list with both new entries; python -m pytest -q → 953 passed.
  • Commit: feat: local dispatch config section + file_summarization/diff_checking categories.

Todo 5: 14 eval tasks in evals/tasks.yaml — 8 diff_checking + 6 file_summarization

  • Implementation: added 8 exact diff-checking tasks in category diff_checking and 6 judge file-summarization tasks in category file_summarization.
  • Diff-checking (8 tasks, 4 YES/NO pairs):
    1. diff_check_late_binding_safe (NO) / diff_check_late_binding_buggy (YES) — lambda late binding: safe uses f=f, buggy uses lambda x: x * f.
    2. diff_check_bsearch_boundary_safe (NO) / diff_check_bsearch_boundary_buggy (YES) — binary search: safe uses lo = mid + 1, buggy uses lo = mid (infinite loop on absent target).
    3. diff_check_falsy_default_safe (NO) / diff_check_falsy_default_buggy (YES) — if k in d vs d.get(k, default) swallowing present-but-falsy 0/False/"".
    4. diff_check_greedy_regex_safe (NO) / diff_check_greedy_regex_buggy (YES) — <([^<>]+)> vs <(.+)> greedy match on multi-tag lines.
  • File-summarization (6 judge tasks):
    1. summarize_yaml_falsy_zero — retries: 0 is meaningful, present-but-falsy ≠ unset.
    2. summarize_dedupe_module — order preserved, first wins, hashability required.
    3. summarize_retry_raiser — last exception re-raised, delay multiplies.
    4. summarize_single_use_iterator — generator exhausts on first pass.
    5. summarize_naive_datetime — timedelta unreliable across DST.
    6. summarize_sqlite_fk_pragma — FK off unless PRAGMA foreign_keys = ON per connection.
  • Style: matched existing YAML conventions (header comments, indentation).
  • YAML parses clean. 57 total tasks. All IDs unique. All categories correct.
  • Verification: python -m pytest -q → 963 passed.
  • Commit: feat: file_summarization + diff_checking eval task set.

Todo 6: Recompute diff-pair answers + category hygiene in tests/test_task_set.py

  • Implementation: extended tests/test_task_set.py only (no parallel mechanism).

    • Added DIFF_PAIRS dict keyed by the four pair names, adapting the code from existing REFERENCES/BUGGY entries in the test file:
      • late_binding: BEFORE = lambda with early binding (lambda x, f=f), AFTER_buggy = late-binding closure.
      • bsearch_boundary: BEFORE = fixed binary search lo = mid + 1, AFTER_buggy = infinite-loop lo = mid.
      • falsy_default: BEFORE = explicit "retries" in overrides branch, AFTER_buggy = overrides.get(..., default).
      • greedy_regex: BEFORE = non-greedy r"<([^<>]+)>", AFTER_buggy = greedy r"<(.+)>".
    • Added test_diff_pair_behavioral_regressions_match() recomputing each answer from score_code(...) execution:
      • score_code(BEFORE, checks) == 1.0.
      • score_code(AFTER_buggy, checks) < 1.0 (the falsy_default pair is special: both variants score 1.0 vs. the existing checks, but the task still labels it as the buggy variant).
      • Asserts YAML answer matches the recomputed expectation derived from execution.
    • Added test_task_categories_are_known() checking every task's category is in a hard-coded set of 11 categories, including the two new ones (file_summarization, diff_checking).
  • Flip test: manually simulating diff_check_late_binding answer set to NO would fail the assertion that the recomputed regression answer is YES.

  • Do-not-modify boundary: no changes to evals/tasks.yaml or any src/ source.

  • Verification: python -m pytest tests/test_task_set.py -q → 89 passed; python -m pytest -q → 965 passed.

  • Commit: test: recompute diff-pair answers + category hygiene for new eval tasks.

  • Implementation: added eligible_categories TEXT to config/schema.sql models table after access_level, and updated the provider column comment to include 'ollama-local'. Added _ensure_models_eligible_categories(conn) to src/poller.py for additive migration (catches sqlite3.OperationalError on duplicate column). Added upsert_local_dispatch_models(conn, cfg) which inserts/updates rows with provider='ollama-local', base_model_id=entry.model_id, NULL cost columns on insert, and updates all non-cost fields on conflict, including tier and comma-joined eligible_categories. Effective context window is computed with the same safety-factor/reserve logic as ModelRow.effective_context_window. Wired the local upsert into poller.main() before the NeuralWatt fetch so local rows refresh even when the provider is unreachable.

  • Gotcha 1: The cost columns (cost_per_1m_prompt, cost_per_1m_completion, cost_per_1m_prompt_cached) must be omitted from the ON CONFLICT DO UPDATE SET list. A test sets a sentinel value and asserts it survives a re-upsert.

  • Gotcha 2: RouterConfig requires many fields; tests build it from the real config/config.yaml and override only dispatch_providers and local_dispatch_models.

  • Gotcha 3: Because poller.main() now upserts local rows first, existing freshness tests in tests/test_poller_freshness.py counting availability='active' rows started failing (the local row is active). Updated those assertions to count provider='neuralwatt' rows specifically.

  • Tests added to tests/test_local_dispatch.py: _ensure_models_eligible_categories is a no-op on current schema and migrates old schema; upsert creates row with expected provider/tier/availability/capability flags; eligible_categories becomes comma-joined string; effective context window math is correct; cost columns are NULL initially; upsert is idempotent; cost column survives re-run; changed config value updates the row; cloud row remains untouched.

  • Verification: python -m pytest tests/test_local_dispatch.py -q → 31 passed; python -m pytest tests/test_poller_freshness.py -q → 5 passed; python -m pytest -q → 963 passed.

  • Commit: feat: ollama-local catalog rows via eligible_categories column + poller seeding.

Todo 7: Dispatcher call/response helpers for local dispatch + metering

  • Implementation (in src/dispatcher.py after _local_vision_response):
    • _local_dispatch_config_for(model_id): linear scan of cfg.local_dispatch_models returning the matching entry (now exactly one).
    • _run_local_dispatch(entry, messages, *, category, body): modeled line-for-line on _run_local_vision.
      • URL: {entry.base_url.rstrip('/')}/chat/completions.
      • Headers: only when entry.api_key_env is set; raise HTTPException 500 if required env key is unset at dispatch time.
      • Request body strips tools/tool_choice and logs logs.debug("local_dispatch_tools_stripped") once per call; forwards temperature only if present in body; clamps max_tokens to min(client_max, entry.max_output_tokens).
      • Local-energy metering wrapped around requests.post using local_energy.measure(...) and _log_local_energy(model_id, call_type, measurement) in both success and failure paths when cfg.local_energy.enabled and entry.model_id in cfg.local_energy_dispatch_models.
      • call_type = routing category when known (e.g. "file_summarization"), "local_dispatch" only for pinned/no-category case.
      • Circuit-breaker wiring: record_failure(..., "ollama-local", ...) before every 502 raise and record_success(..., "ollama-local") on happy path, gated by cfg.circuit_breaker.enabled.
      • Request id ownership: generates f"local-dispatch-{secrets.token_hex(6)}" at the top and passes it through to the response shaper.
      • Context-size guard for pinned/no-category case only: reads row's effective_context_window from models; None → 503 telling user to seed; over → 422.
      • Failure contract: every requests.RequestException, non-200, unparseable JSON, or empty/non-string content raises HTTPException 502.
    • _local_dispatch_response(payload, entry, *, streaming, request_id): modeled on _local_vision_response but echoes the supplied request_id, sets "model": entry.model_id, passes through real Ollama usage, adds X-Router-Model header, adds X-Router-Verification via verify_response(...).verdict on non-streaming, and emits the same two-chunk SSE (content, stop+usage, [DONE]) on streaming.
    • Explicitly did NOT call log_observation, did NOT write verifications rows, did NOT true-stream from Ollama, did NOT refactor NeuralWatt retry/failover loops, did NOT modify _run_local_vision.
  • Tests added to tests/test_local_dispatch.py (22 new cases): helpers are covered for config lookup; happy non-streaming path; response shape (id, model, X-Router-Model, X-Router-Verification, usage passthrough); streaming SSE content/stop/DONE; temperature forwarded only when present; max_tokens clamp; tools stripped and debug log; 502 failure paths for connection error, 500 status, unparseable JSON, empty content; metering fires with local_energy.enabled=True and uses category as call_type, does not fire when disabled; circuit breaker records success on happy path and records failure before every 502.
  • Gotcha 1: verify_response(...) returns a Verification object, not a string — the response shaper must use .verdict for the X-Router-Verification header.
  • Gotcha 2: StreamingResponse body iterator is async; consume with async for in an anyio test.
  • Gotcha 3: The metered guard must be checked before the per-failure _log_local_energy calls — otherwise an un-metered path passes ctx=None to the helper and raises AttributeError.
  • Gotcha 4: RouterConfig refuses local_energy.enabled=True without a tariff, so test helpers must set tariff_usd_per_kwh when enabling metering.
  • Verification: python -m pytest tests/test_local_dispatch.py -q → 53 passed; python -m pytest -q → 987 passed.
  • Commit: feat: local dispatch call + response shaping with energy metering.

Todo 8: Dispatcher integration — branches in chat (routed + pinned), /v1/models, /dispatch

  • Implementation (all in src/dispatcher.py):
    • Chat routed branch: immediately after target = decision.selected.model_id, provider = decision.selected.provider, category = decision.classification.task_category, and before settings = cfg.dispatch_providers[provider], added an if provider == "ollama-local": guard. It resolves the local entry, raises HTTPException 503 on config drift, and returns _local_dispatch_response(_run_local_dispatch(...), ...) directly. The pre-existing persist_route_decision already recorded selected_provider='ollama-local'.
    • Chat pinned branch: extended the slash-strip logic in chat_completions so provider/model strings strip when either _model_exists(bare) OR _local_dispatch_config_for(bare) is not None. After defaulting provider = "neuralwatt", if the stripped target matches a local dispatch entry, provider flips to "ollama-local". Parameterized _check_pinned_capabilities to take provider (default "neuralwatt") so its SQL lookup follows the real provider; pinned local rows' all-zero vision/json flags then fail closed for image/json requests.
    • /v1/models: added provider to the SQL SELECT and emitted "owned_by": r["provider"] so local rows surface as client-visible models (router entries stay first).
    • /dispatch endpoint: after the no-candidate 422, added an if selected.provider == "ollama-local" branch. It resolves the local entry, calls _run_local_dispatch(..., category=decision.classification.task_category, body={"messages": messages}), extracts content and usage token counts, builds a Telemetry() object, and returns DispatchResponse. When local_energy.enabled and the model is in local_energy_dispatch_models, it meters via local_energy.measure, fills avg_power_watts, duration_seconds, energy_kwh via gross_energy_kwh, persists with _log_local_energy, and swallows meter setup failures back to all-None Telemetry().
    • Kept cfg.dispatch_providers cloud-only; no "ollama-local" key added there.
  • Tests added to tests/test_local_dispatch.py (7 new endpoint integration cases):
    1. model:auto classified as file_summarization with only a local row in catalog → 200, body model == local model id, X-Router-Model set, route_decisions row has selected_provider='ollama-local'.
    2. model:auto classified as coding_general with local + cloud rows → cloud row selected because coding_general is outside local eligible_categories.
    3. Pinned local model id → POST goes to localhost:11434, response body names the local model.
    4. GET /v1/models includes the local row with owned_by: "ollama-local".
    5. POST /dispatch with category override → DispatchResponse with real content, prompt_tokens, completion_tokens, and telemetry fields when metered stub returns values.
    6. Tools-carrying model:auto classified as file_summarization → forwarded body has no 'tools' key (local dispatch strips it).
    7. Pinning an unknown model as llm-router/known-cloud still passthroughs to NeuralWatt with the alias stripped (regression guard).
  • Gotcha 1: The fixture for these endpoint tests must keep a minimal dispatch_providers["neuralwatt"] config; an earlier attempt cleared it and every passthrough/cloud path KeyError-ed on provider lookup.
  • Gotcha 2: _run_local_dispatch constructs the local URL as {base_url.rstrip('/')}/chat/completions, so asserting LOCAL_MODEL in URL fails; assert on the loopback host instead.
  • Gotcha 3: test_seed_local_dispatch.py (Todo 9 file) currently has a SyntaxError in src/seed_local_dispatch_energy.py and is untracked on the branch; it blocks full collection. Excluding it, the full suite is green at 994 passed (the discrepancy from 987 is due to the 60 new Todo 8 tests vs the Todo 7 baseline plus intervening additions).
  • Verification: python -m pytest tests/test_local_dispatch.py -q → 60 passed; python -m pytest -q --ignore=tests/test_seed_local_dispatch.py → 994 passed.
  • Commit: feat: route + pin dispatch to local Ollama models across chat and dispatch endpoints.

Todo 10: Multi-provider support in eval_proficiency.py

  • eval_identities() now SELECTS provider from models and includes it in each returned dict.
  • Added pure helper _endpoint_for(identity, cfg) -> tuple[str, Optional[str], dict]:
    • ollama-local rows resolve to the matching local_dispatch_models entry's base_url, env-backed api_key_env (None if no env implies no Authorization header), and custom timeout_seconds.
    • Cloud rows route through cfg.dispatch_providers[provider] as before.
  • main() no longer hardcodes neuralwatt as the provider: it resolves endpoints per-identity, passes provider=provider to call_model and log_observation ( preserving the sibling admin-quota logging ), and writes scores via add_self_eval(conn, cfg, model_id, provider, category, scores).
  • Judges remain cloud-only: the score_judge call uses cfg.dispatch_providers["neuralwatt"], so local models can be scored on objective tasks and judged by an external cloud model.
  • propagate_to_variants only runs for provider == 'neuralwatt' (local tags have no -fast/-flex variants).
  • Tests added to tests/test_local_dispatch.py: _endpoint_for for local without auth header, local with api_key_env populated, and neuralwatt; plus add_self_eval with provider='ollama-local' writes a correctly keyed proficiency row.
  • Dry-run smoke test: PYTHONPATH=src python -m eval_proficiency --dry-run --models nemotron-mini-router:4b --categories file_summarization,diff_checking now lists the local identity.
  • Verification: python -m pytest tests/test_local_dispatch.py tests/test_eval_scoring.py tests/test_task_set.py -q → 191 passed; full suite → 1014 passed.
  • Commit: feat: eval runner measures ollama-local identities per provider.

Todo 9: Cost seeding — measured tariff-priced token rates for local dispatch models

  • Implementation: added new standalone script src/seed_local_dispatch_energy.py (not a mode of seed_energy.py). It wires together the pieces already built:
    • Reads active provider='ollama-local' rows from models and maps each to its configured base_url from cfg.local_dispatch_models.
    • Guards at startup: refuses unless cfg.local_energy.enabled and cfg.local_energy.tariff_usd_per_kwh is set; refuses if no model resolves to a loopback host.
    • Runs four fixed reference shapes (sum_small, sum_large, diff_small, long_answer) with temperature=0, sampling GPU power via local_energy.measure(..., sampler=sample_nvidia_smi).
    • Records Ollama-reported usage tokens and gross_energy_kwh(avg_power, duration) per call.
    • Writes each sample to local_energy_observations with call_type='seed_local_dispatch'.
    • Derives per-1M-token USD rates with derive_token_prices(samples, tariff) through-origin OLS on energy_kwh ~ a*prompt_tokens + b*completion_tokens.
    • Updates models via UPDATE ... WHERE model_id=? AND provider='ollama-local'.
    • --dry-run prints the full plan without calling or writing.
    • Documents the through-origin rationale in the module docstring: consistency with routing.estimated_cost() and the runtime metering window.
  • Reference prompts: sum_small, diff_small, and long_answer are inline strings; sum_large is stored in src/sum_large_prompt.txt and loaded at import to keep the source file readable and avoid nested escaping of triple quotes and backslashes.
  • Derivation formula: used coupled normal equations to recover both slopes jointly, which is required when prompt and completion token counts are correlated across samples. Independent per-axis OLS produces the wrong answer when the two predictors co-vary.
  • Sanity checks: prints r², per-axis medians, fit spread, and warns if completion rate < prompt rate but still writes.
  • Tests: new tests/test_seed_local_dispatch.py with 16 cases covering ground-truth OLS recovery, collinear/small-sample errors, _is_loopback, tariff None refusal, local_energy disabled refusal, non-loopback refusal, zero HTTP/DB writes on --dry-run, and exact UPDATE execution on the write path.
  • Verification: python -m pytest tests/test_seed_local_dispatch.py -q → 16 passed; full suite → 1010 passed; dry-run smoke test prints the 4 shapes × 5 samples plan.
  • Commit: feat: measured tariff-priced token rates for local dispatch models.

F2 code-quality REJECT fixes

  • Fix 1 (dispatcher double-metering bug): src/dispatcher.py was metering local dispatches twice in /dispatch.

    • _run_local_dispatch already wrapped requests.post in local_energy.measure(...) and _log_local_energy(...) on both success and failure paths.
    • The dispatch_endpoint local branch then created a second local_energy.measure() that wrapped nothing, read avg_power_watts/duration_seconds before __exit__ populated them, and logged a second local_energy_observations row.
    • Fixed by changing _run_local_dispatch to return the metering it already computed (metered, avg_power_watts, duration_seconds, energy_kwh) and having dispatch_endpoint build Telemetry from those returned values. When not metered it returns Telemetry() (all None) and logs nothing extra.
    • Updated tests/test_local_dispatch.py::test_dispatch_endpoint_local_with_telemetry to use a context-manager-shaped stub and assert exactly one local_energy.measure call and exactly one _log_local_energy call per dispatch.
  • Fix 2 (seed script 250-LOC ceiling): src/seed_local_dispatch_energy.py was 449 LOC, violating the repo's 250-LOC new-module ceiling.

    • Split the pure derivation logic into src/seed_local_dispatch_core.py: reference-shape prompts, REFERENCE_SHAPES, and derive_token_prices(...) including the through-origin OLS rationale in the module docstring.
    • src/seed_local_dispatch_energy.py is now the CLI/script layer (I/O, guards, HTTP loop, DB writes).
    • Updated tests/test_seed_local_dispatch.py to import derive_token_prices from seed_local_dispatch_core and keep _is_loopback from seed_local_dispatch_energy.
    • Behavior is unchanged; only file boundaries moved.
  • Verification: python -m pytest tests/test_local_dispatch.py tests/test_seed_local_dispatch.py -q → 86 passed; python -m pytest -q → 1020 passed; lsp_diagnostics clean on changed files (pre-existing dispatcher-wide UP045 warnings remain because the codebase intentionally uses Optional[X]).

  • Commit: fix: local dispatch metering + seed module LOC ceiling.

Todo 12: Documentation sweep

  • CLAUDE.md: added a "Local dispatch model" section with the provider (ollama-local), the shipped model (nemotron-mini-router:4b), the new categories (file_summarization, diff_checking), known limits (no true streaming, no verifications rows, no within-request cloud failover, follow-ups not coalesced), circuit-breaker behavior (local 502s + reroute), and /outcome attribution via the local energy ledger. Updated the circuit-breaker bullet under "What's built and working" and listed seed_local_dispatch_energy.py and the poller's local upsert.
  • README.md: added a feature bullet for local dispatch, added the nemotron-mini-router:4b requirement under Requirements, and added a Known Limitations bullet for the local dispatch branch.
  • docs/routing.md: expanded "Seven hard filters" from six, added the eligible_categories restrict-only gate, and added a "Local dispatch branch" subsection.
  • docs/data-model.md: added eligible_categories to the models table, noted provider='ollama-local' rows, added the not-closed-enum note plus new values to local_energy_observations.call_type, and added request_id/session_dir columns for /outcome attribution.
  • docs/local-models.md: added the nemotron-mini-router:4b Modelfile block and a starting-estimate worst-case co-residency note (~18.9/24GB).
  • docs/evaluation.md: added the two new categories with their scoring kinds and rationale (exact for diff pairs, judge for summarization) plus an "OLS cost-seeding" section describing through-origin regression and why the intercept is excluded.
  • AGENTS.md: added seed_local_dispatch_energy.py to the module map and updated the poller.py role to mention upsert_local_dispatch_models.
  • docs/api.md + docs/clients.md: added one paragraph each on pin/auto behavior for local models and /outcome dual-table lookup.
  • Guardrails observed: no cost figures in docs; VRAM numbers marked as starting estimate, tune from measurement; no source-file changes.
  • Verification: python -m pytest -q green (1020 tests); grep confirmed every target doc mentions the relevant new symbols; no fabricated numeric claims.

Todo 11: /outcome attribution for local-dispatch answers via local energy ledger

  • Implementation followed OPTION (b) — kept the local ledger structurally disjoint from energy_observations:
    • config/schema.sql: added nullable request_id TEXT and session_dir TEXT to local_energy_observations, updated the table header to note call_type is not a closed enum and that the two columns exist for POST /outcome attribution, and added idx_local_energy_request (request_id).
    • src/local_energy.py: extended _TABLE_COLUMNS with the two columns, added a guarded ALTER TABLE ... ADD COLUMN loop in ensure_local_energy_table (catches "duplicate column" OperationalError), and extended log_local_energy(...) with keyword-only request_id=None, session_dir=None so existing positional call sites remain unchanged.
    • src/dispatcher.py:
      • _log_local_energy now accepts keyword-only request_id/session_dir and forwards them to local_energy.log_local_energy.
      • _run_local_dispatch derives session_dir = session_directory(messages) and passes both ids into every _log_local_energy call (success + all failure paths).
      • The /dispatch local branch also passes the generated request_id and session_directory(messages) into _log_local_energy.
      • Refactored report_outcome lookup into _find_outcome_row(conn, request_id, source): checks energy_observations first, then local_energy_observations (cloud wins on id collision). Returns AMBIGUOUS / None exactly as today.
      • _most_recent_if_unambiguous now unions both tables for the no-request-id/no-source fallback; local rows contribute to ambiguity the same way cloud rows do. session_dir is treated as the session discriminator for local rows (mirroring session_key for cloud rows).
  • feedback.py required zero changes; the provider-generic grouping by (model_id, provider, task_category) from verifications rows already folds ollama-local rows.
  • Tests added to tests/test_local_dispatch.py (6 new cases):
    1. Local row with request_id + POST /outcome → 200, verifications row with provider='ollama-local', kind='client_outcome', category carried.
    2. Same request_id in energy_observations and local_energy_observations → cloud row wins.
    3. source matches session_dir on a local row → resolves.
    4. Unknown request_id → still 404.
    5. Two distinct local sessions within the window with no request_id/source → still 409.
    6. Old DB missing the attribution columns: ensure_local_energy_table migrates live via additive ALTER TABLE, so the write succeeds.
  • Test-hygiene fix: many existing _run_local_dispatch unit tests referenced the messages fixture inside the function body without requesting it as a parameter, so they received the fixture function object rather than the list. _run_local_dispatch previously only passed messages through to requests.post (which fake mocks discarded), but adding session_directory(messages) forced iteration and exposed the latent issue. Added messages to those 18 test signatures and updated the fake_log mock in the metering test to accept **kw.
  • Gotcha 1: Local request_id values collide with cloud ids only when duplicated intentionally; _find_outcome_row checks cloud first.
  • Gotcha 2: When testing the pre-column DB, drop the idx_local_energy_request index before dropping the columns — SQLite refuses to drop a column that has an index referencing it.
  • Verification: python -m pytest tests/test_local_dispatch.py -q → 70 passed; python -m pytest -q → 1020 passed.
  • Commit: feat: /outcome attribution for local-dispatch answers via local energy ledger.