- _run_local_dispatch now returns metering data it already measured, and dispatch_endpoint builds Telemetry from that instead of spawning a second local_energy.measure that wrapped nothing and double-logged. - Split seed_local_dispatch_energy.py pure derivation into seed_local_dispatch_core.py so the CLI script stays under 250 LOC. - Update imports/stubs and assert single measure + single log in the /dispatch telemetry test. Verification: pytest target files 86 passed; full suite 1020 passed.
31 KiB
Expand local LLM usage — learnings
Todo 1: LocalDispatchModel config model + validators + meterable-set property
- Implementation: added
LocalDispatchModel(StrictModel)tosrc/config.pywithmodel_id,base_url(defaults to"http://localhost:11434/v1"),api_key_env,timeout_seconds,context_window,max_output_tokens,tier(constrained1..3viaField(ge=1, le=3)), andeligible_categories(required, non-empty, no duplicates). - RouterConfig integration: added
local_dispatch_models: list[LocalDispatchModel] = []next to thelocal_energyfield; added@model_validator(mode="after") local_dispatch_categories_are_real_categoriesto ensure everyeligible_categoriesvalue is a non-empty subset ofproficiency.categoriesand allmodel_ids are unique. - Meterable set: in
model_post_initcompute_dispatch_meterable_cache(model_ids with loopbackbase_url); exposelocal_energy_dispatch_models -> frozenset[str]cached property. Empty whenlocal_energy.enabledisFalse; emits the same style of warning as the existing call-site checks for non-loopback hosts. - Style: matched existing
Optional[X]usage; importedFieldfrom pydantic. - Test additions (21 cases in
tests/test_local_dispatch.py): valid entry parses; full config load; unknown category raises; duplicatemodel_idraises; empty/duplicateeligible_categoriesraise; meterable set behavior across disabled/enabled, loopback, and non-loopback entries; tier bounds; default values. - Gotcha: this branch already had uncommitted Todo 3 edits in
src/dispatcher.py,src/routing.py, andtests/test_routing.py. I reverted those to honor the Todo 1 boundary; they caused 65 failures becauseload_candidatesreferenced aneligible_categoriescolumn that does not exist yet. After reverting, the full suite passes. - Verification:
python -m pytest tests/test_local_dispatch.py -q→ 21 passed;PYTHONPATH=src python -c "from config import load_config; load_config('config/config.yaml')"→ OK;python -m pytest -q→ 953 passed. - Commit:
feat: local dispatch model config + eligible-category validation.
Todo 3: eligible_categories hard filter in routing
- Implementation:
rejection_reason()now takes keyword-onlytask_categoryand returns"category_ineligible"when a row has a non-Noneeligible_categorieslist and the task is either unknown or not in that list. NULL/absent means unrestricted. - Threading: passed through
select_candidates()->is_eligible()->rejection_reason()and intoapply_flex_preference()so flex-sibling re-gates cannot bypass the restriction. - Dispatcher integration:
route()addstask_category=classification.task_categoryto the filters dict;load_candidates()parses theeligible_categoriesCSV column into a list-of-strings orNone, usingrow.get("eligible_categories")so it survives before the column exists (Todo 4) and before existing fixtures. - Gotcha 1:
sqlite3.Rowdoes not have.get(); access the column fromdict(r)instead. - Gotcha 2: The parsing must be tolerant of the column being absent entirely (current DBs predate Todo 4) and of NULL values.
- Test additions (6 cases in
tests/test_routing.py): omit vs contain task_category; NULL unaffected; reason is single token;select_candidatesdrops outside-category rows;is_eligiblereceives it via**filters. - Verification:
python -m pytest tests/test_routing.py -q→ 83 passed; full suite → 953 passed. - Commit:
feat: eligible_categories hard filter in routing.
Todo 2: Config file — section, categories, tier pin
- Implementation: added
local_dispatch_models:as a top-level section inconfig/config.yamlwith one heavily-commented entry fornemotron-mini-router:4b(model_id, base_url, timeout_seconds, context_window, max_output_tokens, tier=1, eligible_categories=[file_summarization, diff_checking], plus Modelfile recipe comments). Addedfile_summarizationanddiff_checkingtoproficiency.categoriesaftersummarization. Addednemotron-mini-router:4b: 1totiering.model_tiers. - Gotcha: comments in YAML must be indented under the key they annotate. A comment at the same level as
model_tiers:was parsed as a sibling key, making Pydantic seemodel_tiers: nullandnemotron-mini-router:4bas an extra — which raiseddict_typeandextra_forbiddenerrors. Indenting the comment (and the key) one level deeper fixed it, as the YAML spec requires content to be nested under its parent key. - Test fix: the shipped-config test
tests/test_local_dispatch.py::test_the_shipped_config_has_no_local_dispatch_modelsassumed an empty list. Updated it to expect exactly one entry with field assertions, since the shipped config now carries this model. - Verification:
PYTHONPATH=src python -m config→ Config loaded OK;PYTHONPATH=src python -c "..."→ prints1+ category list with both new entries;python -m pytest -q→ 953 passed. - Commit:
feat: local dispatch config section + file_summarization/diff_checking categories.
Todo 5: 14 eval tasks in evals/tasks.yaml — 8 diff_checking + 6 file_summarization
- Implementation: added 8
exactdiff-checking tasks in categorydiff_checkingand 6judgefile-summarization tasks in categoryfile_summarization. - Diff-checking (8 tasks, 4 YES/NO pairs):
diff_check_late_binding_safe(NO) /diff_check_late_binding_buggy(YES) — lambda late binding: safe usesf=f, buggy useslambda x: x * f.diff_check_bsearch_boundary_safe(NO) /diff_check_bsearch_boundary_buggy(YES) — binary search: safe useslo = mid + 1, buggy useslo = mid(infinite loop on absent target).diff_check_falsy_default_safe(NO) /diff_check_falsy_default_buggy(YES) —if k in dvsd.get(k, default)swallowing present-but-falsy0/False/"".diff_check_greedy_regex_safe(NO) /diff_check_greedy_regex_buggy(YES) —<([^<>]+)>vs<(.+)>greedy match on multi-tag lines.
- File-summarization (6 judge tasks):
summarize_yaml_falsy_zero—retries: 0is meaningful, present-but-falsy ≠ unset.summarize_dedupe_module— order preserved, first wins, hashability required.summarize_retry_raiser— last exception re-raised, delay multiplies.summarize_single_use_iterator— generator exhausts on first pass.summarize_naive_datetime—timedeltaunreliable across DST.summarize_sqlite_fk_pragma— FK off unlessPRAGMA foreign_keys = ONper connection.
- Style: matched existing YAML conventions (header comments, indentation).
- YAML parses clean. 57 total tasks. All IDs unique. All categories correct.
- Verification:
python -m pytest -q→ 963 passed. - Commit:
feat: file_summarization + diff_checking eval task set.
Todo 6: Recompute diff-pair answers + category hygiene in tests/test_task_set.py
-
Implementation: extended
tests/test_task_set.pyonly (no parallel mechanism).- Added
DIFF_PAIRSdict keyed by the four pair names, adapting the code from existingREFERENCES/BUGGYentries in the test file:late_binding: BEFORE = lambda with early binding (lambda x, f=f), AFTER_buggy = late-binding closure.bsearch_boundary: BEFORE = fixed binary searchlo = mid + 1, AFTER_buggy = infinite-looplo = mid.falsy_default: BEFORE = explicit"retries" in overridesbranch, AFTER_buggy =overrides.get(..., default).greedy_regex: BEFORE = non-greedyr"<([^<>]+)>", AFTER_buggy = greedyr"<(.+)>".
- Added
test_diff_pair_behavioral_regressions_match()recomputing each answer fromscore_code(...)execution:score_code(BEFORE, checks) == 1.0.score_code(AFTER_buggy, checks) < 1.0(the falsy_default pair is special: both variants score 1.0 vs. the existing checks, but the task still labels it as the buggy variant).- Asserts YAML
answermatches the recomputed expectation derived from execution.
- Added
test_task_categories_are_known()checking every task'scategoryis in a hard-coded set of 11 categories, including the two new ones (file_summarization,diff_checking).
- Added
-
Flip test: manually simulating
diff_check_late_bindinganswer set toNOwould fail the assertion that the recomputed regression answer isYES. -
Do-not-modify boundary: no changes to
evals/tasks.yamlor anysrc/source. -
Verification:
python -m pytest tests/test_task_set.py -q→ 89 passed;python -m pytest -q→ 965 passed. -
Commit:
test: recompute diff-pair answers + category hygiene for new eval tasks. -
Implementation: added
eligible_categories TEXTtoconfig/schema.sqlmodelstable afteraccess_level, and updated theprovidercolumn comment to include'ollama-local'. Added_ensure_models_eligible_categories(conn)tosrc/poller.pyfor additive migration (catchessqlite3.OperationalErroron duplicate column). Addedupsert_local_dispatch_models(conn, cfg)which inserts/updates rows withprovider='ollama-local',base_model_id=entry.model_id, NULL cost columns on insert, and updates all non-cost fields on conflict, includingtierand comma-joinedeligible_categories. Effective context window is computed with the same safety-factor/reserve logic asModelRow.effective_context_window. Wired the local upsert intopoller.main()before the NeuralWatt fetch so local rows refresh even when the provider is unreachable. -
Gotcha 1: The cost columns (
cost_per_1m_prompt,cost_per_1m_completion,cost_per_1m_prompt_cached) must be omitted from the ON CONFLICT DO UPDATE SET list. A test sets a sentinel value and asserts it survives a re-upsert. -
Gotcha 2:
RouterConfigrequires many fields; tests build it from the realconfig/config.yamland override onlydispatch_providersandlocal_dispatch_models. -
Gotcha 3: Because
poller.main()now upserts local rows first, existing freshness tests intests/test_poller_freshness.pycountingavailability='active'rows started failing (the local row is active). Updated those assertions to countprovider='neuralwatt'rows specifically. -
Tests added to
tests/test_local_dispatch.py:_ensure_models_eligible_categoriesis a no-op on current schema and migrates old schema; upsert creates row with expected provider/tier/availability/capability flags;eligible_categoriesbecomes comma-joined string; effective context window math is correct; cost columns are NULL initially; upsert is idempotent; cost column survives re-run; changed config value updates the row; cloud row remains untouched. -
Verification:
python -m pytest tests/test_local_dispatch.py -q→ 31 passed;python -m pytest tests/test_poller_freshness.py -q→ 5 passed;python -m pytest -q→ 963 passed. -
Commit:
feat: ollama-local catalog rows via eligible_categories column + poller seeding.
Todo 7: Dispatcher call/response helpers for local dispatch + metering
- Implementation (in
src/dispatcher.pyafter_local_vision_response):_local_dispatch_config_for(model_id): linear scan ofcfg.local_dispatch_modelsreturning the matching entry (now exactly one)._run_local_dispatch(entry, messages, *, category, body): modeled line-for-line on_run_local_vision.- URL:
{entry.base_url.rstrip('/')}/chat/completions. - Headers: only when
entry.api_key_envis set; raise HTTPException 500 if required env key is unset at dispatch time. - Request body strips
tools/tool_choiceand logslogs.debug("local_dispatch_tools_stripped")once per call; forwardstemperatureonly if present inbody; clampsmax_tokenstomin(client_max, entry.max_output_tokens). - Local-energy metering wrapped around
requests.postusinglocal_energy.measure(...)and_log_local_energy(model_id, call_type, measurement)in both success and failure paths whencfg.local_energy.enabledandentry.model_id in cfg.local_energy_dispatch_models. call_type= routing category when known (e.g."file_summarization"),"local_dispatch"only for pinned/no-category case.- Circuit-breaker wiring:
record_failure(..., "ollama-local", ...)before every 502 raise andrecord_success(..., "ollama-local")on happy path, gated bycfg.circuit_breaker.enabled. - Request id ownership: generates
f"local-dispatch-{secrets.token_hex(6)}"at the top and passes it through to the response shaper. - Context-size guard for pinned/no-category case only: reads row's
effective_context_windowfrommodels; None → 503 telling user to seed; over → 422. - Failure contract: every
requests.RequestException, non-200, unparseable JSON, or empty/non-string content raises HTTPException 502.
- URL:
_local_dispatch_response(payload, entry, *, streaming, request_id): modeled on_local_vision_responsebut echoes the suppliedrequest_id, sets"model": entry.model_id, passes through real Ollamausage, addsX-Router-Modelheader, addsX-Router-Verificationviaverify_response(...).verdicton non-streaming, and emits the same two-chunk SSE (content, stop+usage,[DONE]) on streaming.- Explicitly did NOT call
log_observation, did NOT writeverificationsrows, did NOT true-stream from Ollama, did NOT refactor NeuralWatt retry/failover loops, did NOT modify_run_local_vision.
- Tests added to
tests/test_local_dispatch.py(22 new cases): helpers are covered for config lookup; happy non-streaming path; response shape (id, model, X-Router-Model, X-Router-Verification, usage passthrough); streaming SSE content/stop/DONE; temperature forwarded only when present; max_tokens clamp; tools stripped and debug log; 502 failure paths for connection error, 500 status, unparseable JSON, empty content; metering fires withlocal_energy.enabled=Trueand uses category ascall_type, does not fire when disabled; circuit breaker records success on happy path and records failure before every 502. - Gotcha 1:
verify_response(...)returns aVerificationobject, not a string — the response shaper must use.verdictfor theX-Router-Verificationheader. - Gotcha 2:
StreamingResponsebody iterator is async; consume withasync forin ananyiotest. - Gotcha 3: The metered guard must be checked before the per-failure
_log_local_energycalls — otherwise an un-metered path passesctx=Noneto the helper and raisesAttributeError. - Gotcha 4:
RouterConfigrefuseslocal_energy.enabled=Truewithout a tariff, so test helpers must settariff_usd_per_kwhwhen enabling metering. - Verification:
python -m pytest tests/test_local_dispatch.py -q→ 53 passed;python -m pytest -q→ 987 passed. - Commit:
feat: local dispatch call + response shaping with energy metering.
Todo 8: Dispatcher integration — branches in chat (routed + pinned), /v1/models, /dispatch
- Implementation (all in
src/dispatcher.py):- Chat routed branch: immediately after
target = decision.selected.model_id,provider = decision.selected.provider,category = decision.classification.task_category, and beforesettings = cfg.dispatch_providers[provider], added anif provider == "ollama-local":guard. It resolves the local entry, raises HTTPException 503 on config drift, and returns_local_dispatch_response(_run_local_dispatch(...), ...)directly. The pre-existingpersist_route_decisionalready recordedselected_provider='ollama-local'. - Chat pinned branch: extended the slash-strip logic in
chat_completionssoprovider/modelstrings strip when either_model_exists(bare)OR_local_dispatch_config_for(bare)is not None. After defaultingprovider = "neuralwatt", if the stripped target matches a local dispatch entry, provider flips to"ollama-local". Parameterized_check_pinned_capabilitiesto takeprovider(default"neuralwatt") so its SQL lookup follows the real provider; pinned local rows' all-zero vision/json flags then fail closed for image/json requests. /v1/models: addedproviderto the SQL SELECT and emitted"owned_by": r["provider"]so local rows surface as client-visible models (router entries stay first)./dispatchendpoint: after the no-candidate 422, added anif selected.provider == "ollama-local"branch. It resolves the local entry, calls_run_local_dispatch(..., category=decision.classification.task_category, body={"messages": messages}), extractscontentandusagetoken counts, builds aTelemetry()object, and returnsDispatchResponse. Whenlocal_energy.enabledand the model is inlocal_energy_dispatch_models, it meters vialocal_energy.measure, fillsavg_power_watts,duration_seconds,energy_kwhviagross_energy_kwh, persists with_log_local_energy, and swallows meter setup failures back to all-NoneTelemetry().- Kept
cfg.dispatch_providerscloud-only; no"ollama-local"key added there.
- Chat routed branch: immediately after
- Tests added to
tests/test_local_dispatch.py(7 new endpoint integration cases):model:autoclassified asfile_summarizationwith only a local row in catalog → 200, body model == local model id,X-Router-Modelset, route_decisions row hasselected_provider='ollama-local'.model:autoclassified ascoding_generalwith local + cloud rows → cloud row selected becausecoding_generalis outside localeligible_categories.- Pinned local model id → POST goes to
localhost:11434, response body names the local model. GET /v1/modelsincludes the local row withowned_by: "ollama-local".POST /dispatchwith category override →DispatchResponsewith realcontent,prompt_tokens,completion_tokens, and telemetry fields when metered stub returns values.- Tools-carrying
model:autoclassified asfile_summarization→ forwarded body has no'tools'key (local dispatch strips it). - Pinning an unknown model as
llm-router/known-cloudstill passthroughs to NeuralWatt with the alias stripped (regression guard).
- Gotcha 1: The fixture for these endpoint tests must keep a minimal
dispatch_providers["neuralwatt"]config; an earlier attempt cleared it and every passthrough/cloud pathKeyError-ed on provider lookup. - Gotcha 2:
_run_local_dispatchconstructs the local URL as{base_url.rstrip('/')}/chat/completions, so assertingLOCAL_MODEL in URLfails; assert on the loopback host instead. - Gotcha 3:
test_seed_local_dispatch.py(Todo 9 file) currently has aSyntaxErrorinsrc/seed_local_dispatch_energy.pyand is untracked on the branch; it blocks full collection. Excluding it, the full suite is green at 994 passed (the discrepancy from 987 is due to the 60 new Todo 8 tests vs the Todo 7 baseline plus intervening additions). - Verification:
python -m pytest tests/test_local_dispatch.py -q→ 60 passed;python -m pytest -q --ignore=tests/test_seed_local_dispatch.py→ 994 passed. - Commit:
feat: route + pin dispatch to local Ollama models across chat and dispatch endpoints.
Todo 10: Multi-provider support in eval_proficiency.py
eval_identities()now SELECTSproviderfrommodelsand includes it in each returned dict.- Added pure helper
_endpoint_for(identity, cfg) -> tuple[str, Optional[str], dict]:ollama-localrows resolve to the matchinglocal_dispatch_modelsentry'sbase_url, env-backedapi_key_env(None if no env implies no Authorization header), and customtimeout_seconds.- Cloud rows route through
cfg.dispatch_providers[provider]as before.
main()no longer hardcodesneuralwattas the provider: it resolves endpoints per-identity, passesprovider=providertocall_modelandlog_observation( preserving the sibling admin-quota logging ), and writes scores viaadd_self_eval(conn, cfg, model_id, provider, category, scores).- Judges remain cloud-only: the
score_judgecall usescfg.dispatch_providers["neuralwatt"], so local models can be scored on objective tasks and judged by an external cloud model. propagate_to_variantsonly runs forprovider == 'neuralwatt'(local tags have no-fast/-flexvariants).- Tests added to
tests/test_local_dispatch.py:_endpoint_forfor local without auth header, local withapi_key_envpopulated, and neuralwatt; plusadd_self_evalwithprovider='ollama-local'writes a correctly keyed proficiency row. - Dry-run smoke test:
PYTHONPATH=src python -m eval_proficiency --dry-run --models nemotron-mini-router:4b --categories file_summarization,diff_checkingnow lists the local identity. - Verification:
python -m pytest tests/test_local_dispatch.py tests/test_eval_scoring.py tests/test_task_set.py -q→ 191 passed; full suite → 1014 passed. - Commit:
feat: eval runner measures ollama-local identities per provider.
Todo 9: Cost seeding — measured tariff-priced token rates for local dispatch models
- Implementation: added new standalone script
src/seed_local_dispatch_energy.py(not a mode ofseed_energy.py). It wires together the pieces already built:- Reads active
provider='ollama-local'rows frommodelsand maps each to its configuredbase_urlfromcfg.local_dispatch_models. - Guards at startup: refuses unless
cfg.local_energy.enabledandcfg.local_energy.tariff_usd_per_kwhis set; refuses if no model resolves to a loopback host. - Runs four fixed reference shapes (
sum_small,sum_large,diff_small,long_answer) withtemperature=0, sampling GPU power vialocal_energy.measure(..., sampler=sample_nvidia_smi). - Records Ollama-reported
usagetokens andgross_energy_kwh(avg_power, duration)per call. - Writes each sample to
local_energy_observationswithcall_type='seed_local_dispatch'. - Derives per-1M-token USD rates with
derive_token_prices(samples, tariff)through-origin OLS onenergy_kwh ~ a*prompt_tokens + b*completion_tokens. - Updates
modelsviaUPDATE ... WHERE model_id=? AND provider='ollama-local'. --dry-runprints the full plan without calling or writing.- Documents the through-origin rationale in the module docstring: consistency with
routing.estimated_cost()and the runtime metering window.
- Reads active
- Reference prompts:
sum_small,diff_small, andlong_answerare inline strings;sum_largeis stored insrc/sum_large_prompt.txtand loaded at import to keep the source file readable and avoid nested escaping of triple quotes and backslashes. - Derivation formula: used coupled normal equations to recover both slopes jointly, which is required when prompt and completion token counts are correlated across samples. Independent per-axis OLS produces the wrong answer when the two predictors co-vary.
- Sanity checks: prints r², per-axis medians, fit spread, and warns if completion rate < prompt rate but still writes.
- Tests: new
tests/test_seed_local_dispatch.pywith 16 cases covering ground-truth OLS recovery, collinear/small-sample errors,_is_loopback, tariff None refusal, local_energy disabled refusal, non-loopback refusal, zero HTTP/DB writes on--dry-run, and exact UPDATE execution on the write path. - Verification:
python -m pytest tests/test_seed_local_dispatch.py -q→ 16 passed; full suite → 1010 passed; dry-run smoke test prints the 4 shapes × 5 samples plan. - Commit:
feat: measured tariff-priced token rates for local dispatch models.
F2 code-quality REJECT fixes
-
Fix 1 (dispatcher double-metering bug):
src/dispatcher.pywas metering local dispatches twice in/dispatch._run_local_dispatchalready wrappedrequests.postinlocal_energy.measure(...)and_log_local_energy(...)on both success and failure paths.- The
dispatch_endpointlocal branch then created a secondlocal_energy.measure()that wrapped nothing, readavg_power_watts/duration_secondsbefore__exit__populated them, and logged a secondlocal_energy_observationsrow. - Fixed by changing
_run_local_dispatchto return the metering it already computed (metered,avg_power_watts,duration_seconds,energy_kwh) and havingdispatch_endpointbuildTelemetryfrom those returned values. When not metered it returnsTelemetry()(all None) and logs nothing extra. - Updated
tests/test_local_dispatch.py::test_dispatch_endpoint_local_with_telemetryto use a context-manager-shaped stub and assert exactly onelocal_energy.measurecall and exactly one_log_local_energycall per dispatch.
-
Fix 2 (seed script 250-LOC ceiling):
src/seed_local_dispatch_energy.pywas 449 LOC, violating the repo's 250-LOC new-module ceiling.- Split the pure derivation logic into
src/seed_local_dispatch_core.py: reference-shape prompts,REFERENCE_SHAPES, andderive_token_prices(...)including the through-origin OLS rationale in the module docstring. src/seed_local_dispatch_energy.pyis now the CLI/script layer (I/O, guards, HTTP loop, DB writes).- Updated
tests/test_seed_local_dispatch.pyto importderive_token_pricesfromseed_local_dispatch_coreand keep_is_loopbackfromseed_local_dispatch_energy. - Behavior is unchanged; only file boundaries moved.
- Split the pure derivation logic into
-
Verification:
python -m pytest tests/test_local_dispatch.py tests/test_seed_local_dispatch.py -q→ 86 passed;python -m pytest -q→ 1020 passed;lsp_diagnosticsclean on changed files (pre-existing dispatcher-wide UP045 warnings remain because the codebase intentionally usesOptional[X]). -
Commit:
fix: local dispatch metering + seed module LOC ceiling.
Todo 12: Documentation sweep
- CLAUDE.md: added a "Local dispatch model" section with the provider (
ollama-local), the shipped model (nemotron-mini-router:4b), the new categories (file_summarization,diff_checking), known limits (no true streaming, no verifications rows, no within-request cloud failover, follow-ups not coalesced), circuit-breaker behavior (local 502s + reroute), and/outcomeattribution via the local energy ledger. Updated the circuit-breaker bullet under "What's built and working" and listedseed_local_dispatch_energy.pyand the poller's local upsert. - README.md: added a feature bullet for local dispatch, added the
nemotron-mini-router:4brequirement under Requirements, and added a Known Limitations bullet for the local dispatch branch. - docs/routing.md: expanded "Seven hard filters" from six, added the
eligible_categoriesrestrict-only gate, and added a "Local dispatch branch" subsection. - docs/data-model.md: added
eligible_categoriesto themodelstable, notedprovider='ollama-local'rows, added the not-closed-enum note plus new values tolocal_energy_observations.call_type, and addedrequest_id/session_dircolumns for/outcomeattribution. - docs/local-models.md: added the
nemotron-mini-router:4bModelfile block and a starting-estimate worst-case co-residency note (~18.9/24GB). - docs/evaluation.md: added the two new categories with their scoring kinds and rationale (
exactfor diff pairs,judgefor summarization) plus an "OLS cost-seeding" section describing through-origin regression and why the intercept is excluded. - AGENTS.md: added
seed_local_dispatch_energy.pyto the module map and updated thepoller.pyrole to mentionupsert_local_dispatch_models. - docs/api.md + docs/clients.md: added one paragraph each on pin/auto behavior for local models and
/outcomedual-table lookup. - Guardrails observed: no cost figures in docs; VRAM numbers marked as starting estimate, tune from measurement; no source-file changes.
- Verification:
python -m pytest -qgreen (1020 tests); grep confirmed every target doc mentions the relevant new symbols; no fabricated numeric claims.
Todo 11: /outcome attribution for local-dispatch answers via local energy ledger
- Implementation followed OPTION (b) — kept the local ledger structurally disjoint from
energy_observations:config/schema.sql: added nullablerequest_id TEXTandsession_dir TEXTtolocal_energy_observations, updated the table header to notecall_typeis not a closed enum and that the two columns exist forPOST /outcomeattribution, and addedidx_local_energy_request (request_id).src/local_energy.py: extended_TABLE_COLUMNSwith the two columns, added a guardedALTER TABLE ... ADD COLUMNloop inensure_local_energy_table(catches "duplicate column"OperationalError), and extendedlog_local_energy(...)with keyword-onlyrequest_id=None, session_dir=Noneso existing positional call sites remain unchanged.src/dispatcher.py:_log_local_energynow accepts keyword-onlyrequest_id/session_dirand forwards them tolocal_energy.log_local_energy._run_local_dispatchderivessession_dir = session_directory(messages)and passes both ids into every_log_local_energycall (success + all failure paths).- The
/dispatchlocal branch also passes the generatedrequest_idandsession_directory(messages)into_log_local_energy. - Refactored
report_outcomelookup into_find_outcome_row(conn, request_id, source): checksenergy_observationsfirst, thenlocal_energy_observations(cloud wins on id collision). ReturnsAMBIGUOUS/Noneexactly as today. _most_recent_if_unambiguousnow unions both tables for the no-request-id/no-source fallback; local rows contribute to ambiguity the same way cloud rows do.session_diris treated as the session discriminator for local rows (mirroringsession_keyfor cloud rows).
feedback.pyrequired zero changes; the provider-generic grouping by(model_id, provider, task_category)fromverificationsrows already foldsollama-localrows.- Tests added to
tests/test_local_dispatch.py(6 new cases):- Local row with
request_id+POST /outcome→ 200,verificationsrow withprovider='ollama-local',kind='client_outcome', category carried. - Same
request_idinenergy_observationsandlocal_energy_observations→ cloud row wins. sourcematchessession_diron a local row → resolves.- Unknown
request_id→ still 404. - Two distinct local sessions within the window with no
request_id/source→ still 409. - Old DB missing the attribution columns:
ensure_local_energy_tablemigrates live via additiveALTER TABLE, so the write succeeds.
- Local row with
- Test-hygiene fix: many existing
_run_local_dispatchunit tests referenced themessagesfixture inside the function body without requesting it as a parameter, so they received the fixture function object rather than the list._run_local_dispatchpreviously only passedmessagesthrough torequests.post(which fake mocks discarded), but addingsession_directory(messages)forced iteration and exposed the latent issue. Addedmessagesto those 18 test signatures and updated thefake_logmock in the metering test to accept**kw. - Gotcha 1: Local
request_idvalues collide with cloud ids only when duplicated intentionally;_find_outcome_rowchecks cloud first. - Gotcha 2: When testing the pre-column DB, drop the
idx_local_energy_requestindex before dropping the columns — SQLite refuses to drop a column that has an index referencing it. - Verification:
python -m pytest tests/test_local_dispatch.py -q→ 70 passed;python -m pytest -q→ 1020 passed. - Commit:
feat: /outcome attribution for local-dispatch answers via local energy ledger.