Files
6krrt/plans/local-dispatch-inert-and-test-coupling.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

13 KiB

Local dispatch is inert, and the tests are coupled to deployment config

Status: done -- local dispatch dormant under the default profile, by design

Status: FINAL — decision-complete except for one question only the user can answer (§1). Findings are from the first live execution of expand-local-llm-usage todo 13 on 2026-09-02/03, against a Quadro RTX 6000 at a measured tariff of $0.159/kWh.

Three of the defects this run found are already fixed and committed in 35a5939. What remains is one design question and two test-hygiene items.

1. The blocker: local dispatch can never be selected

The feature works end to end — catalog row, hard filter, metering, measured pricing — and then never fires, because it cannot win the ranking:

model file_summarization 8k+400 request
deepseek-v4-flash (cloud) 0.95 $0.001232
qwen2.5-coder-router:14b (local) 0.767 $0.000135

The gap is 0.183. objective.quality_tolerance is 0.1. Ranking is quality-first with cost as a tiebreak only inside that band, so the local model's 9.1x cost advantage is never consulted. Verified live:

POST /route {"task_category":"file_summarization","task_tier":1}
  -> selected: deepseek-v4-flash  prof=0.95

The hard filter is fine — coding_general correctly excludes the local row. The problem is purely that quality-first ranking makes the cost axis unreachable for any model more than quality_tolerance below the best.

This is not a bug. It is the documented routing philosophy ("quality is the objective and cost is the tiebreak") doing exactly what it says. The plan built the mechanism without checking that a local model could clear the bar it would be judged against.

The question for the user, and it is genuinely theirs: what is local dispatch for? Three coherent answers, with different implementations:

  • (a) It is a cost/privacy lever the user opts into. Then quality-first is the wrong rule for these categories and it needs an explicit override — e.g. a per-category prefer_local or a cost-ceiling rule that admits the cheapest candidate above a quality FLOOR rather than the best candidate within a tolerance. Do NOT implement by widening quality_tolerance globally: 0.183 would change routing in every other category too, and the tolerance exists to describe measurement noise (2-3 samples), not preference.
  • (b) It should fire only when it is genuinely competitive. Then nothing to build; the feature is correct and dormant until a local model scores within 0.1 of the cloud leader. Say so in the docs so the next reader does not re-diagnose it, and drop the eligible categories to avoid implying otherwise.
  • (c) It is for when the cloud is unavailable (out of credits, offline). Then it is a fallback, not a routing candidate, and belongs on the failover path next to circuit_breaker, not in the ranking.

DECIDED 2026-09-03 by the user: (b) AND (c). They compose cleanly — (b) governs the normal path, (c) the degraded one — so implement both:

  • (b): change nothing in ranking. Local stays an ordinary candidate that wins only if it scores within quality_tolerance of the cloud leader. Today it does not, so it is dormant under the default profile — by design. Say so in docs/ (and in the local_dispatch_models config comment) with the measured numbers, so the next reader does not spend an evening re-diagnosing a working system. Do NOT widen quality_tolerance, and do NOT strip eligible_categories — they are what make (c) safe.

    Word this as "dormant under the default profile", not "dormant". The project is moving toward named profiles (2026-09-03): default stays the feature-rich automatic router, alongside user-named model sets — "locality", "onlycheaps", "bigboybritches". Option (a) above is therefore NOT rejected, only scoped: a per-category local preference is the wrong thing to impose globally and the right thing to put inside a locality profile. A doc that says "local never routes" becomes false the day profiles land, which is exactly the staleness this section exists to prevent.

    Architectural note for whoever builds profiles: auto:batch is already one. It is a named variant of auto that widens the candidate set to admit -flex rows, so auto:locality extends a documented pattern rather than inventing a mechanism. A profile is a named HARD FILTER over candidates — the stage routing.py already owns, pure and injected — not a new ranking path. eligible_categories is the same idea from the model's side. The open design question is whether a profile is only a candidate set or also carries objective overrides (quality_tolerance, max_energy_per_request); a local-ONLY set works fine as a pure filter, but a MIXED set that prefers local needs the override, or it reproduces the inertness diagnosed here.

  • (c): add local fallback to the failover path, so an unavailable cloud degrades to a local answer instead of an error.

(c) design notes — read before implementing

Credit exhaustion is ACCOUNT-level, and that breaks the existing failover. circuit_breaker records a failure only on 5xx and fails over to the next-ranked cloud candidate. An out-of-credits response is a 4xx (402/403/429): it will not trip the breaker, it is not transient, and every cloud row shares one account — so trying the next cloud candidate is guaranteed to fail identically. Walking the whole candidate list before giving up wastes the user's latency to reach a foregone conclusion.

So the fallback trigger is NOT "5xx after retries". It is closer to: a non-retryable, account-level upstream refusal, or no cloud candidate left.

The exact signature is UNKNOWN and must not be invented. The credits ran out on 2026-09-02/03 and the router journal recorded no upstream failure at all — no non-200 /v1/chat/completions, no traceback. We therefore do not know whether NeuralWatt returns 402, 403, or 429, nor its body. First task: add logging that captures the full status line and body of any non-2xx upstream response, so the next occurrence is diagnosable. Until that signature is observed, implement the trigger conservatively (see below) rather than hard-coding a guessed status code.

Conservative trigger that does not need the signature: fall back to local when the dispatch loop has exhausted every cloud candidate, OR when the upstream returns any 4xx that is not attributable to the request itself (i.e. not 400/404/422, which mean the request was malformed and would fail locally too).

Respect eligible_categories even in fallback. It is tempting to serve any category when the cloud is down — "something beats nothing". It does not, for code: this model scored 0/4 at detecting bugs and FABRICATED on 3 of 6 summaries at 4B, and even at 14b it is a summarization-grade model. A confidently wrong code answer is worse than an honest 502. Serve the eligible categories; fail loudly for the rest.

Follow the local_vision precedent, which is the same shape. _run_local_vision / _local_vision_response already implement "cloud cannot serve this, answer locally, return a normal OpenAI-shaped response (including a stream-wrapped variant)". Reuse that structure rather than inventing a second one; CLAUDE.md documents its budget guards and its deliberate choice to fail through to the ordinary error when the local call fails, both of which apply here unchanged.

Mark the decision row. A fallback answer must be distinguishable in route_decisions from a chosen one, or the proficiency loop will read degraded-mode traffic as evidence that local won on merit. Reuse or extend the existing kind values (local_vision is the precedent) rather than silently recording it as chat.

Note the measured asymmetry, which bears on (a): local is 26x cheaper on prompt tokens ($0.0054 vs $0.14 per 1M) but only 1.2x cheaper on completions ($0.2286 vs $0.28). Since prompts dominate real traffic, local wins most on exactly the large-context summarization it is assigned — the cost case is strongest precisely where it is currently blocked.

2. Seven tests fail, and none of them is a product defect

Reproduce: pytest -n auto -q with the current config/config.yaml. With local_energy.enabled: false the count drops to 2, which isolates the two causes cleanly.

2a. Five tests assume local_energy is disabled (test-design coupling)

tests/test_metrics_endpoint.py:27 does CFG = load_config(str(ROOT / "config" / "config.yaml")) — it loads the REAL deployment config. The tests then assert:

assert data["local_energy"] is None

which only holds while the feature is off. Enabling a documented, supported, default-off feature therefore turns the suite red without any code changing.

The suite's correctness must not depend on deployment config. This is the mirror of the rule CLAUDE.md already states in the other direction ("never point config/config.yaml at test fixtures"). Fix by monkeypatching cfg.local_energy.enabled to a known value in the affected tests, or by asserting against the loaded config rather than a hard-coded None. Do NOT fix it by turning the feature back off.

Affected (7 total across both causes; confirm the exact list by running): tests/test_metrics_endpoint.py, tests/test_admin_snapshot.py.

2b. Two tests hardcode the old dispatch model name

tests/test_local_dispatch.py::test_the_shipped_config_has_one_dispatch_model asserts nemotron-mini-router:4b. The shipped model is now qwen2.5-coder-router:14b (see §3). Update the expected value.

Consider whether that test should assert a specific model id at all, or only the shape (exactly one entry, with eligible_categories set). Pinning the id means every deliberate model change is a test failure — which is arguably the point, but it should be a decision rather than an accident.

3. config/config.yaml is deliberately uncommitted

The working tree holds three changes, and they must NOT be committed as one:

change commit?
local_energy.tariff_usd_per_kwh: 0.159 NO — user's electricity bill
local_energy.enabled: true NO — user's deployment choice
dispatch/classifier/verifier model -> qwen2.5-coder-router:14b YES
tiering.model_tiers pin for the new model YES

The model selection is a project decision backed by measurement and belongs in the repo. The tariff is personal data; expand-local-llm-usage todo 13 says so explicitly. Split the file when committing.

4. Why the model changed from the plan's choice

The plan specified nemotron-mini:4b. Measurement disqualified it:

  • diff_checking: answered "no" to 8/8 tasks, and to 3/3 blatant probes including return a + b -> return a - b. A detector that never fires.
  • file_summarization: 0.25, and it FABRICATED on 3 of 6 — "falsely claims new exception", "wrongly assumes FK enforcement".

Candidates benched on file_summarization (n=6, judge-scored):

model score
deepseek-v4-flash (cloud reference) 0.95
qwen2.5-coder-router:14b 0.767
qwen3-router:8b 0.75 (n=4; 2 truncated on reasoning trace)
qwen2.5-coder-router:7b 0.667
nemotron-mini-router:4b 0.25

14b vs 7b is 0.6 of a task on n=6 — inside noise. The 14b was chosen because it avoided the FABRICATION the 7b made on the sqlite-FK task, not for the 0.10.

It also replaced mistral-nemo-router:12b as classifier and verifier, on a 14-case bench replicating CLAUDE.md's original methodology (cold load excluded — the first run did not exclude it and was misleading):

14b mistral-nemo:12b
category correct 11/14 10/14
hard failures 0/14 1/14
latency mean 1.07s 1.25s
latency MAX 1.14s 6.10s

Accuracy is a tie (one case). The tail is not: mistral-nemo still shows the documented runaway mode — 6.1s spent to return no parseable JSON, degrading that request to source: "fallback" (tier 2, general_chat).

diff_checking cannot rank models and should not be used to. It has 8 tasks over only 4 scenarios, so one task moves a score 0.125 and the entire observed spread (0.50-0.75) is two tasks. At temperature: 0 re-running is byte-identical, so more samples means more TASKS, not more runs. Either expand it to ~12 scenarios or stop treating its numbers as a ranking.

Non-goals

  • Do not widen objective.quality_tolerance to make §1 go away.
  • Do not disable local_energy to make §2a go away.
  • Do not commit the tariff.
  • Do not re-run the seed step expecting different numbers; 35a5939 fixed it and the fit was clean (r^2 = 0.9985).

Success criteria

  • §1 answered by the user and implemented, or explicitly deferred with the reason recorded in docs/ so it is not re-diagnosed.
  • pytest -n auto -q green with local_energy.enabled: true AND with it false.
  • A test proves the suite is insensitive to local_energy.enabled.
  • config.yaml's model change committed; tariff and enabled still local.
  • docs/ states that local dispatch is quality-gated and, if §1(b) is chosen, that it is dormant by design.