plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
13 KiB
Local dispatch is inert, and the tests are coupled to deployment config
Status: done -- local dispatch dormant under the default profile, by design
Status: FINAL — decision-complete except for one question only the user can
answer (§1). Findings are from the first live execution of
expand-local-llm-usage todo 13 on 2026-09-02/03, against a Quadro RTX 6000
at a measured tariff of $0.159/kWh.
Three of the defects this run found are already fixed and committed in
35a5939. What remains is one design question and two test-hygiene items.
1. The blocker: local dispatch can never be selected
The feature works end to end — catalog row, hard filter, metering, measured pricing — and then never fires, because it cannot win the ranking:
| model | file_summarization |
8k+400 request |
|---|---|---|
deepseek-v4-flash (cloud) |
0.95 | $0.001232 |
qwen2.5-coder-router:14b (local) |
0.767 | $0.000135 |
The gap is 0.183. objective.quality_tolerance is 0.1. Ranking is
quality-first with cost as a tiebreak only inside that band, so the local
model's 9.1x cost advantage is never consulted. Verified live:
POST /route {"task_category":"file_summarization","task_tier":1}
-> selected: deepseek-v4-flash prof=0.95
The hard filter is fine — coding_general correctly excludes the local row.
The problem is purely that quality-first ranking makes the cost axis
unreachable for any model more than quality_tolerance below the best.
This is not a bug. It is the documented routing philosophy ("quality is the objective and cost is the tiebreak") doing exactly what it says. The plan built the mechanism without checking that a local model could clear the bar it would be judged against.
The question for the user, and it is genuinely theirs: what is local dispatch for? Three coherent answers, with different implementations:
- (a) It is a cost/privacy lever the user opts into. Then quality-first is
the wrong rule for these categories and it needs an explicit override — e.g.
a per-category
prefer_localor a cost-ceiling rule that admits the cheapest candidate above a quality FLOOR rather than the best candidate within a tolerance. Do NOT implement by wideningquality_toleranceglobally: 0.183 would change routing in every other category too, and the tolerance exists to describe measurement noise (2-3 samples), not preference. - (b) It should fire only when it is genuinely competitive. Then nothing to build; the feature is correct and dormant until a local model scores within 0.1 of the cloud leader. Say so in the docs so the next reader does not re-diagnose it, and drop the eligible categories to avoid implying otherwise.
- (c) It is for when the cloud is unavailable (out of credits, offline).
Then it is a fallback, not a routing candidate, and belongs on the failover
path next to
circuit_breaker, not in the ranking.
DECIDED 2026-09-03 by the user: (b) AND (c). They compose cleanly — (b) governs the normal path, (c) the degraded one — so implement both:
-
(b): change nothing in ranking. Local stays an ordinary candidate that wins only if it scores within
quality_toleranceof the cloud leader. Today it does not, so it is dormant under the default profile — by design. Say so indocs/(and in thelocal_dispatch_modelsconfig comment) with the measured numbers, so the next reader does not spend an evening re-diagnosing a working system. Do NOT widenquality_tolerance, and do NOT stripeligible_categories— they are what make (c) safe.Word this as "dormant under the default profile", not "dormant". The project is moving toward named profiles (2026-09-03):
defaultstays the feature-rich automatic router, alongside user-named model sets — "locality", "onlycheaps", "bigboybritches". Option (a) above is therefore NOT rejected, only scoped: a per-category local preference is the wrong thing to impose globally and the right thing to put inside alocalityprofile. A doc that says "local never routes" becomes false the day profiles land, which is exactly the staleness this section exists to prevent.Architectural note for whoever builds profiles:
auto:batchis already one. It is a named variant ofautothat widens the candidate set to admit-flexrows, soauto:localityextends a documented pattern rather than inventing a mechanism. A profile is a named HARD FILTER over candidates — the stagerouting.pyalready owns, pure and injected — not a new ranking path.eligible_categoriesis the same idea from the model's side. The open design question is whether a profile is only a candidate set or also carries objective overrides (quality_tolerance,max_energy_per_request); a local-ONLY set works fine as a pure filter, but a MIXED set that prefers local needs the override, or it reproduces the inertness diagnosed here. -
(c): add local fallback to the failover path, so an unavailable cloud degrades to a local answer instead of an error.
(c) design notes — read before implementing
Credit exhaustion is ACCOUNT-level, and that breaks the existing failover.
circuit_breaker records a failure only on 5xx and fails over to the
next-ranked cloud candidate. An out-of-credits response is a 4xx
(402/403/429): it will not trip the breaker, it is not transient, and every
cloud row shares one account — so trying the next cloud candidate is
guaranteed to fail identically. Walking the whole candidate list before giving
up wastes the user's latency to reach a foregone conclusion.
So the fallback trigger is NOT "5xx after retries". It is closer to: a non-retryable, account-level upstream refusal, or no cloud candidate left.
The exact signature is UNKNOWN and must not be invented. The credits ran
out on 2026-09-02/03 and the router journal recorded no upstream failure at
all — no non-200 /v1/chat/completions, no traceback. We therefore do not know
whether NeuralWatt returns 402, 403, or 429, nor its body. First task: add
logging that captures the full status line and body of any non-2xx upstream
response, so the next occurrence is diagnosable. Until that signature is
observed, implement the trigger conservatively (see below) rather than
hard-coding a guessed status code.
Conservative trigger that does not need the signature: fall back to local when the dispatch loop has exhausted every cloud candidate, OR when the upstream returns any 4xx that is not attributable to the request itself (i.e. not 400/404/422, which mean the request was malformed and would fail locally too).
Respect eligible_categories even in fallback. It is tempting to serve any
category when the cloud is down — "something beats nothing". It does not, for
code: this model scored 0/4 at detecting bugs and FABRICATED on 3 of 6
summaries at 4B, and even at 14b it is a summarization-grade model. A
confidently wrong code answer is worse than an honest 502. Serve the eligible
categories; fail loudly for the rest.
Follow the local_vision precedent, which is the same shape.
_run_local_vision / _local_vision_response already implement "cloud cannot
serve this, answer locally, return a normal OpenAI-shaped response (including a
stream-wrapped variant)". Reuse that structure rather than inventing a second
one; CLAUDE.md documents its budget guards and its deliberate choice to fail
through to the ordinary error when the local call fails, both of which apply
here unchanged.
Mark the decision row. A fallback answer must be distinguishable in
route_decisions from a chosen one, or the proficiency loop will read
degraded-mode traffic as evidence that local won on merit. Reuse or extend the
existing kind values (local_vision is the precedent) rather than silently
recording it as chat.
Note the measured asymmetry, which bears on (a): local is 26x cheaper on prompt tokens ($0.0054 vs $0.14 per 1M) but only 1.2x cheaper on completions ($0.2286 vs $0.28). Since prompts dominate real traffic, local wins most on exactly the large-context summarization it is assigned — the cost case is strongest precisely where it is currently blocked.
2. Seven tests fail, and none of them is a product defect
Reproduce: pytest -n auto -q with the current config/config.yaml.
With local_energy.enabled: false the count drops to 2, which isolates the
two causes cleanly.
2a. Five tests assume local_energy is disabled (test-design coupling)
tests/test_metrics_endpoint.py:27 does
CFG = load_config(str(ROOT / "config" / "config.yaml")) — it loads the REAL
deployment config. The tests then assert:
assert data["local_energy"] is None
which only holds while the feature is off. Enabling a documented, supported, default-off feature therefore turns the suite red without any code changing.
The suite's correctness must not depend on deployment config. This is the
mirror of the rule CLAUDE.md already states in the other direction ("never
point config/config.yaml at test fixtures"). Fix by monkeypatching
cfg.local_energy.enabled to a known value in the affected tests, or by
asserting against the loaded config rather than a hard-coded None. Do NOT fix
it by turning the feature back off.
Affected (7 total across both causes; confirm the exact list by running):
tests/test_metrics_endpoint.py, tests/test_admin_snapshot.py.
2b. Two tests hardcode the old dispatch model name
tests/test_local_dispatch.py::test_the_shipped_config_has_one_dispatch_model
asserts nemotron-mini-router:4b. The shipped model is now
qwen2.5-coder-router:14b (see §3). Update the expected value.
Consider whether that test should assert a specific model id at all, or only the shape (exactly one entry, with eligible_categories set). Pinning the id means every deliberate model change is a test failure — which is arguably the point, but it should be a decision rather than an accident.
3. config/config.yaml is deliberately uncommitted
The working tree holds three changes, and they must NOT be committed as one:
| change | commit? |
|---|---|
local_energy.tariff_usd_per_kwh: 0.159 |
NO — user's electricity bill |
local_energy.enabled: true |
NO — user's deployment choice |
dispatch/classifier/verifier model -> qwen2.5-coder-router:14b |
YES |
tiering.model_tiers pin for the new model |
YES |
The model selection is a project decision backed by measurement and belongs in
the repo. The tariff is personal data; expand-local-llm-usage todo 13 says so
explicitly. Split the file when committing.
4. Why the model changed from the plan's choice
The plan specified nemotron-mini:4b. Measurement disqualified it:
diff_checking: answered "no" to 8/8 tasks, and to 3/3 blatant probes includingreturn a + b->return a - b. A detector that never fires.file_summarization: 0.25, and it FABRICATED on 3 of 6 — "falsely claims new exception", "wrongly assumes FK enforcement".
Candidates benched on file_summarization (n=6, judge-scored):
| model | score |
|---|---|
deepseek-v4-flash (cloud reference) |
0.95 |
qwen2.5-coder-router:14b |
0.767 |
qwen3-router:8b |
0.75 (n=4; 2 truncated on reasoning trace) |
qwen2.5-coder-router:7b |
0.667 |
nemotron-mini-router:4b |
0.25 |
14b vs 7b is 0.6 of a task on n=6 — inside noise. The 14b was chosen because it avoided the FABRICATION the 7b made on the sqlite-FK task, not for the 0.10.
It also replaced mistral-nemo-router:12b as classifier and verifier, on a
14-case bench replicating CLAUDE.md's original methodology (cold load
excluded — the first run did not exclude it and was misleading):
| 14b | mistral-nemo:12b | |
|---|---|---|
| category correct | 11/14 | 10/14 |
| hard failures | 0/14 | 1/14 |
| latency mean | 1.07s | 1.25s |
| latency MAX | 1.14s | 6.10s |
Accuracy is a tie (one case). The tail is not: mistral-nemo still shows the
documented runaway mode — 6.1s spent to return no parseable JSON, degrading
that request to source: "fallback" (tier 2, general_chat).
diff_checking cannot rank models and should not be used to. It has 8
tasks over only 4 scenarios, so one task moves a score 0.125 and the entire
observed spread (0.50-0.75) is two tasks. At temperature: 0 re-running is
byte-identical, so more samples means more TASKS, not more runs. Either expand
it to ~12 scenarios or stop treating its numbers as a ranking.
Non-goals
- Do not widen
objective.quality_toleranceto make §1 go away. - Do not disable
local_energyto make §2a go away. - Do not commit the tariff.
- Do not re-run the seed step expecting different numbers;
35a5939fixed it and the fit was clean (r^2 = 0.9985).
Success criteria
- §1 answered by the user and implemented, or explicitly deferred with the
reason recorded in
docs/so it is not re-diagnosed. pytest -n auto -qgreen withlocal_energy.enabled: trueAND with it false.- A test proves the suite is insensitive to
local_energy.enabled. config.yaml's model change committed; tariff andenabledstill local.docs/states that local dispatch is quality-gated and, if §1(b) is chosen, that it is dormant by design.