Files
6krrt/plans/local-dispatch-inert-and-test-coupling.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

261 lines
13 KiB
Markdown

# Local dispatch is inert, and the tests are coupled to deployment config
Status: done -- local dispatch dormant under the default profile, by design
**Status: FINAL — decision-complete except for one question only the user can
answer (§1).** Findings are from the first live execution of
`expand-local-llm-usage` todo 13 on 2026-09-02/03, against a Quadro RTX 6000
at a measured tariff of $0.159/kWh.
Three of the defects this run found are already fixed and committed in
`35a5939`. What remains is one design question and two test-hygiene items.
## 1. The blocker: local dispatch can never be selected
The feature works end to end — catalog row, hard filter, metering, measured
pricing — and then never fires, because it cannot win the ranking:
| model | `file_summarization` | 8k+400 request |
|---|---|---|
| `deepseek-v4-flash` (cloud) | **0.95** | $0.001232 |
| `qwen2.5-coder-router:14b` (local) | **0.767** | **$0.000135** |
The gap is **0.183**. `objective.quality_tolerance` is **0.1**. Ranking is
quality-first with cost as a tiebreak only *inside* that band, so the local
model's **9.1x** cost advantage is never consulted. Verified live:
```
POST /route {"task_category":"file_summarization","task_tier":1}
-> selected: deepseek-v4-flash prof=0.95
```
The hard filter is fine — `coding_general` correctly excludes the local row.
The problem is purely that quality-first ranking makes the cost axis
unreachable for any model more than `quality_tolerance` below the best.
**This is not a bug.** It is the documented routing philosophy ("quality is the
objective and cost is the tiebreak") doing exactly what it says. The plan built
the mechanism without checking that a local model could clear the bar it would
be judged against.
**The question for the user, and it is genuinely theirs:** what is local
dispatch *for*? Three coherent answers, with different implementations:
- **(a) It is a cost/privacy lever the user opts into.** Then quality-first is
the wrong rule for these categories and it needs an explicit override — e.g.
a per-category `prefer_local` or a cost-ceiling rule that admits the cheapest
candidate above a quality FLOOR rather than the best candidate within a
tolerance. Do NOT implement by widening `quality_tolerance` globally: 0.183
would change routing in every other category too, and the tolerance exists to
describe measurement noise (2-3 samples), not preference.
- **(b) It should fire only when it is genuinely competitive.** Then nothing to
build; the feature is correct and dormant until a local model scores within
0.1 of the cloud leader. Say so in the docs so the next reader does not
re-diagnose it, and drop the eligible categories to avoid implying otherwise.
- **(c) It is for when the cloud is unavailable** (out of credits, offline).
Then it is a fallback, not a routing candidate, and belongs on the failover
path next to `circuit_breaker`, not in the ranking.
**DECIDED 2026-09-03 by the user: (b) AND (c).** They compose cleanly — (b)
governs the normal path, (c) the degraded one — so implement both:
- **(b): change nothing in ranking.** Local stays an ordinary candidate that
wins only if it scores within `quality_tolerance` of the cloud leader. Today
it does not, so it is dormant **under the default profile** — by design. Say
so in `docs/` (and in the `local_dispatch_models` config comment) with the
measured numbers, so the next reader does not spend an evening re-diagnosing
a working system. Do NOT widen `quality_tolerance`, and do NOT strip
`eligible_categories` — they are what make (c) safe.
**Word this as "dormant under the default profile", not "dormant".** The
project is moving toward named profiles (2026-09-03): `default` stays the
feature-rich automatic router, alongside user-named model sets — "locality",
"onlycheaps", "bigboybritches". Option (a) above is therefore NOT rejected,
only scoped: a per-category local preference is the wrong thing to impose
globally and the right thing to put inside a `locality` profile. A doc that
says "local never routes" becomes false the day profiles land, which is
exactly the staleness this section exists to prevent.
Architectural note for whoever builds profiles: **`auto:batch` is already
one.** It is a named variant of `auto` that widens the candidate set to admit
`-flex` rows, so `auto:locality` extends a documented pattern rather than
inventing a mechanism. A profile is a named HARD FILTER over candidates —
the stage `routing.py` already owns, pure and injected — not a new ranking
path. `eligible_categories` is the same idea from the model's side. The open
design question is whether a profile is only a candidate set or also carries
objective overrides (`quality_tolerance`, `max_energy_per_request`); a
local-ONLY set works fine as a pure filter, but a MIXED set that prefers
local needs the override, or it reproduces the inertness diagnosed here.
- **(c): add local fallback to the failover path**, so an unavailable cloud
degrades to a local answer instead of an error.
### (c) design notes — read before implementing
**Credit exhaustion is ACCOUNT-level, and that breaks the existing failover.**
`circuit_breaker` records a failure only on **5xx** and fails over to the
**next-ranked cloud candidate**. An out-of-credits response is a **4xx**
(402/403/429): it will not trip the breaker, it is not transient, and every
cloud row shares one account — so trying the next cloud candidate is
guaranteed to fail identically. Walking the whole candidate list before giving
up wastes the user's latency to reach a foregone conclusion.
So the fallback trigger is NOT "5xx after retries". It is closer to:
*a non-retryable, account-level upstream refusal, or no cloud candidate left*.
**The exact signature is UNKNOWN and must not be invented.** The credits ran
out on 2026-09-02/03 and the router journal recorded no upstream failure at
all — no non-200 `/v1/chat/completions`, no traceback. We therefore do not know
whether NeuralWatt returns 402, 403, or 429, nor its body. **First task: add
logging that captures the full status line and body of any non-2xx upstream
response**, so the next occurrence is diagnosable. Until that signature is
observed, implement the trigger conservatively (see below) rather than
hard-coding a guessed status code.
Conservative trigger that does not need the signature: fall back to local when
the dispatch loop has exhausted every cloud candidate, OR when the upstream
returns any 4xx that is not attributable to the request itself (i.e. not
400/404/422, which mean the request was malformed and would fail locally too).
**Respect `eligible_categories` even in fallback.** It is tempting to serve any
category when the cloud is down — "something beats nothing". It does not, for
code: this model scored **0/4** at detecting bugs and FABRICATED on 3 of 6
summaries at 4B, and even at 14b it is a summarization-grade model. A
confidently wrong code answer is worse than an honest 502. Serve the eligible
categories; fail loudly for the rest.
**Follow the `local_vision` precedent, which is the same shape.**
`_run_local_vision` / `_local_vision_response` already implement "cloud cannot
serve this, answer locally, return a normal OpenAI-shaped response (including a
stream-wrapped variant)". Reuse that structure rather than inventing a second
one; `CLAUDE.md` documents its budget guards and its deliberate choice to fail
through to the ordinary error when the local call fails, both of which apply
here unchanged.
**Mark the decision row.** A fallback answer must be distinguishable in
`route_decisions` from a chosen one, or the proficiency loop will read
degraded-mode traffic as evidence that local won on merit. Reuse or extend the
existing `kind` values (`local_vision` is the precedent) rather than silently
recording it as `chat`.
Note the measured asymmetry, which bears on (a): local is **26x** cheaper on
prompt tokens ($0.0054 vs $0.14 per 1M) but only **1.2x** cheaper on
completions ($0.2286 vs $0.28). Since prompts dominate real traffic, local wins
most on exactly the large-context summarization it is assigned — the cost case
is strongest precisely where it is currently blocked.
## 2. Seven tests fail, and none of them is a product defect
Reproduce: `pytest -n auto -q` with the current `config/config.yaml`.
With `local_energy.enabled: false` the count drops to 2, which isolates the
two causes cleanly.
### 2a. Five tests assume `local_energy` is disabled (test-design coupling)
`tests/test_metrics_endpoint.py:27` does
`CFG = load_config(str(ROOT / "config" / "config.yaml"))` — it loads the REAL
deployment config. The tests then assert:
```python
assert data["local_energy"] is None
```
which only holds while the feature is off. Enabling a documented, supported,
default-off feature therefore turns the suite red without any code changing.
**The suite's correctness must not depend on deployment config.** This is the
mirror of the rule `CLAUDE.md` already states in the other direction ("never
point `config/config.yaml` at test fixtures"). Fix by monkeypatching
`cfg.local_energy.enabled` to a known value in the affected tests, or by
asserting against the loaded config rather than a hard-coded `None`. Do NOT fix
it by turning the feature back off.
Affected (7 total across both causes; confirm the exact list by running):
`tests/test_metrics_endpoint.py`, `tests/test_admin_snapshot.py`.
### 2b. Two tests hardcode the old dispatch model name
`tests/test_local_dispatch.py::test_the_shipped_config_has_one_dispatch_model`
asserts `nemotron-mini-router:4b`. The shipped model is now
`qwen2.5-coder-router:14b` (see §3). Update the expected value.
Consider whether that test should assert a specific model id at all, or only
the *shape* (exactly one entry, with eligible_categories set). Pinning the id
means every deliberate model change is a test failure — which is arguably the
point, but it should be a decision rather than an accident.
## 3. `config/config.yaml` is deliberately uncommitted
The working tree holds three changes, and they must NOT be committed as one:
| change | commit? |
|---|---|
| `local_energy.tariff_usd_per_kwh: 0.159` | **NO** — user's electricity bill |
| `local_energy.enabled: true` | **NO** — user's deployment choice |
| dispatch/classifier/verifier model -> `qwen2.5-coder-router:14b` | **YES** |
| `tiering.model_tiers` pin for the new model | **YES** |
The model selection is a project decision backed by measurement and belongs in
the repo. The tariff is personal data; `expand-local-llm-usage` todo 13 says so
explicitly. Split the file when committing.
## 4. Why the model changed from the plan's choice
The plan specified `nemotron-mini:4b`. Measurement disqualified it:
- `diff_checking`: answered "no" to **8/8** tasks, and to **3/3** blatant probes
including `return a + b` -> `return a - b`. A detector that never fires.
- `file_summarization`: **0.25**, and it FABRICATED on 3 of 6 — "falsely claims
new exception", "wrongly assumes FK enforcement".
Candidates benched on `file_summarization` (n=6, judge-scored):
| model | score |
|---|---|
| `deepseek-v4-flash` (cloud reference) | 0.95 |
| `qwen2.5-coder-router:14b` | **0.767** |
| `qwen3-router:8b` | 0.75 (n=4; 2 truncated on reasoning trace) |
| `qwen2.5-coder-router:7b` | 0.667 |
| `nemotron-mini-router:4b` | 0.25 |
14b vs 7b is 0.6 of a task on n=6 — **inside noise**. The 14b was chosen because
it avoided the FABRICATION the 7b made on the sqlite-FK task, not for the 0.10.
It also replaced `mistral-nemo-router:12b` as classifier and verifier, on a
14-case bench replicating `CLAUDE.md`'s original methodology (cold load
excluded — the first run did not exclude it and was misleading):
| | 14b | mistral-nemo:12b |
|---|---|---|
| category correct | 11/14 | 10/14 |
| **hard failures** | **0/14** | **1/14** |
| latency mean | 1.07s | 1.25s |
| **latency MAX** | **1.14s** | **6.10s** |
Accuracy is a tie (one case). The tail is not: mistral-nemo still shows the
documented runaway mode — 6.1s spent to return no parseable JSON, degrading
that request to `source: "fallback"` (tier 2, `general_chat`).
**`diff_checking` cannot rank models and should not be used to.** It has 8
tasks over only 4 scenarios, so one task moves a score 0.125 and the entire
observed spread (0.50-0.75) is two tasks. At `temperature: 0` re-running is
byte-identical, so more samples means more TASKS, not more runs. Either expand
it to ~12 scenarios or stop treating its numbers as a ranking.
## Non-goals
- Do not widen `objective.quality_tolerance` to make §1 go away.
- Do not disable `local_energy` to make §2a go away.
- Do not commit the tariff.
- Do not re-run the seed step expecting different numbers; `35a5939` fixed it
and the fit was clean (r^2 = 0.9985).
## Success criteria
- §1 answered by the user and implemented, or explicitly deferred with the
reason recorded in `docs/` so it is not re-diagnosed.
- `pytest -n auto -q` green with `local_energy.enabled: true` AND with it false.
- A test proves the suite is insensitive to `local_energy.enabled`.
- `config.yaml`'s model change committed; tariff and `enabled` still local.
- `docs/` states that local dispatch is quality-gated and, if §1(b) is chosen,
that it is dormant by design.