plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
261 lines
13 KiB
Markdown
261 lines
13 KiB
Markdown
# Local dispatch is inert, and the tests are coupled to deployment config
|
|
|
|
Status: done -- local dispatch dormant under the default profile, by design
|
|
|
|
**Status: FINAL — decision-complete except for one question only the user can
|
|
answer (§1).** Findings are from the first live execution of
|
|
`expand-local-llm-usage` todo 13 on 2026-09-02/03, against a Quadro RTX 6000
|
|
at a measured tariff of $0.159/kWh.
|
|
|
|
Three of the defects this run found are already fixed and committed in
|
|
`35a5939`. What remains is one design question and two test-hygiene items.
|
|
|
|
## 1. The blocker: local dispatch can never be selected
|
|
|
|
The feature works end to end — catalog row, hard filter, metering, measured
|
|
pricing — and then never fires, because it cannot win the ranking:
|
|
|
|
| model | `file_summarization` | 8k+400 request |
|
|
|---|---|---|
|
|
| `deepseek-v4-flash` (cloud) | **0.95** | $0.001232 |
|
|
| `qwen2.5-coder-router:14b` (local) | **0.767** | **$0.000135** |
|
|
|
|
The gap is **0.183**. `objective.quality_tolerance` is **0.1**. Ranking is
|
|
quality-first with cost as a tiebreak only *inside* that band, so the local
|
|
model's **9.1x** cost advantage is never consulted. Verified live:
|
|
|
|
```
|
|
POST /route {"task_category":"file_summarization","task_tier":1}
|
|
-> selected: deepseek-v4-flash prof=0.95
|
|
```
|
|
|
|
The hard filter is fine — `coding_general` correctly excludes the local row.
|
|
The problem is purely that quality-first ranking makes the cost axis
|
|
unreachable for any model more than `quality_tolerance` below the best.
|
|
|
|
**This is not a bug.** It is the documented routing philosophy ("quality is the
|
|
objective and cost is the tiebreak") doing exactly what it says. The plan built
|
|
the mechanism without checking that a local model could clear the bar it would
|
|
be judged against.
|
|
|
|
**The question for the user, and it is genuinely theirs:** what is local
|
|
dispatch *for*? Three coherent answers, with different implementations:
|
|
|
|
- **(a) It is a cost/privacy lever the user opts into.** Then quality-first is
|
|
the wrong rule for these categories and it needs an explicit override — e.g.
|
|
a per-category `prefer_local` or a cost-ceiling rule that admits the cheapest
|
|
candidate above a quality FLOOR rather than the best candidate within a
|
|
tolerance. Do NOT implement by widening `quality_tolerance` globally: 0.183
|
|
would change routing in every other category too, and the tolerance exists to
|
|
describe measurement noise (2-3 samples), not preference.
|
|
- **(b) It should fire only when it is genuinely competitive.** Then nothing to
|
|
build; the feature is correct and dormant until a local model scores within
|
|
0.1 of the cloud leader. Say so in the docs so the next reader does not
|
|
re-diagnose it, and drop the eligible categories to avoid implying otherwise.
|
|
- **(c) It is for when the cloud is unavailable** (out of credits, offline).
|
|
Then it is a fallback, not a routing candidate, and belongs on the failover
|
|
path next to `circuit_breaker`, not in the ranking.
|
|
|
|
**DECIDED 2026-09-03 by the user: (b) AND (c).** They compose cleanly — (b)
|
|
governs the normal path, (c) the degraded one — so implement both:
|
|
|
|
- **(b): change nothing in ranking.** Local stays an ordinary candidate that
|
|
wins only if it scores within `quality_tolerance` of the cloud leader. Today
|
|
it does not, so it is dormant **under the default profile** — by design. Say
|
|
so in `docs/` (and in the `local_dispatch_models` config comment) with the
|
|
measured numbers, so the next reader does not spend an evening re-diagnosing
|
|
a working system. Do NOT widen `quality_tolerance`, and do NOT strip
|
|
`eligible_categories` — they are what make (c) safe.
|
|
|
|
**Word this as "dormant under the default profile", not "dormant".** The
|
|
project is moving toward named profiles (2026-09-03): `default` stays the
|
|
feature-rich automatic router, alongside user-named model sets — "locality",
|
|
"onlycheaps", "bigboybritches". Option (a) above is therefore NOT rejected,
|
|
only scoped: a per-category local preference is the wrong thing to impose
|
|
globally and the right thing to put inside a `locality` profile. A doc that
|
|
says "local never routes" becomes false the day profiles land, which is
|
|
exactly the staleness this section exists to prevent.
|
|
|
|
Architectural note for whoever builds profiles: **`auto:batch` is already
|
|
one.** It is a named variant of `auto` that widens the candidate set to admit
|
|
`-flex` rows, so `auto:locality` extends a documented pattern rather than
|
|
inventing a mechanism. A profile is a named HARD FILTER over candidates —
|
|
the stage `routing.py` already owns, pure and injected — not a new ranking
|
|
path. `eligible_categories` is the same idea from the model's side. The open
|
|
design question is whether a profile is only a candidate set or also carries
|
|
objective overrides (`quality_tolerance`, `max_energy_per_request`); a
|
|
local-ONLY set works fine as a pure filter, but a MIXED set that prefers
|
|
local needs the override, or it reproduces the inertness diagnosed here.
|
|
- **(c): add local fallback to the failover path**, so an unavailable cloud
|
|
degrades to a local answer instead of an error.
|
|
|
|
### (c) design notes — read before implementing
|
|
|
|
**Credit exhaustion is ACCOUNT-level, and that breaks the existing failover.**
|
|
`circuit_breaker` records a failure only on **5xx** and fails over to the
|
|
**next-ranked cloud candidate**. An out-of-credits response is a **4xx**
|
|
(402/403/429): it will not trip the breaker, it is not transient, and every
|
|
cloud row shares one account — so trying the next cloud candidate is
|
|
guaranteed to fail identically. Walking the whole candidate list before giving
|
|
up wastes the user's latency to reach a foregone conclusion.
|
|
|
|
So the fallback trigger is NOT "5xx after retries". It is closer to:
|
|
*a non-retryable, account-level upstream refusal, or no cloud candidate left*.
|
|
|
|
**The exact signature is UNKNOWN and must not be invented.** The credits ran
|
|
out on 2026-09-02/03 and the router journal recorded no upstream failure at
|
|
all — no non-200 `/v1/chat/completions`, no traceback. We therefore do not know
|
|
whether NeuralWatt returns 402, 403, or 429, nor its body. **First task: add
|
|
logging that captures the full status line and body of any non-2xx upstream
|
|
response**, so the next occurrence is diagnosable. Until that signature is
|
|
observed, implement the trigger conservatively (see below) rather than
|
|
hard-coding a guessed status code.
|
|
|
|
Conservative trigger that does not need the signature: fall back to local when
|
|
the dispatch loop has exhausted every cloud candidate, OR when the upstream
|
|
returns any 4xx that is not attributable to the request itself (i.e. not
|
|
400/404/422, which mean the request was malformed and would fail locally too).
|
|
|
|
**Respect `eligible_categories` even in fallback.** It is tempting to serve any
|
|
category when the cloud is down — "something beats nothing". It does not, for
|
|
code: this model scored **0/4** at detecting bugs and FABRICATED on 3 of 6
|
|
summaries at 4B, and even at 14b it is a summarization-grade model. A
|
|
confidently wrong code answer is worse than an honest 502. Serve the eligible
|
|
categories; fail loudly for the rest.
|
|
|
|
**Follow the `local_vision` precedent, which is the same shape.**
|
|
`_run_local_vision` / `_local_vision_response` already implement "cloud cannot
|
|
serve this, answer locally, return a normal OpenAI-shaped response (including a
|
|
stream-wrapped variant)". Reuse that structure rather than inventing a second
|
|
one; `CLAUDE.md` documents its budget guards and its deliberate choice to fail
|
|
through to the ordinary error when the local call fails, both of which apply
|
|
here unchanged.
|
|
|
|
**Mark the decision row.** A fallback answer must be distinguishable in
|
|
`route_decisions` from a chosen one, or the proficiency loop will read
|
|
degraded-mode traffic as evidence that local won on merit. Reuse or extend the
|
|
existing `kind` values (`local_vision` is the precedent) rather than silently
|
|
recording it as `chat`.
|
|
|
|
Note the measured asymmetry, which bears on (a): local is **26x** cheaper on
|
|
prompt tokens ($0.0054 vs $0.14 per 1M) but only **1.2x** cheaper on
|
|
completions ($0.2286 vs $0.28). Since prompts dominate real traffic, local wins
|
|
most on exactly the large-context summarization it is assigned — the cost case
|
|
is strongest precisely where it is currently blocked.
|
|
|
|
## 2. Seven tests fail, and none of them is a product defect
|
|
|
|
Reproduce: `pytest -n auto -q` with the current `config/config.yaml`.
|
|
With `local_energy.enabled: false` the count drops to 2, which isolates the
|
|
two causes cleanly.
|
|
|
|
### 2a. Five tests assume `local_energy` is disabled (test-design coupling)
|
|
|
|
`tests/test_metrics_endpoint.py:27` does
|
|
`CFG = load_config(str(ROOT / "config" / "config.yaml"))` — it loads the REAL
|
|
deployment config. The tests then assert:
|
|
|
|
```python
|
|
assert data["local_energy"] is None
|
|
```
|
|
|
|
which only holds while the feature is off. Enabling a documented, supported,
|
|
default-off feature therefore turns the suite red without any code changing.
|
|
|
|
**The suite's correctness must not depend on deployment config.** This is the
|
|
mirror of the rule `CLAUDE.md` already states in the other direction ("never
|
|
point `config/config.yaml` at test fixtures"). Fix by monkeypatching
|
|
`cfg.local_energy.enabled` to a known value in the affected tests, or by
|
|
asserting against the loaded config rather than a hard-coded `None`. Do NOT fix
|
|
it by turning the feature back off.
|
|
|
|
Affected (7 total across both causes; confirm the exact list by running):
|
|
`tests/test_metrics_endpoint.py`, `tests/test_admin_snapshot.py`.
|
|
|
|
### 2b. Two tests hardcode the old dispatch model name
|
|
|
|
`tests/test_local_dispatch.py::test_the_shipped_config_has_one_dispatch_model`
|
|
asserts `nemotron-mini-router:4b`. The shipped model is now
|
|
`qwen2.5-coder-router:14b` (see §3). Update the expected value.
|
|
|
|
Consider whether that test should assert a specific model id at all, or only
|
|
the *shape* (exactly one entry, with eligible_categories set). Pinning the id
|
|
means every deliberate model change is a test failure — which is arguably the
|
|
point, but it should be a decision rather than an accident.
|
|
|
|
## 3. `config/config.yaml` is deliberately uncommitted
|
|
|
|
The working tree holds three changes, and they must NOT be committed as one:
|
|
|
|
| change | commit? |
|
|
|---|---|
|
|
| `local_energy.tariff_usd_per_kwh: 0.159` | **NO** — user's electricity bill |
|
|
| `local_energy.enabled: true` | **NO** — user's deployment choice |
|
|
| dispatch/classifier/verifier model -> `qwen2.5-coder-router:14b` | **YES** |
|
|
| `tiering.model_tiers` pin for the new model | **YES** |
|
|
|
|
The model selection is a project decision backed by measurement and belongs in
|
|
the repo. The tariff is personal data; `expand-local-llm-usage` todo 13 says so
|
|
explicitly. Split the file when committing.
|
|
|
|
## 4. Why the model changed from the plan's choice
|
|
|
|
The plan specified `nemotron-mini:4b`. Measurement disqualified it:
|
|
|
|
- `diff_checking`: answered "no" to **8/8** tasks, and to **3/3** blatant probes
|
|
including `return a + b` -> `return a - b`. A detector that never fires.
|
|
- `file_summarization`: **0.25**, and it FABRICATED on 3 of 6 — "falsely claims
|
|
new exception", "wrongly assumes FK enforcement".
|
|
|
|
Candidates benched on `file_summarization` (n=6, judge-scored):
|
|
|
|
| model | score |
|
|
|---|---|
|
|
| `deepseek-v4-flash` (cloud reference) | 0.95 |
|
|
| `qwen2.5-coder-router:14b` | **0.767** |
|
|
| `qwen3-router:8b` | 0.75 (n=4; 2 truncated on reasoning trace) |
|
|
| `qwen2.5-coder-router:7b` | 0.667 |
|
|
| `nemotron-mini-router:4b` | 0.25 |
|
|
|
|
14b vs 7b is 0.6 of a task on n=6 — **inside noise**. The 14b was chosen because
|
|
it avoided the FABRICATION the 7b made on the sqlite-FK task, not for the 0.10.
|
|
|
|
It also replaced `mistral-nemo-router:12b` as classifier and verifier, on a
|
|
14-case bench replicating `CLAUDE.md`'s original methodology (cold load
|
|
excluded — the first run did not exclude it and was misleading):
|
|
|
|
| | 14b | mistral-nemo:12b |
|
|
|---|---|---|
|
|
| category correct | 11/14 | 10/14 |
|
|
| **hard failures** | **0/14** | **1/14** |
|
|
| latency mean | 1.07s | 1.25s |
|
|
| **latency MAX** | **1.14s** | **6.10s** |
|
|
|
|
Accuracy is a tie (one case). The tail is not: mistral-nemo still shows the
|
|
documented runaway mode — 6.1s spent to return no parseable JSON, degrading
|
|
that request to `source: "fallback"` (tier 2, `general_chat`).
|
|
|
|
**`diff_checking` cannot rank models and should not be used to.** It has 8
|
|
tasks over only 4 scenarios, so one task moves a score 0.125 and the entire
|
|
observed spread (0.50-0.75) is two tasks. At `temperature: 0` re-running is
|
|
byte-identical, so more samples means more TASKS, not more runs. Either expand
|
|
it to ~12 scenarios or stop treating its numbers as a ranking.
|
|
|
|
## Non-goals
|
|
|
|
- Do not widen `objective.quality_tolerance` to make §1 go away.
|
|
- Do not disable `local_energy` to make §2a go away.
|
|
- Do not commit the tariff.
|
|
- Do not re-run the seed step expecting different numbers; `35a5939` fixed it
|
|
and the fit was clean (r^2 = 0.9985).
|
|
|
|
## Success criteria
|
|
|
|
- §1 answered by the user and implemented, or explicitly deferred with the
|
|
reason recorded in `docs/` so it is not re-diagnosed.
|
|
- `pytest -n auto -q` green with `local_energy.enabled: true` AND with it false.
|
|
- A test proves the suite is insensitive to `local_energy.enabled`.
|
|
- `config.yaml`'s model change committed; tariff and `enabled` still local.
|
|
- `docs/` states that local dispatch is quality-gated and, if §1(b) is chosen,
|
|
that it is dormant by design.
|