# Local dispatch is inert, and the tests are coupled to deployment config Status: done -- local dispatch dormant under the default profile, by design **Status: FINAL — decision-complete except for one question only the user can answer (§1).** Findings are from the first live execution of `expand-local-llm-usage` todo 13 on 2026-09-02/03, against a Quadro RTX 6000 at a measured tariff of $0.159/kWh. Three of the defects this run found are already fixed and committed in `35a5939`. What remains is one design question and two test-hygiene items. ## 1. The blocker: local dispatch can never be selected The feature works end to end — catalog row, hard filter, metering, measured pricing — and then never fires, because it cannot win the ranking: | model | `file_summarization` | 8k+400 request | |---|---|---| | `deepseek-v4-flash` (cloud) | **0.95** | $0.001232 | | `qwen2.5-coder-router:14b` (local) | **0.767** | **$0.000135** | The gap is **0.183**. `objective.quality_tolerance` is **0.1**. Ranking is quality-first with cost as a tiebreak only *inside* that band, so the local model's **9.1x** cost advantage is never consulted. Verified live: ``` POST /route {"task_category":"file_summarization","task_tier":1} -> selected: deepseek-v4-flash prof=0.95 ``` The hard filter is fine — `coding_general` correctly excludes the local row. The problem is purely that quality-first ranking makes the cost axis unreachable for any model more than `quality_tolerance` below the best. **This is not a bug.** It is the documented routing philosophy ("quality is the objective and cost is the tiebreak") doing exactly what it says. The plan built the mechanism without checking that a local model could clear the bar it would be judged against. **The question for the user, and it is genuinely theirs:** what is local dispatch *for*? Three coherent answers, with different implementations: - **(a) It is a cost/privacy lever the user opts into.** Then quality-first is the wrong rule for these categories and it needs an explicit override — e.g. a per-category `prefer_local` or a cost-ceiling rule that admits the cheapest candidate above a quality FLOOR rather than the best candidate within a tolerance. Do NOT implement by widening `quality_tolerance` globally: 0.183 would change routing in every other category too, and the tolerance exists to describe measurement noise (2-3 samples), not preference. - **(b) It should fire only when it is genuinely competitive.** Then nothing to build; the feature is correct and dormant until a local model scores within 0.1 of the cloud leader. Say so in the docs so the next reader does not re-diagnose it, and drop the eligible categories to avoid implying otherwise. - **(c) It is for when the cloud is unavailable** (out of credits, offline). Then it is a fallback, not a routing candidate, and belongs on the failover path next to `circuit_breaker`, not in the ranking. **DECIDED 2026-09-03 by the user: (b) AND (c).** They compose cleanly — (b) governs the normal path, (c) the degraded one — so implement both: - **(b): change nothing in ranking.** Local stays an ordinary candidate that wins only if it scores within `quality_tolerance` of the cloud leader. Today it does not, so it is dormant **under the default profile** — by design. Say so in `docs/` (and in the `local_dispatch_models` config comment) with the measured numbers, so the next reader does not spend an evening re-diagnosing a working system. Do NOT widen `quality_tolerance`, and do NOT strip `eligible_categories` — they are what make (c) safe. **Word this as "dormant under the default profile", not "dormant".** The project is moving toward named profiles (2026-09-03): `default` stays the feature-rich automatic router, alongside user-named model sets — "locality", "onlycheaps", "bigboybritches". Option (a) above is therefore NOT rejected, only scoped: a per-category local preference is the wrong thing to impose globally and the right thing to put inside a `locality` profile. A doc that says "local never routes" becomes false the day profiles land, which is exactly the staleness this section exists to prevent. Architectural note for whoever builds profiles: **`auto:batch` is already one.** It is a named variant of `auto` that widens the candidate set to admit `-flex` rows, so `auto:locality` extends a documented pattern rather than inventing a mechanism. A profile is a named HARD FILTER over candidates — the stage `routing.py` already owns, pure and injected — not a new ranking path. `eligible_categories` is the same idea from the model's side. The open design question is whether a profile is only a candidate set or also carries objective overrides (`quality_tolerance`, `max_energy_per_request`); a local-ONLY set works fine as a pure filter, but a MIXED set that prefers local needs the override, or it reproduces the inertness diagnosed here. - **(c): add local fallback to the failover path**, so an unavailable cloud degrades to a local answer instead of an error. ### (c) design notes — read before implementing **Credit exhaustion is ACCOUNT-level, and that breaks the existing failover.** `circuit_breaker` records a failure only on **5xx** and fails over to the **next-ranked cloud candidate**. An out-of-credits response is a **4xx** (402/403/429): it will not trip the breaker, it is not transient, and every cloud row shares one account — so trying the next cloud candidate is guaranteed to fail identically. Walking the whole candidate list before giving up wastes the user's latency to reach a foregone conclusion. So the fallback trigger is NOT "5xx after retries". It is closer to: *a non-retryable, account-level upstream refusal, or no cloud candidate left*. **The exact signature is UNKNOWN and must not be invented.** The credits ran out on 2026-09-02/03 and the router journal recorded no upstream failure at all — no non-200 `/v1/chat/completions`, no traceback. We therefore do not know whether NeuralWatt returns 402, 403, or 429, nor its body. **First task: add logging that captures the full status line and body of any non-2xx upstream response**, so the next occurrence is diagnosable. Until that signature is observed, implement the trigger conservatively (see below) rather than hard-coding a guessed status code. Conservative trigger that does not need the signature: fall back to local when the dispatch loop has exhausted every cloud candidate, OR when the upstream returns any 4xx that is not attributable to the request itself (i.e. not 400/404/422, which mean the request was malformed and would fail locally too). **Respect `eligible_categories` even in fallback.** It is tempting to serve any category when the cloud is down — "something beats nothing". It does not, for code: this model scored **0/4** at detecting bugs and FABRICATED on 3 of 6 summaries at 4B, and even at 14b it is a summarization-grade model. A confidently wrong code answer is worse than an honest 502. Serve the eligible categories; fail loudly for the rest. **Follow the `local_vision` precedent, which is the same shape.** `_run_local_vision` / `_local_vision_response` already implement "cloud cannot serve this, answer locally, return a normal OpenAI-shaped response (including a stream-wrapped variant)". Reuse that structure rather than inventing a second one; `CLAUDE.md` documents its budget guards and its deliberate choice to fail through to the ordinary error when the local call fails, both of which apply here unchanged. **Mark the decision row.** A fallback answer must be distinguishable in `route_decisions` from a chosen one, or the proficiency loop will read degraded-mode traffic as evidence that local won on merit. Reuse or extend the existing `kind` values (`local_vision` is the precedent) rather than silently recording it as `chat`. Note the measured asymmetry, which bears on (a): local is **26x** cheaper on prompt tokens ($0.0054 vs $0.14 per 1M) but only **1.2x** cheaper on completions ($0.2286 vs $0.28). Since prompts dominate real traffic, local wins most on exactly the large-context summarization it is assigned — the cost case is strongest precisely where it is currently blocked. ## 2. Seven tests fail, and none of them is a product defect Reproduce: `pytest -n auto -q` with the current `config/config.yaml`. With `local_energy.enabled: false` the count drops to 2, which isolates the two causes cleanly. ### 2a. Five tests assume `local_energy` is disabled (test-design coupling) `tests/test_metrics_endpoint.py:27` does `CFG = load_config(str(ROOT / "config" / "config.yaml"))` — it loads the REAL deployment config. The tests then assert: ```python assert data["local_energy"] is None ``` which only holds while the feature is off. Enabling a documented, supported, default-off feature therefore turns the suite red without any code changing. **The suite's correctness must not depend on deployment config.** This is the mirror of the rule `CLAUDE.md` already states in the other direction ("never point `config/config.yaml` at test fixtures"). Fix by monkeypatching `cfg.local_energy.enabled` to a known value in the affected tests, or by asserting against the loaded config rather than a hard-coded `None`. Do NOT fix it by turning the feature back off. Affected (7 total across both causes; confirm the exact list by running): `tests/test_metrics_endpoint.py`, `tests/test_admin_snapshot.py`. ### 2b. Two tests hardcode the old dispatch model name `tests/test_local_dispatch.py::test_the_shipped_config_has_one_dispatch_model` asserts `nemotron-mini-router:4b`. The shipped model is now `qwen2.5-coder-router:14b` (see §3). Update the expected value. Consider whether that test should assert a specific model id at all, or only the *shape* (exactly one entry, with eligible_categories set). Pinning the id means every deliberate model change is a test failure — which is arguably the point, but it should be a decision rather than an accident. ## 3. `config/config.yaml` is deliberately uncommitted The working tree holds three changes, and they must NOT be committed as one: | change | commit? | |---|---| | `local_energy.tariff_usd_per_kwh: 0.159` | **NO** — user's electricity bill | | `local_energy.enabled: true` | **NO** — user's deployment choice | | dispatch/classifier/verifier model -> `qwen2.5-coder-router:14b` | **YES** | | `tiering.model_tiers` pin for the new model | **YES** | The model selection is a project decision backed by measurement and belongs in the repo. The tariff is personal data; `expand-local-llm-usage` todo 13 says so explicitly. Split the file when committing. ## 4. Why the model changed from the plan's choice The plan specified `nemotron-mini:4b`. Measurement disqualified it: - `diff_checking`: answered "no" to **8/8** tasks, and to **3/3** blatant probes including `return a + b` -> `return a - b`. A detector that never fires. - `file_summarization`: **0.25**, and it FABRICATED on 3 of 6 — "falsely claims new exception", "wrongly assumes FK enforcement". Candidates benched on `file_summarization` (n=6, judge-scored): | model | score | |---|---| | `deepseek-v4-flash` (cloud reference) | 0.95 | | `qwen2.5-coder-router:14b` | **0.767** | | `qwen3-router:8b` | 0.75 (n=4; 2 truncated on reasoning trace) | | `qwen2.5-coder-router:7b` | 0.667 | | `nemotron-mini-router:4b` | 0.25 | 14b vs 7b is 0.6 of a task on n=6 — **inside noise**. The 14b was chosen because it avoided the FABRICATION the 7b made on the sqlite-FK task, not for the 0.10. It also replaced `mistral-nemo-router:12b` as classifier and verifier, on a 14-case bench replicating `CLAUDE.md`'s original methodology (cold load excluded — the first run did not exclude it and was misleading): | | 14b | mistral-nemo:12b | |---|---|---| | category correct | 11/14 | 10/14 | | **hard failures** | **0/14** | **1/14** | | latency mean | 1.07s | 1.25s | | **latency MAX** | **1.14s** | **6.10s** | Accuracy is a tie (one case). The tail is not: mistral-nemo still shows the documented runaway mode — 6.1s spent to return no parseable JSON, degrading that request to `source: "fallback"` (tier 2, `general_chat`). **`diff_checking` cannot rank models and should not be used to.** It has 8 tasks over only 4 scenarios, so one task moves a score 0.125 and the entire observed spread (0.50-0.75) is two tasks. At `temperature: 0` re-running is byte-identical, so more samples means more TASKS, not more runs. Either expand it to ~12 scenarios or stop treating its numbers as a ranking. ## Non-goals - Do not widen `objective.quality_tolerance` to make §1 go away. - Do not disable `local_energy` to make §2a go away. - Do not commit the tariff. - Do not re-run the seed step expecting different numbers; `35a5939` fixed it and the fit was clean (r^2 = 0.9985). ## Success criteria - §1 answered by the user and implemented, or explicitly deferred with the reason recorded in `docs/` so it is not re-diagnosed. - `pytest -n auto -q` green with `local_energy.enabled: true` AND with it false. - A test proves the suite is insensitive to `local_energy.enabled`. - `config.yaml`'s model change committed; tariff and `enabled` still local. - `docs/` states that local dispatch is quality-gated and, if §1(b) is chosen, that it is dormant by design.