# Local LLM Model Router — project brief `README.md` is the concise front door. Deep-dive module-by-module reference docs live in **`docs/`**. `design/local-llm-model-router.md` holds the architecture and rationale, including parts still unbuilt. This file is the working state + immediate next steps, and is the one to trust on what is currently true. ## NORTH STAR GUIDELINES Four rules that outrank local cleverness. Each exists because it was broken first and the breakage was expensive to find. When a change conflicts with one of these, the change is wrong — not the rule. ### 1. Every config knob is reachable from the admin page Every scalar knob in `config.yaml` gets an admin control — runtime, persisted, or both as its mechanism warrants — **or** a recorded decision saying why it must not. The absence of a control has to be a decision someone made, never an oversight nobody noticed. `objective.credit_attenuation.enabled` is the model exception: deliberately absent, because enabling it must be a config edit plus a restart. That is a recorded choice, not a gap. The classifier section is now fully accounted for under the gate. Seven classifier scalars (`context_framing`, `cooldown_seconds`, `fallback_tier`, `fallback_category`, `max_input_chars`, `degraded_warn_min`, `degraded_warn_threshold`) have a runtime control, a persisted control, or both (`fallback_category` and `max_input_chars` are persisted-only). Every other classifier scalar is either card-backed (`_CARD_BACKED_PATHS` in `src/admin.py`: mode, cloud primary/fallback, encoder, and local-decision fields) or excused in `DELIBERATELY_NOT_IN_ADMIN` (timeout, temperature, `max_output_tokens`, `encoder.tier_from_features`). Outside-gate knobs are still pending. Sections not yet reached by the portal (database, dispatch providers/settings, local dispatch models, profiles, local energy, tiers/tiering/proficiency/context, and deployment wiring inside classifier/verification/local vision) are tracked in `plans/deferred-knobs.md`. The first control added under any of those sections will drag every scalar under it into scope at once, per the coverage test's clause 1. **Why:** Wave 2 shipped `incumbent_cache_pricing` and `incumbent_challenger_cache_rate` with no control at all. Nobody decided that; it just never came up. The dial's entire purpose is tuning from neutral to full *without reverting code* — and hand-editing a tracked file and restarting is a loop nobody walks, so the knob's rationale evaporated on contact with reality. Enforced by `tests/test_admin_knob_coverage.py`, which fails naming the knob. Honour it; do not add a `DELIBERATELY_NOT_IN_ADMIN` entry to silence it unless the reason is true. ### 2. Fix classifier latency by improving the classifier, never by reusing a stale decision Classification latency is a real problem and the answer is a faster or better classifier — a smaller model, a warmer process, a cheaper backend. The answer is **never** to let one classification stand in for later, different work. A stale label does not merely add noise. It silently redirects money: the whole point of this router is sending each task to the model that fits *that task*, and a replayed label routes the next task to whatever fitted the last one. **Why:** the session cache reached **96.6% of all classifications**. One real classification drove **107 consecutive turns** across 14 minutes on 2026-09-16. It was introduced to avoid paying a 2-5s round trip per turn, which was a reasonable trade in isolation and became the dominant path without anyone choosing that. Two costs, and the second is worse than the first: - **Routing** decides on what the session was doing when it started, not what this turn is. A mid-session pivot — docs, then bug fixes — routes the bug fixes on the docs label. - **Proficiency** is trained on those labels, because `POST /outcome` attributes to `(model, task_category)`. That is part of how `deepseek-v4-flash` came to hold a saturated 1.000 on `file_summarization` while failing 50% of them on real traffic. It also made a working classifier look broken: replaying a handful of session labels across hundreds of turns produced "pages and pages of `diff_checking`", which read as a classifier stuck on one label when the underlying classifications were a reasonable mix. If latency forces a cache, that is a **measured, time-boxed concession with an expiry condition written next to it** — not a default. ### 3. Check worktrees and other agents' work before touching any file Before editing, run `git worktree list` and check for other agents or sessions working in the same tree. Do not assume a worktree is yours. When parallel work is unavoidable, isolate it — a separate worktree per agent — and commit by explicit path, never `git add -A`. **Why:** two agents were once launched into the *same* worktree while each was told the other was elsewhere. It survived only because the files were disjoint and one of them committed by explicit path rather than sweeping the tree. The second agent's test count was also measured against a tree containing the first's uncommitted work, so its "verified green" was not a clean signal. Related: a `router.db` inside a worktree is a **stale copy**, not the live database. The live one is `/home/alee/Sources/6krrt/router.db`. Check mtimes before measuring anything, and open it read-only: `sqlite3.connect("file:...?mode=ro", uri=True)` — the `sqlite3` CLI here does not accept `-uri`. ### 4. Waste is surfaced in the admin portal, and stopping it is one click When the router can see money being wasted, it shows the operator where they already look, with the evidence and the lever to stop it side by side. A detector that only writes a log line, or a fix that needs a config edit, a restart or a long table scan, has not met this rule. "Waste" here means **spend with no concrete change landing**: an agent session looping, re-reading, or retrying the same failure, and a model that keeps producing such sessions. It does **not** mean steady spend. A healthy agent run can burn for hours, and a spend-rate alarm cannot tell the two apart. In practice: - **Stalled sessions are visible** in the portal with their evidence: turns and $ since the last landed change, and the top repeated target. They also reach the operator when nobody is watching (desktop alert). - **A model that keeps producing them is visible** as a per-model rollup, with **Block model** beside the evidence. The block records its reason, shows in a Blocked list, and is one click to undo. - **Automatic responses come after the visible one,** never instead of it, and every automatic action shows up in the same place. **Why:** incident #8 (`docs/incidents.md`). Agent sessions looped for hours on 2026-09-25/26: one worker read the same file 61 times, a planner re-read a spec 68 times its length, and a model confabulated truncation that was not there. Every existing check stayed green. It was caught only by a human, or a Claude session, reading opencode's session store by hand. Pulling the model took a trip through a dropdown whose vocabulary is catalog `deprecated`. The detection existed nowhere, and the lever existed only for someone who already knew. ## What this is A router that uses a local model (served via Ollama) to classify incoming coding/documentation tasks — category, tier, required context size — and dispatch each task to the best-fit open-weight model on **Neuralwatt Cloud**, under a per-request cost ceiling, tiebroken by price, and ranked by category-level expected pass rate. Every measurement in this file was taken on one deployment against one provider account. They are recorded because the reasoning is worth more than the conclusion, but treat them as observations with a date on them, not as constants — the catalog, prices, grid intensity and pool load all move. When a number here decides something, re-run the measurement before trusting it. Neuralwatt remains the primary provider. OpenRouter is back as an opt-in provider gated by the `provider_model_allowlist` table — a provider with `require_allowlist: true` is ignored unless its model_id is explicitly allowlisted, and any previously-upserted row that drops off the list is marked `deprecated` on the next poll. The `provider` column and the `(model_id, provider)` primary key are unchanged, so this adds no migration. ## Stack - Python (chosen over Rust — this is I/O-bound against provider APIs, not CPU-bound; iteration speed on the scoring/weighting logic matters more than raw execution speed at this scale) - SQLite for the decision table - Ollama for local classification, via any OpenAI-compatible endpoint — `localhost:11434/v1`, or an Ollama on another machine across a VPN (`classifier.base_url`) - FastAPI for the dispatcher service ## Billing is per-kWh, not per-token — and neither is what scoring uses **Measured against the live API, 2026-08-11.** Neuralwatt bills a flat **$8.00 per kWh** and the catalog's `input_per_million` / `output_per_million` prices are not what this account is charged. Confirmed across five models; `cost_usd / energy_kwh` came back 8.00 every time: | model | list $/1M out | completion tokens | billed USD | kWh | $/kWh | |---|---|---|---|---|---| | deepseek-v4-flash | 0.28 | 600 | 5.00e-06 | 5.88e-07 | 8.50\* | | gemma-4-31b | 0.42 | 540 | 3.80e-04 | 4.75e-05 | 8.00 | | qwen3.6-35b-fast | 1.15 | 600 | 2.53e-04 | 3.16e-05 | 8.01 | | kimi-k2.7-code-fast | 4.00 | 600 | 2.25e-03 | 2.81e-04 | 8.00 | | kimi-k3-fast | 15.00 | 539 | 2.17e-04 | 2.71e-05 | 8.00 | \* rounding — billed cost is quantized to ~1e-06. The precise rule, validated against all 65 samples of the reference sweep (61/65 within 2%; the 4 outliers are microdollar rounding, not misses): ``` cost_usd = min( $8.00/kWh x energy_kwh , 3 x list token price ) ``` The ceiling bound in only 2 of 65 samples, both `deepseek-v4-flash` energy spikes, and matched to the cent: 31 prompt x $0.14/1M + 400 completion x $0.28/1M = 1.1634e-04, x3 = 3.4902e-04, billed 0.000349. **List price ranks models backwards.** Not approximately — invertedly: | | list $/1M | actually billed | gCO2eq | |---|---|---|---| | `deepseek-v4-flash` | **0.28** | 9.80e-05 | 8.90e-04 | | `gemma-4-31b` | 0.42 | **5.02e-05** | **2.32e-04** | deepseek lists 33% cheaper, costs 95% more, and emits 284% more carbon. **But cost and eco are NOT the same axis** — the tempting simplification, and it is wrong. Cost tracks energy, but carbon is energy x the serving region's grid intensity, and models run in different regions: | grid | gCO2/kWh | models | |---|---|---| | `FI` | ~49-50 | most of the catalog; varies by time | | `FI` (reported) | 475 | `glm-5.2-fast`, `glm-5.2-flex` | | `US-MIDA-PJM` | ~442 | the `kimi-k3` family | A 13.6x spread, so the two axes disagree: `glm-5.2-fast` is the 2nd cheapest model and only the 6th cleanest; `kimi-k3-flex` draws 3.7x *less* energy than `kimi-k2.7-code` while emitting 3.6x *more* carbon. Weighting them separately is load-bearing, and `tests/test_routing.py` pins it. **Superseded — cost no longer comes from the sweep at all.** `cost` was the *median* measured USD over the reference sweep. That was measured to be WRONG for real traffic, because the reference workload is the wrong shape. The sweep sends a 400-token prompt with a 400-token completion. Real agent traffic is a 150,000-token prompt with a ~400-token completion and ~92% cache hits (token-weighted over 50 sessions and 40.7M tokens on 2026-08-23; the figure was 84% when measured on 2.2M tokens of earlier traffic). The attribution ratio moves with prompt size, so the ranking inverts: | workload | winner | |---|---| | reference sweep (400/400) | `glm-5.2-fast`, 3.2x cheaper | | realistic (70k prompt, short answer) | **`deepseek-v4-flash`, 5.0x cheaper** | Same two models, opposite answer. `glm-5.2-fast` sits at attribution 0.006 on a toy prompt and 0.50 on a 70k one — it batches beautifully on small prompts and badly on real ones. `deepseek-v4-flash` barely moves (0.21 -> 0.25). So `routing.estimated_cost` prices each request from **catalog token prices, scaled to that request's actual shape** (prompt size, assumed completion length, `assumed_cache_rate`). List price is not what gets billed, but billing is capped at 3x list, so it tracks the real ordering and bounds it — and on the one case that was checked live it agrees with the measurement in direction and magnitude (7.8x predicted vs 5.0x measured). It is also free, needs no sweep, and refreshes whenever the poller runs. `objective.plan_kwh_per_period` is a planning figure only: per-request traffic is **never** refused for exceeding it — it **gates nothing**. Overage is billed against the account's credit balance (`allowance_remaining_usd` from the provider). The `/admin/api/snapshot` endpoint reports balance, estimated burn rate, and projected runway **per provider** inside `quota.accounts[]` — a list of per-provider billing shapes (`metered_plan`, `prepaid_credit`, `self_hosted`, or `unmetered`) with `plan`, `pool`, `burn`, and `credit` blocks as appropriate. `quota.spend` aggregates provider spend and a list-price estimate. The old flat keys and `by_provider`/`total_balance_usd` shape were removed; every consumer was updated in the same change, so there are no deprecated aliases. Three signals said `deepseek-v4-flash` — catalog token price (7.8x cheaper), NeuralWatt's own published per-request energy (~10x lower), and a live 70k measurement (5.0x cheaper). Only the 400-token benchmark disagreed. Trust the workload you actually run. `eco` still comes from the sweep's median gCO2eq, and is still not an objective. `flex_cost_multiplier` is gone: a flex row's measured cost already is its flex cost. **Open, and worth knowing:** NeuralWatt's model cards publish *gross* energy (~1.99e-04 kWh for deepseek, ~1.91e-03 for GLM), while the billed figure is gross x attribution. GLM burns roughly 7x more actual electricity per request and charges ~5x less, because far more tenants share its GPUs. Anything built on `eco` inherits that inversion — the attributed carbon figure answers "what is my share", not "what was burned". ## Energy attribution: signal that looks like noise Billed energy decomposes exactly: ``` energy_kwh = avg_power_watts x duration_seconds x attribution_ratio ``` `attribution_ratio` is the request's share of a shared multi-tenant GPU pool. Up close it looks like pure noise — eight rapid identical calls to one model spanned 20x in billed energy, correlating **+0.997** with the ratio while power and duration held steady. Two sweeps of the same 13 models with the same prompt disagreed by up to 36x. Scoring on the pre-attribution product (`power x duration`) was tried, and it is **wrong**. Across the sweep: | | spread | |---|---| | median attribution, **between** models | **750x** | | typical spread **within** one model | **1.8x** | The ratios are quantized (0.001, 0.25, 0.5, 0.75) — that is serving concurrency, a stable per-model property, not weather. A model whose GPUs carry far more concurrent requests genuinely costs less per request, and that is most of the real cost difference in the catalog: `deepseek-v4-flash` bills ~1000x under its share of pool gross. Stripping attribution discards a 750x real signal to suppress a 1.8x one. So scoring reads the attributed figures, and the **median** absorbs what noise remains. A split-half check on the 7-sample sweep (median of first three vs last four) shows that working: - **10 of 13 models agree within 1.4x** — stable enough to route on - **3 do not**: `kimi-k2.7-code-fast` (29x), `kimi-k3` (14x), `glm-5.2-flex` (2.2x). Those need more samples before their position is trustworthy. `dispatcher.gross_energy_kwh` remains as a diagnostic on the identity, not a scoring input. ### Attribution drifts across hours, so sampling must too Within about 30 minutes the billed figures reproduce (0.3-1.1x on a spot-check). Across hours they do not: between two sweeps, `deepseek-v4-flash` moved roughly 50x and `qwen3.6-35b` about 7x the other way — enough to **invert their cost ranking**. Attribution tracks pool load, and pool load tracks time of day. More samples inside one sweep does not fix this; it measures one moment more precisely. Coverage across time does. `load_candidates` already takes the median over ALL `seed_reference` rows, so repeated sweeps accumulate into a median-across-time for free — hence `llm-router-seed.timer`, which runs a small sweep every 6 hours. Until several sweeps have accumulated, treat the eco ordering as provisional. A single sweep's ranking is one sample of a moving quantity. **And none have accumulated since 6e729ad.** That commit moved `log_observation`'s trailing arguments to keyword-only without updating `seed_energy.py`, so every timer run since spent one billed completion and then died on `TypeError` — which is not a `RequestException`, so the per-sample `except` did not catch it. Fixed, and the sweep now has an offline end-to-end test, but the accumulation this section describes starts from the next run rather than from months of history. ## What's built and working - `config/schema.sql` — `models`, `proficiency`, `energy_observations`; applies cleanly (`sqlite3 router.db < config/schema.sql`). See [data-model](docs/data-model.md). - `poller.py` — fetches Neuralwatt's catalog (public, unauthenticated), normalizes, upserts, marks stale. Verified live: 14 routable models. See [data-model](docs/data-model.md). For providers with `require_allowlist: true` (`openrouter` in the base config), the poller filters the fetched catalog against `provider_model_allowlist` before upsert and prunes existing rows that are no longer on the list to `deprecated` — see the #45 OpenRouter opt-in allowlist section below. - `config/config.yaml` / `src/config.py` — weights, thresholds, provider settings, Pydantic-validated. [architecture](docs/architecture.md). - `scoring.py` — one `normalize_inverted` (cost and eco normalize identically) + the weighted composite. [routing](docs/routing.md). - `seed_energy.py` — reference task × N per model → `energy_observations` tagged `seed_reference`; makes `cost`/`eco` real. `--samples 5` = 65 calls, under a cent. [architecture](docs/architecture.md). - `tiering.py` / `tier.py` — pure tier resolver + DB pass. Why tier on `reasoning_default_enabled`, cheapness-not-ceiling, `tier1_context_max`: [routing#tiering](docs/routing.md#tiering). - `routing.py` — pure hard filters + ranking, plus the request-side capability gates (fail-closed asymmetry). [routing](docs/routing.md). - `routing.py` `rank_candidates` incumbency: prices the session's last chat model at its measured cache rate and every challenger at the `objective.incumbent_challenger_cache_rate` dial, behind a load-bearing `min(dial, incumbent_rate)` clamp; gated off by default (`incumbent_cache_pricing: false`), tunable from off to full in config. [routing#incumbency-and-cache-pricing](docs/routing.md#incumbency-and-cache-pricing). - `circuit_breaker.py` — passive availability skip on a 5xx (cooldown + backoff, clears on next success, no poller), on by default. Eval harness deliberately stays outside it (isolation): [routing#circuit-breaker](docs/routing.md#circuit-breaker--circuit_breakerpy). Covers `ollama-local` too: a local outage raises 502 on the first request and the breaker excludes the dead local row on the next one, so traffic reroutes to cloud candidates. - `dispatcher.py` — FastAPI service: `GET /health`, `POST /route` (no provider call), `POST /dispatch`, OpenAI-compatible `/v1/models` + `/v1/chat/completions`, SSE `GET /events/decisions`. [api](docs/api.md). On an account-level cloud refusal/exhaustion, eligible routed requests degrade to the local dispatch model instead of surfacing the cloud error (see the Local dispatch model section below). - `proficiency.py` / `proficiency_store.py` / `proficiency_outcome.py` — blend leaderboard + self-eval into a benchmark prior, accumulate client outcomes, and recompute expected pass rates; the only write paths to `proficiency`, so `blended_score`/`source` never drift. [architecture](docs/architecture.md). - `context_prune.py` — relevance-based stage trimming only tool results once over `budget_tokens`, before any paid token ships. See [pinch](docs/pinch.md) for `budget_tokens`; see also the `protected_max_chars` note there if you are changing how much prefix context is guarded. - `feedback.py` — folds `POST /outcome` client reports into `proficiency.outcome_score` via `add_outcome()`. Structural and `local_llm` verdicts are diagnostics only; `POST /outcome` is the posterior. [verification](docs/verification.md). `--dry-run` no longer just describes the fold, it **projects** it: `feedback_preview.py` copies the DB into memory, runs the real `add_outcome` against the copy, and reports the per-row before/after, opening the source `mode=ro` so a preview cannot write to what it is previewing. `deploy/llm-router-feedback.{service,timer}` gives this loop the timer it never had — and ships **not enabled**, because the fold is irreversible and the first one against an accumulated backlog is an operator decision. See `deploy/README.md`. - `exploration.py` — epsilon-greedy exploration chooser; injected RNG, no mutable state. [routing](docs/routing.md). - `seed_local_dispatch_energy.py` — standalone reference-shape sweep for `ollama-local` rows; derives per-token USD rates through the user's tariff and OLS on measured GPU draw. [architecture](docs/architecture.md). - `poller.py` — also seeds/updates `provider='ollama-local'` rows from `config.yaml` each poll so local rows stay current even when NeuralWatt is unreachable. - `logs.py` — per-request trace id (ContextVar), logfmt, journald priority prefixes; `logs.bind()` survives StreamingResponse generators. [operations](docs/operations.md). - `metrics.py` / `GET /metrics` — read-only observability; takes `(conn, cfg)`, never imports `dispatcher`. Also carries the three detectors added after the incidents below: capability sub-ceilings, the reactive rejection detector, and the classifier-degradation share. [api](docs/api.md). - `tui.py` — Textual dashboard over `/metrics` + `/events/decisions`; live feed, category→model panel, detail popup; data layer split into `tui_model.py`. The decision table leads with a `time` column and carries `profile` plus an `E` flag for exploratory picks; the quota panel now shows per-account billing shapes (`metered_plan`, `prepaid_credit`, etc.) with plan/pool/burn/credit blocks, spend aggregates, and an alarm line. [architecture](docs/architecture.md). - `tests/test_tui_schema_drift.py` — the tripwire that keeps the two honest. A new `route_decisions` column must be registered as surfaced or deliberately-not, or the test fails **naming the column**. Five columns had already reached the schema without reaching the dashboard; `ROUTE_DECISIONS_COLUMNS` in `tests/test_route_decisions.py` had itself drifted. - `tests/test_tui_warnings.py` — the same idea for warnings. Every class `/metrics` can emit must render in `#warnings-panel`, and every emitted warning must be registered — the second failing with the RAW text, because the point is that nobody knew the class existed. **Its fixture is a coupled system**: adding a seed can silence an existing class (a small-context seed once killed the escalation hazard by dragging the p95 down), which is why both directions are asserted. - `router_cli.py` — one-shot `/route` probe (no spend), raw JSON with `--json`. [api](docs/api.md). - `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — loopback `/admin` portal: dashboard, models overrides, decisions log, profiles, a read-only proficiency matrix (`GET /admin/api/proficiency`) that distinguishes a measured score from an inherited one, and controls (including the Local Compute and `classifier.mode` cards). [admin-portal](docs/admin-portal.md). Provider management includes list/detail/update/delete endpoints (`GET /admin/api/providers`, `GET /admin/api/providers/{name}`, `POST /admin/api/providers/{name}`, `DELETE /admin/api/providers/{name}`) plus per-provider allowlist endpoints (`GET/POST /admin/api/providers/{name}/allowlist`, `DELETE /admin/api/providers/{name}/allowlist/{model_id}`); the providers page shows `require_allowlist` and links to the allowlist editor. - `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md). - `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below. - `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted. - `progress_detect.py` — loop-detection signals over a window of per-session probe calls: duplicate-bulk (`dup_min`), top-similarity (`top_min`/`top_min_ro`), slow-progress (`cum_min`) and coverage (`cover_min`) heuristics, gated on `min_calls`. See [watchdog](docs/watchdog.md). - `watchdog.py` — the per-session watchdog loop: every ~5 minutes it judges each session with a tool call since the last tick, on its full history, and writes a quiet `no_opencode` tick when opencode is not running. See [watchdog](docs/watchdog.md). - `notifier.py` — alert fan-out: desktop/`notify-send` plus per-channel `min_severity` and a rate limit. See [watchdog](docs/watchdog.md). - `watchdog_store.py` — `watchdog_ticks`, `watchdog_verdicts`, `watchdog_alerts`, `watchdog_channel_settings` (four tables + indexes). See [watchdog](docs/watchdog.md). - `router-link.js` — the opencode plugin; exported as a factory with `parentCache` as a property, because opencode 1.18.x rejects the whole plugin when any export is not a function. See [watchdog](docs/watchdog.md). - `tests/` — 2414 tests across 108 files, offline, verified on Python 3.10 and 3.14. [README](README.md). - Agent guardrails — `scripts/verify_commit.py` (is this commit good?), `scripts/oc_dispatch_audit.py` (audit an orchestrator session's dispatches), and the opencode plugin `deploy/opencode-plugin/guardrails.js` whose `tool.execute.before` hook blocks rule-breaking tool calls. [agent-guardrails](docs/agent-guardrails.md). ## #45 — OpenRouter is an opt-in allowlist provider OpenRouter used to be ingested whole, then removed, and is now back — but only as an opt-in provider. The base config sets `openrouter.require_allowlist: true` and ships a short seed allowlist. Models on that seed list are upserted and kept active; anything else in the OpenRouter catalog is filtered out before upsert and any previously-active OpenRouter row that is not on the list is marked `deprecated` on the next poll. This is deliberately different from the old ingest-everything behavior. The previous approach once pulled in a non-chat model (`lyria/...`) that returned HTTP 404 on dispatch because the endpoint expected chat completions. The router had paid for the classification, selected the model, and then failed on the provider call. Allowlist-gating prevents that class of failure by default: if a model id has not been reviewed and explicitly added, the router acts as if it does not exist. The seed allowlist is short and has firm exclusions. It does NOT include: - `x-ai/*` (Grok) - `openai/*` - `anthropic/*` Those exclusions are non-negotiable. They are not "currently excluded" or planned for future inclusion; they are deliberately absent from the seed list. Adding one requires editing both the seed allowlist and this file. The admin portal exposes the allowlist under `/admin`: the providers page shows which providers require one, and each provider row links to an allowlist editor where entries can be added or removed. The underlying four API endpoints are `GET /admin/api/providers/{name}/allowlist`, `POST /admin/api/providers/{name}/allowlist`, `DELETE /admin/api/providers/{name}/allowlist/{model_id}`, and the providers page itself surfaces `require_allowlist` with a link to the allowlist editor. Neuralwatt remains the primary, ungated provider. The allowlist behavior only fires for providers with `require_allowlist: true`. ## Proficiency: category now changes routing `proficiency_score` is the ONLY category-dependent term in the ranking, so until this table had data, `task_category` could not change a decision at all — the classifier computed it, the router paid ~10s for it, and then it made no difference. It does now. Two categories were added for local dispatch: `file_summarization` and `diff_checking`; see [evaluation](docs/evaluation.md). The score is also no longer a raw benchmark level: it has been converted into an **expected pass rate on real traffic**, calibrated against 1,059 client-reported outcomes and shrunk with a 20-pseudo-observation prior so thin data does not dominate. `proficiency` now holds 141 rows: | source | count | meaning | |---|---|---| | `outcome_blended` | 33 | Fresh per-model outcome evidence | | `outcome_prior` | 76 | Trafficked-sibling rows inheriting the peer-rate prior | | `self_eval_thin` | 32 | Cold categories (summarization, translation); benchmark preserved verbatim | A further 33 (model, category) pairs have direct outcome samples. The biggest evidence gains went to `deepseek-v4-flash` (notably `coding_general`, `coding_refactor`, and `general_chat`), `kimi-k2.7-code` (across 7 categories), and `qwen3.6-35b`. Sweeping 9 categories x 3 tiers currently returns **5 distinct winners** at both 50k and 120k of context. | context | winners over 27 decisions | |---|---| | 50k | `qwen3.6-35b` (10), `gemma-4-31b` (7), `deepseek-v4-flash` (5), `kimi-k3` (3), `kimi-k3-fast` (2) | | 120k | `kimi-k2.7-code` (10), `gemma-4-31b` (7), `deepseek-v4-flash` (5), `kimi-k3` (3), `kimi-k3-fast` (2) | **This spread is recent, and how it got here is the useful part.** For a long time all 27 decisions returned ONE model, and that was the correct answer at the time rather than a bug: with cost and eco both populated, `qwen3.6-35b` was Pareto-dominant — cheapest AND cleanest in the routable set, while scoring within `quality_tolerance` of the best. No defensible weighting picks anything else out of that. Two corrections widened it, and neither was a tuning change: - **Cost stopped being a benchmark average.** It is now priced per request from catalog prices scaled to the request's shape, so the ranking depends on the workload instead of on a 400-token reference sweep that no real traffic resembles. - **Tier stopped being inferred from price.** `deepseek-v4-flash` was pinned to tier 1 for being cheap, which excluded it from every tier-2 request regardless of what any score said. A third shift is under way: the outcome backlog has been spent, so the score now reflects real pass/fail reports rather than the benchmark alone. That changes the numbers; it does not change the rule. Quality is still the objective and cost is still the tiebreak within `quality_tolerance`. Note what changes between the two rows above: only the leader, and only because of the hard context filter. That is the filter working, not the scoring disagreeing with itself. **If you see one model win everything again, check for dominance before reaching for config.** One winner is a legitimate outcome. The levers, if a genuinely different balance is wanted, are `objective.quality_tolerance` (how large an expected-success-rate gap must be before it outranks a cost saving) or `objective.max_energy_per_request` (a hard ceiling). There is no weight to tune. ### What the task set actually found **The benchmark could not discriminate these models on coding.** Every row scored exactly 1.00 on `coding_general`, `coding_refactor` and `debugging` — and that is after the tasks were deliberately hardened with touching intervals, full semver, present-but-falsy defaults, late-binding closures and a binary search that infinite-loops. Every model in this catalog is simply good at that class of problem, so cost decides coding routes, which is the right outcome. **Real traffic broke one of those ties, which the benchmark never could.** `coding_general` now spans 0.86-1.00: `glm-5.2-fast` fell to 0.862 over 29 samples folded in by `feedback.py` from an actual agent session, and crossed `self_eval_min_samples` on the way, so it reads `self_eval` rather than `self_eval_thin`. That is the intended shape of this system — the 43-task benchmark establishes a floor, and your own traffic is what refines it. `coding_refactor` and `debugging` are still flat at 1.00, awaiting the same treatment. **Benchmark-sourced hardening landed, and it broke both remaining ties.** The task set grew from 23 to 43: 8 BFCL tool tasks, 6 CRUXEval-O exact tasks (after the `score_exact` literal-eval fix), and three Exercism refactor/debug pairs. The CRUXEval-O rows split `coding_general` into a 0-1 mix across models — five of six now fail at least one model — and the Exercism refactor rows moved `coding_refactor` off its flat 1.00: bowling 0.25-1.00, dominoes 0.10-1.00, affine 0.56-1.00. The debug pairs split `debugging` the same way (bowling 0.80-1.00, dominoes a sharp 0/1 split). Two follow-ups recorded, not deleted: `debug_affine_coprime` and most of the BFCL tasks sat flat at 1.00, so they do not discriminate. **A 1.00 can also be a sampling artifact, and `docs_writing` was one.** At 2 samples per model the category read 0.70-1.00 with a model at the ceiling, and the router paid for that ceiling: `kimi-k3-fast` won every docs route. Six more benchmark passes moved every score and left NOTHING at 1.00: | model | n=2 | n=11-14 | |---|---|---| | `kimi-k3` | 0.85 | **0.973** | | `kimi-k2.7-code` | 0.85 | 0.886 | | `deepseek-v4-flash` | 0.80 | 0.864 | | `kimi-k3-fast` | **1.00** | 0.864 | | `qwen3.6-35b` | 0.85 | 0.800 | | `gemma-4-31b` | 0.85 | 0.786 | The winner moved to `kimi-k2.7-code`, **3.2x cheaper** at 50k of context ($0.0441 -> $0.0136), with no config change — `kimi-k3` scores higher but sits inside `quality_tolerance`, so cost breaks the tie. `deepseek-v4-flash` ($0.0024) misses the band by 0.009, which is the kind of margin the tolerance exists to describe rather than a verdict. **The whole spread rests on one rubric line, though.** `docs_function` is effectively saturated — 1.00 on nine of every ten samples — and nearly every `docs_gotcha` deduction is the same omission: the model documents that order is preserved, that the first occurrence is kept, and what `key` does, then never says elements must be hashable. That is real discrimination, since it is a real property of the function, but one sentence is deciding a category. Treat this ordering as thinner than n=14 makes it look. **The self-judging guard costs sample density, and it shows up here.** Most models reached n=14; `kimi-k3` and `kimi-k3-fast` reached only 11, because those two are the ones diverted to the alternate judge `qwen3.6-35b`, which returns unparseable JSON more often than `kimi-k3` does. The guard is still right — a model grading its own family is worse than a thinner sample — but the alternate judges should be picked for parseability, not just for being someone else. The reflex when a category looks flat is to reach for `quality_tolerance`. Neither tie broken so far was broken that way: `coding_general` opened up when `feedback.py` folded in real traffic, and `docs_writing` opened up on six more benchmark passes. Both were samples, not settings. `coding_refactor` and `debugging` are still flat at 1.00 on 2-3 samples each — which is now a state this project has mistaken for a measurement once. Current spread by category, widest first: | category | spread | |---|---| | `tool_use_agentic` | 0.33 - 1.00 | | `summarization` | 0.60 - 1.00 | | `reasoning_math` | 0.67 - 1.00 | | `docs_writing` | 0.66 - 0.97 | | `general_chat` | 0.80 - 1.00 | | `translation` | 0.85 - 1.00 | | `coding_general` | 0.86 - 1.00 | | `coding_refactor`, `debugging` | flat at 1.00 | **What does discriminate is tool use, arithmetic traps, and prose.** `deepseek-v4-flash` scores 1.00 on all three coding categories yet **0.33 on `tool_use_agentic`** and 0.67 on `reasoning_math`. Verified live, not an artifact: given "It is 1:20pm and my meeting starts at 3pm, how many minutes away?" — both times supplied — it calls *two* tools rather than subtracting. It over-reaches for tools, which is exactly the failure mode that matters in an agent loop. The router now avoids it for those categories while still picking it for coding. Most rows still read `source='self_eval_thin'` (118 of 132): real measurement, but below `self_eval_min_samples` at 2-3 tasks per category per run. The 14 that have crossed it are all `docs_writing`, from the six extra passes above. Two paths thicken it, and they are complementary — re-run `eval_proficiency.py` to accumulate benchmark samples, or just use the router and let `feedback.py` fold in real outcomes. Both fold into a running mean rather than replacing, so samples add up across runs. ### A score is only as fresh as the row it was copied to Proficiency is a property of the weights, not the queue, so the eval harness scores one row per family and `propagate_to_variants` copies the result onto the serving variants — `kimi-k3-flex` gets `kimi-k3`'s number, because no benchmark rates a `-flex` row separately. That copy used to happen **exactly once per variant, ever.** The guard skipped any row with `self_eval_samples > 0`, meaning "measured directly, do not overwrite" — but inheritance copies the sample count too, so after the first propagation an inherited row was indistinguishable from a measured one and was never refreshed again. `kimi-k3-flex` sat at 0.85/n=2 while `kimi-k3` moved to 0.973/n=11. `proficiency.inherited_from` records the provenance that was missing, and the **migration** was the delicate half, not the fix: `ADD COLUMN` gives every existing row NULL, which reads as "measured here", so shipping the guard alone would have permanently frozen the exact rows it exists to unfreeze. The backfill infers provenance from the harness's own selection rule rather than guessing — `eval_identities` only ever evaluates standard rows plus flex rows with **no** standard equivalent, so a flex row that has one was never a candidate for direct evaluation, whatever its sample count claims. Everything else keeps NULL, which fails safe: NULL means "do not overwrite", so no real measurement can be lost to a wrong guess. Confirmed on the live database, and on the catalog's one genuine exception — `glm-5.2` is canary, so `glm-5.2-flex` is the routable row the harness scores directly, and its NULL is correct. ### Harness bugs this shook out Three separate defects, each of which scored the rig rather than the model, and each caught by reading per-task detail rather than the summary: - **Token budget.** `max_tokens` was shared between a reasoning model's trace and its answer. At 1200, qwen3.6-35b spent ~4,200 characters thinking and returned an EMPTY content field, scoring 0.00 on tasks it can plainly do. Now 24000, clamped per model (gemma-4-31b caps at 16384), and `finish_reason: length` skips the sample instead of scoring it. - **One leading space.** kimi-k2.7-code returns `" def f(...)"`, which becomes IndentationError once the harness prepends its imports — 0.00 across all nine coding tasks for a model with "code" in its name. - **Judge failures scored as model failures.** 44% of judge calls returned unparseable output (the judge is itself a reasoning model and leaks its thinking despite `response_format`). Each was recorded as 0.0. Now the JSON is extracted from surrounding prose and an unusable reply yields no sample. `tests/test_task_set.py` exists so this stops happening: it implements a reference solution for every `code` task and asserts it passes every check, recomputes every `exact` answer (one by brute force), and confirms each refactor target already passes its own checks while each debugging target fails. It immediately caught a check where the expected value was simply wrong — which would have docked every model on a task and been indistinguishable from genuine difficulty. ## Tool competence is read from the request, not guessed at Neither local classifier can identify agentic work. Asked to label six unambiguous tool-use prompts ("read the config then update the manifest", "run the tests and fix what fails"), `qwen3.5` got 2/6 and `mistral-nemo` 1/6 — and `mistral-nemo`'s misses collapse to `general_chat`, which is also the configured `fallback_category`, so qwen3.5's crashes land in the same place. That mattered because `tool_use_agentic` has the widest proficiency spread in the table (0.33-1.00) and `deepseek-v4-flash` — the current winner on coding — sits at the bottom of it. **The fix was not a better classifier.** Whether tools are on the table is stated in the request: every agent client sends a `tools` array, and `chat_completions` never looked at it. Reading it is exact and free. It is applied as a **hard filter**, not a category override, and the distinction is load-bearing. The question is not "is this task agentic" but "can this model be trusted with tools that exist". The recorded failure is precisely the second one: `deepseek-v4-flash` was given a *non*-agentic prompt ("it is 1:20pm and my meeting is at 3pm, how many minutes away?", both times supplied) and called two tools rather than subtracting. A model that over-reaches is a hazard on every request where tools are available, whatever a classifier would have labelled the task. So `routing.min_tool_proficiency` drops any candidate whose measured `tool_use_agentic` score is below it, but only when the request carries tools: | request | winner on `coding_general` @ 50k | |---|---| | no tools | `deepseek-v4-flash` ($0.0024) | | tools present | `qwen3.6-35b` ($0.0041) | Verified live through `/v1/chat/completions` with identical bodies differing only by the `tools` array. The cost of safety here is 1.7x on that route, paid only where tools exist. **It is currently set to `null`, i.e. OFF**, deliberately and pending experiment. opencode sends `tools` on essentially every request, so with the filter on, `deepseek-v4-flash` is excluded from ordinary agent traffic and its ~7x cost advantage goes unused; with it off, that advantage applies and a model measured at 0.33 on tool use handles requests where tools are on the table. Which is right is an empirical question and the benchmark cannot answer it — the 0.33 comes from 3 tasks. What settles it is `POST /outcome`: run with the filter off, let real pass/fail reports accumulate, and compare `deepseek-v4-flash`'s `tool_use_agentic` proficiency before and after. That is the one signal here that knows whether the work actually worked, and `feedback.py` folds client outcomes in both directions, so success counts too. **Update 2026-09-15: that accumulation path is closed.** The classifier no longer emits `tool_use_agentic` — per-turn classification collapsed onto it (31 of 31 consecutive live turns, all routed to the slowest model in the catalog) — so no new outcome attributes to the category and its scores are frozen at today's values. Accepted, not fixed; the reasoning and the rejected alternatives are in [docs/routing.md](docs/routing.md), "The classifier's labels are not the proficiency scoring axis". The filter experiment now either finds a different signal or reads frozen data. 0.5 sits in the empty band between the only two values the catalog holds (0.33 and 1.00), so it is not fitted to either. A model with **no** measured tool score is unproven rather than proven bad and is not dropped — the same rule as the tier-1 context gate. Config load refuses a `routing.tool_use_category` that is not a real category, because a name matching nothing yields NULL for every row and NULL means "do not disqualify": the filter would silently stop filtering. ## Tier is an iteration budget, not just a floor A tier used to mean only "do not route below this". It now also buys corrective attempts after a verification failure: | tier | batch | interactive | |---|---|---| | 1 | 0 retries | 0 | | 2 | 1 | 1 | | 3 | 2 | 1 | Interactive is capped below its tier because every retry doubles time-to-answer, and in interactive use latency **is** a quality loss. Retries are matched to the failure, since the causes differ: - **truncated** — raise the token budget on the same model; a different one would run out too. If there is no cap to raise, the model's own output ceiling is the wall, so escalate to a candidate that can emit more. - **malformed** — more tokens will not make unparseable output parse, so escalate to the next-ranked candidate. - **ok / unverifiable** — buy nothing. Retrying `unverifiable` would burn quota across the majority of prose traffic for no signal. `escalation.preemptive_on_low_confidence` is now **off by default**. Bumping the tier because the classifier was unsure pays frontier prices before anything has gone wrong; spending after a check has actually failed is better on both mandates — the cheap attempt usually succeeds, and when it fails you have evidence rather than a hunch. ## The only ground truth: `POST /outcome` Everything else the router records is a proxy. Structural checks know whether code *parses*. The local checker guesses whether prose *looks* right. Neither knows whether the answer did the job — the client does, because it ran the tests. ```bash # id comes from the completion body, or any stream chunk curl -s localhost:8080/outcome -H 'content-type: application/json' \ -d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}' ``` Two things make this the highest-value signal available: - **It is the only quality signal that survives streaming.** A retry cannot reach a streamed response because the bytes are already gone; a report arrives afterwards and works either way. Every agent client streams. - **Its successes count.** `feedback.py` folds `ok: true` and `ok: false` into `proficiency.outcome_score` via `add_outcome()`. Structural and `local_llm` verdicts are now diagnostics only; they no longer move the score, because a structural failure is not a verified client outcome. Client outcomes are what calibrate the routing score. The benchmark is a prior; `POST /outcome` is the posterior. `request_id` is written back to `route_decisions.request_id` on both the streamed and buffered paths, so the report joins cleanly to the decision that produced it. An unknown `request_id` returns 404 rather than being quietly accepted — a client whose reports go nowhere should find out. ## Verification: what local compute is actually good for Local inference is a poor substitute for cloud completions here — 7-82x the energy, ~670x the carbon on Michigan's grid, and slower (6.2s vs 1.4-2.0s). But it is very good at stopping a cloud completion from being wasted, and completions are where all the money is: fitted on real traffic, a completion token costs **201x** a prompt token. | check | cost | what it catches | |---|---|---| | structural (`verification.py`) | **free** | truncation, malformed code/JSON/YAML, empty answers | | local LLM (Ollama) | ~$6.9e-05, ~6s | refusals, wrong-question answers, incoherence | Structural checks run on every response and never execute the code — they parse it. The local LLM check runs only on answers above `verification.min_completion_tokens`, in the background after the client has its response, because it costs ~15% of a median 193-token answer and only pays above ~600 tokens. ### An agent turn is not a prose answer Both checkers had mirror halves of the same blind spot, and real traffic is what found it. On the first genuine agent session (63 completions, a shipped feature, 349 passing tests, clean mypy and ruff): | checker | called it | actually | |---|---|---| | structural | `malformed: empty response` x29 | turns that ended in a tool call | | local LLM | `cuts off mid-sentence` x8 of 9 | same turns, judged from the other side | A turn that calls a tool has empty or half-finished text **by design**. Both paths now take `has_tool_calls` and return `unverifiable`; in `verify_response` that check outranks even `finish_reason == 'length'`, because stopping mid-sentence at a call boundary is not a budget overrun. `worth_local_check` declines outright, which also stops paying ~6s of local inference to mis-grade a tool call. Had `feedback.py` run before this, it would have applied ~12 false failures to the models that had just shipped the feature. That is the **fourth** harness bug in this project that would have scored the rig rather than the model, and the first caught by real traffic instead of a synthetic test. The pre-fix rows are kept but set `model_attributable = 0`, so the record survives without steering routing. **`client_capped` was also over-applied.** It marked *every* verdict non-attributable whenever the client set `max_tokens` — and opencode always does — so real failures were invisible to feedback. A client's token cap explains a `truncated` verdict and nothing else; it is now scoped to exactly that. `feedback.py` folds observed failures into `proficiency`, so routing learns from your traffic rather than only the 43-task benchmark. Only **failures** are folded in: a structural 'ok' means the code parsed, not that it was correct, and recording those as 1.0 would flatten every score toward the ceiling. Failures the model did not cause — a client's own tight `max_tokens` truncating the answer — are recorded but excluded. ## The classifier is the latency floor Every routed request pays a full local classification round-trip before an upstream token is requested — the classifier is the latency floor. Four settings keep it usable: `max_retries=0` (SDK retry → 3x wall-clock), `max_output_tokens: 1024` (bounds reasoning), `max_input_chars: 8000` (doesn't need the document; head+tail clamp), `fallback_tier`/`fallback_category` (mid tier, not 502/503), plus `temperature: 0` (deterministic tier). "Local" means your hardware, not this machine: bind to the VPN address, not `0.0.0.0` (Ollama has no auth). Full measurements: [docs/local-models.md](docs/local-models.md#the-classifier-is-the-latency-floor). ### A classifier failure no longer collapses to one fixed guess `fallback_tier`/`fallback_category` above is now the LAST step, not the only one. When the local classifier fails, `classify()` walks a cascade: | # | step | cost | |---|---|---| | 1 | this session's cached classification, **staleness ignored** | free | | 2 | this session's history in `route_decisions` | free | | 3 | `classifier.cloud_fallback`, if configured | money | | 4 | `fallback_tier` / `fallback_category` | free | Steps 1 and 2 reuse a *real* classification rather than guessing. Step 2 is what survives a restart, when the in-memory cache is gone but the session's decisions are still on disk. Both filter degraded sources, so one outage's guess cannot propagate through every later turn of a session and end up looking like a measurement. **Step 3 is absent by default and the whole feature costs nothing until it is configured.** Three things bound it once it is: - `classifier.cooldown_seconds` (30) is a **global** backoff, so a sustained outage buys one cloud attempt per window across ALL requests, not one per request; - an account-level provider refusal suppresses step 3 for the same window — an out-of-credit account makes that call a guaranteed wasted request; - the `session_cache.put()` gate admits `classifier_cloud`, bounding a session to one cloud call per staleness window instead of one per turn. `session_stale` and `session_history` are deliberately NOT cached: borrowed answers must not renew a staleness clock they never earned. **The trap this had to avoid, and it is the reason to read this section.** A degraded classification is recorded with `task_category = general_chat` — which is a *fully scored* category (13 proficiency rows) as well as the configured `fallback_category`. Without a guard, `POST /outcome` reports on mislabelled outage traffic fold into `proficiency(model, general_chat)` and drag real scores toward whatever happened to be flowing while the classifier was down. So `report_outcome` marks outcomes of degraded-source decisions `model_attributable = 0` — kept as a record, excluded from folding, the same treatment the tool-call false-failures got. `feedback.py` is untouched; its existing `AND model_attributable = 1` already does the work. Attributable: `classifier`, `cached`, `override`, `classifier_cloud`. Not: `fallback`, `session_stale`, `session_history`. The asymmetry is deliberate — a degraded source is good enough to route one visibly-flagged request, but a proficiency score is consulted by every future request, so attribution must not trust more than routing does. Unknown decision rows **fail open**; an over-applied exclusion already starved feedback once (`client_capped`). `/metrics` warns when the degraded share of the last 24h crosses `classifier.degraded_warn_threshold` over at least `degraded_warn_min` decisions. A survivable failure is exactly the kind that goes unnoticed for weeks. **Why a decline is now recorded, and what the numbers said (2026-10-04).** Under `local_decision` on the 4b the classifier declines about a third of agent turns on purpose: 769 of 2,557 (30%) that day, 8% the day before, against 0% for `local_llm` and `local_encoder`. 22.5% were `session_history` replays and 7.6% the static `general_chat` guess, and `session_history` has no age bound (p99 67 minutes, max 100). None of that could be tuned from the database, because `route_decisions.confidence` is the override's hard-coded 1.0 on 99.9% of chat rows. `classifier_confidence`, `classifier_coverage` and `classifier_reject` now hold the classifier's own numbers and the reason (codes in [data-model](docs/data-model.md#what-the-classifier-did-and-why-its-answer-was-not-used)), and the degraded-share warning lists the reasons it saw. The warning's default threshold (0.5) never fires at that rate, so the live overlay sets 0.2. `degraded_warn_min` and `degraded_warn_threshold` have no admin control, and that is a deferred gap, not a decision: the `classifier` section is outside the knob-coverage gate because it has its own card, and adding any `classifier.*` knob to the generic registries would pull every other classifier scalar into the gate at once. Do it in one pass when classification is reworked. ### Which implementation is PRIMARY is now a config choice `classifier.mode` in the live deployment is currently `local_encoder`, set in `config.local.yaml` with `device: cpu` on the classifier host. That is the mode answering real traffic right now. The explicit caveat is that a CPU-vs-CUDA latency and confidence comparison on the live classifier host is still pending; until that measurement exists, flipping the project default to `local_encoder` is not decided. **Turning that mode on for real found three real bugs in one afternoon (2026-09-06), each a fresh instance of this project's own recurring lesson — verify against the live system, not the plan.** First, the config-load validator for `confidence_threshold` didn't exist yet: the admin UI saved a raw `80` (meant as 80%) straight into the overlay with no conversion, which would have made every real confidence score read as below-threshold on the next restart (`classify_zero_shot` returns `[0.0, 1.0]`; no probability exceeds 1.0). Caught before the restart, not after — see `confidence_threshold` in [local-models](docs/local-models.md) for the full incident and the fix. Second, once that was corrected and the service actually restarted, it crash-looped twice more before coming up clean: the shipped default model (`MoritzLaurer/deberta-v3-base-zeroshot-v2`) had become gated on HuggingFace sometime after this project picked it (401 on an unauthenticated GET of its own model page), and separately `HF_HOME`'s default cache path falls outside this service's `ProtectHome=read-only` sandbox exception — both are now fixed (switched to `facebook/bart-large-mnli`, `HF_HOME` redirected into the repo). Third, and most consequential: the very first real classifications measured only 5 of 9 test categories correct, because `classify_zero_shot` was passing raw config identifiers like `tool_use_agentic` and `diff_checking` directly as zero-shot candidate labels — HF's pipeline scores a label against a hypothesis template ("This example is {}."), and an underscored code token is not a sentence the model's NLI training ever saw. Mapping each category to a natural- language description before scoring, plus `multi_label=True` (the pipeline's single-label default forces every candidate to compete for the same probability mass), brought that to 8 of 9 correct with confidence scores 0.77-0.999 on the hits — the one remaining miss scored below the configured threshold and correctly fell through to the safe fallback rather than mis-routing. `classifier.mode` (`local_llm` default, `cloud_llm`, `local_encoder`, `local_decision`) picks what answers a classification request — a peer concept to the cascade above, **not a replacement for it**. Whichever mode is primary, a failure still walks the exact same cascade (stale session → session history → `cloud_fallback` → the static guess), unmodified. - **`cloud_llm`** makes a cloud model the primary attempt, not just the cascade's post-failure backup. Either pin one (`classifier.cloud_primary`, same shape as `cloud_fallback`) or set `cloud_primary_auto: true` to resolve the cheapest currently-routable model **live** against the catalog (`routing.cheapest_classifier_candidate`, priced for the classifier's own short-prompt/short-completion shape — not the task's). A success here records `source="classifier"`, deliberately the *same* string `local_llm`'s success uses, not `"classifier_cloud"` — that string means specifically "the cascade's backup step fired" and feeds the `/metrics` degradation-share warning above as a degraded signal. An intentionally configured primary succeeding is not degraded. - **`local_encoder`** classifies with a small, non-generative zero-shot model instead of an LLM — structurally immune to the runaway-reasoning failure mode documented above, since there is no generation to run away. Zero-shot rather than fine-tuned: this router never stores raw task text anywhere, so there is no training corpus without a new, separate opt-in capture feature (not built). Only produces `task_category`; `task_tier` falls back to `fallback_tier` — a real limitation, not a bug. A below-threshold confidence is treated as a failure and cascades exactly like a local-LLM parse failure would. - **`local_decision`** asks a small generative model (configured in `classifier.decision`) to pick a category from a set of natural-language descriptions, then optionally classifies tier and — when enabled in its decision config — runs the same A/B confidence check on the chosen label. Requires a `classifier.decision:` block in config (confidence threshold, base URL, model name); it is an opt-in overlay mode, not the default, and the config load will refuse to start without it when the mode is selected. Only produces `task_category` and an optional `task_tier` when the decision block's `tier_classification.enabled` is true. Below-threshold confidence cascades like a parse failure would. - **Neither is gated by `local_compute.enabled`** (gaming mode, below) the way `local_llm` is: `cloud_llm` never touches local hardware, and `local_encoder` is small enough to run on CPU, so neither competes for the GPU gaming mode exists to free up. Configured via `config.yaml` (global default) and overridable per-machine in `config.local.yaml` — the existing overlay, not a new mechanism — or through the admin portal's Classifier card, which reports the **live** resolved primary for `cloud_primary_auto` rather than echoing the config value (see [admin-portal](docs/admin-portal.md)). The default in `config.yaml` is still `local_llm`. `local_encoder` is intentionally an opt-in per-deployment choice rather than the repository default until the pending CPU-vs-CUDA comparison on the live classifier host is available. `local_decision` is also opt-in: it requires a `classifier.decision:` block with its own base URL, model, and confidence thresholds — the config load refuses to start without it when the mode is selected. ## Local dispatch model A second local model can now be dispatched directly for specific categories. `qwen2.5-coder-router:14b` is configured as a tier-1 local row with `provider='ollama-local'`, gated by `models.eligible_categories` (`file_summarization` and `diff_checking`). The poller refreshes the row each run; `seed_local_dispatch_energy.py` derives its price from measured GPU draw and the user's tariff. Routing treats a local row like any other candidate once the category filter admits it, and the circuit breaker excludes it on a local failure so the next request reroutes to cloud candidates. Known limitations of the local dispatch branch right now: - **No true streaming.** The response is shaped into an SSE stream, but the local answer is generated before any bytes leave the router. - **No verification rows.** Structural and local-LLM checks run but are not written to `verifications` for local answers. - **No within-request cloud failover into the upload path is gone.** An eligible routed request degrades to the local dispatch model when the cloud account refuses or is exhausted (a fallback, not a preference), so the local row is no longer a dead-end before a client retry. The degraded answer is recorded as `kind='local_dispatch_fallback'`. - **Follow-ups are not special-cased.** A pinned or auto-routed follow-up to the same local model works, but nothing caches the loaded model between turns. Dormant under the default profile by design — see docs/routing.md § Local dispatch branch. `POST /outcome` now attributes through the local energy ledger too. Local rows include `request_id` and `session_dir` in `local_energy_observations`, so a client report on a local answer resolves to the same `(model_id, provider, task_category)` provider-agnostic record as a cloud one. ## Routing notes Ranking is quality-first, cost as a tiebreak; cost is never allowed to override a real quality gap. The optional `objective.credit_attenuation` block extends that tiebreak without changing it: when the block is enabled, a per-provider multiplier is applied to a candidate's comparison cost only, producing an `effective_cost` that breaks ties. The multiplier is derived from the provider's polled account balance (the `balance_url` path, such as OpenRouter), so a low prepaid balance can nudge a near-tie toward a healthier provider. The logged `est_cost_usd` and the decision history stay as raw catalog estimates. The multiplier is 1.0 for providers whose balance comes from per-completion `allowance_remaining_usd` telemetry (NeuralWatt), so normal overage readings do not bias routing. Two semantics matter when reading the numbers. `total_balance_usd` is a sum of heterogeneous provider-reported readings: OpenRouter's prepaid credits plus NeuralWatt's overage allowance, which normally reads near -$0.004. It can be negative and it is not a single spendable figure. `credit_attenuation.enabled` deliberately lives only in the config file; it is absent from the admin persisted-config allowlist and from provider edits. Turning it on or off requires editing `config/config.yaml` and `systemctl --user restart llm-router.service`, because the dispatcher's `cfg` binds at import time. ## What's NOT built yet — pick up here Built: session-directory attribution, the local energy ledger, local model dispatch, admin profiles/proficiency/gaming-mode, and the configurable classifier backend (`classifier.mode`, including its admin card — all listed under "What's built and working" above). **As of 2026-09-05, PRs #26-#36 landed** the gitignored config overlay, admin profile CRUD writing to it, the provider literal cleanup, quota balance/burn/runway, capability-aware ceiling and rejection warnings, the TUI schema catch-up, the classifier fallback cascade, and the admin portal uplift (proficiency page, profiles duplicate/coverage fix, gaming mode, the classifier-backoff bug fix). `classifier.mode` (this document's own section above) is a further, independent addition on top of that. `plans/multi-provider-support.md` is PARKED on provider selection — Z.ai was the recommendation and is no longer settled; the coupling surface in it is measured and still valid. Known follow-ups recorded but not specced, both small: - The TUI decision table renders the literal `"None"` in the `ctx` cell when `required_context_tokens` is absent — the same defect the `profile` cell was written to avoid. See `.omo/notepads/tui-overhaul/issues.md`. - A NULL `required_context_tokens` raises `TypeError` inside the demand-ceiling comparison, and the column is nullable. Latent only: zero such rows exist today, checked on the live DB. The items below remain open. 1. **Fine-tuning `local_encoder` on real traffic.** Scoped, not built: `classifier.training_capture.enabled` (opt-in, off by default — a deliberate reversal of "never store task text", so it must be impossible to enable by accident), a `classifier_training_samples` table gated the same way `report_outcome` already filters proficiency (only rows whose `classification_source` is in the attributable set), and a `train_local_encoder.py` script matching `eval_proficiency.py`'s conventions. `local_encoder.py` currently ships zero-shot only. 2. **Leaderboard priors are unfilled.** `leaderboards.yaml` ships empty on purpose — inventing plausible-looking benchmark numbers would put fabricated data straight into routing, the same failure as the provider's `static_fallback` carbon constant this project already excludes. Until real sourced figures go in, a newly listed NeuralWatt family has no prior and relies entirely on self-eval accumulating. `python leaderboard.py --check` lists what is missing. 3. **Sampling depth for three models — now `eco`-only.** 7 samples/model gives split-half agreement within 1.4x for 10 of 13, but `kimi-k2.7-code-fast` (29x), `kimi-k3` (14x) and `glm-5.2-flex` (2.2x) are still unsettled. This no longer touches cost, which is priced per-request from the catalog, so it only affects `eco` — which is not an objective. Low priority unless eco comes back. 4. **Retry does not reach streaming.** The iteration budget (`iteration.py`) retries after a failed check, but only on the non-streaming path — once bytes have gone to the client there is nothing to take back. Buffering to fix that would cost streaming itself, a worse trade for interactive work. `POST /outcome` is the answer for streamed traffic: it arrives afterwards, so it works identically either way. 5. **`local_encoder` noise isolation — built for the confirmed shapes; residuals below.** The raw task's fenced code blocks, `Tool result:`-shaped lines, and closed `` spans are now stripped by `_isolate_task_text` (pure, stdlib-only, deterministic) before the fit + embed pass — `classify_zero_shot` runs isolate → fit → prefix, so cleaning happens first and a noisy task often fits the token window outright. Measured on the real `BAAI/bge-large-en-v1.5` (2026-09-19): 8 clean one-sentence tasks scored 8/8, but the same instructions wrapped in that noise scored 2/8 with the truncation fix already in place — and tail-biased windowing ALONE also scored 2/8 on long noisy pairs, so the fit does not subsume isolation. Post-isolation: 8/8 on short and long noisy pairs, clean-vs-noisy pair consistency 8/8 + 8/8; `'[code]'`/`'[elided]'` placeholder tokens measured worse than pure removal (7/8, 6/8 on long pairs) and were rejected; a 20%-ratio floor guard measured harmful (3/8 — it reverts exactly the short noisy inputs isolation exists to fix) in favor of an absolute 24-char floor that only catches near-all-code inputs. Offline regression: a noise tripwire (same instruction bare vs wrapped must classify identically) fails against pre-isolation code and passes after. Still not built: attention-masking de-weighting as an alternative to stripping (it would preserve the noise tokens' presence without letting them dominate); unclosed `` tags and un-fenced diff hunks are left in place (only closed-tag spans, fenced blocks, and marker-prefixed lines are stripped); and a code-grounded instruction whose pasted snippet is the subject can still land on a near-category — measured 3/4 on a 4-task grounded set, the residual miss being description similarity ("refactor this helper" + code → `debugging`), not noise dominance. ## Gaming mode, and the backoff that used to do nothing **The classifier circuit breaker did not break the circuit.** `_last_classifier_failure` was written by `_record_failure()` and read by nothing — `_classify_cascade` gated only its *cloud* step, and on a different timestamp. So the router re-dialled a known-dead local classifier on every request. A stopped Ollama refuses immediately and costs little; a **hung** one, or a VPN-bound one that black-holes, costs the full 120s `timeout_seconds` per request for as long as the outage lasts. `_classifier_backoff_active()` is now the read, consulted **before** the client is constructed. One detail there is load-bearing: recording the failure lives in `classify()`'s exception handlers, NOT in `_classify_cascade`. The cascade is walked for reasons other than a fresh failure, and if those re-stamped the clock, every request during an outage would push the deadline forward and the local classifier would never be re-probed while traffic flowed — a permanent outage wearing a circuit breaker's clothes. **`local_compute.enabled` (default true) is the outer gate over local hardware.** Turn it off when you stop Ollama for a game and the router *skips* every local call rather than discovering the outage one timeout at a time: classifier, `/health` probe, local verification, local-vision fallback, and local dispatch rows (dropped in `load_candidates`, so a local row is never picked and then 503'd). `/v1/models` stops listing local rows, and an explicit pin gets a 503 naming the flag instead of NeuralWatt's unknown-model 400. ONE flag the code reads, **not** a macro writing five keys — a macro is hard to undo cleanly, drifts the moment a sixth call site appears, and leaves nobody able to answer "why isn't the classifier running?" from one place. `verification.local_llm_enabled`, `local_vision.enabled` and `local_energy.enabled` keep their own meanings; this ANDs over them. **It REFUSES to engage without `classifier.cloud_fallback`** — 409 on the runtime knob, a validation error at config load. Skipping the local classifier does not make classification remote; without a cloud classifier it stops classifying, and every request falls through to a static guess recorded as `general_chat`, a fully scored category indistinguishable from a real classification afterwards. A refusal, not a warning, because a warning is what nobody reads while their game is loading. Nothing auto-writes the block. Cascade steps 1 and 2 still run **ahead** of the cloud call: a stale session classification is free and was a real classification of that same session, so paying to re-derive an answer already held is spending money for nothing. **A latent substring bug fell out of the profiles work.** SQLite stores `eligible_categories` as a comma-joined string, and `task_category not in ","` is a SUBSTRING test — so a row eligible only for `file_summarization` also admitted `summarization`. Latent on main (the category-less probe short-circuits before the compare) and live the moment anything probes per category. `routing.parse_eligible_categories` is now the single parser `dispatcher.load_candidates` and `admin.py`'s probe both use. ## Known open questions - Answered: cost and eco stay separate axes — grid intensity spans 13.6x across the catalog, so they rank models differently. - Answered: the GLM rows reporting `grid_id: FI` at 475 gCO2/kWh were `carbon_source: static_fallback` — a substituted constant, not a measurement. They are now excluded from eco rather than trusted. Still worth asking NeuralWatt why the fallback keeps the original `grid_id`, since that is what made it look like a real regional difference. - Three models still fail a split-half stability check at 7 samples. Is the instability real (variable serving conditions) or an artifact of when the sweep ran? Re-sweeping at a different hour would tell. - Answered, and the question no longer parses: tier-1 composites used to sit within 0.009 of each other because min-max normalization compressed them. There is no composite any more — ranking is quality first, cost as the tiebreak inside `quality_tolerance` — so nothing normalizes and nothing compresses. - Answered: the eval set exists (`evals/tasks.yaml`, 43 tasks, four scoring kinds) and `tests/test_task_set.py` keeps it honest. The benchmark-sourced rows now split the coding categories, but two tasks are still flat at 1.00 (`debug_affine_coprime` and most of the BFCL set) and either need hardening again or should be conceded as non-discriminating. **Try samples before hardening.** `docs_writing` looked flat at the top too, and six more passes spread it 0.66-0.97 without touching a task; two samples per model is not enough to tell a saturated task from an unsampled one. - How much context-assembly (RAG-style retrieval) belongs in the classifier step vs. a separate pre-step? Leaning decoupled, undecided. - Should `eco_score` use real-time grid carbon intensity per request or a stable per-model average? Currently the latter, from the reference sweep. `grid_carbon_intensity` and `grid_id` are logged per observation, so this stays answerable from data without a re-run. ## Config is strict: an unknown key is an error Pydantic ignores extra keys by default, which means a typo or a misplaced setting loads cleanly, does nothing, and still looks configured. Every config model now inherits `StrictModel` (`extra="forbid"`), so both of these fail at load rather than silently: ``` verification.max_input_chars # right key, wrong section routing.min_tool_proficency # sic ``` This is not hypothetical. `max_input_chars` shipped into the `verification:` block instead of `classifier:` and was accepted and discarded — it happened to match the code default, so behaviour was correct and the file was a lie. Editing it would have done nothing. The corollary worth keeping: **every knob belongs in `config.yaml`, not only in a Pydantic default.** A default the file never mentions is invisible to anyone tuning it. `classifier.outcome_attribution_window_seconds` was removed in the same pass — it was declared, never read, and shadowed the `verification` one that actually is. ## Setup Full install steps (venv, deps, config, first run) in [README ## Installation](README.md#installation). Host-local deployment values go in `config/config.local.yaml` (gitignored overlay) — `classifier.model`/`base_url`, `local_energy.*`, host-specific URLs. General defaults in `config/config.yaml` stay shareable. `objective.plan_kwh_per_period` in the README config table. Model tags (`num_ctx`) + `verification.model` same-tag note in [docs/local-models.md](docs/local-models.md). Requirements are pinned — bump deliberately (README). ## Run as a service `deploy/` holds the dispatcher's systemd **user** unit plus a timer and a oneshot service each for the poller, the seed sweep, the backup, the offsite sync and the feedback fold, and one drop-in for a *system* Ollama — see `deploy/README.md` for install and operation. Every timer there is enabled on install except `llm-router-feedback.timer`, which is not, on purpose. In short: ```bash echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env cp deploy/llm-router*.{service,timer} ~/.config/systemd/user/ systemctl --user daemon-reload systemctl --user enable --now llm-router.service llm-router-poller.timer ``` The dispatcher binds `127.0.0.1:8080`. **The poller timer is load-bearing, not housekeeping** — but not for the reason this section used to give, and the correction matters because it inverts which failure to watch for. `mark_stale` runs only *inside* `poller.main()`, and `main()` returns early on a `RequestException` — **before** `upsert` and **before** `mark_stale`. So a stopped timer or a provider outage marks nothing: the catalog freezes at last-known-good and the router keeps routing on prices that may be weeks old. The failure is **silent and open**, not loud and closed. Nothing surfaces it, because a frozen row still reads `availability = 'active'`. The path that *can* empty the candidate set is narrower and is not governed by the timer at all. `fetch_neuralwatt` reads `payload.get("data", [])` with no floor on row count, so a 200 response carrying an empty or truncated `data` array — a partial provider outage, a schema change, an auth path degrading to an empty list — clears `raise_for_status()`, upserts nothing, and then lets `mark_stale` run anyway. Three days of that and every row is stale and `exclude_stale: true` leaves zero candidates for everything. `stale_after_days: 3` only sets the length of that fuse; it does not arm or disarm it. Against a 2-hourly poll it is 36 successful polls of margin, and recovery is automatic — `upsert` writes `availability = excluded.availability`, so one good poll flips every stale row back to active. The fix is a sanity floor on the fetch, not a larger number. **A second path empties the candidate set, and it bit on 2026-09-01.** Admin availability overrides are not governed by the poller at all. Deprecating the seven expensive models through `/admin` collapsed tier 3's context ceiling from 782,324 to **94,196** — while tiers 1 and 2 stayed at 782,324 — so every tier-3 request above 94k returned 422 with nothing warning anywhere. It surfaced ~19 hours later as an agent failing mid-task on an opaque error. **The obvious check for this is wrong, and the reason is worth remembering.** Tempting: warn when a higher tier's context ceiling sits below a lower tier's. But `ceiling(T)` is the max `effective_context_window` over models with `tier >= T`, and tier is a capability *floor*, so the eligible set shrinks monotonically as T rises — `ceiling(1) >= ceiling(2) >= ceiling(3)` is a theorem, true of every catalog. Such a warning fires always and means nothing. What actually failed is that a tier's ceiling dropped below what that tier is *asked* to serve, which is only knowable from traffic: compare `ceiling(T)` against the observed `required_context_tokens` for decisions classified at tier T. That is silent on all three tiers today and fires on the outage state (94,196 vs an observed max of 268,168). See `plans/catalog-staleness-and-poller-failure-modes.md` §4.4. **This recurred on 2026-09-04 through a dimension the detector did not model, and both halves of the fix are now in `metrics.py`.** Admin deprecations took out `kimi-k3*` — the only vision-capable rows with enough context — so a 242,486-token image request 422'd while every existing check stayed silent, because the *all-models* tier-1 ceiling was still 782,324. The vision-capable ceiling had collapsed to 192,500. - **Predictive:** `capability_ceilings` / `capability_demand_warnings` compute `vision` and `json_mode` sub-ceilings and compare each against demand actually observed for requests carrying images / requesting JSON. Two extra series, not a bucket per capability combination. - **Reactive:** `rejection_warnings` watches `route_decisions` for rows with `selected_model IS NULL`. This is the more valuable half and the simpler one — it catches the *next* dimension nobody predicted, at the cost of firing after the first failure rather than before. Two details in the reactive detector are load-bearing and easy to undo by accident. It groups by `(task_tier, digit-normalized reason)` using the **structured column**, because normalizing digits alone merges `tier >= 1` and `tier >= 3` rejections into one group and hides whether the broadest or the frontier candidate set went empty. And the signal is **novelty OR rate**, never mere presence: measured on the live DB, routine rejections run ~3/hr while the 2026-09-04 incident was only n=2 — *below* the noise floor — so no single count threshold can both catch it and stay quiet. A group absent from the 24h baseline warns at n≥2; a familiar group warns at the configured count. Zero rejections warn about nothing: a genuinely impossible request SHOULD 422. The service holds a billable API key and has **no auth of its own**. Loopback bind is the only thing standing between the open internet and your allowance; add auth before widening `--host`. The same applies to an Ollama shared over a VPN — it has no auth either, so `deploy/ollama-over-vpn.conf` binds it to the VPN address rather than `0.0.0.0`, which would publish it on whatever network the client happens to be on. ### When the router goes unreachable, start at docs/incidents.md Eight incidents so far, nearly all sharing one shape: a change that looked local to the router silently degraded the agent depending on it, and none announced itself as a router problem. **`docs/incidents.md` carries the full write-ups plus a symptom -> one-line-check table**; read it rather than re-deriving a diagnosis. #8 is the exception worth knowing before an unattended agent run: the router worked perfectly while agent workers looped for hours with no progress, and no check noticed, because every check watched spend or availability rather than whether changes landed (`plans/no-progress-detection.md`). Two conventions from those incidents that bind every session, and so stay here: - **8080 is production, always.** It is baked into `opencode.json`, the systemd unit, every curl example here, and the admin frontend's own fetches. A throwaway instance (manual iteration, Playwright smoke tests, anything that is not "use the real router") binds **8081**. Never send a kill signal to a process matched by name or port rather than by a PID you started yourself -- `Restart=always` will fight you, and on this repo it may be your own model access. - **Never point `config/config.yaml` at test fixtures.** It is the file the live service reads. Pass a different config file, monkeypatch `cfg.database.path` in-process, or use a temp copy. Recovery for an unreachable-but-`active` service is `systemctl --user restart llm-router.service` -- a hung process was never in a tracked stop job, so this issues a fresh cycle systemd does enforce a timeout on. The watchdog runs as its own systemd **timer** (`llm-router-watchdog.timer`), separate from the dispatcher, so a stuck router cannot silence the thing meant to notice it is stuck. Install and enable it with `systemctl --user enable --now llm-router-watchdog.timer` (the timer's unit file ships pointing at a placeholder home path and must be `sed`-repointed to the real one first), or run it once by hand with the `--once` flag. See [watchdog](docs/watchdog.md) for the signals it fires on, the alert lifecycle, and the known limits. ## Pointing a coding agent at it The `/v1` endpoints are OpenAI-compatible, so any normal client works — opencode, an SDK, plain curl. Repo-local `opencode.json` is already wired up, so running `opencode` from a clone of this repo routes by default. For global use, merge `provider.llm-router` into `~/.config/opencode/opencode.json`. | model name | behavior | |---|---| | `auto` | router picks; flex rows excluded so nothing is held during peak | | `auto:batch` | router picks; flex rows admitted, for overnight/async work | | any real model id | dispatched as asked, still logged | Streaming is proxied chunk by chunk rather than buffered, so tokens still render as they arrive. NeuralWatt emits its energy and cost blocks as SSE **comment** lines (`: energy {...}`) before `data: [DONE]` — ordinary clients ignore comments, so the stream passes through untouched while the router reads the telemetry on the way past. Without that, streamed calls would log no energy at all, which is most of the point of this project. ## Try it Copy-pasteable `curl` examples (`/route`, `/dispatch`, `/v1` models + chat) and the classifier-skip overrides (`task_category`, `task_tier`, `required_context_tokens`) live in [README ## Usage](README.md#usage). Inspect what a dispatch cost/burned with the `sqlite3` query in [docs/operations.md](docs/operations.md).