Files
6krrt/CLAUDE.md
adlee-was-taken c639d60859 feat(feedback): a dry run that projects the fold, and the timer it never had
`POST /outcome` is the only ground truth this router has, and feedback.py is
the one loop with no timer -- poller, seed sweep, backup and offsite all have
one. Live state on 2026-09-15: `proficiency` last written 2026-09-10, with 404
unapplied attributable outcomes and thousands of decisions routed off the stale
scores in between.

The fold is irreversible. add_outcome accumulates into a running mean and
recompute_category re-derives every row in the category from the new peer rate;
neither keeps the pre-fold value anywhere, and verifications.applied_at means a
second run will not redo the work either. So the deliverable is a preview plus
units that ship unstarted, not an automatic fold.

`--dry-run` now projects instead of describing. feedback_preview copies the
database into memory, runs the REAL add_outcome against the copy, and diffs the
two proficiency snapshots. It does not re-implement the empirical-Bayes
conversion -- a second implementation would drift, and a confidently wrong
forecast of an irreversible action is the worst failure available here.

Two kinds of movement come out, and the second is the surprise: `direct` rows
carry new outcomes of their own; `ripple` rows carry none and move anyway,
because the whole category is re-derived against a peer rate the new evidence
just changed. On the live backlog 404 samples across 17 pairs move 17 rows
directly and 363 by ripple, so ripple is digested per category and `--csv`
carries every row.

The units are named and hardened like the poller/seed pair, and take no
EnvironmentFile and no network-online.target because the fold makes no HTTP
request of any kind. The .service has no [Install] section, so it cannot be
enabled on its own -- "not enabled by default" is structural rather than a
README promise. 12h cadence because the fold is exactly additive: ten folds of
five land on the same numbers as one fold of fifty, so cadence caps staleness
and batch size and cannot change where the scores end up.

Three corrections to what was in the tree:

- The timer carried `Persistent=true`, which systemd.timer(5) says "only has an
  effect on timers configured with OnCalendar=". This timer is monotonic, so
  the line bought nothing; OnBootSec is the real catch-up. The sibling poller
  and seed timers carry the same inert line -- noted, not fixed in passing.
- `--csv` without `--dry-run` was accepted and discarded, and the run it was
  silently dropped from is the irreversible one. It is now refused.
- preview() copied the whole 41 MB database into memory just to print coverage,
  then project() copied it again. Coverage only reads, so it takes a mode=ro
  handle instead.

Tests: 2068 -> 2087. The load-bearing one asserts a dry run leaves the database
byte-identical -- proficiency rows, verifications.applied_at, and the file's
sha256 -- while still reporting the 8-sample delta it would apply. Verified
non-vacuous by handing copy_database the real connection and watching it fail.
A second test folds for real onto an identical copy and demands the projection
match every score, which is what stops the preview drifting from the store.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-15 22:34:16 -04:00

75 KiB
Raw Permalink Blame History

Local LLM Model Router — project brief

README.md is the concise front door. Deep-dive module-by-module reference docs live in docs/. design/local-llm-model-router.md holds the architecture and rationale, including parts still unbuilt. This file is the working state + immediate next steps, and is the one to trust on what is currently true.

NORTH STAR GUIDELINES

Three rules that outrank local cleverness. Each exists because it was broken first and the breakage was expensive to find. When a change conflicts with one of these, the change is wrong — not the rule.

1. Every config knob is reachable from the admin page

Every scalar knob in config.yaml gets an admin control — runtime, persisted, or both as its mechanism warrants — or a recorded decision saying why it must not. The absence of a control has to be a decision someone made, never an oversight nobody noticed.

objective.credit_attenuation.enabled is the model exception: deliberately absent, because enabling it must be a config edit plus a restart. That is a recorded choice, not a gap.

Why: Wave 2 shipped incumbent_cache_pricing and incumbent_challenger_cache_rate with no control at all. Nobody decided that; it just never came up. The dial's entire purpose is tuning from neutral to full without reverting code — and hand-editing a tracked file and restarting is a loop nobody walks, so the knob's rationale evaporated on contact with reality.

Enforced by tests/test_admin_knob_coverage.py, which fails naming the knob. Honour it; do not add a DELIBERATELY_NOT_IN_ADMIN entry to silence it unless the reason is true.

2. Fix classifier latency by improving the classifier, never by reusing a

stale decision

Classification latency is a real problem and the answer is a faster or better classifier — a smaller model, a warmer process, a cheaper backend. The answer is never to let one classification stand in for later, different work.

A stale label does not merely add noise. It silently redirects money: the whole point of this router is sending each task to the model that fits that task, and a replayed label routes the next task to whatever fitted the last one.

Why: the session cache reached 96.6% of all classifications. One real classification drove 107 consecutive turns across 14 minutes on 2026-09-16. It was introduced to avoid paying a 2-5s round trip per turn, which was a reasonable trade in isolation and became the dominant path without anyone choosing that.

Two costs, and the second is worse than the first:

  • Routing decides on what the session was doing when it started, not what this turn is. A mid-session pivot — docs, then bug fixes — routes the bug fixes on the docs label.
  • Proficiency is trained on those labels, because POST /outcome attributes to (model, task_category). That is part of how deepseek-v4-flash came to hold a saturated 1.000 on file_summarization while failing 50% of them on real traffic.

It also made a working classifier look broken: replaying a handful of session labels across hundreds of turns produced "pages and pages of diff_checking", which read as a classifier stuck on one label when the underlying classifications were a reasonable mix.

If latency forces a cache, that is a measured, time-boxed concession with an expiry condition written next to it — not a default.

3. Check worktrees and other agents' work before touching any file

Before editing, run git worktree list and check for other agents or sessions working in the same tree. Do not assume a worktree is yours. When parallel work is unavoidable, isolate it — a separate worktree per agent — and commit by explicit path, never git add -A.

Why: two agents were once launched into the same worktree while each was told the other was elsewhere. It survived only because the files were disjoint and one of them committed by explicit path rather than sweeping the tree. The second agent's test count was also measured against a tree containing the first's uncommitted work, so its "verified green" was not a clean signal.

Related: a router.db inside a worktree is a stale copy, not the live database. The live one is /home/alee/Sources/6krrt/router.db. Check mtimes before measuring anything, and open it read-only: sqlite3.connect("file:...?mode=ro", uri=True) — the sqlite3 CLI here does not accept -uri.

What this is

A router that uses a local model (served via Ollama) to classify incoming coding/documentation tasks — category, tier, required context size — and dispatch each task to the best-fit open-weight model on Neuralwatt Cloud, under a per-request cost ceiling, tiebroken by price, and ranked by category-level expected pass rate.

Every measurement in this file was taken on one deployment against one provider account. They are recorded because the reasoning is worth more than the conclusion, but treat them as observations with a date on them, not as constants — the catalog, prices, grid intensity and pool load all move. When a number here decides something, re-run the measurement before trusting it.

Neuralwatt remains the primary provider. OpenRouter is back as an opt-in provider gated by the provider_model_allowlist table — a provider with require_allowlist: true is ignored unless its model_id is explicitly allowlisted, and any previously-upserted row that drops off the list is marked deprecated on the next poll. The provider column and the (model_id, provider) primary key are unchanged, so this adds no migration.

Stack

  • Python (chosen over Rust — this is I/O-bound against provider APIs, not CPU-bound; iteration speed on the scoring/weighting logic matters more than raw execution speed at this scale)
  • SQLite for the decision table
  • Ollama for local classification, via any OpenAI-compatible endpoint — localhost:11434/v1, or an Ollama on another machine across a VPN (classifier.base_url)
  • FastAPI for the dispatcher service

Billing is per-kWh, not per-token — and neither is what scoring uses

Measured against the live API, 2026-08-11. Neuralwatt bills a flat $8.00 per kWh and the catalog's input_per_million / output_per_million prices are not what this account is charged. Confirmed across five models; cost_usd / energy_kwh came back 8.00 every time:

model list $/1M out completion tokens billed USD kWh $/kWh
deepseek-v4-flash 0.28 600 5.00e-06 5.88e-07 8.50*
gemma-4-31b 0.42 540 3.80e-04 4.75e-05 8.00
qwen3.6-35b-fast 1.15 600 2.53e-04 3.16e-05 8.01
kimi-k2.7-code-fast 4.00 600 2.25e-03 2.81e-04 8.00
kimi-k3-fast 15.00 539 2.17e-04 2.71e-05 8.00

* rounding — billed cost is quantized to ~1e-06.

The precise rule, validated against all 65 samples of the reference sweep (61/65 within 2%; the 4 outliers are microdollar rounding, not misses):

cost_usd = min( $8.00/kWh x energy_kwh ,  3 x list token price )

The ceiling bound in only 2 of 65 samples, both deepseek-v4-flash energy spikes, and matched to the cent: 31 prompt x $0.14/1M + 400 completion x $0.28/1M = 1.1634e-04, x3 = 3.4902e-04, billed 0.000349.

List price ranks models backwards. Not approximately — invertedly:

list $/1M actually billed gCO2eq
deepseek-v4-flash 0.28 9.80e-05 8.90e-04
gemma-4-31b 0.42 5.02e-05 2.32e-04

deepseek lists 33% cheaper, costs 95% more, and emits 284% more carbon.

But cost and eco are NOT the same axis — the tempting simplification, and it is wrong. Cost tracks energy, but carbon is energy x the serving region's grid intensity, and models run in different regions:

grid gCO2/kWh models
FI ~49-50 most of the catalog; varies by time
FI (reported) 475 glm-5.2-fast, glm-5.2-flex
US-MIDA-PJM ~442 the kimi-k3 family

A 13.6x spread, so the two axes disagree: glm-5.2-fast is the 2nd cheapest model and only the 6th cleanest; kimi-k3-flex draws 3.7x less energy than kimi-k2.7-code while emitting 3.6x more carbon. Weighting them separately is load-bearing, and tests/test_routing.py pins it.

Superseded — cost no longer comes from the sweep at all. cost was the median measured USD over the reference sweep. That was measured to be WRONG for real traffic, because the reference workload is the wrong shape.

The sweep sends a 400-token prompt with a 400-token completion. Real agent traffic is a 150,000-token prompt with a ~400-token completion and ~92% cache hits (token-weighted over 50 sessions and 40.7M tokens on 2026-08-23; the figure was 84% when measured on 2.2M tokens of earlier traffic). The attribution ratio moves with prompt size, so the ranking inverts:

workload winner
reference sweep (400/400) glm-5.2-fast, 3.2x cheaper
realistic (70k prompt, short answer) deepseek-v4-flash, 5.0x cheaper

Same two models, opposite answer. glm-5.2-fast sits at attribution 0.006 on a toy prompt and 0.50 on a 70k one — it batches beautifully on small prompts and badly on real ones. deepseek-v4-flash barely moves (0.21 -> 0.25).

So routing.estimated_cost prices each request from catalog token prices, scaled to that request's actual shape (prompt size, assumed completion length, assumed_cache_rate). List price is not what gets billed, but billing is capped at 3x list, so it tracks the real ordering and bounds it — and on the one case that was checked live it agrees with the measurement in direction and magnitude (7.8x predicted vs 5.0x measured). It is also free, needs no sweep, and refreshes whenever the poller runs.

objective.plan_kwh_per_period is a planning figure only: per-request traffic is never refused for exceeding it — it gates nothing. Overage is billed against the account's credit balance (allowance_remaining_usd from the provider). The /admin/api/snapshot endpoint reports balance, estimated burn rate, and projected runway per provider inside quota.accounts[] — a list of per-provider billing shapes (metered_plan, prepaid_credit, self_hosted, or unmetered) with plan, pool, burn, and credit blocks as appropriate. quota.spend aggregates provider spend and a list-price estimate. The old flat keys and by_provider/total_balance_usd shape were removed; every consumer was updated in the same change, so there are no deprecated aliases.

Three signals said deepseek-v4-flash — catalog token price (7.8x cheaper), NeuralWatt's own published per-request energy (~10x lower), and a live 70k measurement (5.0x cheaper). Only the 400-token benchmark disagreed. Trust the workload you actually run.

eco still comes from the sweep's median gCO2eq, and is still not an objective. flex_cost_multiplier is gone: a flex row's measured cost already is its flex cost.

Open, and worth knowing: NeuralWatt's model cards publish gross energy (~1.99e-04 kWh for deepseek, ~1.91e-03 for GLM), while the billed figure is gross x attribution. GLM burns roughly 7x more actual electricity per request and charges ~5x less, because far more tenants share its GPUs. Anything built on eco inherits that inversion — the attributed carbon figure answers "what is my share", not "what was burned".

Energy attribution: signal that looks like noise

Billed energy decomposes exactly:

energy_kwh = avg_power_watts x duration_seconds x attribution_ratio

attribution_ratio is the request's share of a shared multi-tenant GPU pool. Up close it looks like pure noise — eight rapid identical calls to one model spanned 20x in billed energy, correlating +0.997 with the ratio while power and duration held steady. Two sweeps of the same 13 models with the same prompt disagreed by up to 36x.

Scoring on the pre-attribution product (power x duration) was tried, and it is wrong. Across the sweep:

spread
median attribution, between models 750x
typical spread within one model 1.8x

The ratios are quantized (0.001, 0.25, 0.5, 0.75) — that is serving concurrency, a stable per-model property, not weather. A model whose GPUs carry far more concurrent requests genuinely costs less per request, and that is most of the real cost difference in the catalog: deepseek-v4-flash bills ~1000x under its share of pool gross. Stripping attribution discards a 750x real signal to suppress a 1.8x one.

So scoring reads the attributed figures, and the median absorbs what noise remains. A split-half check on the 7-sample sweep (median of first three vs last four) shows that working:

  • 10 of 13 models agree within 1.4x — stable enough to route on
  • 3 do not: kimi-k2.7-code-fast (29x), kimi-k3 (14x), glm-5.2-flex (2.2x). Those need more samples before their position is trustworthy.

dispatcher.gross_energy_kwh remains as a diagnostic on the identity, not a scoring input.

Attribution drifts across hours, so sampling must too

Within about 30 minutes the billed figures reproduce (0.3-1.1x on a spot-check). Across hours they do not: between two sweeps, deepseek-v4-flash moved roughly 50x and qwen3.6-35b about 7x the other way — enough to invert their cost ranking. Attribution tracks pool load, and pool load tracks time of day.

More samples inside one sweep does not fix this; it measures one moment more precisely. Coverage across time does. load_candidates already takes the median over ALL seed_reference rows, so repeated sweeps accumulate into a median-across-time for free — hence llm-router-seed.timer, which runs a small sweep every 6 hours.

Until several sweeps have accumulated, treat the eco ordering as provisional. A single sweep's ranking is one sample of a moving quantity.

And none have accumulated since 6e729ad. That commit moved log_observation's trailing arguments to keyword-only without updating seed_energy.py, so every timer run since spent one billed completion and then died on TypeError — which is not a RequestException, so the per-sample except did not catch it. Fixed, and the sweep now has an offline end-to-end test, but the accumulation this section describes starts from the next run rather than from months of history.

What's built and working

  • config/schema.sql — models, proficiency, energy_observations; applies cleanly (sqlite3 router.db < config/schema.sql). See data-model.
  • poller.py — fetches Neuralwatt's catalog (public, unauthenticated), normalizes, upserts, marks stale. Verified live: 14 routable models. See data-model. For providers with require_allowlist: true (openrouter in the base config), the poller filters the fetched catalog against provider_model_allowlist before upsert and prunes existing rows that are no longer on the list to deprecated — see the #45 OpenRouter opt-in allowlist section below.
  • config/config.yaml / src/config.py — weights, thresholds, provider settings, Pydantic-validated. architecture.
  • scoring.py — one normalize_inverted (cost and eco normalize identically) + the weighted composite. routing.
  • seed_energy.py — reference task × N per model → energy_observations tagged seed_reference; makes cost/eco real. --samples 5 = 65 calls, under a cent. architecture.
  • tiering.py / tier.py — pure tier resolver + DB pass. Why tier on reasoning_default_enabled, cheapness-not-ceiling, tier1_context_max: routing#tiering.
  • routing.py — pure hard filters + ranking, plus the request-side capability gates (fail-closed asymmetry). routing.
  • routing.py rank_candidates incumbency: prices the session's last chat model at its measured cache rate and every challenger at the objective.incumbent_challenger_cache_rate dial, behind a load-bearing min(dial, incumbent_rate) clamp; gated off by default (incumbent_cache_pricing: false), tunable from off to full in config. routing#incumbency-and-cache-pricing.
  • circuit_breaker.py — passive availability skip on a 5xx (cooldown + backoff, clears on next success, no poller), on by default. Eval harness deliberately stays outside it (isolation): routing#circuit-breaker. Covers ollama-local too: a local outage raises 502 on the first request and the breaker excludes the dead local row on the next one, so traffic reroutes to cloud candidates.
  • dispatcher.py — FastAPI service: GET /health, POST /route (no provider call), POST /dispatch, OpenAI-compatible /v1/models + /v1/chat/completions, SSE GET /events/decisions. api. On an account-level cloud refusal/exhaustion, eligible routed requests degrade to the local dispatch model instead of surfacing the cloud error (see the Local dispatch model section below).
  • proficiency.py / proficiency_store.py / proficiency_outcome.py — blend leaderboard + self-eval into a benchmark prior, accumulate client outcomes, and recompute expected pass rates; the only write paths to proficiency, so blended_score/source never drift. architecture.
  • context_prune.py — relevance-based stage trimming only tool results once over budget_tokens, before any paid token ships. See pinch for budget_tokens; see also the protected_max_chars note there if you are changing how much prefix context is guarded.
  • feedback.py — folds POST /outcome client reports into proficiency.outcome_score via add_outcome(). Structural and local_llm verdicts are diagnostics only; POST /outcome is the posterior. verification. --dry-run no longer just describes the fold, it projects it: feedback_preview.py copies the DB into memory, runs the real add_outcome against the copy, and reports the per-row before/after, opening the source mode=ro so a preview cannot write to what it is previewing. deploy/llm-router-feedback.{service,timer} gives this loop the timer it never had — and ships not enabled, because the fold is irreversible and the first one against an accumulated backlog is an operator decision. See deploy/README.md.
  • exploration.py — epsilon-greedy exploration chooser; injected RNG, no mutable state. routing.
  • seed_local_dispatch_energy.py — standalone reference-shape sweep for ollama-local rows; derives per-token USD rates through the user's tariff and OLS on measured GPU draw. architecture.
  • poller.py — also seeds/updates provider='ollama-local' rows from config.yaml each poll so local rows stay current even when NeuralWatt is unreachable.
  • logs.py — per-request trace id (ContextVar), logfmt, journald priority prefixes; logs.bind() survives StreamingResponse generators. operations.
  • metrics.py / GET /metrics — read-only observability; takes (conn, cfg), never imports dispatcher. Also carries the three detectors added after the incidents below: capability sub-ceilings, the reactive rejection detector, and the classifier-degradation share. api.
  • tui.py — Textual dashboard over /metrics + /events/decisions; live feed, category→model panel, detail popup; data layer split into tui_model.py. The decision table leads with a time column and carries profile plus an E flag for exploratory picks; the quota panel now shows per-account billing shapes (metered_plan, prepaid_credit, etc.) with plan/pool/burn/credit blocks, spend aggregates, and an alarm line. architecture.
  • tests/test_tui_schema_drift.py — the tripwire that keeps the two honest. A new route_decisions column must be registered as surfaced or deliberately-not, or the test fails naming the column. Five columns had already reached the schema without reaching the dashboard; ROUTE_DECISIONS_COLUMNS in tests/test_route_decisions.py had itself drifted.
  • tests/test_tui_warnings.py — the same idea for warnings. Every class /metrics can emit must render in #warnings-panel, and every emitted warning must be registered — the second failing with the RAW text, because the point is that nobody knew the class existed. Its fixture is a coupled system: adding a seed can silence an existing class (a small-context seed once killed the escalation hazard by dragging the p95 down), which is why both directions are asserted.
  • router_cli.py — one-shot /route probe (no spend), raw JSON with --json. api.
  • admin.py / config/admin_schema.sql / admin/frontend/*.html — loopback /admin portal: dashboard, models overrides, decisions log, profiles, a read-only proficiency matrix (GET /admin/api/proficiency) that distinguishes a measured score from an inherited one, and controls (including the Local Compute and classifier.mode cards). admin-portal. Provider management includes list/detail/update/delete endpoints (GET /admin/api/providers, GET /admin/api/providers/{name}, POST /admin/api/providers/{name}, DELETE /admin/api/providers/{name}) plus per-provider allowlist endpoints (GET/POST /admin/api/providers/{name}/allowlist, DELETE /admin/api/providers/{name}/allowlist/{model_id}); the providers page shows require_allowlist and links to the allowlist editor.
  • local_encoder.py — zero-shot category classification via a non-generative encoder, backing classifier.mode: local_encoder. transformers/torch imported lazily; a deployment that never selects the mode needs neither installed. local-models.
  • provider_model_allowlist table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become deprecated on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below.
  • config.py / DispatchProvider.require_allowlist — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted.
  • tests/ — 1155 tests across 40+ files, offline, verified on Python 3.10 and 3.14. README.

#45 — OpenRouter is an opt-in allowlist provider

OpenRouter used to be ingested whole, then removed, and is now back — but only as an opt-in provider. The base config sets openrouter.require_allowlist: true and ships a short seed allowlist. Models on that seed list are upserted and kept active; anything else in the OpenRouter catalog is filtered out before upsert and any previously-active OpenRouter row that is not on the list is marked deprecated on the next poll.

This is deliberately different from the old ingest-everything behavior. The previous approach once pulled in a non-chat model (lyria/...) that returned HTTP 404 on dispatch because the endpoint expected chat completions. The router had paid for the classification, selected the model, and then failed on the provider call. Allowlist-gating prevents that class of failure by default: if a model id has not been reviewed and explicitly added, the router acts as if it does not exist.

The seed allowlist is short and has firm exclusions. It does NOT include:

  • x-ai/* (Grok)
  • openai/*
  • anthropic/*

Those exclusions are non-negotiable. They are not "currently excluded" or planned for future inclusion; they are deliberately absent from the seed list. Adding one requires editing both the seed allowlist and this file.

The admin portal exposes the allowlist under /admin: the providers page shows which providers require one, and each provider row links to an allowlist editor where entries can be added or removed. The underlying four API endpoints are GET /admin/api/providers/{name}/allowlist, POST /admin/api/providers/{name}/allowlist, DELETE /admin/api/providers/{name}/allowlist/{model_id}, and the providers page itself surfaces require_allowlist with a link to the allowlist editor.

Neuralwatt remains the primary, ungated provider. The allowlist behavior only fires for providers with require_allowlist: true.

Proficiency: category now changes routing

proficiency_score is the ONLY category-dependent term in the ranking, so until this table had data, task_category could not change a decision at all — the classifier computed it, the router paid ~10s for it, and then it made no difference. It does now. Two categories were added for local dispatch: file_summarization and diff_checking; see evaluation. The score is also no longer a raw benchmark level: it has been converted into an expected pass rate on real traffic, calibrated against 1,059 client-reported outcomes and shrunk with a 20-pseudo-observation prior so thin data does not dominate.

proficiency now holds 141 rows:

source count meaning
outcome_blended 33 Fresh per-model outcome evidence
outcome_prior 76 Trafficked-sibling rows inheriting the peer-rate prior
self_eval_thin 32 Cold categories (summarization, translation); benchmark preserved verbatim

A further 33 (model, category) pairs have direct outcome samples. The biggest evidence gains went to deepseek-v4-flash (notably coding_general, coding_refactor, and general_chat), kimi-k2.7-code (across 7 categories), and qwen3.6-35b.

Sweeping 9 categories x 3 tiers currently returns 5 distinct winners at both 50k and 120k of context.

context winners over 27 decisions
50k qwen3.6-35b (10), gemma-4-31b (7), deepseek-v4-flash (5), kimi-k3 (3), kimi-k3-fast (2)
120k kimi-k2.7-code (10), gemma-4-31b (7), deepseek-v4-flash (5), kimi-k3 (3), kimi-k3-fast (2)

This spread is recent, and how it got here is the useful part. For a long time all 27 decisions returned ONE model, and that was the correct answer at the time rather than a bug: with cost and eco both populated, qwen3.6-35b was Pareto-dominant — cheapest AND cleanest in the routable set, while scoring within quality_tolerance of the best. No defensible weighting picks anything else out of that.

Two corrections widened it, and neither was a tuning change:

  • Cost stopped being a benchmark average. It is now priced per request from catalog prices scaled to the request's shape, so the ranking depends on the workload instead of on a 400-token reference sweep that no real traffic resembles.
  • Tier stopped being inferred from price. deepseek-v4-flash was pinned to tier 1 for being cheap, which excluded it from every tier-2 request regardless of what any score said.

A third shift is under way: the outcome backlog has been spent, so the score now reflects real pass/fail reports rather than the benchmark alone. That changes the numbers; it does not change the rule. Quality is still the objective and cost is still the tiebreak within quality_tolerance.

Note what changes between the two rows above: only the leader, and only because of the hard context filter. That is the filter working, not the scoring disagreeing with itself.

If you see one model win everything again, check for dominance before reaching for config. One winner is a legitimate outcome. The levers, if a genuinely different balance is wanted, are objective.quality_tolerance (how large an expected-success-rate gap must be before it outranks a cost saving) or objective.max_energy_per_request (a hard ceiling). There is no weight to tune.

What the task set actually found

The benchmark could not discriminate these models on coding. Every row scored exactly 1.00 on coding_general, coding_refactor and debugging — and that is after the tasks were deliberately hardened with touching intervals, full semver, present-but-falsy defaults, late-binding closures and a binary search that infinite-loops. Every model in this catalog is simply good at that class of problem, so cost decides coding routes, which is the right outcome.

Real traffic broke one of those ties, which the benchmark never could. coding_general now spans 0.86-1.00: glm-5.2-fast fell to 0.862 over 29 samples folded in by feedback.py from an actual agent session, and crossed self_eval_min_samples on the way, so it reads self_eval rather than self_eval_thin. That is the intended shape of this system — the 43-task benchmark establishes a floor, and your own traffic is what refines it. coding_refactor and debugging are still flat at 1.00, awaiting the same treatment.

Benchmark-sourced hardening landed, and it broke both remaining ties. The task set grew from 23 to 43: 8 BFCL tool tasks, 6 CRUXEval-O exact tasks (after the score_exact literal-eval fix), and three Exercism refactor/debug pairs. The CRUXEval-O rows split coding_general into a 0-1 mix across models — five of six now fail at least one model — and the Exercism refactor rows moved coding_refactor off its flat 1.00: bowling 0.25-1.00, dominoes 0.10-1.00, affine 0.56-1.00. The debug pairs split debugging the same way (bowling 0.80-1.00, dominoes a sharp 0/1 split). Two follow-ups recorded, not deleted: debug_affine_coprime and most of the BFCL tasks sat flat at 1.00, so they do not discriminate.

A 1.00 can also be a sampling artifact, and docs_writing was one. At 2 samples per model the category read 0.70-1.00 with a model at the ceiling, and the router paid for that ceiling: kimi-k3-fast won every docs route. Six more benchmark passes moved every score and left NOTHING at 1.00:

model n=2 n=11-14
kimi-k3 0.85 0.973
kimi-k2.7-code 0.85 0.886
deepseek-v4-flash 0.80 0.864
kimi-k3-fast 1.00 0.864
qwen3.6-35b 0.85 0.800
gemma-4-31b 0.85 0.786

The winner moved to kimi-k2.7-code, 3.2x cheaper at 50k of context ($0.0441 -> $0.0136), with no config change — kimi-k3 scores higher but sits inside quality_tolerance, so cost breaks the tie. deepseek-v4-flash ($0.0024) misses the band by 0.009, which is the kind of margin the tolerance exists to describe rather than a verdict.

The whole spread rests on one rubric line, though. docs_function is effectively saturated — 1.00 on nine of every ten samples — and nearly every docs_gotcha deduction is the same omission: the model documents that order is preserved, that the first occurrence is kept, and what key does, then never says elements must be hashable. That is real discrimination, since it is a real property of the function, but one sentence is deciding a category. Treat this ordering as thinner than n=14 makes it look.

The self-judging guard costs sample density, and it shows up here. Most models reached n=14; kimi-k3 and kimi-k3-fast reached only 11, because those two are the ones diverted to the alternate judge qwen3.6-35b, which returns unparseable JSON more often than kimi-k3 does. The guard is still right — a model grading its own family is worse than a thinner sample — but the alternate judges should be picked for parseability, not just for being someone else.

The reflex when a category looks flat is to reach for quality_tolerance. Neither tie broken so far was broken that way: coding_general opened up when feedback.py folded in real traffic, and docs_writing opened up on six more benchmark passes. Both were samples, not settings. coding_refactor and debugging are still flat at 1.00 on 2-3 samples each — which is now a state this project has mistaken for a measurement once.

Current spread by category, widest first:

category spread
tool_use_agentic 0.33 - 1.00
summarization 0.60 - 1.00
reasoning_math 0.67 - 1.00
docs_writing 0.66 - 0.97
general_chat 0.80 - 1.00
translation 0.85 - 1.00
coding_general 0.86 - 1.00
coding_refactor, debugging flat at 1.00

What does discriminate is tool use, arithmetic traps, and prose. deepseek-v4-flash scores 1.00 on all three coding categories yet 0.33 on tool_use_agentic and 0.67 on reasoning_math. Verified live, not an artifact: given "It is 1:20pm and my meeting starts at 3pm, how many minutes away?" — both times supplied — it calls two tools rather than subtracting. It over-reaches for tools, which is exactly the failure mode that matters in an agent loop. The router now avoids it for those categories while still picking it for coding.

Most rows still read source='self_eval_thin' (118 of 132): real measurement, but below self_eval_min_samples at 2-3 tasks per category per run. The 14 that have crossed it are all docs_writing, from the six extra passes above. Two paths thicken it, and they are complementary — re-run eval_proficiency.py to accumulate benchmark samples, or just use the router and let feedback.py fold in real outcomes. Both fold into a running mean rather than replacing, so samples add up across runs.

A score is only as fresh as the row it was copied to

Proficiency is a property of the weights, not the queue, so the eval harness scores one row per family and propagate_to_variants copies the result onto the serving variants — kimi-k3-flex gets kimi-k3's number, because no benchmark rates a -flex row separately.

That copy used to happen exactly once per variant, ever. The guard skipped any row with self_eval_samples > 0, meaning "measured directly, do not overwrite" — but inheritance copies the sample count too, so after the first propagation an inherited row was indistinguishable from a measured one and was never refreshed again. kimi-k3-flex sat at 0.85/n=2 while kimi-k3 moved to 0.973/n=11.

proficiency.inherited_from records the provenance that was missing, and the migration was the delicate half, not the fix: ADD COLUMN gives every existing row NULL, which reads as "measured here", so shipping the guard alone would have permanently frozen the exact rows it exists to unfreeze. The backfill infers provenance from the harness's own selection rule rather than guessing — eval_identities only ever evaluates standard rows plus flex rows with no standard equivalent, so a flex row that has one was never a candidate for direct evaluation, whatever its sample count claims. Everything else keeps NULL, which fails safe: NULL means "do not overwrite", so no real measurement can be lost to a wrong guess.

Confirmed on the live database, and on the catalog's one genuine exception — glm-5.2 is canary, so glm-5.2-flex is the routable row the harness scores directly, and its NULL is correct.

Harness bugs this shook out

Three separate defects, each of which scored the rig rather than the model, and each caught by reading per-task detail rather than the summary:

  • Token budget. max_tokens was shared between a reasoning model's trace and its answer. At 1200, qwen3.6-35b spent ~4,200 characters thinking and returned an EMPTY content field, scoring 0.00 on tasks it can plainly do. Now 24000, clamped per model (gemma-4-31b caps at 16384), and finish_reason: length skips the sample instead of scoring it.
  • One leading space. kimi-k2.7-code returns " def f(...)", which becomes IndentationError once the harness prepends its imports — 0.00 across all nine coding tasks for a model with "code" in its name.
  • Judge failures scored as model failures. 44% of judge calls returned unparseable output (the judge is itself a reasoning model and leaks its thinking despite response_format). Each was recorded as 0.0. Now the JSON is extracted from surrounding prose and an unusable reply yields no sample.

tests/test_task_set.py exists so this stops happening: it implements a reference solution for every code task and asserts it passes every check, recomputes every exact answer (one by brute force), and confirms each refactor target already passes its own checks while each debugging target fails. It immediately caught a check where the expected value was simply wrong — which would have docked every model on a task and been indistinguishable from genuine difficulty.

Tool competence is read from the request, not guessed at

Neither local classifier can identify agentic work. Asked to label six unambiguous tool-use prompts ("read the config then update the manifest", "run the tests and fix what fails"), qwen3.5 got 2/6 and mistral-nemo 1/6 — and mistral-nemo's misses collapse to general_chat, which is also the configured fallback_category, so qwen3.5's crashes land in the same place.

That mattered because tool_use_agentic has the widest proficiency spread in the table (0.33-1.00) and deepseek-v4-flash — the current winner on coding — sits at the bottom of it.

The fix was not a better classifier. Whether tools are on the table is stated in the request: every agent client sends a tools array, and chat_completions never looked at it. Reading it is exact and free.

It is applied as a hard filter, not a category override, and the distinction is load-bearing. The question is not "is this task agentic" but "can this model be trusted with tools that exist". The recorded failure is precisely the second one: deepseek-v4-flash was given a non-agentic prompt ("it is 1:20pm and my meeting is at 3pm, how many minutes away?", both times supplied) and called two tools rather than subtracting. A model that over-reaches is a hazard on every request where tools are available, whatever a classifier would have labelled the task.

So routing.min_tool_proficiency drops any candidate whose measured tool_use_agentic score is below it, but only when the request carries tools:

request winner on coding_general @ 50k
no tools deepseek-v4-flash ($0.0024)
tools present qwen3.6-35b ($0.0041)

Verified live through /v1/chat/completions with identical bodies differing only by the tools array. The cost of safety here is 1.7x on that route, paid only where tools exist.

It is currently set to null, i.e. OFF, deliberately and pending experiment. opencode sends tools on essentially every request, so with the filter on, deepseek-v4-flash is excluded from ordinary agent traffic and its ~7x cost advantage goes unused; with it off, that advantage applies and a model measured at 0.33 on tool use handles requests where tools are on the table. Which is right is an empirical question and the benchmark cannot answer it — the 0.33 comes from 3 tasks.

What settles it is POST /outcome: run with the filter off, let real pass/fail reports accumulate, and compare deepseek-v4-flash's tool_use_agentic proficiency before and after. That is the one signal here that knows whether the work actually worked, and feedback.py folds client outcomes in both directions, so success counts too.

Update 2026-09-15: that accumulation path is closed. The classifier no longer emits tool_use_agentic — per-turn classification collapsed onto it (31 of 31 consecutive live turns, all routed to the slowest model in the catalog) — so no new outcome attributes to the category and its scores are frozen at today's values. Accepted, not fixed; the reasoning and the rejected alternatives are in docs/routing.md, "The classifier's labels are not the proficiency scoring axis". The filter experiment now either finds a different signal or reads frozen data.

0.5 sits in the empty band between the only two values the catalog holds (0.33 and 1.00), so it is not fitted to either. A model with no measured tool score is unproven rather than proven bad and is not dropped — the same rule as the tier-1 context gate. Config load refuses a routing.tool_use_category that is not a real category, because a name matching nothing yields NULL for every row and NULL means "do not disqualify": the filter would silently stop filtering.

Tier is an iteration budget, not just a floor

A tier used to mean only "do not route below this". It now also buys corrective attempts after a verification failure:

tier batch interactive
1 0 retries 0
2 1 1
3 2 1

Interactive is capped below its tier because every retry doubles time-to-answer, and in interactive use latency is a quality loss.

Retries are matched to the failure, since the causes differ:

  • truncated — raise the token budget on the same model; a different one would run out too. If there is no cap to raise, the model's own output ceiling is the wall, so escalate to a candidate that can emit more.
  • malformed — more tokens will not make unparseable output parse, so escalate to the next-ranked candidate.
  • ok / unverifiable — buy nothing. Retrying unverifiable would burn quota across the majority of prose traffic for no signal.

escalation.preemptive_on_low_confidence is now off by default. Bumping the tier because the classifier was unsure pays frontier prices before anything has gone wrong; spending after a check has actually failed is better on both mandates — the cheap attempt usually succeeds, and when it fails you have evidence rather than a hunch.

The only ground truth: POST /outcome

Everything else the router records is a proxy. Structural checks know whether code parses. The local checker guesses whether prose looks right. Neither knows whether the answer did the job — the client does, because it ran the tests.

# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
  -d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'

Two things make this the highest-value signal available:

  • It is the only quality signal that survives streaming. A retry cannot reach a streamed response because the bytes are already gone; a report arrives afterwards and works either way. Every agent client streams.
  • Its successes count. feedback.py folds ok: true and ok: false into proficiency.outcome_score via add_outcome(). Structural and local_llm verdicts are now diagnostics only; they no longer move the score, because a structural failure is not a verified client outcome.

Client outcomes are what calibrate the routing score. The benchmark is a prior; POST /outcome is the posterior.

request_id is written back to route_decisions.request_id on both the streamed and buffered paths, so the report joins cleanly to the decision that produced it. An unknown request_id returns 404 rather than being quietly accepted — a client whose reports go nowhere should find out.

Verification: what local compute is actually good for

Local inference is a poor substitute for cloud completions here — 7-82x the energy, ~670x the carbon on Michigan's grid, and slower (6.2s vs 1.4-2.0s). But it is very good at stopping a cloud completion from being wasted, and completions are where all the money is: fitted on real traffic, a completion token costs 201x a prompt token.

check cost what it catches
structural (verification.py) free truncation, malformed code/JSON/YAML, empty answers
local LLM (Ollama) ~$6.9e-05, ~6s refusals, wrong-question answers, incoherence

Structural checks run on every response and never execute the code — they parse it. The local LLM check runs only on answers above verification.min_completion_tokens, in the background after the client has its response, because it costs ~15% of a median 193-token answer and only pays above ~600 tokens.

An agent turn is not a prose answer

Both checkers had mirror halves of the same blind spot, and real traffic is what found it. On the first genuine agent session (63 completions, a shipped feature, 349 passing tests, clean mypy and ruff):

checker called it actually
structural malformed: empty response x29 turns that ended in a tool call
local LLM cuts off mid-sentence x8 of 9 same turns, judged from the other side

A turn that calls a tool has empty or half-finished text by design. Both paths now take has_tool_calls and return unverifiable; in verify_response that check outranks even finish_reason == 'length', because stopping mid-sentence at a call boundary is not a budget overrun. worth_local_check declines outright, which also stops paying ~6s of local inference to mis-grade a tool call.

Had feedback.py run before this, it would have applied ~12 false failures to the models that had just shipped the feature. That is the fourth harness bug in this project that would have scored the rig rather than the model, and the first caught by real traffic instead of a synthetic test. The pre-fix rows are kept but set model_attributable = 0, so the record survives without steering routing.

client_capped was also over-applied. It marked every verdict non-attributable whenever the client set max_tokens — and opencode always does — so real failures were invisible to feedback. A client's token cap explains a truncated verdict and nothing else; it is now scoped to exactly that.

feedback.py folds observed failures into proficiency, so routing learns from your traffic rather than only the 43-task benchmark. Only failures are folded in: a structural 'ok' means the code parsed, not that it was correct, and recording those as 1.0 would flatten every score toward the ceiling. Failures the model did not cause — a client's own tight max_tokens truncating the answer — are recorded but excluded.

The classifier is the latency floor

Every routed request pays a full local classification round-trip before an upstream token is requested — the classifier is the latency floor. Four settings keep it usable: max_retries=0 (SDK retry → 3x wall-clock), max_output_tokens: 1024 (bounds reasoning), max_input_chars: 8000 (doesn't need the document; head+tail clamp), fallback_tier/fallback_category (mid tier, not 502/503), plus temperature: 0 (deterministic tier). "Local" means your hardware, not this machine: bind to the VPN address, not 0.0.0.0 (Ollama has no auth). Full measurements: docs/local-models.md.

A classifier failure no longer collapses to one fixed guess

fallback_tier/fallback_category above is now the LAST step, not the only one. When the local classifier fails, classify() walks a cascade:

# step cost
1 this session's cached classification, staleness ignored free
2 this session's history in route_decisions free
3 classifier.cloud_fallback, if configured money
4 fallback_tier / fallback_category free

Steps 1 and 2 reuse a real classification rather than guessing. Step 2 is what survives a restart, when the in-memory cache is gone but the session's decisions are still on disk. Both filter degraded sources, so one outage's guess cannot propagate through every later turn of a session and end up looking like a measurement.

Step 3 is absent by default and the whole feature costs nothing until it is configured. Three things bound it once it is:

  • classifier.cooldown_seconds (30) is a global backoff, so a sustained outage buys one cloud attempt per window across ALL requests, not one per request;
  • an account-level provider refusal suppresses step 3 for the same window — an out-of-credit account makes that call a guaranteed wasted request;
  • the session_cache.put() gate admits classifier_cloud, bounding a session to one cloud call per staleness window instead of one per turn. session_stale and session_history are deliberately NOT cached: borrowed answers must not renew a staleness clock they never earned.

The trap this had to avoid, and it is the reason to read this section. A degraded classification is recorded with task_category = general_chat — which is a fully scored category (13 proficiency rows) as well as the configured fallback_category. Without a guard, POST /outcome reports on mislabelled outage traffic fold into proficiency(model, general_chat) and drag real scores toward whatever happened to be flowing while the classifier was down. So report_outcome marks outcomes of degraded-source decisions model_attributable = 0 — kept as a record, excluded from folding, the same treatment the tool-call false-failures got. feedback.py is untouched; its existing AND model_attributable = 1 already does the work.

Attributable: classifier, cached, override, classifier_cloud. Not: fallback, session_stale, session_history. The asymmetry is deliberate — a degraded source is good enough to route one visibly-flagged request, but a proficiency score is consulted by every future request, so attribution must not trust more than routing does. Unknown decision rows fail open; an over-applied exclusion already starved feedback once (client_capped).

/metrics warns when the degraded share of the last 24h crosses classifier.degraded_warn_threshold over at least degraded_warn_min decisions. A survivable failure is exactly the kind that goes unnoticed for weeks.

Which implementation is PRIMARY is now a config choice

classifier.mode in the live deployment is currently local_encoder, set in config.local.yaml with device: cuda on the classifier host. That is the mode answering real traffic right now. The explicit caveat is that a CPU-vs-CUDA latency and confidence comparison on the live classifier host is still pending; until that measurement exists, flipping the project default to local_encoder is not decided.

Turning that mode on for real found three real bugs in one afternoon (2026-09-06), each a fresh instance of this project's own recurring lesson — verify against the live system, not the plan. First, the config-load validator for confidence_threshold didn't exist yet: the admin UI saved a raw 80 (meant as 80%) straight into the overlay with no conversion, which would have made every real confidence score read as below-threshold on the next restart (classify_zero_shot returns [0.0, 1.0]; no probability exceeds 1.0). Caught before the restart, not after — see confidence_threshold in local-models for the full incident and the fix. Second, once that was corrected and the service actually restarted, it crash-looped twice more before coming up clean: the shipped default model (MoritzLaurer/deberta-v3-base-zeroshot-v2) had become gated on HuggingFace sometime after this project picked it (401 on an unauthenticated GET of its own model page), and separately HF_HOME's default cache path falls outside this service's ProtectHome=read-only sandbox exception — both are now fixed (switched to facebook/bart-large-mnli, HF_HOME redirected into the repo). Third, and most consequential: the very first real classifications measured only 5 of 9 test categories correct, because classify_zero_shot was passing raw config identifiers like tool_use_agentic and diff_checking directly as zero-shot candidate labels — HF's pipeline scores a label against a hypothesis template ("This example is {}."), and an underscored code token is not a sentence the model's NLI training ever saw. Mapping each category to a natural- language description before scoring, plus multi_label=True (the pipeline's single-label default forces every candidate to compete for the same probability mass), brought that to 8 of 9 correct with confidence scores 0.77-0.999 on the hits — the one remaining miss scored below the configured threshold and correctly fell through to the safe fallback rather than mis-routing.

classifier.mode (local_llm default, cloud_llm, local_encoder) picks what answers a classification request — a peer concept to the cascade above, not a replacement for it. Whichever mode is primary, a failure still walks the exact same cascade (stale session → session history → cloud_fallback → the static guess), unmodified.

  • cloud_llm makes a cloud model the primary attempt, not just the cascade's post-failure backup. Either pin one (classifier.cloud_primary, same shape as cloud_fallback) or set cloud_primary_auto: true to resolve the cheapest currently-routable model live against the catalog (routing.cheapest_classifier_candidate, priced for the classifier's own short-prompt/short-completion shape — not the task's). A success here records source="classifier", deliberately the same string local_llm's success uses, not "classifier_cloud" — that string means specifically "the cascade's backup step fired" and feeds the /metrics degradation-share warning above as a degraded signal. An intentionally configured primary succeeding is not degraded.
  • local_encoder classifies with a small, non-generative zero-shot model instead of an LLM — structurally immune to the runaway-reasoning failure mode documented above, since there is no generation to run away. Zero-shot rather than fine-tuned: this router never stores raw task text anywhere, so there is no training corpus without a new, separate opt-in capture feature (not built). Only produces task_category; task_tier falls back to fallback_tier — a real limitation, not a bug. A below-threshold confidence is treated as a failure and cascades exactly like a local-LLM parse failure would.
  • Neither is gated by local_compute.enabled (gaming mode, below) the way local_llm is: cloud_llm never touches local hardware, and local_encoder is small enough to run on CPU, so neither competes for the GPU gaming mode exists to free up.

Configured via config.yaml (global default) and overridable per-machine in config.local.yaml — the existing overlay, not a new mechanism — or through the admin portal's Classifier card, which reports the live resolved primary for cloud_primary_auto rather than echoing the config value (see admin-portal).

The default in config.yaml is still local_llm. local_encoder is intentionally an opt-in per-deployment choice rather than the repository default until the pending CPU-vs-CUDA comparison on the live classifier host is available.

Local dispatch model

A second local model can now be dispatched directly for specific categories. qwen2.5-coder-router:14b is configured as a tier-1 local row with provider='ollama-local', gated by models.eligible_categories (file_summarization and diff_checking). The poller refreshes the row each run; seed_local_dispatch_energy.py derives its price from measured GPU draw and the user's tariff. Routing treats a local row like any other candidate once the category filter admits it, and the circuit breaker excludes it on a local failure so the next request reroutes to cloud candidates.

Known limitations of the local dispatch branch right now:

  • No true streaming. The response is shaped into an SSE stream, but the local answer is generated before any bytes leave the router.
  • No verification rows. Structural and local-LLM checks run but are not written to verifications for local answers.
  • No within-request cloud failover into the upload path is gone. An eligible routed request degrades to the local dispatch model when the cloud account refuses or is exhausted (a fallback, not a preference), so the local row is no longer a dead-end before a client retry. The degraded answer is recorded as kind='local_dispatch_fallback'.
  • Follow-ups are not special-cased. A pinned or auto-routed follow-up to the same local model works, but nothing caches the loaded model between turns.

Dormant under the default profile by design — see docs/routing.md § Local dispatch branch.

POST /outcome now attributes through the local energy ledger too. Local rows include request_id and session_dir in local_energy_observations, so a client report on a local answer resolves to the same (model_id, provider, task_category) provider-agnostic record as a cloud one.

Routing notes

Ranking is quality-first, cost as a tiebreak; cost is never allowed to override a real quality gap. The optional objective.credit_attenuation block extends that tiebreak without changing it: when the block is enabled, a per-provider multiplier is applied to a candidate's comparison cost only, producing an effective_cost that breaks ties. The multiplier is derived from the provider's polled account balance (the balance_url path, such as OpenRouter), so a low prepaid balance can nudge a near-tie toward a healthier provider. The logged est_cost_usd and the decision history stay as raw catalog estimates. The multiplier is 1.0 for providers whose balance comes from per-completion allowance_remaining_usd telemetry (NeuralWatt), so normal overage readings do not bias routing.

Two semantics matter when reading the numbers. total_balance_usd is a sum of heterogeneous provider-reported readings: OpenRouter's prepaid credits plus NeuralWatt's overage allowance, which normally reads near -$0.004. It can be negative and it is not a single spendable figure. credit_attenuation.enabled deliberately lives only in the config file; it is absent from the admin persisted-config allowlist and from provider edits. Turning it on or off requires editing config/config.yaml and systemctl --user restart llm-router.service, because the dispatcher's cfg binds at import time.

What's NOT built yet — pick up here

Built: session-directory attribution, the local energy ledger, local model dispatch, admin profiles/proficiency/gaming-mode, and the configurable classifier backend (classifier.mode, including its admin card — all listed under "What's built and working" above).

As of 2026-09-05, PRs #26-#36 landed the gitignored config overlay, admin profile CRUD writing to it, the provider literal cleanup, quota balance/burn/runway, capability-aware ceiling and rejection warnings, the TUI schema catch-up, the classifier fallback cascade, and the admin portal uplift (proficiency page, profiles duplicate/coverage fix, gaming mode, the classifier-backoff bug fix). classifier.mode (this document's own section above) is a further, independent addition on top of that. plans/multi-provider-support.md is PARKED on provider selection — Z.ai was the recommendation and is no longer settled; the coupling surface in it is measured and still valid.

Known follow-ups recorded but not specced, both small:

  • The TUI decision table renders the literal "None" in the ctx cell when required_context_tokens is absent — the same defect the profile cell was written to avoid. See .omo/notepads/tui-overhaul/issues.md.
  • A NULL required_context_tokens raises TypeError inside the demand-ceiling comparison, and the column is nullable. Latent only: zero such rows exist today, checked on the live DB.

The items below remain open.

  1. Fine-tuning local_encoder on real traffic. Scoped, not built: classifier.training_capture.enabled (opt-in, off by default — a deliberate reversal of "never store task text", so it must be impossible to enable by accident), a classifier_training_samples table gated the same way report_outcome already filters proficiency (only rows whose classification_source is in the attributable set), and a train_local_encoder.py script matching eval_proficiency.py's conventions. local_encoder.py currently ships zero-shot only.

  2. Leaderboard priors are unfilled. leaderboards.yaml ships empty on purpose — inventing plausible-looking benchmark numbers would put fabricated data straight into routing, the same failure as the provider's static_fallback carbon constant this project already excludes. Until real sourced figures go in, a newly listed NeuralWatt family has no prior and relies entirely on self-eval accumulating. python leaderboard.py --check lists what is missing.

  3. Sampling depth for three models — now eco-only. 7 samples/model gives split-half agreement within 1.4x for 10 of 13, but kimi-k2.7-code-fast (29x), kimi-k3 (14x) and glm-5.2-flex (2.2x) are still unsettled. This no longer touches cost, which is priced per-request from the catalog, so it only affects eco — which is not an objective. Low priority unless eco comes back.

  4. Retry does not reach streaming. The iteration budget (iteration.py) retries after a failed check, but only on the non-streaming path — once bytes have gone to the client there is nothing to take back. Buffering to fix that would cost streaming itself, a worse trade for interactive work. POST /outcome is the answer for streamed traffic: it arrives afterwards, so it works identically either way.

Gaming mode, and the backoff that used to do nothing

The classifier circuit breaker did not break the circuit. _last_classifier_failure was written by _record_failure() and read by nothing — _classify_cascade gated only its cloud step, and on a different timestamp. So the router re-dialled a known-dead local classifier on every request. A stopped Ollama refuses immediately and costs little; a hung one, or a VPN-bound one that black-holes, costs the full 120s timeout_seconds per request for as long as the outage lasts. _classifier_backoff_active() is now the read, consulted before the client is constructed.

One detail there is load-bearing: recording the failure lives in classify()'s exception handlers, NOT in _classify_cascade. The cascade is walked for reasons other than a fresh failure, and if those re-stamped the clock, every request during an outage would push the deadline forward and the local classifier would never be re-probed while traffic flowed — a permanent outage wearing a circuit breaker's clothes.

local_compute.enabled (default true) is the outer gate over local hardware. Turn it off when you stop Ollama for a game and the router skips every local call rather than discovering the outage one timeout at a time: classifier, /health probe, local verification, local-vision fallback, and local dispatch rows (dropped in load_candidates, so a local row is never picked and then 503'd). /v1/models stops listing local rows, and an explicit pin gets a 503 naming the flag instead of NeuralWatt's unknown-model 400.

ONE flag the code reads, not a macro writing five keys — a macro is hard to undo cleanly, drifts the moment a sixth call site appears, and leaves nobody able to answer "why isn't the classifier running?" from one place. verification.local_llm_enabled, local_vision.enabled and local_energy.enabled keep their own meanings; this ANDs over them.

It REFUSES to engage without classifier.cloud_fallback — 409 on the runtime knob, a validation error at config load. Skipping the local classifier does not make classification remote; without a cloud classifier it stops classifying, and every request falls through to a static guess recorded as general_chat, a fully scored category indistinguishable from a real classification afterwards. A refusal, not a warning, because a warning is what nobody reads while their game is loading. Nothing auto-writes the block.

Cascade steps 1 and 2 still run ahead of the cloud call: a stale session classification is free and was a real classification of that same session, so paying to re-derive an answer already held is spending money for nothing.

A latent substring bug fell out of the profiles work. SQLite stores eligible_categories as a comma-joined string, and task_category not in "<a>,<b>" is a SUBSTRING test — so a row eligible only for file_summarization also admitted summarization. Latent on main (the category-less probe short-circuits before the compare) and live the moment anything probes per category. routing.parse_eligible_categories is now the single parser dispatcher.load_candidates and admin.py's probe both use.

Known open questions

  • Answered: cost and eco stay separate axes — grid intensity spans 13.6x across the catalog, so they rank models differently.
  • Answered: the GLM rows reporting grid_id: FI at 475 gCO2/kWh were carbon_source: static_fallback — a substituted constant, not a measurement. They are now excluded from eco rather than trusted. Still worth asking NeuralWatt why the fallback keeps the original grid_id, since that is what made it look like a real regional difference.
  • Three models still fail a split-half stability check at 7 samples. Is the instability real (variable serving conditions) or an artifact of when the sweep ran? Re-sweeping at a different hour would tell.
  • Answered, and the question no longer parses: tier-1 composites used to sit within 0.009 of each other because min-max normalization compressed them. There is no composite any more — ranking is quality first, cost as the tiebreak inside quality_tolerance — so nothing normalizes and nothing compresses.
  • Answered: the eval set exists (evals/tasks.yaml, 43 tasks, four scoring kinds) and tests/test_task_set.py keeps it honest. The benchmark-sourced rows now split the coding categories, but two tasks are still flat at 1.00 (debug_affine_coprime and most of the BFCL set) and either need hardening again or should be conceded as non-discriminating. Try samples before hardening. docs_writing looked flat at the top too, and six more passes spread it 0.66-0.97 without touching a task; two samples per model is not enough to tell a saturated task from an unsampled one.
  • How much context-assembly (RAG-style retrieval) belongs in the classifier step vs. a separate pre-step? Leaning decoupled, undecided.
  • Should eco_score use real-time grid carbon intensity per request or a stable per-model average? Currently the latter, from the reference sweep. grid_carbon_intensity and grid_id are logged per observation, so this stays answerable from data without a re-run.

Config is strict: an unknown key is an error

Pydantic ignores extra keys by default, which means a typo or a misplaced setting loads cleanly, does nothing, and still looks configured. Every config model now inherits StrictModel (extra="forbid"), so both of these fail at load rather than silently:

verification.max_input_chars     # right key, wrong section
routing.min_tool_proficency      # sic

This is not hypothetical. max_input_chars shipped into the verification: block instead of classifier: and was accepted and discarded — it happened to match the code default, so behaviour was correct and the file was a lie. Editing it would have done nothing.

The corollary worth keeping: every knob belongs in config.yaml, not only in a Pydantic default. A default the file never mentions is invisible to anyone tuning it. classifier.outcome_attribution_window_seconds was removed in the same pass — it was declared, never read, and shadowed the verification one that actually is.

Setup

Full install steps (venv, deps, config, first run) in README ## Installation. Host-local deployment values go in config/config.local.yaml (gitignored overlay) — classifier.model/base_url, local_energy.*, host-specific URLs. General defaults in config/config.yaml stay shareable. objective.plan_kwh_per_period in the README config table. Model tags (num_ctx) + verification.model same-tag note in docs/local-models.md. Requirements are pinned — bump deliberately (README).

Run as a service

deploy/ holds the dispatcher's systemd user unit plus a timer and a oneshot service each for the poller, the seed sweep, the backup, the offsite sync and the feedback fold, and one drop-in for a system Ollama — see deploy/README.md for install and operation. Every timer there is enabled on install except llm-router-feedback.timer, which is not, on purpose. In short:

echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env
cp deploy/llm-router*.{service,timer} ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now llm-router.service llm-router-poller.timer

The dispatcher binds 127.0.0.1:8080. The poller timer is load-bearing, not housekeeping — but not for the reason this section used to give, and the correction matters because it inverts which failure to watch for.

mark_stale runs only inside poller.main(), and main() returns early on a RequestException — before upsert and before mark_stale. So a stopped timer or a provider outage marks nothing: the catalog freezes at last-known-good and the router keeps routing on prices that may be weeks old. The failure is silent and open, not loud and closed. Nothing surfaces it, because a frozen row still reads availability = 'active'.

The path that can empty the candidate set is narrower and is not governed by the timer at all. fetch_neuralwatt reads payload.get("data", []) with no floor on row count, so a 200 response carrying an empty or truncated data array — a partial provider outage, a schema change, an auth path degrading to an empty list — clears raise_for_status(), upserts nothing, and then lets mark_stale run anyway. Three days of that and every row is stale and exclude_stale: true leaves zero candidates for everything.

stale_after_days: 3 only sets the length of that fuse; it does not arm or disarm it. Against a 2-hourly poll it is 36 successful polls of margin, and recovery is automatic — upsert writes availability = excluded.availability, so one good poll flips every stale row back to active. The fix is a sanity floor on the fetch, not a larger number.

A second path empties the candidate set, and it bit on 2026-09-01. Admin availability overrides are not governed by the poller at all. Deprecating the seven expensive models through /admin collapsed tier 3's context ceiling from 782,324 to 94,196 — while tiers 1 and 2 stayed at 782,324 — so every tier-3 request above 94k returned 422 with nothing warning anywhere. It surfaced ~19 hours later as an agent failing mid-task on an opaque error.

The obvious check for this is wrong, and the reason is worth remembering. Tempting: warn when a higher tier's context ceiling sits below a lower tier's. But ceiling(T) is the max effective_context_window over models with tier >= T, and tier is a capability floor, so the eligible set shrinks monotonically as T rises — ceiling(1) >= ceiling(2) >= ceiling(3) is a theorem, true of every catalog. Such a warning fires always and means nothing. What actually failed is that a tier's ceiling dropped below what that tier is asked to serve, which is only knowable from traffic: compare ceiling(T) against the observed required_context_tokens for decisions classified at tier T. That is silent on all three tiers today and fires on the outage state (94,196 vs an observed max of 268,168). See plans/catalog-staleness-and-poller-failure-modes.md §4.4.

This recurred on 2026-09-04 through a dimension the detector did not model, and both halves of the fix are now in metrics.py. Admin deprecations took out kimi-k3* — the only vision-capable rows with enough context — so a 242,486-token image request 422'd while every existing check stayed silent, because the all-models tier-1 ceiling was still 782,324. The vision-capable ceiling had collapsed to 192,500.

  • Predictive: capability_ceilings / capability_demand_warnings compute vision and json_mode sub-ceilings and compare each against demand actually observed for requests carrying images / requesting JSON. Two extra series, not a bucket per capability combination.
  • Reactive: rejection_warnings watches route_decisions for rows with selected_model IS NULL. This is the more valuable half and the simpler one — it catches the next dimension nobody predicted, at the cost of firing after the first failure rather than before.

Two details in the reactive detector are load-bearing and easy to undo by accident. It groups by (task_tier, digit-normalized reason) using the structured column, because normalizing digits alone merges tier >= 1 and tier >= 3 rejections into one group and hides whether the broadest or the frontier candidate set went empty. And the signal is novelty OR rate, never mere presence: measured on the live DB, routine rejections run ~3/hr while the 2026-09-04 incident was only n=2 — below the noise floor — so no single count threshold can both catch it and stay quiet. A group absent from the 24h baseline warns at n≥2; a familiar group warns at the configured count. Zero rejections warn about nothing: a genuinely impossible request SHOULD 422.

The service holds a billable API key and has no auth of its own. Loopback bind is the only thing standing between the open internet and your allowance; add auth before widening --host.

The same applies to an Ollama shared over a VPN — it has no auth either, so deploy/ollama-over-vpn.conf binds it to the VPN address rather than 0.0.0.0, which would publish it on whatever network the client happens to be on.

When the router goes unreachable, start at docs/incidents.md

Four incidents so far, all sharing one shape: a change that looked local to the router silently degraded the agent depending on it, and none announced itself as a router problem. docs/incidents.md carries the full write-ups plus a symptom -> one-line-check table; read it rather than re-deriving a diagnosis.

Two conventions from those incidents that bind every session, and so stay here:

  • 8080 is production, always. It is baked into opencode.json, the systemd unit, every curl example here, and the admin frontend's own fetches. A throwaway instance (manual iteration, Playwright smoke tests, anything that is not "use the real router") binds 8081. Never send a kill signal to a process matched by name or port rather than by a PID you started yourself -- Restart=always will fight you, and on this repo it may be your own model access.
  • Never point config/config.yaml at test fixtures. It is the file the live service reads. Pass a different config file, monkeypatch cfg.database.path in-process, or use a temp copy.

Recovery for an unreachable-but-active service is systemctl --user restart llm-router.service -- a hung process was never in a tracked stop job, so this issues a fresh cycle systemd does enforce a timeout on.

Pointing a coding agent at it

The /v1 endpoints are OpenAI-compatible, so any normal client works — opencode, an SDK, plain curl. Repo-local opencode.json is already wired up, so running opencode from a clone of this repo routes by default. For global use, merge provider.llm-router into ~/.config/opencode/opencode.json.

model name behavior
auto router picks; flex rows excluded so nothing is held during peak
auto:batch router picks; flex rows admitted, for overnight/async work
any real model id dispatched as asked, still logged

Streaming is proxied chunk by chunk rather than buffered, so tokens still render as they arrive. NeuralWatt emits its energy and cost blocks as SSE comment lines (: energy {...}) before data: [DONE] — ordinary clients ignore comments, so the stream passes through untouched while the router reads the telemetry on the way past. Without that, streamed calls would log no energy at all, which is most of the point of this project.

Try it

Copy-pasteable curl examples (/route, /dispatch, /v1 models + chat) and the classifier-skip overrides (task_category, task_tier, required_context_tokens) live in README ## Usage. Inspect what a dispatch cost/burned with the sqlite3 query in docs/operations.md.