classify_zero_shot embedded the RAW task text: a prose instruction wrapped in realistic agent-session noise (fenced code block, "Tool result:" line, <system-reminder> tag) mean-pooled the noise tokens with equal weight, and the verdict scored the noise. Confirmed on the real BAAI/bge-large-en-v1.5: 8 clean one-sentence tasks scored 8/8, the same instructions wrapped scored 2/8 on SHORT inputs far under any truncation limit — a second bug, distinct from the fixed 512-token collapse; tail-biased windowing alone also measured 2/8 on long noisy pairs, so the fit does not subsume isolation. New _isolate_task_text() (pure, stdlib-only re, deterministic) strips fenced code blocks, tool result/output/call lines, and closed <system-reminder> spans before the embed pass; classify_zero_shot now runs isolate -> fit -> prefix. Measured decisions: pure removal beats '[code]'/'[elided]' placeholders (8/8 vs 7/8, 6/8 on long noisy pairs — the placeholder token itself pulls toward code categories); inline code spans stay (file names in instructions are signal); a 20%-ratio floor guard measured harmful (3/8 — it reverts exactly the short noisy inputs) in favor of an absolute 24-char floor that only falls back on near-all-code inputs. Post-isolation: 8/8 on short and long noisy pairs, clean-vs-noisy pair consistency 8/8 + 8/8; code-grounded instructions improved 2/4 -> 3/4 (residual miss is description similarity, not noise). Regression: a noise tripwire (same instruction bare vs wrapped must classify identically) fails against pre-isolation code (empirically confirmed: docs_writing -> coding_general) and passes after; the existing truncation tripwires still pass. CLAUDE.md open item #5 updated to reflect what is built and what is not (attention-masking de-weighting, unclosed tags / un-fenced diff hunks, grounded-task residual). Full suite 2170 green. Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
77 KiB
Local LLM Model Router — project brief
README.md is the concise front door. Deep-dive module-by-module reference
docs live in docs/.
design/local-llm-model-router.md holds the architecture and rationale,
including parts still unbuilt. This file is the working state + immediate
next steps, and is the one to trust on what is currently true.
NORTH STAR GUIDELINES
Three rules that outrank local cleverness. Each exists because it was broken first and the breakage was expensive to find. When a change conflicts with one of these, the change is wrong — not the rule.
1. Every config knob is reachable from the admin page
Every scalar knob in config.yaml gets an admin control — runtime, persisted,
or both as its mechanism warrants — or a recorded decision saying why it
must not. The absence of a control has to be a decision someone made, never an
oversight nobody noticed.
objective.credit_attenuation.enabled is the model exception: deliberately
absent, because enabling it must be a config edit plus a restart. That is a
recorded choice, not a gap.
Why: Wave 2 shipped incumbent_cache_pricing and
incumbent_challenger_cache_rate with no control at all. Nobody decided that;
it just never came up. The dial's entire purpose is tuning from neutral to full
without reverting code — and hand-editing a tracked file and restarting is a
loop nobody walks, so the knob's rationale evaporated on contact with reality.
Enforced by tests/test_admin_knob_coverage.py, which fails naming the knob.
Honour it; do not add a DELIBERATELY_NOT_IN_ADMIN entry to silence it unless
the reason is true.
2. Fix classifier latency by improving the classifier, never by reusing a
stale decision
Classification latency is a real problem and the answer is a faster or better classifier — a smaller model, a warmer process, a cheaper backend. The answer is never to let one classification stand in for later, different work.
A stale label does not merely add noise. It silently redirects money: the whole point of this router is sending each task to the model that fits that task, and a replayed label routes the next task to whatever fitted the last one.
Why: the session cache reached 96.6% of all classifications. One real classification drove 107 consecutive turns across 14 minutes on 2026-09-16. It was introduced to avoid paying a 2-5s round trip per turn, which was a reasonable trade in isolation and became the dominant path without anyone choosing that.
Two costs, and the second is worse than the first:
- Routing decides on what the session was doing when it started, not what this turn is. A mid-session pivot — docs, then bug fixes — routes the bug fixes on the docs label.
- Proficiency is trained on those labels, because
POST /outcomeattributes to(model, task_category). That is part of howdeepseek-v4-flashcame to hold a saturated 1.000 onfile_summarizationwhile failing 50% of them on real traffic.
It also made a working classifier look broken: replaying a handful of session
labels across hundreds of turns produced "pages and pages of diff_checking",
which read as a classifier stuck on one label when the underlying
classifications were a reasonable mix.
If latency forces a cache, that is a measured, time-boxed concession with an expiry condition written next to it — not a default.
3. Check worktrees and other agents' work before touching any file
Before editing, run git worktree list and check for other agents or sessions
working in the same tree. Do not assume a worktree is yours. When parallel work
is unavoidable, isolate it — a separate worktree per agent — and commit by
explicit path, never git add -A.
Why: two agents were once launched into the same worktree while each was told the other was elsewhere. It survived only because the files were disjoint and one of them committed by explicit path rather than sweeping the tree. The second agent's test count was also measured against a tree containing the first's uncommitted work, so its "verified green" was not a clean signal.
Related: a router.db inside a worktree is a stale copy, not the live
database. The live one is /home/alee/Sources/6krrt/router.db. Check mtimes
before measuring anything, and open it read-only:
sqlite3.connect("file:...?mode=ro", uri=True) — the sqlite3 CLI here does
not accept -uri.
What this is
A router that uses a local model (served via Ollama) to classify incoming coding/documentation tasks — category, tier, required context size — and dispatch each task to the best-fit open-weight model on Neuralwatt Cloud, under a per-request cost ceiling, tiebroken by price, and ranked by category-level expected pass rate.
Every measurement in this file was taken on one deployment against one provider account. They are recorded because the reasoning is worth more than the conclusion, but treat them as observations with a date on them, not as constants — the catalog, prices, grid intensity and pool load all move. When a number here decides something, re-run the measurement before trusting it.
Neuralwatt remains the primary provider. OpenRouter is back as an opt-in
provider gated by the provider_model_allowlist table — a provider with
require_allowlist: true is ignored unless its model_id is explicitly
allowlisted, and any previously-upserted row that drops off the list is
marked deprecated on the next poll. The provider column and the
(model_id, provider) primary key are unchanged, so this adds no migration.
Stack
- Python (chosen over Rust — this is I/O-bound against provider APIs, not CPU-bound; iteration speed on the scoring/weighting logic matters more than raw execution speed at this scale)
- SQLite for the decision table
- Ollama for local classification, via any OpenAI-compatible endpoint —
localhost:11434/v1, or an Ollama on another machine across a VPN (classifier.base_url) - FastAPI for the dispatcher service
Billing is per-kWh, not per-token — and neither is what scoring uses
Measured against the live API, 2026-08-11. Neuralwatt bills a flat
$8.00 per kWh and the catalog's input_per_million /
output_per_million prices are not what this account is charged. Confirmed
across five models; cost_usd / energy_kwh came back 8.00 every time:
| model | list $/1M out | completion tokens | billed USD | kWh | $/kWh |
|---|---|---|---|---|---|
| deepseek-v4-flash | 0.28 | 600 | 5.00e-06 | 5.88e-07 | 8.50* |
| gemma-4-31b | 0.42 | 540 | 3.80e-04 | 4.75e-05 | 8.00 |
| qwen3.6-35b-fast | 1.15 | 600 | 2.53e-04 | 3.16e-05 | 8.01 |
| kimi-k2.7-code-fast | 4.00 | 600 | 2.25e-03 | 2.81e-04 | 8.00 |
| kimi-k3-fast | 15.00 | 539 | 2.17e-04 | 2.71e-05 | 8.00 |
* rounding — billed cost is quantized to ~1e-06.
The precise rule, validated against all 65 samples of the reference sweep (61/65 within 2%; the 4 outliers are microdollar rounding, not misses):
cost_usd = min( $8.00/kWh x energy_kwh , 3 x list token price )
The ceiling bound in only 2 of 65 samples, both deepseek-v4-flash energy
spikes, and matched to the cent: 31 prompt x $0.14/1M + 400 completion x
$0.28/1M = 1.1634e-04, x3 = 3.4902e-04, billed 0.000349.
List price ranks models backwards. Not approximately — invertedly:
| list $/1M | actually billed | gCO2eq | |
|---|---|---|---|
deepseek-v4-flash |
0.28 | 9.80e-05 | 8.90e-04 |
gemma-4-31b |
0.42 | 5.02e-05 | 2.32e-04 |
deepseek lists 33% cheaper, costs 95% more, and emits 284% more carbon.
But cost and eco are NOT the same axis — the tempting simplification, and it is wrong. Cost tracks energy, but carbon is energy x the serving region's grid intensity, and models run in different regions:
| grid | gCO2/kWh | models |
|---|---|---|
FI |
~49-50 | most of the catalog; varies by time |
FI (reported) |
475 | glm-5.2-fast, glm-5.2-flex |
US-MIDA-PJM |
~442 | the kimi-k3 family |
A 13.6x spread, so the two axes disagree: glm-5.2-fast is the 2nd cheapest
model and only the 6th cleanest; kimi-k3-flex draws 3.7x less energy than
kimi-k2.7-code while emitting 3.6x more carbon. Weighting them separately
is load-bearing, and tests/test_routing.py pins it.
Superseded — cost no longer comes from the sweep at all. cost was the
median measured USD over the reference sweep. That was measured to be WRONG
for real traffic, because the reference workload is the wrong shape.
The sweep sends a 400-token prompt with a 400-token completion. Real agent traffic is a 150,000-token prompt with a ~400-token completion and ~92% cache hits (token-weighted over 50 sessions and 40.7M tokens on 2026-08-23; the figure was 84% when measured on 2.2M tokens of earlier traffic). The attribution ratio moves with prompt size, so the ranking inverts:
| workload | winner |
|---|---|
| reference sweep (400/400) | glm-5.2-fast, 3.2x cheaper |
| realistic (70k prompt, short answer) | deepseek-v4-flash, 5.0x cheaper |
Same two models, opposite answer. glm-5.2-fast sits at attribution 0.006 on
a toy prompt and 0.50 on a 70k one — it batches beautifully on small prompts
and badly on real ones. deepseek-v4-flash barely moves (0.21 -> 0.25).
So routing.estimated_cost prices each request from catalog token prices,
scaled to that request's actual shape (prompt size, assumed completion
length, assumed_cache_rate). List price is not what gets billed, but billing
is capped at 3x list, so it tracks the real ordering and bounds it — and on
the one case that was checked live it agrees with the measurement in direction
and magnitude (7.8x predicted vs 5.0x measured). It is also free, needs no
sweep, and refreshes whenever the poller runs.
objective.plan_kwh_per_period is a planning figure only: per-request traffic
is never refused for exceeding it — it gates nothing. Overage is billed
against the account's credit balance (allowance_remaining_usd from the
provider). The /admin/api/snapshot endpoint reports balance, estimated burn
rate, and projected runway per provider inside quota.accounts[] — a
list of per-provider billing shapes (metered_plan, prepaid_credit,
self_hosted, or unmetered) with plan, pool, burn, and credit
blocks as appropriate. quota.spend aggregates provider spend and a
list-price estimate. The old flat keys and by_provider/total_balance_usd
shape were removed; every consumer was updated in the same change, so there
are no deprecated aliases.
Three signals said deepseek-v4-flash — catalog token price (7.8x cheaper),
NeuralWatt's own published per-request energy (~10x lower), and a live 70k
measurement (5.0x cheaper). Only the 400-token benchmark disagreed. Trust the
workload you actually run.
eco still comes from the sweep's median gCO2eq, and is still not an
objective. flex_cost_multiplier is gone: a flex row's measured cost already
is its flex cost.
Open, and worth knowing: NeuralWatt's model cards publish gross energy
(~1.99e-04 kWh for deepseek, ~1.91e-03 for GLM), while the billed figure is
gross x attribution. GLM burns roughly 7x more actual electricity per request
and charges ~5x less, because far more tenants share its GPUs. Anything built
on eco inherits that inversion — the attributed carbon figure answers "what
is my share", not "what was burned".
Energy attribution: signal that looks like noise
Billed energy decomposes exactly:
energy_kwh = avg_power_watts x duration_seconds x attribution_ratio
attribution_ratio is the request's share of a shared multi-tenant GPU pool.
Up close it looks like pure noise — eight rapid identical calls to one model
spanned 20x in billed energy, correlating +0.997 with the ratio while
power and duration held steady. Two sweeps of the same 13 models with the
same prompt disagreed by up to 36x.
Scoring on the pre-attribution product (power x duration) was tried, and it
is wrong. Across the sweep:
| spread | |
|---|---|
| median attribution, between models | 750x |
| typical spread within one model | 1.8x |
The ratios are quantized (0.001, 0.25, 0.5, 0.75) — that is serving
concurrency, a stable per-model property, not weather. A model whose GPUs
carry far more concurrent requests genuinely costs less per request, and
that is most of the real cost difference in the catalog: deepseek-v4-flash
bills ~1000x under its share of pool gross. Stripping attribution discards a
750x real signal to suppress a 1.8x one.
So scoring reads the attributed figures, and the median absorbs what noise remains. A split-half check on the 7-sample sweep (median of first three vs last four) shows that working:
- 10 of 13 models agree within 1.4x — stable enough to route on
- 3 do not:
kimi-k2.7-code-fast(29x),kimi-k3(14x),glm-5.2-flex(2.2x). Those need more samples before their position is trustworthy.
dispatcher.gross_energy_kwh remains as a diagnostic on the identity, not a
scoring input.
Attribution drifts across hours, so sampling must too
Within about 30 minutes the billed figures reproduce (0.3-1.1x on a
spot-check). Across hours they do not: between two sweeps,
deepseek-v4-flash moved roughly 50x and qwen3.6-35b about 7x the other
way — enough to invert their cost ranking. Attribution tracks pool load,
and pool load tracks time of day.
More samples inside one sweep does not fix this; it measures one moment more
precisely. Coverage across time does. load_candidates already takes the
median over ALL seed_reference rows, so repeated sweeps accumulate into a
median-across-time for free — hence llm-router-seed.timer, which runs a
small sweep every 6 hours.
Until several sweeps have accumulated, treat the eco ordering as provisional. A single sweep's ranking is one sample of a moving quantity.
And none have accumulated since 6e729ad. That commit moved
log_observation's trailing arguments to keyword-only without updating
seed_energy.py, so every timer run since spent one billed completion and then
died on TypeError — which is not a RequestException, so the per-sample
except did not catch it. Fixed, and the sweep now has an offline end-to-end
test, but the accumulation this section describes starts from the next run
rather than from months of history.
What's built and working
config/schema.sql—models,proficiency,energy_observations; applies cleanly (sqlite3 router.db < config/schema.sql). See data-model.poller.py— fetches Neuralwatt's catalog (public, unauthenticated), normalizes, upserts, marks stale. Verified live: 14 routable models. See data-model. For providers withrequire_allowlist: true(openrouterin the base config), the poller filters the fetched catalog againstprovider_model_allowlistbefore upsert and prunes existing rows that are no longer on the list todeprecated— see the #45 OpenRouter opt-in allowlist section below.config/config.yaml/src/config.py— weights, thresholds, provider settings, Pydantic-validated. architecture.scoring.py— onenormalize_inverted(cost and eco normalize identically) + the weighted composite. routing.seed_energy.py— reference task × N per model →energy_observationstaggedseed_reference; makescost/ecoreal.--samples 5= 65 calls, under a cent. architecture.tiering.py/tier.py— pure tier resolver + DB pass. Why tier onreasoning_default_enabled, cheapness-not-ceiling,tier1_context_max: routing#tiering.routing.py— pure hard filters + ranking, plus the request-side capability gates (fail-closed asymmetry). routing.routing.pyrank_candidatesincumbency: prices the session's last chat model at its measured cache rate and every challenger at theobjective.incumbent_challenger_cache_ratedial, behind a load-bearingmin(dial, incumbent_rate)clamp; gated off by default (incumbent_cache_pricing: false), tunable from off to full in config. routing#incumbency-and-cache-pricing.circuit_breaker.py— passive availability skip on a 5xx (cooldown + backoff, clears on next success, no poller), on by default. Eval harness deliberately stays outside it (isolation): routing#circuit-breaker. Coversollama-localtoo: a local outage raises 502 on the first request and the breaker excludes the dead local row on the next one, so traffic reroutes to cloud candidates.dispatcher.py— FastAPI service:GET /health,POST /route(no provider call),POST /dispatch, OpenAI-compatible/v1/models+/v1/chat/completions, SSEGET /events/decisions. api. On an account-level cloud refusal/exhaustion, eligible routed requests degrade to the local dispatch model instead of surfacing the cloud error (see the Local dispatch model section below).proficiency.py/proficiency_store.py/proficiency_outcome.py— blend leaderboard + self-eval into a benchmark prior, accumulate client outcomes, and recompute expected pass rates; the only write paths toproficiency, soblended_score/sourcenever drift. architecture.context_prune.py— relevance-based stage trimming only tool results once overbudget_tokens, before any paid token ships. See pinch forbudget_tokens; see also theprotected_max_charsnote there if you are changing how much prefix context is guarded.feedback.py— foldsPOST /outcomeclient reports intoproficiency.outcome_scoreviaadd_outcome(). Structural andlocal_llmverdicts are diagnostics only;POST /outcomeis the posterior. verification.--dry-runno longer just describes the fold, it projects it:feedback_preview.pycopies the DB into memory, runs the realadd_outcomeagainst the copy, and reports the per-row before/after, opening the sourcemode=roso a preview cannot write to what it is previewing.deploy/llm-router-feedback.{service,timer}gives this loop the timer it never had — and ships not enabled, because the fold is irreversible and the first one against an accumulated backlog is an operator decision. Seedeploy/README.md.exploration.py— epsilon-greedy exploration chooser; injected RNG, no mutable state. routing.seed_local_dispatch_energy.py— standalone reference-shape sweep forollama-localrows; derives per-token USD rates through the user's tariff and OLS on measured GPU draw. architecture.poller.py— also seeds/updatesprovider='ollama-local'rows fromconfig.yamleach poll so local rows stay current even when NeuralWatt is unreachable.logs.py— per-request trace id (ContextVar), logfmt, journald priority prefixes;logs.bind()survives StreamingResponse generators. operations.metrics.py/GET /metrics— read-only observability; takes(conn, cfg), never importsdispatcher. Also carries the three detectors added after the incidents below: capability sub-ceilings, the reactive rejection detector, and the classifier-degradation share. api.tui.py— Textual dashboard over/metrics+/events/decisions; live feed, category→model panel, detail popup; data layer split intotui_model.py. The decision table leads with atimecolumn and carriesprofileplus anEflag for exploratory picks; the quota panel now shows per-account billing shapes (metered_plan,prepaid_credit, etc.) with plan/pool/burn/credit blocks, spend aggregates, and an alarm line. architecture.tests/test_tui_schema_drift.py— the tripwire that keeps the two honest. A newroute_decisionscolumn must be registered as surfaced or deliberately-not, or the test fails naming the column. Five columns had already reached the schema without reaching the dashboard;ROUTE_DECISIONS_COLUMNSintests/test_route_decisions.pyhad itself drifted.tests/test_tui_warnings.py— the same idea for warnings. Every class/metricscan emit must render in#warnings-panel, and every emitted warning must be registered — the second failing with the RAW text, because the point is that nobody knew the class existed. Its fixture is a coupled system: adding a seed can silence an existing class (a small-context seed once killed the escalation hazard by dragging the p95 down), which is why both directions are asserted.router_cli.py— one-shot/routeprobe (no spend), raw JSON with--json. api.admin.py/config/admin_schema.sql/admin/frontend/*.html— loopback/adminportal: dashboard, models overrides, decisions log, profiles, a read-only proficiency matrix (GET /admin/api/proficiency) that distinguishes a measured score from an inherited one, and controls (including the Local Compute andclassifier.modecards). admin-portal. Provider management includes list/detail/update/delete endpoints (GET /admin/api/providers,GET /admin/api/providers/{name},POST /admin/api/providers/{name},DELETE /admin/api/providers/{name}) plus per-provider allowlist endpoints (GET/POST /admin/api/providers/{name}/allowlist,DELETE /admin/api/providers/{name}/allowlist/{model_id}); the providers page showsrequire_allowlistand links to the allowlist editor.local_encoder.py— zero-shot category classification via a non-generative encoder, backingclassifier.mode: local_encoder.transformers/torchimported lazily; a deployment that never selects the mode needs neither installed. local-models.provider_model_allowlisttable — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist becomedeprecatedon the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below.config.py/DispatchProvider.require_allowlist— Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted.tests/— 1155 tests across 40+ files, offline, verified on Python 3.10 and 3.14. README.
#45 — OpenRouter is an opt-in allowlist provider
OpenRouter used to be ingested whole, then removed, and is now back — but only
as an opt-in provider. The base config sets openrouter.require_allowlist: true
and ships a short seed allowlist. Models on that seed list are upserted and
kept active; anything else in the OpenRouter catalog is filtered out before
upsert and any previously-active OpenRouter row that is not on the list is
marked deprecated on the next poll.
This is deliberately different from the old ingest-everything behavior. The
previous approach once pulled in a non-chat model (lyria/...) that returned
HTTP 404 on dispatch because the endpoint expected chat completions. The router
had paid for the classification, selected the model, and then failed on the
provider call. Allowlist-gating prevents that class of failure by default: if a
model id has not been reviewed and explicitly added, the router acts as if it
does not exist.
The seed allowlist is short and has firm exclusions. It does NOT include:
x-ai/*(Grok)openai/*anthropic/*
Those exclusions are non-negotiable. They are not "currently excluded" or planned for future inclusion; they are deliberately absent from the seed list. Adding one requires editing both the seed allowlist and this file.
The admin portal exposes the allowlist under /admin: the providers page shows
which providers require one, and each provider row links to an allowlist editor
where entries can be added or removed. The underlying four API endpoints are
GET /admin/api/providers/{name}/allowlist,
POST /admin/api/providers/{name}/allowlist,
DELETE /admin/api/providers/{name}/allowlist/{model_id}, and the providers
page itself surfaces require_allowlist with a link to the allowlist editor.
Neuralwatt remains the primary, ungated provider. The allowlist behavior only
fires for providers with require_allowlist: true.
Proficiency: category now changes routing
proficiency_score is the ONLY category-dependent term in the ranking, so
until this table had data, task_category could not change a decision at
all — the classifier computed it, the router paid ~10s for it, and then it
made no difference. It does now. Two categories were added for local dispatch:
file_summarization and diff_checking; see evaluation. The score is also no longer a raw benchmark
level: it has been converted into an expected pass rate on real traffic,
calibrated against 1,059 client-reported outcomes and shrunk with a
20-pseudo-observation prior so thin data does not dominate.
proficiency now holds 141 rows:
| source | count | meaning |
|---|---|---|
outcome_blended |
33 | Fresh per-model outcome evidence |
outcome_prior |
76 | Trafficked-sibling rows inheriting the peer-rate prior |
self_eval_thin |
32 | Cold categories (summarization, translation); benchmark preserved verbatim |
A further 33 (model, category) pairs have direct outcome samples. The biggest
evidence gains went to deepseek-v4-flash (notably coding_general,
coding_refactor, and general_chat), kimi-k2.7-code (across 7
categories), and qwen3.6-35b.
Sweeping 9 categories x 3 tiers currently returns 5 distinct winners at both 50k and 120k of context.
| context | winners over 27 decisions |
|---|---|
| 50k | qwen3.6-35b (10), gemma-4-31b (7), deepseek-v4-flash (5), kimi-k3 (3), kimi-k3-fast (2) |
| 120k | kimi-k2.7-code (10), gemma-4-31b (7), deepseek-v4-flash (5), kimi-k3 (3), kimi-k3-fast (2) |
This spread is recent, and how it got here is the useful part. For a long
time all 27 decisions returned ONE model, and that was the correct answer at
the time rather than a bug: with cost and eco both populated, qwen3.6-35b
was Pareto-dominant — cheapest AND cleanest in the routable set, while
scoring within quality_tolerance of the best. No defensible weighting picks
anything else out of that.
Two corrections widened it, and neither was a tuning change:
- Cost stopped being a benchmark average. It is now priced per request from catalog prices scaled to the request's shape, so the ranking depends on the workload instead of on a 400-token reference sweep that no real traffic resembles.
- Tier stopped being inferred from price.
deepseek-v4-flashwas pinned to tier 1 for being cheap, which excluded it from every tier-2 request regardless of what any score said.
A third shift is under way: the outcome backlog has been spent, so the score
now reflects real pass/fail reports rather than the benchmark alone. That
changes the numbers; it does not change the rule. Quality is still the
objective and cost is still the tiebreak within quality_tolerance.
Note what changes between the two rows above: only the leader, and only because of the hard context filter. That is the filter working, not the scoring disagreeing with itself.
If you see one model win everything again, check for dominance before
reaching for config. One winner is a legitimate outcome. The levers, if a
genuinely different balance is wanted, are objective.quality_tolerance
(how large an expected-success-rate gap must be before it outranks a cost
saving) or objective.max_energy_per_request (a hard ceiling). There is no
weight to tune.
What the task set actually found
The benchmark could not discriminate these models on coding. Every row
scored exactly 1.00 on coding_general, coding_refactor and debugging —
and that is after the tasks were deliberately hardened with touching
intervals, full semver, present-but-falsy defaults, late-binding closures and
a binary search that infinite-loops. Every model in this catalog is simply
good at that class of problem, so cost decides coding routes, which is the
right outcome.
Real traffic broke one of those ties, which the benchmark never could.
coding_general now spans 0.86-1.00: glm-5.2-fast fell to 0.862 over 29
samples folded in by feedback.py from an actual agent session, and crossed
self_eval_min_samples on the way, so it reads self_eval rather than
self_eval_thin. That is the intended shape of this system — the 43-task
benchmark establishes a floor, and your own traffic is what refines it.
coding_refactor and debugging are still flat at 1.00, awaiting the same
treatment.
Benchmark-sourced hardening landed, and it broke both remaining ties. The
task set grew from 23 to 43: 8 BFCL tool tasks, 6 CRUXEval-O exact tasks (after
the score_exact literal-eval fix), and three Exercism refactor/debug pairs.
The CRUXEval-O rows split coding_general into a 0-1 mix across models — five
of six now fail at least one model — and the Exercism refactor rows moved
coding_refactor off its flat 1.00: bowling 0.25-1.00, dominoes 0.10-1.00,
affine 0.56-1.00. The debug pairs split debugging the same way (bowling
0.80-1.00, dominoes a sharp 0/1 split). Two follow-ups recorded, not deleted:
debug_affine_coprime and most of the BFCL tasks sat flat at 1.00, so they do
not discriminate.
A 1.00 can also be a sampling artifact, and docs_writing was one. At 2
samples per model the category read 0.70-1.00 with a model at the ceiling, and
the router paid for that ceiling: kimi-k3-fast won every docs route. Six more
benchmark passes moved every score and left NOTHING at 1.00:
| model | n=2 | n=11-14 |
|---|---|---|
kimi-k3 |
0.85 | 0.973 |
kimi-k2.7-code |
0.85 | 0.886 |
deepseek-v4-flash |
0.80 | 0.864 |
kimi-k3-fast |
1.00 | 0.864 |
qwen3.6-35b |
0.85 | 0.800 |
gemma-4-31b |
0.85 | 0.786 |
The winner moved to kimi-k2.7-code, 3.2x cheaper at 50k of context
($0.0441 -> $0.0136), with no config change — kimi-k3 scores higher but sits
inside quality_tolerance, so cost breaks the tie. deepseek-v4-flash
($0.0024) misses the band by 0.009, which is the kind of margin the tolerance
exists to describe rather than a verdict.
The whole spread rests on one rubric line, though. docs_function is
effectively saturated — 1.00 on nine of every ten samples — and nearly every
docs_gotcha deduction is the same omission: the model documents that order is
preserved, that the first occurrence is kept, and what key does, then never
says elements must be hashable. That is real discrimination, since it is a real
property of the function, but one sentence is deciding a category. Treat this
ordering as thinner than n=14 makes it look.
The self-judging guard costs sample density, and it shows up here. Most
models reached n=14; kimi-k3 and kimi-k3-fast reached only 11, because
those two are the ones diverted to the alternate judge qwen3.6-35b, which
returns unparseable JSON more often than kimi-k3 does. The guard is still
right — a model grading its own family is worse than a thinner sample — but
the alternate judges should be picked for parseability, not just for being
someone else.
The reflex when a category looks flat is to reach for quality_tolerance.
Neither tie broken so far was broken that way: coding_general opened up when
feedback.py folded in real traffic, and docs_writing opened up on six more
benchmark passes. Both were samples, not settings. coding_refactor and
debugging are still flat at 1.00 on 2-3 samples each — which is now a state
this project has mistaken for a measurement once.
Current spread by category, widest first:
| category | spread |
|---|---|
tool_use_agentic |
0.33 - 1.00 |
summarization |
0.60 - 1.00 |
reasoning_math |
0.67 - 1.00 |
docs_writing |
0.66 - 0.97 |
general_chat |
0.80 - 1.00 |
translation |
0.85 - 1.00 |
coding_general |
0.86 - 1.00 |
coding_refactor, debugging |
flat at 1.00 |
What does discriminate is tool use, arithmetic traps, and prose.
deepseek-v4-flash scores 1.00 on all three coding categories yet 0.33 on
tool_use_agentic and 0.67 on reasoning_math. Verified live, not an
artifact: given "It is 1:20pm and my meeting starts at 3pm, how many minutes
away?" — both times supplied — it calls two tools rather than subtracting.
It over-reaches for tools, which is exactly the failure mode that matters in
an agent loop. The router now avoids it for those categories while still
picking it for coding.
Most rows still read source='self_eval_thin' (118 of 132): real
measurement, but below self_eval_min_samples at 2-3 tasks per category per
run. The 14 that have crossed it are all docs_writing, from the six extra
passes above. Two paths thicken it, and they are complementary — re-run
eval_proficiency.py to accumulate benchmark samples, or just use the router
and let feedback.py fold in real outcomes. Both fold into a running mean
rather than replacing, so samples add up across runs.
A score is only as fresh as the row it was copied to
Proficiency is a property of the weights, not the queue, so the eval harness
scores one row per family and propagate_to_variants copies the result onto
the serving variants — kimi-k3-flex gets kimi-k3's number, because no
benchmark rates a -flex row separately.
That copy used to happen exactly once per variant, ever. The guard skipped
any row with self_eval_samples > 0, meaning "measured directly, do not
overwrite" — but inheritance copies the sample count too, so after the first
propagation an inherited row was indistinguishable from a measured one and was
never refreshed again. kimi-k3-flex sat at 0.85/n=2 while kimi-k3 moved to
0.973/n=11.
proficiency.inherited_from records the provenance that was missing, and the
migration was the delicate half, not the fix: ADD COLUMN gives every
existing row NULL, which reads as "measured here", so shipping the guard alone
would have permanently frozen the exact rows it exists to unfreeze. The
backfill infers provenance from the harness's own selection rule rather than
guessing — eval_identities only ever evaluates standard rows plus flex rows
with no standard equivalent, so a flex row that has one was never a
candidate for direct evaluation, whatever its sample count claims. Everything
else keeps NULL, which fails safe: NULL means "do not overwrite", so no real
measurement can be lost to a wrong guess.
Confirmed on the live database, and on the catalog's one genuine exception —
glm-5.2 is canary, so glm-5.2-flex is the routable row the harness scores
directly, and its NULL is correct.
Harness bugs this shook out
Three separate defects, each of which scored the rig rather than the model, and each caught by reading per-task detail rather than the summary:
- Token budget.
max_tokenswas shared between a reasoning model's trace and its answer. At 1200, qwen3.6-35b spent ~4,200 characters thinking and returned an EMPTY content field, scoring 0.00 on tasks it can plainly do. Now 24000, clamped per model (gemma-4-31b caps at 16384), andfinish_reason: lengthskips the sample instead of scoring it. - One leading space. kimi-k2.7-code returns
" def f(...)", which becomes IndentationError once the harness prepends its imports — 0.00 across all nine coding tasks for a model with "code" in its name. - Judge failures scored as model failures. 44% of judge calls returned
unparseable output (the judge is itself a reasoning model and leaks its
thinking despite
response_format). Each was recorded as 0.0. Now the JSON is extracted from surrounding prose and an unusable reply yields no sample.
tests/test_task_set.py exists so this stops happening: it implements a
reference solution for every code task and asserts it passes every check,
recomputes every exact answer (one by brute force), and confirms each
refactor target already passes its own checks while each debugging target
fails. It immediately caught a check where the expected value was simply
wrong — which would have docked every model on a task and been
indistinguishable from genuine difficulty.
Tool competence is read from the request, not guessed at
Neither local classifier can identify agentic work. Asked to label six
unambiguous tool-use prompts ("read the config then update the manifest",
"run the tests and fix what fails"), qwen3.5 got 2/6 and mistral-nemo
1/6 — and mistral-nemo's misses collapse to general_chat, which is also
the configured fallback_category, so qwen3.5's crashes land in the same
place.
That mattered because tool_use_agentic has the widest proficiency spread in
the table (0.33-1.00) and deepseek-v4-flash — the current winner on coding —
sits at the bottom of it.
The fix was not a better classifier. Whether tools are on the table is
stated in the request: every agent client sends a tools array, and
chat_completions never looked at it. Reading it is exact and free.
It is applied as a hard filter, not a category override, and the
distinction is load-bearing. The question is not "is this task agentic" but
"can this model be trusted with tools that exist". The recorded failure is
precisely the second one: deepseek-v4-flash was given a non-agentic prompt
("it is 1:20pm and my meeting is at 3pm, how many minutes away?", both times
supplied) and called two tools rather than subtracting. A model that
over-reaches is a hazard on every request where tools are available, whatever
a classifier would have labelled the task.
So routing.min_tool_proficiency drops any candidate whose measured
tool_use_agentic score is below it, but only when the request carries tools:
| request | winner on coding_general @ 50k |
|---|---|
| no tools | deepseek-v4-flash ($0.0024) |
| tools present | qwen3.6-35b ($0.0041) |
Verified live through /v1/chat/completions with identical bodies differing
only by the tools array. The cost of safety here is 1.7x on that route,
paid only where tools exist.
It is currently set to null, i.e. OFF, deliberately and pending
experiment. opencode sends tools on essentially every request, so with the
filter on, deepseek-v4-flash is excluded from ordinary agent traffic and its
~7x cost advantage goes unused; with it off, that advantage applies and a
model measured at 0.33 on tool use handles requests where tools are on the
table. Which is right is an empirical question and the benchmark cannot
answer it — the 0.33 comes from 3 tasks.
What settles it is POST /outcome: run with the filter off, let real pass/fail
reports accumulate, and compare deepseek-v4-flash's tool_use_agentic
proficiency before and after. That is the one signal here that knows whether
the work actually worked, and feedback.py folds client outcomes in both
directions, so success counts too.
Update 2026-09-15: that accumulation path is closed. The classifier no
longer emits tool_use_agentic — per-turn classification collapsed onto it
(31 of 31 consecutive live turns, all routed to the slowest model in the
catalog) — so no new outcome attributes to the category and its scores are
frozen at today's values. Accepted, not fixed; the reasoning and the rejected
alternatives are in docs/routing.md, "The classifier's
labels are not the proficiency scoring axis". The filter experiment now either
finds a different signal or reads frozen data.
0.5 sits in the empty band between the only two values the catalog holds
(0.33 and 1.00), so it is not fitted to either. A model with no measured
tool score is unproven rather than proven bad and is not dropped — the same
rule as the tier-1 context gate. Config load refuses a
routing.tool_use_category that is not a real category, because a name
matching nothing yields NULL for every row and NULL means "do not
disqualify": the filter would silently stop filtering.
Tier is an iteration budget, not just a floor
A tier used to mean only "do not route below this". It now also buys corrective attempts after a verification failure:
| tier | batch | interactive |
|---|---|---|
| 1 | 0 retries | 0 |
| 2 | 1 | 1 |
| 3 | 2 | 1 |
Interactive is capped below its tier because every retry doubles time-to-answer, and in interactive use latency is a quality loss.
Retries are matched to the failure, since the causes differ:
- truncated — raise the token budget on the same model; a different one would run out too. If there is no cap to raise, the model's own output ceiling is the wall, so escalate to a candidate that can emit more.
- malformed — more tokens will not make unparseable output parse, so escalate to the next-ranked candidate.
- ok / unverifiable — buy nothing. Retrying
unverifiablewould burn quota across the majority of prose traffic for no signal.
escalation.preemptive_on_low_confidence is now off by default. Bumping
the tier because the classifier was unsure pays frontier prices before
anything has gone wrong; spending after a check has actually failed is better
on both mandates — the cheap attempt usually succeeds, and when it fails you
have evidence rather than a hunch.
The only ground truth: POST /outcome
Everything else the router records is a proxy. Structural checks know whether code parses. The local checker guesses whether prose looks right. Neither knows whether the answer did the job — the client does, because it ran the tests.
# id comes from the completion body, or any stream chunk
curl -s localhost:8080/outcome -H 'content-type: application/json' \
-d '{"request_id":"chatcmpl-...","ok":false,"detail":"tests failed"}'
Two things make this the highest-value signal available:
- It is the only quality signal that survives streaming. A retry cannot reach a streamed response because the bytes are already gone; a report arrives afterwards and works either way. Every agent client streams.
- Its successes count.
feedback.pyfoldsok: trueandok: falseintoproficiency.outcome_scoreviaadd_outcome(). Structural andlocal_llmverdicts are now diagnostics only; they no longer move the score, because a structural failure is not a verified client outcome.
Client outcomes are what calibrate the routing score. The benchmark is a
prior; POST /outcome is the posterior.
request_id is written back to route_decisions.request_id on both the
streamed and buffered paths, so the report joins cleanly to the decision that
produced it. An unknown request_id returns 404 rather than being quietly
accepted — a client whose reports go nowhere should find out.
Verification: what local compute is actually good for
Local inference is a poor substitute for cloud completions here — 7-82x the energy, ~670x the carbon on Michigan's grid, and slower (6.2s vs 1.4-2.0s). But it is very good at stopping a cloud completion from being wasted, and completions are where all the money is: fitted on real traffic, a completion token costs 201x a prompt token.
| check | cost | what it catches |
|---|---|---|
structural (verification.py) |
free | truncation, malformed code/JSON/YAML, empty answers |
| local LLM (Ollama) | ~$6.9e-05, ~6s | refusals, wrong-question answers, incoherence |
Structural checks run on every response and never execute the code — they
parse it. The local LLM check runs only on answers above
verification.min_completion_tokens, in the background after the client has
its response, because it costs ~15% of a median 193-token answer and only
pays above ~600 tokens.
An agent turn is not a prose answer
Both checkers had mirror halves of the same blind spot, and real traffic is what found it. On the first genuine agent session (63 completions, a shipped feature, 349 passing tests, clean mypy and ruff):
| checker | called it | actually |
|---|---|---|
| structural | malformed: empty response x29 |
turns that ended in a tool call |
| local LLM | cuts off mid-sentence x8 of 9 |
same turns, judged from the other side |
A turn that calls a tool has empty or half-finished text by design. Both
paths now take has_tool_calls and return unverifiable; in verify_response
that check outranks even finish_reason == 'length', because stopping
mid-sentence at a call boundary is not a budget overrun. worth_local_check
declines outright, which also stops paying ~6s of local inference to
mis-grade a tool call.
Had feedback.py run before this, it would have applied ~12 false failures to
the models that had just shipped the feature. That is the fourth harness
bug in this project that would have scored the rig rather than the model, and
the first caught by real traffic instead of a synthetic test. The pre-fix rows
are kept but set model_attributable = 0, so the record survives without
steering routing.
client_capped was also over-applied. It marked every verdict
non-attributable whenever the client set max_tokens — and opencode always
does — so real failures were invisible to feedback. A client's token cap
explains a truncated verdict and nothing else; it is now scoped to exactly
that.
feedback.py folds observed failures into proficiency, so routing learns
from your traffic rather than only the 43-task benchmark. Only failures
are folded in: a structural 'ok' means the code parsed, not that it was
correct, and recording those as 1.0 would flatten every score toward the
ceiling. Failures the model did not cause — a client's own tight max_tokens
truncating the answer — are recorded but excluded.
The classifier is the latency floor
Every routed request pays a full local classification round-trip before an
upstream token is requested — the classifier is the latency floor. Four settings keep it usable:
max_retries=0 (SDK retry → 3x wall-clock), max_output_tokens: 1024 (bounds reasoning), max_input_chars: 8000 (doesn't need the document; head+tail clamp), fallback_tier/fallback_category (mid tier, not 502/503), plus temperature: 0 (deterministic tier). "Local" means your hardware, not this machine: bind to the VPN address, not 0.0.0.0 (Ollama has no auth).
Full measurements: docs/local-models.md.
A classifier failure no longer collapses to one fixed guess
fallback_tier/fallback_category above is now the LAST step, not the only
one. When the local classifier fails, classify() walks a cascade:
| # | step | cost |
|---|---|---|
| 1 | this session's cached classification, staleness ignored | free |
| 2 | this session's history in route_decisions |
free |
| 3 | classifier.cloud_fallback, if configured |
money |
| 4 | fallback_tier / fallback_category |
free |
Steps 1 and 2 reuse a real classification rather than guessing. Step 2 is what survives a restart, when the in-memory cache is gone but the session's decisions are still on disk. Both filter degraded sources, so one outage's guess cannot propagate through every later turn of a session and end up looking like a measurement.
Step 3 is absent by default and the whole feature costs nothing until it is configured. Three things bound it once it is:
classifier.cooldown_seconds(30) is a global backoff, so a sustained outage buys one cloud attempt per window across ALL requests, not one per request;- an account-level provider refusal suppresses step 3 for the same window — an out-of-credit account makes that call a guaranteed wasted request;
- the
session_cache.put()gate admitsclassifier_cloud, bounding a session to one cloud call per staleness window instead of one per turn.session_staleandsession_historyare deliberately NOT cached: borrowed answers must not renew a staleness clock they never earned.
The trap this had to avoid, and it is the reason to read this section. A
degraded classification is recorded with task_category = general_chat —
which is a fully scored category (13 proficiency rows) as well as the
configured fallback_category. Without a guard, POST /outcome reports on
mislabelled outage traffic fold into proficiency(model, general_chat) and
drag real scores toward whatever happened to be flowing while the classifier
was down. So report_outcome marks outcomes of degraded-source decisions
model_attributable = 0 — kept as a record, excluded from folding, the same
treatment the tool-call false-failures got. feedback.py is untouched; its
existing AND model_attributable = 1 already does the work.
Attributable: classifier, cached, override, classifier_cloud. Not:
fallback, session_stale, session_history. The asymmetry is deliberate —
a degraded source is good enough to route one visibly-flagged request, but a
proficiency score is consulted by every future request, so attribution must
not trust more than routing does. Unknown decision rows fail open; an
over-applied exclusion already starved feedback once (client_capped).
/metrics warns when the degraded share of the last 24h crosses
classifier.degraded_warn_threshold over at least degraded_warn_min
decisions. A survivable failure is exactly the kind that goes unnoticed for
weeks.
Which implementation is PRIMARY is now a config choice
classifier.mode in the live deployment is currently local_encoder, set in
config.local.yaml with device: cuda on the classifier host. That is the
mode answering real traffic right now. The explicit caveat is that a
CPU-vs-CUDA latency and confidence comparison on the live classifier host is
still pending; until that measurement exists, flipping the project default to
local_encoder is not decided.
Turning that mode on for real found three real bugs in one afternoon
(2026-09-06), each a fresh instance of this project's own recurring lesson —
verify against the live system, not the plan. First, the config-load
validator for confidence_threshold didn't exist yet: the admin UI saved a
raw 80 (meant as 80%) straight into the overlay with no conversion, which
would have made every real confidence score read as below-threshold on the
next restart (classify_zero_shot returns [0.0, 1.0]; no probability
exceeds 1.0). Caught before the restart, not after — see confidence_threshold
in local-models for the full incident and the fix.
Second, once that was corrected and the service actually restarted, it
crash-looped twice more before coming up clean: the shipped default model
(MoritzLaurer/deberta-v3-base-zeroshot-v2) had become gated on HuggingFace
sometime after this project picked it (401 on an unauthenticated GET of its
own model page), and separately HF_HOME's default cache path falls outside
this service's ProtectHome=read-only sandbox exception — both are now fixed
(switched to facebook/bart-large-mnli, HF_HOME redirected into the repo).
Third, and most consequential: the very first real classifications measured
only 5 of 9 test categories correct, because classify_zero_shot was passing
raw config identifiers like tool_use_agentic and diff_checking directly as
zero-shot candidate labels — HF's pipeline scores a label against a hypothesis
template ("This example is {}."), and an underscored code token is not a
sentence the model's NLI training ever saw. Mapping each category to a natural-
language description before scoring, plus multi_label=True (the pipeline's
single-label default forces every candidate to compete for the same
probability mass), brought that to 8 of 9 correct with confidence scores
0.77-0.999 on the hits — the one remaining miss scored below the configured
threshold and correctly fell through to the safe fallback rather than
mis-routing.
classifier.mode (local_llm default, cloud_llm, local_encoder) picks
what answers a classification request — a peer concept to the cascade above,
not a replacement for it. Whichever mode is primary, a failure still
walks the exact same cascade (stale session → session history →
cloud_fallback → the static guess), unmodified.
cloud_llmmakes a cloud model the primary attempt, not just the cascade's post-failure backup. Either pin one (classifier.cloud_primary, same shape ascloud_fallback) or setcloud_primary_auto: trueto resolve the cheapest currently-routable model live against the catalog (routing.cheapest_classifier_candidate, priced for the classifier's own short-prompt/short-completion shape — not the task's). A success here recordssource="classifier", deliberately the same stringlocal_llm's success uses, not"classifier_cloud"— that string means specifically "the cascade's backup step fired" and feeds the/metricsdegradation-share warning above as a degraded signal. An intentionally configured primary succeeding is not degraded.local_encoderclassifies with a small, non-generative zero-shot model instead of an LLM — structurally immune to the runaway-reasoning failure mode documented above, since there is no generation to run away. Zero-shot rather than fine-tuned: this router never stores raw task text anywhere, so there is no training corpus without a new, separate opt-in capture feature (not built). Only producestask_category;task_tierfalls back tofallback_tier— a real limitation, not a bug. A below-threshold confidence is treated as a failure and cascades exactly like a local-LLM parse failure would.- Neither is gated by
local_compute.enabled(gaming mode, below) the waylocal_llmis:cloud_llmnever touches local hardware, andlocal_encoderis small enough to run on CPU, so neither competes for the GPU gaming mode exists to free up.
Configured via config.yaml (global default) and overridable per-machine
in config.local.yaml — the existing overlay, not a new mechanism — or
through the admin portal's Classifier card, which reports the live
resolved primary for cloud_primary_auto rather than echoing the config
value (see admin-portal).
The default in config.yaml is still local_llm. local_encoder is
intentionally an opt-in per-deployment choice rather than the repository
default until the pending CPU-vs-CUDA comparison on the live classifier host is
available.
Local dispatch model
A second local model can now be dispatched directly for specific categories.
qwen2.5-coder-router:14b is configured as a tier-1 local row with
provider='ollama-local', gated by models.eligible_categories
(file_summarization and diff_checking). The poller refreshes the row each
run; seed_local_dispatch_energy.py derives its price from measured GPU draw
and the user's tariff. Routing treats a local row like any other candidate
once the category filter admits it, and the circuit breaker excludes it on a
local failure so the next request reroutes to cloud candidates.
Known limitations of the local dispatch branch right now:
- No true streaming. The response is shaped into an SSE stream, but the local answer is generated before any bytes leave the router.
- No verification rows. Structural and local-LLM checks run but are not
written to
verificationsfor local answers. - No within-request cloud failover into the upload path is gone. An eligible
routed request degrades to the local dispatch model when the cloud account
refuses or is exhausted (a fallback, not a preference), so the local row is
no longer a dead-end before a client retry. The degraded answer is recorded
as
kind='local_dispatch_fallback'. - Follow-ups are not special-cased. A pinned or auto-routed follow-up to the same local model works, but nothing caches the loaded model between turns.
Dormant under the default profile by design — see docs/routing.md § Local dispatch branch.
POST /outcome now attributes through the local energy ledger too. Local rows
include request_id and session_dir in local_energy_observations, so a
client report on a local answer resolves to the same (model_id, provider, task_category) provider-agnostic record as a cloud one.
Routing notes
Ranking is quality-first, cost as a tiebreak; cost is never allowed to override
a real quality gap. The optional objective.credit_attenuation block extends
that tiebreak without changing it: when the block is enabled, a per-provider
multiplier is applied to a candidate's comparison cost only, producing an
effective_cost that breaks ties. The multiplier is derived from the provider's
polled account balance (the balance_url path, such as OpenRouter), so a low
prepaid balance can nudge a near-tie toward a healthier provider. The logged
est_cost_usd and the decision history stay as raw catalog estimates. The
multiplier is 1.0 for providers whose balance comes from per-completion
allowance_remaining_usd telemetry (NeuralWatt), so normal overage readings do
not bias routing.
Two semantics matter when reading the numbers. total_balance_usd is a sum of
heterogeneous provider-reported readings: OpenRouter's prepaid credits plus
NeuralWatt's overage allowance, which normally reads near -$0.004. It can be
negative and it is not a single spendable figure. credit_attenuation.enabled
deliberately lives only in the config file; it is absent from the admin
persisted-config allowlist and from provider edits. Turning it on or off
requires editing config/config.yaml and systemctl --user restart llm-router.service, because the dispatcher's cfg binds at import time.
What's NOT built yet — pick up here
Built: session-directory attribution, the local energy ledger, local model
dispatch, admin profiles/proficiency/gaming-mode, and the configurable
classifier backend (classifier.mode, including its admin card — all listed
under "What's built and working" above).
As of 2026-09-05, PRs #26-#36 landed the gitignored config overlay, admin
profile CRUD writing to it, the provider literal cleanup, quota
balance/burn/runway, capability-aware ceiling and rejection warnings, the TUI
schema catch-up, the classifier fallback cascade, and the admin portal uplift
(proficiency page, profiles duplicate/coverage fix, gaming mode, the
classifier-backoff bug fix). classifier.mode (this document's own section
above) is a further, independent addition on top of that.
plans/multi-provider-support.md is PARKED on provider selection — Z.ai was
the recommendation and is no longer settled; the coupling surface in it is
measured and still valid.
Known follow-ups recorded but not specced, both small:
- The TUI decision table renders the literal
"None"in thectxcell whenrequired_context_tokensis absent — the same defect theprofilecell was written to avoid. See.omo/notepads/tui-overhaul/issues.md. - A NULL
required_context_tokensraisesTypeErrorinside the demand-ceiling comparison, and the column is nullable. Latent only: zero such rows exist today, checked on the live DB.
The items below remain open.
-
Fine-tuning
local_encoderon real traffic. Scoped, not built:classifier.training_capture.enabled(opt-in, off by default — a deliberate reversal of "never store task text", so it must be impossible to enable by accident), aclassifier_training_samplestable gated the same wayreport_outcomealready filters proficiency (only rows whoseclassification_sourceis in the attributable set), and atrain_local_encoder.pyscript matchingeval_proficiency.py's conventions.local_encoder.pycurrently ships zero-shot only. -
Leaderboard priors are unfilled.
leaderboards.yamlships empty on purpose — inventing plausible-looking benchmark numbers would put fabricated data straight into routing, the same failure as the provider'sstatic_fallbackcarbon constant this project already excludes. Until real sourced figures go in, a newly listed NeuralWatt family has no prior and relies entirely on self-eval accumulating.python leaderboard.py --checklists what is missing. -
Sampling depth for three models — now
eco-only. 7 samples/model gives split-half agreement within 1.4x for 10 of 13, butkimi-k2.7-code-fast(29x),kimi-k3(14x) andglm-5.2-flex(2.2x) are still unsettled. This no longer touches cost, which is priced per-request from the catalog, so it only affectseco— which is not an objective. Low priority unless eco comes back. -
Retry does not reach streaming. The iteration budget (
iteration.py) retries after a failed check, but only on the non-streaming path — once bytes have gone to the client there is nothing to take back. Buffering to fix that would cost streaming itself, a worse trade for interactive work.POST /outcomeis the answer for streamed traffic: it arrives afterwards, so it works identically either way. -
local_encodernoise isolation — built for the confirmed shapes; residuals below. The raw task's fenced code blocks,Tool result:-shaped lines, and closed<system-reminder>spans are now stripped by_isolate_task_text(pure, stdlib-only, deterministic) before the fit + embed pass —classify_zero_shotruns isolate → fit → prefix, so cleaning happens first and a noisy task often fits the token window outright. Measured on the realBAAI/bge-large-en-v1.5(2026-09-19): 8 clean one-sentence tasks scored 8/8, but the same instructions wrapped in that noise scored 2/8 with the truncation fix already in place — and tail-biased windowing ALONE also scored 2/8 on long noisy pairs, so the fit does not subsume isolation. Post-isolation: 8/8 on short and long noisy pairs, clean-vs-noisy pair consistency 8/8 + 8/8;'[code]'/'[elided]'placeholder tokens measured worse than pure removal (7/8, 6/8 on long pairs) and were rejected; a 20%-ratio floor guard measured harmful (3/8 — it reverts exactly the short noisy inputs isolation exists to fix) in favor of an absolute 24-char floor that only catches near-all-code inputs. Offline regression: a noise tripwire (same instruction bare vs wrapped must classify identically) fails against pre-isolation code and passes after. Still not built: attention-masking de-weighting as an alternative to stripping (it would preserve the noise tokens' presence without letting them dominate); unclosed<system-reminder>tags and un-fenced diff hunks are left in place (only closed-tag spans, fenced blocks, and marker-prefixed lines are stripped); and a code-grounded instruction whose pasted snippet is the subject can still land on a near-category — measured 3/4 on a 4-task grounded set, the residual miss being description similarity ("refactor this helper" + code →debugging), not noise dominance.
Gaming mode, and the backoff that used to do nothing
The classifier circuit breaker did not break the circuit.
_last_classifier_failure was written by _record_failure() and read by
nothing — _classify_cascade gated only its cloud step, and on a different
timestamp. So the router re-dialled a known-dead local classifier on every
request. A stopped Ollama refuses immediately and costs little; a hung one,
or a VPN-bound one that black-holes, costs the full 120s timeout_seconds per
request for as long as the outage lasts. _classifier_backoff_active() is now
the read, consulted before the client is constructed.
One detail there is load-bearing: recording the failure lives in classify()'s
exception handlers, NOT in _classify_cascade. The cascade is walked for
reasons other than a fresh failure, and if those re-stamped the clock, every
request during an outage would push the deadline forward and the local
classifier would never be re-probed while traffic flowed — a permanent outage
wearing a circuit breaker's clothes.
local_compute.enabled (default true) is the outer gate over local
hardware. Turn it off when you stop Ollama for a game and the router skips
every local call rather than discovering the outage one timeout at a time:
classifier, /health probe, local verification, local-vision fallback, and
local dispatch rows (dropped in load_candidates, so a local row is never
picked and then 503'd). /v1/models stops listing local rows, and an explicit
pin gets a 503 naming the flag instead of NeuralWatt's unknown-model 400.
ONE flag the code reads, not a macro writing five keys — a macro is hard to
undo cleanly, drifts the moment a sixth call site appears, and leaves nobody
able to answer "why isn't the classifier running?" from one place.
verification.local_llm_enabled, local_vision.enabled and
local_energy.enabled keep their own meanings; this ANDs over them.
It REFUSES to engage without classifier.cloud_fallback — 409 on the
runtime knob, a validation error at config load. Skipping the local classifier
does not make classification remote; without a cloud classifier it stops
classifying, and every request falls through to a static guess recorded as
general_chat, a fully scored category indistinguishable from a real
classification afterwards. A refusal, not a warning, because a warning is what
nobody reads while their game is loading. Nothing auto-writes the block.
Cascade steps 1 and 2 still run ahead of the cloud call: a stale session classification is free and was a real classification of that same session, so paying to re-derive an answer already held is spending money for nothing.
A latent substring bug fell out of the profiles work. SQLite stores
eligible_categories as a comma-joined string, and
task_category not in "<a>,<b>" is a SUBSTRING test — so a row eligible only
for file_summarization also admitted summarization. Latent on main (the
category-less probe short-circuits before the compare) and live the moment
anything probes per category. routing.parse_eligible_categories is now the
single parser dispatcher.load_candidates and admin.py's probe both use.
Known open questions
- Answered: cost and eco stay separate axes — grid intensity spans 13.6x across the catalog, so they rank models differently.
- Answered: the GLM rows reporting
grid_id: FIat 475 gCO2/kWh werecarbon_source: static_fallback— a substituted constant, not a measurement. They are now excluded from eco rather than trusted. Still worth asking NeuralWatt why the fallback keeps the originalgrid_id, since that is what made it look like a real regional difference. - Three models still fail a split-half stability check at 7 samples. Is the instability real (variable serving conditions) or an artifact of when the sweep ran? Re-sweeping at a different hour would tell.
- Answered, and the question no longer parses: tier-1 composites used to sit
within 0.009 of each other because min-max normalization compressed them.
There is no composite any more — ranking is quality first, cost as the
tiebreak inside
quality_tolerance— so nothing normalizes and nothing compresses. - Answered: the eval set exists (
evals/tasks.yaml, 43 tasks, four scoring kinds) andtests/test_task_set.pykeeps it honest. The benchmark-sourced rows now split the coding categories, but two tasks are still flat at 1.00 (debug_affine_coprimeand most of the BFCL set) and either need hardening again or should be conceded as non-discriminating. Try samples before hardening.docs_writinglooked flat at the top too, and six more passes spread it 0.66-0.97 without touching a task; two samples per model is not enough to tell a saturated task from an unsampled one. - How much context-assembly (RAG-style retrieval) belongs in the classifier step vs. a separate pre-step? Leaning decoupled, undecided.
- Should
eco_scoreuse real-time grid carbon intensity per request or a stable per-model average? Currently the latter, from the reference sweep.grid_carbon_intensityandgrid_idare logged per observation, so this stays answerable from data without a re-run.
Config is strict: an unknown key is an error
Pydantic ignores extra keys by default, which means a typo or a misplaced
setting loads cleanly, does nothing, and still looks configured. Every config
model now inherits StrictModel (extra="forbid"), so both of these fail at
load rather than silently:
verification.max_input_chars # right key, wrong section
routing.min_tool_proficency # sic
This is not hypothetical. max_input_chars shipped into the verification:
block instead of classifier: and was accepted and discarded — it happened to
match the code default, so behaviour was correct and the file was a lie.
Editing it would have done nothing.
The corollary worth keeping: every knob belongs in config.yaml, not only
in a Pydantic default. A default the file never mentions is invisible to
anyone tuning it. classifier.outcome_attribution_window_seconds was removed
in the same pass — it was declared, never read, and shadowed the
verification one that actually is.
Setup
Full install steps (venv, deps, config, first run) in README ## Installation. Host-local deployment values go in config/config.local.yaml (gitignored overlay) — classifier.model/base_url, local_energy.*, host-specific URLs. General defaults in config/config.yaml stay shareable. objective.plan_kwh_per_period in the README config table. Model tags (num_ctx) + verification.model same-tag note in docs/local-models.md. Requirements are pinned — bump deliberately (README).
Run as a service
deploy/ holds the dispatcher's systemd user unit plus a timer and a
oneshot service each for the poller, the seed sweep, the backup, the offsite
sync and the feedback fold, and one drop-in for a system Ollama — see
deploy/README.md for install and operation. Every timer there is enabled on
install except llm-router-feedback.timer, which is not, on purpose. In short:
echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env
cp deploy/llm-router*.{service,timer} ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now llm-router.service llm-router-poller.timer
The dispatcher binds 127.0.0.1:8080. The poller timer is load-bearing, not
housekeeping — but not for the reason this section used to give, and the
correction matters because it inverts which failure to watch for.
mark_stale runs only inside poller.main(), and main() returns early on
a RequestException — before upsert and before mark_stale. So a
stopped timer or a provider outage marks nothing: the catalog freezes at
last-known-good and the router keeps routing on prices that may be weeks old.
The failure is silent and open, not loud and closed. Nothing surfaces it,
because a frozen row still reads availability = 'active'.
The path that can empty the candidate set is narrower and is not governed by
the timer at all. fetch_neuralwatt reads payload.get("data", []) with no
floor on row count, so a 200 response carrying an empty or truncated data
array — a partial provider outage, a schema change, an auth path degrading to
an empty list — clears raise_for_status(), upserts nothing, and then lets
mark_stale run anyway. Three days of that and every row is stale and
exclude_stale: true leaves zero candidates for everything.
stale_after_days: 3 only sets the length of that fuse; it does not arm or
disarm it. Against a 2-hourly poll it is 36 successful polls of margin, and
recovery is automatic — upsert writes availability = excluded.availability,
so one good poll flips every stale row back to active. The fix is a sanity
floor on the fetch, not a larger number.
A second path empties the candidate set, and it bit on 2026-09-01. Admin
availability overrides are not governed by the poller at all. Deprecating the
seven expensive models through /admin collapsed tier 3's context ceiling from
782,324 to 94,196 — while tiers 1 and 2 stayed at 782,324 — so every tier-3
request above 94k returned 422 with nothing warning anywhere. It surfaced ~19
hours later as an agent failing mid-task on an opaque error.
The obvious check for this is wrong, and the reason is worth remembering.
Tempting: warn when a higher tier's context ceiling sits below a lower tier's.
But ceiling(T) is the max effective_context_window over models with
tier >= T, and tier is a capability floor, so the eligible set shrinks
monotonically as T rises — ceiling(1) >= ceiling(2) >= ceiling(3) is a
theorem, true of every catalog. Such a warning fires always and means nothing.
What actually failed is that a tier's ceiling dropped below what that tier is
asked to serve, which is only knowable from traffic: compare ceiling(T)
against the observed required_context_tokens for decisions classified at tier
T. That is silent on all three tiers today and fires on the outage state
(94,196 vs an observed max of 268,168). See
plans/catalog-staleness-and-poller-failure-modes.md §4.4.
This recurred on 2026-09-04 through a dimension the detector did not model,
and both halves of the fix are now in metrics.py. Admin deprecations took
out kimi-k3* — the only vision-capable rows with enough context — so a
242,486-token image request 422'd while every existing check stayed silent,
because the all-models tier-1 ceiling was still 782,324. The vision-capable
ceiling had collapsed to 192,500.
- Predictive:
capability_ceilings/capability_demand_warningscomputevisionandjson_modesub-ceilings and compare each against demand actually observed for requests carrying images / requesting JSON. Two extra series, not a bucket per capability combination. - Reactive:
rejection_warningswatchesroute_decisionsfor rows withselected_model IS NULL. This is the more valuable half and the simpler one — it catches the next dimension nobody predicted, at the cost of firing after the first failure rather than before.
Two details in the reactive detector are load-bearing and easy to undo by
accident. It groups by (task_tier, digit-normalized reason) using the
structured column, because normalizing digits alone merges tier >= 1
and tier >= 3 rejections into one group and hides whether the broadest or
the frontier candidate set went empty. And the signal is novelty OR rate,
never mere presence: measured on the live DB, routine rejections run ~3/hr
while the 2026-09-04 incident was only n=2 — below the noise floor — so no
single count threshold can both catch it and stay quiet. A group absent from
the 24h baseline warns at n≥2; a familiar group warns at the configured
count. Zero rejections warn about nothing: a genuinely impossible request
SHOULD 422.
The service holds a billable API key and has no auth of its own. Loopback
bind is the only thing standing between the open internet and your allowance;
add auth before widening --host.
The same applies to an Ollama shared over a VPN — it has no auth either, so
deploy/ollama-over-vpn.conf binds it to the VPN address rather than
0.0.0.0, which would publish it on whatever network the client happens to
be on.
When the router goes unreachable, start at docs/incidents.md
Four incidents so far, all sharing one shape: a change that looked local to the
router silently degraded the agent depending on it, and none announced itself as
a router problem. docs/incidents.md carries the full write-ups plus a
symptom -> one-line-check table; read it rather than re-deriving a diagnosis.
Two conventions from those incidents that bind every session, and so stay here:
- 8080 is production, always. It is baked into
opencode.json, the systemd unit, every curl example here, and the admin frontend's own fetches. A throwaway instance (manual iteration, Playwright smoke tests, anything that is not "use the real router") binds 8081. Never send a kill signal to a process matched by name or port rather than by a PID you started yourself --Restart=alwayswill fight you, and on this repo it may be your own model access. - Never point
config/config.yamlat test fixtures. It is the file the live service reads. Pass a different config file, monkeypatchcfg.database.pathin-process, or use a temp copy.
Recovery for an unreachable-but-active service is
systemctl --user restart llm-router.service -- a hung process was never in a
tracked stop job, so this issues a fresh cycle systemd does enforce a timeout on.
Pointing a coding agent at it
The /v1 endpoints are OpenAI-compatible, so any normal client works —
opencode, an SDK, plain curl. Repo-local opencode.json is already wired up,
so running opencode from a clone of this repo routes by default. For global
use, merge provider.llm-router into ~/.config/opencode/opencode.json.
| model name | behavior |
|---|---|
auto |
router picks; flex rows excluded so nothing is held during peak |
auto:batch |
router picks; flex rows admitted, for overnight/async work |
| any real model id | dispatched as asked, still logged |
Streaming is proxied chunk by chunk rather than buffered, so tokens still
render as they arrive. NeuralWatt emits its energy and cost blocks as SSE
comment lines (: energy {...}) before data: [DONE] — ordinary clients
ignore comments, so the stream passes through untouched while the router
reads the telemetry on the way past. Without that, streamed calls would log
no energy at all, which is most of the point of this project.
Try it
Copy-pasteable curl examples (/route, /dispatch, /v1 models + chat) and
the classifier-skip overrides (task_category, task_tier,
required_context_tokens) live in README ## Usage. Inspect
what a dispatch cost/burned with the sqlite3 query in
docs/operations.md.