# A 10,000ft review: what this router needs to be a more useful tool Status: reference -- a diagnosis, not a queue item; the executable work it feeds is scoped in `plans/token-waste-waves.md`. **Measured on the LIVE database** at `/home/alee/Sources/6krrt/router.db` (read-only), 2026-09-13, at repo state `26398d7`. An earlier draft of this file measured a stale 11:47 copy of `router.db` sitting inside the `profiles-header-fix` worktree and drew a wrong headline from it; that copy is ~10 hours behind and missing the traffic that matters most. If you are measuring from a worktree, check the mtime first. Audience: the operator, and whoever scopes the next round of work. Deliberately does not repeat the item-level work already scoped in `code-analysis-refactoring-opportunities.md`. --- ## The headline: the router works, and it just proved it On 2026-09-06 NeuralWatt billed **$19.85 in one day** across 4,129 decisions. What followed was the most consequential routing change in the project's history: OpenRouter came back as an allowlist-gated provider and traffic moved. | day | decisions | billed | |---|---|---| | 2026-09-06 | 4,129 | **$19.85** | | 2026-09-09 | 594 | $0.32 | | 2026-09-10 | 722 | $0.25 | | 2026-09-11 | 755 | $0.25 | | 2026-09-12 | 116 | $0.40 | | 2026-09-13 | 154 | $0.21 | Volume held at 600-750 decisions a day. Billed spend fell roughly **50x**. That is the system doing precisely what it exists to do, and it is the strongest evidence in the database that the thing is useful. Two qualifications, and they are where the rest of this document lives: - **The human did the adaptation, not the router.** The migration was an operator decision — add a provider, write an allowlist, let ranking follow. The router has no automatic response to a cost shock; `plan_kwh_per_period` gates nothing and `max_energy_per_request` is null, so every cost mechanism in the system is advisory. It reported the spike well. It did nothing about it. - **The migration moved 73% of traffic onto a provider the router cannot measure.** That is the single most actionable finding here, and it is new since the migration. ## The thesis This project is unusually good at recording *why* a decision was made, and has no mechanism for noticing when a recorded reason has expired. | setting | the comment's own premise | what is true now | |---|---|---| | `quality_tolerance: 0.1` | "scores currently rest on 2-3 samples per category, so a 0.05 gap is indistinguishable from sampling variation... **Narrow it as samples accumulate**" | 46 rows per category, all outcome-backed; `outcome_blended` rows average 51.8 client samples, max 460 | | cost-as-tiebreak, not an objective | "60% of every decision adjudicated fractions of a cent (**all real traffic to date totals $0.07**)" | **$118.79** billed. 1,697x the figure the whole argument rests on | | `assumed_cache_rate: 0.917` | measured once, 2026-08-23, over 40.7M tokens | `cached_prompt_tokens` is NULL in **all 35,086 rows**. The router receives this field and drops it | | cost feedback generally | catalog estimates "track the real ordering" because billing is capped at 3x list | true for NeuralWatt, untestable for OpenRouter: **95.7% of its rows have NULL `cost_usd`** | Four settings, four expiry conditions met, nothing re-fired. That pattern is worth fixing before any individual number. The documentation is a real strength — the reasoning is preserved, measurements are dated, caveats are honest. What is missing is anything that *acts* when a premise the code depends on stops holding. --- ## 1. The provider that now carries the traffic is the one the router cannot see Recent picks, since 2026-09-09 (2,341 decisions): | provider | model | n | |---|---|---| | openrouter | `xiaomi/mimo-v2.5` | 744 | | openrouter | `deepseek/deepseek-v4-flash` | 659 | | neuralwatt | `deepseek-v4-flash` | 266 | | neuralwatt | `qwen3.6-35b` | 247 | | openrouter | `nvidia/nemotron-3-nano-omni-...:free` | 178 | | neuralwatt | `qwen3.6-35b-fast` | 119 | | openrouter | `z-ai/glm-5.3-flash` | 95 | | openrouter | `xiaomi/mimo-v2.5-pro` | 26 | **OpenRouter is 72.8% of current decisions.** Telemetry coverage in `energy_observations`: | provider | rows | NULL `cost_usd` | NULL `duration_seconds` | |---|---|---|---| | neuralwatt | 31,188 | 66 | 66 | | **openrouter** | **3,838** | **3,673 (95.7%)** | **3,838 (100%)** | | ollama-local | 60 | 60 | 60 | So on the majority of current traffic the router has no billed cost to calibrate its estimator against and no completion time at all. Every feedback loop that makes this system smarter than a static config is blind on the provider it now prefers. This inverts the priority of everything below: recalibrating the cost model (item 3) cannot work for 73% of traffic until this is fixed. OpenRouter returns usage and generation-cost data; it is an ingestion gap, not a provider limitation. ## 2. `feedback.py` is the only loop without a timer `proficiency` was written in a single batch at **2026-09-10T00:44:55** — every row, one instant, three days ago. Since then: | | | |---|---| | client outcomes applied | 2,224 | | client outcomes **unapplied** | **95** | And the unapplied 95 are exactly the current traffic: | model | provider | unapplied | of which failed | |---|---|---|---| | `xiaomi/mimo-v2.5` | openrouter | 41 | 9 | | `deepseek/deepseek-v4-flash` | openrouter | 36 | 15 | | `deepseek-v4-flash` | neuralwatt | 17 | 3 | | `xiaomi/mimo-v2.5-pro` | openrouter | 1 | 0 | Meanwhile **every OpenRouter model's proficiency row is `outcome_prior` with `outcome_samples = 0`** — an inherited peer-rate prior, not a measurement. `xiaomi/mimo-v2.5` is the most-routed model in the catalog right now (744 picks) and scores 0.749 / 0.770 / 0.894 on coding_general / coding_refactor / tool_use_agentic entirely on inheritance. Its 41 real outcomes (78% pass) are sitting in `verifications` unapplied. `deploy/` ships timers for the poller, the seed sweep, backup and offsite sync, and `~/.config/systemd/user` additionally has a baseline-report timer. There is **no feedback timer**. The loop the docs call "the only ground truth" is the only one that has to be cranked by hand. This is the cheapest high-value fix in the document: a oneshot service and a timer, matching the four that already exist. ## 3. The cost estimator is wrong in a model-dependent way `routing.estimated_cost` breaks every tie, and the ranking reduces to it whenever candidates land in the same quality band. Per request, on rows with a real billed figure: | model | provider | est µ$ | billed µ$ | est/billed | |---|---|---|---|---| | gemma-4-31b | neuralwatt | 3,822 | 6,041 | **0.63x** | | kimi-k2.7-code-fast | neuralwatt | 28,623 | 21,747 | 1.32x | | kimi-k3 | neuralwatt | 87,910 | 42,388 | 2.07x | | deepseek-v4-flash | neuralwatt | 4,154 | 1,796 | 2.31x | | glm-5.3 | neuralwatt | 48,232 | 20,056 | 2.40x | | `deepseek/deepseek-v4-flash` | openrouter | 3,830 | 1,383 | 2.77x | | kimi-k2.7-code | neuralwatt | 22,253 | 3,906 | 5.70x | | qwen3.6-35b | neuralwatt | 4,091 | 484 | 8.46x | | `qwen3.6-35b-fast` | neuralwatt | 3,773 | 282 | 13.37x | | `xiaomi/mimo-v2.5` | openrouter | 9,689 | 715 | **13.55x** | Total: $401.99 estimated against $118.79 billed. The scale error does not matter; the **21x spread in the error** does, because it reorders candidates. `gemma-4-31b` is the estimator's cheapest model in the catalog and billing's most expensive per request of the ten. The router's current favourite, `xiaomi/mimo-v2.5`, is the one it overestimates most — it is winning on quality *despite* the cost model, not because of it. The cause is already written down in `CLAUDE.md`, two sections apart, never reconciled: attribution ratio spans **750x between models** and is "most of the real cost difference in the catalog", while `estimated_cost` prices from catalog token prices, which contain **no attribution term at all**. Cost moved off measurement to fix a real problem — a 400-token sweep cannot price a 150k-token workload — and in the move discarded the term the project had already proven was dominant. Third recurrence of one pattern: list price ranks models backwards; the reference sweep ranks them backwards for real traffic; catalog pricing misorders them because it omits pool concurrency. **The fix needs no sweep.** 31,119 NeuralWatt rows carry a real billed figure, joinable to decisions by `request_id`. A per-model correction factor over a trailing window, refreshed by the poller, recalibrates against the bill and keeps the per-request shape sensitivity that motivated the move. For OpenRouter it needs item 1 first. ## 4. Switching models mid-session costs 2.5x, and the estimator prices it as free No sticky routing, no concept of the previously-used model. On turns with prompts over 20k tokens and a real billed figure: | turn | n | avg prompt | µ$ / prompt token | spend | |---|---|---|---|---| | same model as previous | 11,246 | 100,994 | 0.037 | $47.91 | | **switched model** | 858 | 90,309 | **0.091** | **$9.68** | 2.46x, on *smaller* prompts, so the effect is understated. Switch turns are 7.1% of turns and **16.8% of joined spend**. The mechanism is not mysterious: a switch lands a ~100k-token prompt on a provider that has never seen it, so the prefix cache is cold and ~92% of the prompt is suddenly billed fresh. But `rank_candidates` prices every candidate at `assumed_cache_rate: 0.917`, including the one it is about to switch to, where the true rate is 0. Pricing the incumbent at the assumed rate and every challenger at 0 is a small change that pays for itself, and it makes the router's own decision an input to its own cost model for the first time. ## 5. Latency is the axis the operator feels and the one nothing optimizes NeuralWatt, where it is recorded at all: | model | avg s | max s | tok/s | |---|---|---|---| | gemma-4-31b | **22.49** | **1,251.1** | 22.6 | | glm-5.3 | 7.45 | 163.7 | 143.0 | | qwen3.6-35b | 4.63 | 172.5 | 99.4 | | kimi-k2.7-code | 3.65 | 226.9 | 93.5 | | deepseek-v4-flash-flex | 1.90 | 14.8 | 210.5 | An 11.8x spread in the mean and a 21-minute worst case. `latency_tolerance` exists but only gates `-flex` rows; `duration_seconds` is read by nothing in routing. And per item 1 it is **not collected at all** for OpenRouter or `ollama-local` — so for 73% of current traffic there is not even a number to ignore. `gemma-4-31b` is where every signal disagrees at once: the estimator's cheapest, billing's most expensive per request, and 6x slower than the field. ## 6. 96.6% of requests are never classified The classifier is the conceptual center of the system — three backends, a four-step failure cascade, its own circuit breaker, a degraded-share warning in `/metrics`, an attribution-exclusion rule in `report_outcome`, and a fine-tuning roadmap. Last five days: | `classification_source` | share | |---|---| | `cached` | **96.6%** | | `classifier` | 3.2% (avg 1,129 ms) | | `override` / null | 0.3% | The consequence is not only that the machinery is oversized. `task_category` on a given turn is **whatever the session was doing at the last cache refresh**, not a classification of that turn — one session ran 3,724 turns and cycled through all nine categories. So ~96% of turns carry a borrowed label, and `POST /outcome` attributes results to `(model, task_category)`. The quality table is largely trained on labels computed for a different turn, and nothing measures that noise. ## 7. The feedback loop worked, and the tolerance band was never narrowed `CLAUDE.md` is stale here in the good direction. The categories it records as "flat at 1.00 awaiting real traffic" are nothing of the kind: | category | rows | spread | outcome-backed | |---|---|---|---| | coding_general | 46 | 0.375 - 0.843 | 46/46 | | coding_refactor | 46 | 0.423 - 0.846 | 46/46 | | debugging | 46 | 0.307 - 0.836 | 46/46 | | tool_use_agentic | 46 | 0.298 - 0.918 | 46/46 | | reasoning_math | 46 | 0.000 - 0.956 | 46/46 | | summarization | 46 | 0.250 - 0.556 | 46/46 | | docs_writing | 46 | 0.607 - 0.949 | 46/46 | 2,319 client outcomes have arrived. The design bet — a 43-task benchmark is the prior, real traffic is the posterior — paid off comprehensively. Note also that scores came *down* as evidence thickened (`coding_general` topped out at 0.947 three days ago, 0.843 now) and that `summarization` now caps at 0.556: nothing in the catalog is good at it, which no benchmark had revealed. `quality_tolerance` is still 0.1, on a comment that says to narrow it as samples accumulate. Two things compound: the top band is 0.1 wide, so models differing by up to 10 points of **outcome-backed pass rate** are treated as equal; and the tiebreak inside that band is the estimator from item 3, wrong by up to 13.6x and misordered by 21x. The system trades up to 10 points of measured quality for a cost saving it cannot compute correctly. Caveat in the other direction: 379 of 474 proficiency rows are `outcome_prior` with zero direct samples, including every OpenRouter row. The band is too wide for the rows that *are* measured and arguably right for the ones that are not, which is an argument for making the band depend on sample depth rather than picking a new constant. ## 8. Structural verification is a 99% no-op on this workload | kind | verdict | n | |---|---|---| | structural | **unverifiable** | **28,958** | | structural | malformed | 218 | | structural | ok | 64 | | structural | truncated | 23 | | local_llm | ok | 480 | | local_llm | malformed | 47 | | client_outcome | succeeded | 1,840 | | client_outcome | failed | 479 | 98.96% of structural verdicts are `unverifiable`, and that is *correct* behavior: agent turns end in tool calls, and `has_tool_calls` short-circuits both checkers, which is the documented fix to a real false-failure incident. The conclusion not drawn is that after that fix the free structural checker checks nothing on the only workload this router serves. 305 of 29,263 verdicts were substantive. All of the subsystem's value is in `client_outcome`, the channel added last — and 527 `local_llm` verdicts have never been applied to anything. ## 9. The router has no concept of what a model is, only numbers on an id On 2026-09-06, before the allowlist existed, the ranker selected `google/lyria-3-clip-preview` — a music generation model — 38 times for coding, refactoring, debugging and summarization, **17 of them as the ranked winner rather than an exploration gamble**. No `energy_observations` rows exist for any: all 38 failed at dispatch. The allowlist fixed the symptom the same day and is the right pragmatic gate. But it is manual curation standing in for a structural absence: every filter in the system is a threshold on a metric, and none asks whether a candidate is a chat model. A new catalog row arrives with no proficiency — deliberately treated as unproven rather than bad — and is immediately eligible to win. --- ## Paradigms being shoehorned ### A. "Classify, then dispatch" is a single-request frame on a conversation The unit of work is a session: hundreds to thousands of turns sharing a monotonically growing prefix, ~92% of which is cached. The router models each turn as an independent classify-and-route problem, then bolts on a session cache to make that affordable — and the cache is now 96.6% of the answer. Inverting the frame dissolves several items at once. If the session is the unit and a turn is a delta: classification becomes "has the task changed?", cheap and usually no, instead of a ~1.1s full-taxonomy inference; the cache stops being an optimization and becomes the model, making its staleness window a real parameter rather than a cost hack; switching models becomes a visibly expensive act (item 4); and `min_tool_proficiency` becomes expressible, because "can this model be trusted with tools" is a session-level property — which is how the docs already describe it. ### B. Quality as a per-(model, category) scalar is a leaderboard paradigm One number per model per category, inherited from benchmarks. The project has already found twice that this cannot hold what it learns. `deepseek-v4-flash` scored 1.00 on all three coding categories and 0.33 on `tool_use_agentic`, and the documented reason is a *failure mode*: given both times in "it is 1:20pm and my meeting is at 3pm", it calls two tools instead of subtracting. That is not a lower skill level on a category; it is a disposition that is hazardous whenever tools are on the table, whatever the task is. Flattening a mode into a category score forced a second mechanism to express the real rule — `routing.min_tool_proficiency`, a hard filter — and that filter is `null` because the flattening made it unusable: opencode sends `tools` on nearly every request, so switching it on excludes the cheap model from all traffic. The paradigm mismatch is what left a known hazard ungated. The inheritance structure shows the same strain from the other side: 379 rows are peer-rate priors carrying a number with zero evidence behind it, and nothing downstream distinguishes a 0.749 that was measured from a 0.749 that was copied — except `source`, which the ranker does not read. ### C. Energy and carbon are structurally central and functionally vestigial `eco` is "not an objective." `max_energy_per_request` is null. `plan_kwh_per_period` gates nothing. The project's own conclusion is that billed energy "ranks how busy the provider was, not how efficient the model is." Around that sit: two columns and a comment block in `energy_observations`, a BTU column kept "purely for comedic dashboard value", `seed_energy.py`, a 6-hourly timer whose accumulation is the stated prerequisite for trusting the eco ordering, `seed_local_dispatch_energy.py`, grid-intensity and `carbon_source` logging, open follow-up item 3, and a large share of a 71KB `CLAUDE.md`. The measurement work was worth doing — the per-kWh billing discovery and the attribution decomposition are the most valuable findings in the project, and item 3 above is an argument for using them *more*. But they belong to the cost model now, not to an eco objective nothing optimizes. And the provider now carrying 73% of traffic reports no energy at all, which makes the framing actively misleading about what the router can see. --- ## Scale check | | | |---|---| | Python | 55,711 lines; `dispatcher.py` 4,974, `admin.py` 2,498, `metrics.py` 2,474 | | tests | 1,155 | | config | 129 leaf keys across 22 sections | | `CLAUDE.md` | 71 KB | | `plans/` | 62 documents | | purpose | route one operator's coding agent | The engineering quality is high in ways that are rare and worth keeping: the schema-drift and warnings tripwires, the `inherited_from` migration that fails safe on NULL, the refusal to fabricate leaderboard priors, the four recorded harness bugs that scored the rig rather than the model. None of that is the problem. The problem is the absence of any forcing function that retires machinery, or that re-fires a premise — 62 plan documents and no mechanism that closes one out. --- ## What it needs, ranked 1. **A `feedback.py` timer.** One oneshot service and one timer, matching the four that already exist. The ground-truth loop is the only one cranked by hand, `proficiency` is three days stale, and the 95 unapplied outcomes are exactly the models carrying today's traffic. Cheapest item here by a wide margin. 2. **Capture OpenRouter cost and latency telemetry.** 95.7% of its rows have NULL `cost_usd`, 100% have NULL `duration_seconds`, and it is 73% of current decisions. This is a prerequisite for items 3 and 5, not a parallel task. 3. **Calibrate `estimated_cost` against the 31,119 billed rows already in the DB.** A per-model correction factor over a trailing window. Fixes the 21x misordering without a sweep and without abandoning per-request shape sensitivity. 4. **Price the switch.** Charge challengers a cold-cache rate in the tiebreak. ~17% of spend sits on 7% of turns, and it makes the router's own decision an input to its own cost model. 5. **Put latency in the objective.** A p50/p95 per model and a latency term or ceiling for `INTERACTIVE` requests. Needs item 2 for most of the catalog. 6. **Make `quality_tolerance` depend on sample depth** rather than replacing one constant with another. 95 of 474 rows are measured, 379 are inherited priors; a band that is right for one is wrong for the other. 7. **An automatic response to a cost shock.** The 09-06 spike was handled well and handled by a human. The minimum honest version: a period spend ceiling that narrows the candidate set as it is approached, then refuses with a clear reason. Not "the router is unusable without it" — it demonstrably is not — but it is the difference between a tool that reports a problem and a tool that responds to one. 8. **Decide what the classifier is for, given 96.6% cached.** Either accept it as session-level and simplify the machinery to match, or make per-turn classification cheap enough to actually run. The current state pays the architecture cost of per-turn classification and gets session-level labels. 9. **Retire the eco objective explicitly.** Move the energy and attribution findings into the cost model where they are load-bearing; archive the rest — the seed timer, the eco scoring path, the BTU column, open item 3. 10. **A premise-expiry mechanism.** The thesis. Every setting whose comment says "revisit when X" should have X as a check that fails loudly — a test, a `/metrics` warning, a line in the poller. Four have expired silently. Items 1-5 are independent and small, and 1 and 2 should land first because everything downstream reads what they produce. Items 6, 8, 9 and 10 are decisions rather than tasks, and should be made before more machinery is added to the subsystems they touch.