Adds router_wall_seconds and router_ttft_seconds to energy_observations: - router_wall_seconds: time.monotonic() from just before the accepted candidate's POST/connection-open to the complete response body (buffered) or the last byte forwarded (streaming). Re-marked per candidate so failover time is excluded — a dead model's 30s stall is not charged to the healthy one that replaced it. - router_ttft_seconds: streaming-only. First delta carrying non-empty content or a tool_calls fragment, excluding the role-only opening delta. NULL on buffered rows (not applicable) and on streams that produced no output token (a broken upstream — not the same as a literal 0). Separate from duration_seconds (provider's reported serving time) on purpose: OpenRouter reports duration_seconds on 0 of 3,850 rows, while the router can always measure its own clock. The two quantities are stored independently and never written into each other. log_observation() defaults both new params to None so seed_energy.py and eval_proficiency.py pass unchanged — a reprise of 6e729ad's bug where new keyword-only arguments killed the seed timer. Also: - docs/data-model.md: document the new columns, span definitions, and the distinction from duration_seconds - config/schema.sql: full CREATE TABLE declaration - plans/token-waste-waves.md: status update for Wave 1 - plans/ten-thousand-foot-review.md: companion diagnosis
22 KiB
A 10,000ft review: what this router needs to be a more useful tool
Status: reference -- a diagnosis, not a queue item; the executable work it
feeds is scoped in plans/token-waste-waves.md.
Measured on the LIVE database at /home/alee/Sources/6krrt/router.db
(read-only), 2026-09-13, at repo state 26398d7. An earlier draft of this file
measured a stale 11:47 copy of router.db sitting inside the
profiles-header-fix worktree and drew a wrong headline from it; that copy is
~10 hours behind and missing the traffic that matters most. If you are
measuring from a worktree, check the mtime first.
Audience: the operator, and whoever scopes the next round of work. Deliberately
does not repeat the item-level work already scoped in
code-analysis-refactoring-opportunities.md.
The headline: the router works, and it just proved it
On 2026-09-06 NeuralWatt billed $19.85 in one day across 4,129 decisions. What followed was the most consequential routing change in the project's history: OpenRouter came back as an allowlist-gated provider and traffic moved.
| day | decisions | billed |
|---|---|---|
| 2026-09-06 | 4,129 | $19.85 |
| 2026-09-09 | 594 | $0.32 |
| 2026-09-10 | 722 | $0.25 |
| 2026-09-11 | 755 | $0.25 |
| 2026-09-12 | 116 | $0.40 |
| 2026-09-13 | 154 | $0.21 |
Volume held at 600-750 decisions a day. Billed spend fell roughly 50x. That is the system doing precisely what it exists to do, and it is the strongest evidence in the database that the thing is useful.
Two qualifications, and they are where the rest of this document lives:
- The human did the adaptation, not the router. The migration was an
operator decision — add a provider, write an allowlist, let ranking follow.
The router has no automatic response to a cost shock;
plan_kwh_per_periodgates nothing andmax_energy_per_requestis null, so every cost mechanism in the system is advisory. It reported the spike well. It did nothing about it. - The migration moved 73% of traffic onto a provider the router cannot measure. That is the single most actionable finding here, and it is new since the migration.
The thesis
This project is unusually good at recording why a decision was made, and has no mechanism for noticing when a recorded reason has expired.
| setting | the comment's own premise | what is true now |
|---|---|---|
quality_tolerance: 0.1 |
"scores currently rest on 2-3 samples per category, so a 0.05 gap is indistinguishable from sampling variation... Narrow it as samples accumulate" | 46 rows per category, all outcome-backed; outcome_blended rows average 51.8 client samples, max 460 |
| cost-as-tiebreak, not an objective | "60% of every decision adjudicated fractions of a cent (all real traffic to date totals $0.07)" | $118.79 billed. 1,697x the figure the whole argument rests on |
assumed_cache_rate: 0.917 |
measured once, 2026-08-23, over 40.7M tokens | cached_prompt_tokens is NULL in all 35,086 rows. The router receives this field and drops it |
| cost feedback generally | catalog estimates "track the real ordering" because billing is capped at 3x list | true for NeuralWatt, untestable for OpenRouter: 95.7% of its rows have NULL cost_usd |
Four settings, four expiry conditions met, nothing re-fired. That pattern is worth fixing before any individual number. The documentation is a real strength — the reasoning is preserved, measurements are dated, caveats are honest. What is missing is anything that acts when a premise the code depends on stops holding.
1. The provider that now carries the traffic is the one the router cannot see
Recent picks, since 2026-09-09 (2,341 decisions):
| provider | model | n |
|---|---|---|
| openrouter | xiaomi/mimo-v2.5 |
744 |
| openrouter | deepseek/deepseek-v4-flash |
659 |
| neuralwatt | deepseek-v4-flash |
266 |
| neuralwatt | qwen3.6-35b |
247 |
| openrouter | nvidia/nemotron-3-nano-omni-...:free |
178 |
| neuralwatt | qwen3.6-35b-fast |
119 |
| openrouter | z-ai/glm-5.3-flash |
95 |
| openrouter | xiaomi/mimo-v2.5-pro |
26 |
OpenRouter is 72.8% of current decisions. Telemetry coverage in
energy_observations:
| provider | rows | NULL cost_usd |
NULL duration_seconds |
|---|---|---|---|
| neuralwatt | 31,188 | 66 | 66 |
| openrouter | 3,838 | 3,673 (95.7%) | 3,838 (100%) |
| ollama-local | 60 | 60 | 60 |
So on the majority of current traffic the router has no billed cost to calibrate its estimator against and no completion time at all. Every feedback loop that makes this system smarter than a static config is blind on the provider it now prefers.
This inverts the priority of everything below: recalibrating the cost model (item 3) cannot work for 73% of traffic until this is fixed. OpenRouter returns usage and generation-cost data; it is an ingestion gap, not a provider limitation.
2. feedback.py is the only loop without a timer
proficiency was written in a single batch at 2026-09-10T00:44:55 — every
row, one instant, three days ago. Since then:
| client outcomes applied | 2,224 |
| client outcomes unapplied | 95 |
And the unapplied 95 are exactly the current traffic:
| model | provider | unapplied | of which failed |
|---|---|---|---|
xiaomi/mimo-v2.5 |
openrouter | 41 | 9 |
deepseek/deepseek-v4-flash |
openrouter | 36 | 15 |
deepseek-v4-flash |
neuralwatt | 17 | 3 |
xiaomi/mimo-v2.5-pro |
openrouter | 1 | 0 |
Meanwhile every OpenRouter model's proficiency row is outcome_prior with
outcome_samples = 0 — an inherited peer-rate prior, not a measurement.
xiaomi/mimo-v2.5 is the most-routed model in the catalog right now (744
picks) and scores 0.749 / 0.770 / 0.894 on coding_general / coding_refactor /
tool_use_agentic entirely on inheritance. Its 41 real outcomes (78% pass) are
sitting in verifications unapplied.
deploy/ ships timers for the poller, the seed sweep, backup and offsite sync,
and ~/.config/systemd/user additionally has a baseline-report timer. There is
no feedback timer. The loop the docs call "the only ground truth" is the
only one that has to be cranked by hand.
This is the cheapest high-value fix in the document: a oneshot service and a timer, matching the four that already exist.
3. The cost estimator is wrong in a model-dependent way
routing.estimated_cost breaks every tie, and the ranking reduces to it
whenever candidates land in the same quality band. Per request, on rows with a
real billed figure:
| model | provider | est µ$ | billed µ$ | est/billed |
|---|---|---|---|---|
| gemma-4-31b | neuralwatt | 3,822 | 6,041 | 0.63x |
| kimi-k2.7-code-fast | neuralwatt | 28,623 | 21,747 | 1.32x |
| kimi-k3 | neuralwatt | 87,910 | 42,388 | 2.07x |
| deepseek-v4-flash | neuralwatt | 4,154 | 1,796 | 2.31x |
| glm-5.3 | neuralwatt | 48,232 | 20,056 | 2.40x |
deepseek/deepseek-v4-flash |
openrouter | 3,830 | 1,383 | 2.77x |
| kimi-k2.7-code | neuralwatt | 22,253 | 3,906 | 5.70x |
| qwen3.6-35b | neuralwatt | 4,091 | 484 | 8.46x |
qwen3.6-35b-fast |
neuralwatt | 3,773 | 282 | 13.37x |
xiaomi/mimo-v2.5 |
openrouter | 9,689 | 715 | 13.55x |
Total: $401.99 estimated against $118.79 billed. The scale error does not
matter; the 21x spread in the error does, because it reorders candidates.
gemma-4-31b is the estimator's cheapest model in the catalog and billing's
most expensive per request of the ten. The router's current favourite,
xiaomi/mimo-v2.5, is the one it overestimates most — it is winning on
quality despite the cost model, not because of it.
The cause is already written down in CLAUDE.md, two sections apart, never
reconciled: attribution ratio spans 750x between models and is "most of
the real cost difference in the catalog", while estimated_cost prices from
catalog token prices, which contain no attribution term at all.
Cost moved off measurement to fix a real problem — a 400-token sweep cannot price a 150k-token workload — and in the move discarded the term the project had already proven was dominant. Third recurrence of one pattern: list price ranks models backwards; the reference sweep ranks them backwards for real traffic; catalog pricing misorders them because it omits pool concurrency.
The fix needs no sweep. 31,119 NeuralWatt rows carry a real billed figure,
joinable to decisions by request_id. A per-model correction factor over a
trailing window, refreshed by the poller, recalibrates against the bill and
keeps the per-request shape sensitivity that motivated the move. For OpenRouter
it needs item 1 first.
4. Switching models mid-session costs 2.5x, and the estimator prices it as free
No sticky routing, no concept of the previously-used model. On turns with prompts over 20k tokens and a real billed figure:
| turn | n | avg prompt | µ$ / prompt token | spend |
|---|---|---|---|---|
| same model as previous | 11,246 | 100,994 | 0.037 | $47.91 |
| switched model | 858 | 90,309 | 0.091 | $9.68 |
2.46x, on smaller prompts, so the effect is understated. Switch turns are 7.1% of turns and 16.8% of joined spend.
The mechanism is not mysterious: a switch lands a ~100k-token prompt on a
provider that has never seen it, so the prefix cache is cold and ~92% of the
prompt is suddenly billed fresh. But rank_candidates prices every candidate
at assumed_cache_rate: 0.917, including the one it is about to switch to,
where the true rate is 0.
Pricing the incumbent at the assumed rate and every challenger at 0 is a small change that pays for itself, and it makes the router's own decision an input to its own cost model for the first time.
5. Latency is the axis the operator feels and the one nothing optimizes
NeuralWatt, where it is recorded at all:
| model | avg s | max s | tok/s |
|---|---|---|---|
| gemma-4-31b | 22.49 | 1,251.1 | 22.6 |
| glm-5.3 | 7.45 | 163.7 | 143.0 |
| qwen3.6-35b | 4.63 | 172.5 | 99.4 |
| kimi-k2.7-code | 3.65 | 226.9 | 93.5 |
| deepseek-v4-flash-flex | 1.90 | 14.8 | 210.5 |
An 11.8x spread in the mean and a 21-minute worst case. latency_tolerance
exists but only gates -flex rows; duration_seconds is read by nothing in
routing. And per item 1 it is not collected at all for OpenRouter or
ollama-local — so for 73% of current traffic there is not even a number to
ignore.
gemma-4-31b is where every signal disagrees at once: the estimator's
cheapest, billing's most expensive per request, and 6x slower than the field.
6. 96.6% of requests are never classified
The classifier is the conceptual center of the system — three backends, a
four-step failure cascade, its own circuit breaker, a degraded-share warning in
/metrics, an attribution-exclusion rule in report_outcome, and a
fine-tuning roadmap. Last five days:
classification_source |
share |
|---|---|
cached |
96.6% |
classifier |
3.2% (avg 1,129 ms) |
override / null |
0.3% |
The consequence is not only that the machinery is oversized. task_category on
a given turn is whatever the session was doing at the last cache refresh,
not a classification of that turn — one session ran 3,724 turns and cycled
through all nine categories. So ~96% of turns carry a borrowed label, and
POST /outcome attributes results to (model, task_category). The quality
table is largely trained on labels computed for a different turn, and nothing
measures that noise.
7. The feedback loop worked, and the tolerance band was never narrowed
CLAUDE.md is stale here in the good direction. The categories it records as
"flat at 1.00 awaiting real traffic" are nothing of the kind:
| category | rows | spread | outcome-backed |
|---|---|---|---|
| coding_general | 46 | 0.375 - 0.843 | 46/46 |
| coding_refactor | 46 | 0.423 - 0.846 | 46/46 |
| debugging | 46 | 0.307 - 0.836 | 46/46 |
| tool_use_agentic | 46 | 0.298 - 0.918 | 46/46 |
| reasoning_math | 46 | 0.000 - 0.956 | 46/46 |
| summarization | 46 | 0.250 - 0.556 | 46/46 |
| docs_writing | 46 | 0.607 - 0.949 | 46/46 |
2,319 client outcomes have arrived. The design bet — a 43-task benchmark is the
prior, real traffic is the posterior — paid off comprehensively. Note also that
scores came down as evidence thickened (coding_general topped out at 0.947
three days ago, 0.843 now) and that summarization now caps at 0.556: nothing
in the catalog is good at it, which no benchmark had revealed.
quality_tolerance is still 0.1, on a comment that says to narrow it as
samples accumulate. Two things compound: the top band is 0.1 wide, so models
differing by up to 10 points of outcome-backed pass rate are treated as
equal; and the tiebreak inside that band is the estimator from item 3, wrong by
up to 13.6x and misordered by 21x. The system trades up to 10 points of
measured quality for a cost saving it cannot compute correctly.
Caveat in the other direction: 379 of 474 proficiency rows are outcome_prior
with zero direct samples, including every OpenRouter row. The band is too wide
for the rows that are measured and arguably right for the ones that are not,
which is an argument for making the band depend on sample depth rather than
picking a new constant.
8. Structural verification is a 99% no-op on this workload
| kind | verdict | n |
|---|---|---|
| structural | unverifiable | 28,958 |
| structural | malformed | 218 |
| structural | ok | 64 |
| structural | truncated | 23 |
| local_llm | ok | 480 |
| local_llm | malformed | 47 |
| client_outcome | succeeded | 1,840 |
| client_outcome | failed | 479 |
98.96% of structural verdicts are unverifiable, and that is correct
behavior: agent turns end in tool calls, and has_tool_calls short-circuits
both checkers, which is the documented fix to a real false-failure incident.
The conclusion not drawn is that after that fix the free structural checker
checks nothing on the only workload this router serves. 305 of 29,263 verdicts
were substantive. All of the subsystem's value is in client_outcome, the
channel added last — and 527 local_llm verdicts have never been applied to
anything.
9. The router has no concept of what a model is, only numbers on an id
On 2026-09-06, before the allowlist existed, the ranker selected
google/lyria-3-clip-preview — a music generation model — 38 times for coding,
refactoring, debugging and summarization, 17 of them as the ranked winner
rather than an exploration gamble. No energy_observations rows exist for
any: all 38 failed at dispatch.
The allowlist fixed the symptom the same day and is the right pragmatic gate. But it is manual curation standing in for a structural absence: every filter in the system is a threshold on a metric, and none asks whether a candidate is a chat model. A new catalog row arrives with no proficiency — deliberately treated as unproven rather than bad — and is immediately eligible to win.
Paradigms being shoehorned
A. "Classify, then dispatch" is a single-request frame on a conversation
The unit of work is a session: hundreds to thousands of turns sharing a monotonically growing prefix, ~92% of which is cached. The router models each turn as an independent classify-and-route problem, then bolts on a session cache to make that affordable — and the cache is now 96.6% of the answer.
Inverting the frame dissolves several items at once. If the session is the unit
and a turn is a delta: classification becomes "has the task changed?", cheap
and usually no, instead of a ~1.1s full-taxonomy inference; the cache stops
being an optimization and becomes the model, making its staleness window a real
parameter rather than a cost hack; switching models becomes a visibly expensive
act (item 4); and min_tool_proficiency becomes expressible, because "can this
model be trusted with tools" is a session-level property — which is how the
docs already describe it.
B. Quality as a per-(model, category) scalar is a leaderboard paradigm
One number per model per category, inherited from benchmarks. The project has already found twice that this cannot hold what it learns.
deepseek-v4-flash scored 1.00 on all three coding categories and 0.33 on
tool_use_agentic, and the documented reason is a failure mode: given both
times in "it is 1:20pm and my meeting is at 3pm", it calls two tools instead of
subtracting. That is not a lower skill level on a category; it is a disposition
that is hazardous whenever tools are on the table, whatever the task is.
Flattening a mode into a category score forced a second mechanism to express
the real rule — routing.min_tool_proficiency, a hard filter — and that filter
is null because the flattening made it unusable: opencode sends tools on
nearly every request, so switching it on excludes the cheap model from all
traffic. The paradigm mismatch is what left a known hazard ungated.
The inheritance structure shows the same strain from the other side: 379 rows
are peer-rate priors carrying a number with zero evidence behind it, and
nothing downstream distinguishes a 0.749 that was measured from a 0.749 that
was copied — except source, which the ranker does not read.
C. Energy and carbon are structurally central and functionally vestigial
eco is "not an objective." max_energy_per_request is null.
plan_kwh_per_period gates nothing. The project's own conclusion is that
billed energy "ranks how busy the provider was, not how efficient the model
is."
Around that sit: two columns and a comment block in energy_observations, a
BTU column kept "purely for comedic dashboard value", seed_energy.py, a
6-hourly timer whose accumulation is the stated prerequisite for trusting the
eco ordering, seed_local_dispatch_energy.py, grid-intensity and
carbon_source logging, open follow-up item 3, and a large share of a 71KB
CLAUDE.md.
The measurement work was worth doing — the per-kWh billing discovery and the attribution decomposition are the most valuable findings in the project, and item 3 above is an argument for using them more. But they belong to the cost model now, not to an eco objective nothing optimizes. And the provider now carrying 73% of traffic reports no energy at all, which makes the framing actively misleading about what the router can see.
Scale check
| Python | 55,711 lines; dispatcher.py 4,974, admin.py 2,498, metrics.py 2,474 |
| tests | 1,155 |
| config | 129 leaf keys across 22 sections |
CLAUDE.md |
71 KB |
plans/ |
62 documents |
| purpose | route one operator's coding agent |
The engineering quality is high in ways that are rare and worth keeping: the
schema-drift and warnings tripwires, the inherited_from migration that fails
safe on NULL, the refusal to fabricate leaderboard priors, the four recorded
harness bugs that scored the rig rather than the model. None of that is the
problem. The problem is the absence of any forcing function that retires
machinery, or that re-fires a premise — 62 plan documents and no mechanism that
closes one out.
What it needs, ranked
-
A
feedback.pytimer. One oneshot service and one timer, matching the four that already exist. The ground-truth loop is the only one cranked by hand,proficiencyis three days stale, and the 95 unapplied outcomes are exactly the models carrying today's traffic. Cheapest item here by a wide margin. -
Capture OpenRouter cost and latency telemetry. 95.7% of its rows have NULL
cost_usd, 100% have NULLduration_seconds, and it is 73% of current decisions. This is a prerequisite for items 3 and 5, not a parallel task. -
Calibrate
estimated_costagainst the 31,119 billed rows already in the DB. A per-model correction factor over a trailing window. Fixes the 21x misordering without a sweep and without abandoning per-request shape sensitivity. -
Price the switch. Charge challengers a cold-cache rate in the tiebreak. ~17% of spend sits on 7% of turns, and it makes the router's own decision an input to its own cost model.
-
Put latency in the objective. A p50/p95 per model and a latency term or ceiling for
INTERACTIVErequests. Needs item 2 for most of the catalog. -
Make
quality_tolerancedepend on sample depth rather than replacing one constant with another. 95 of 474 rows are measured, 379 are inherited priors; a band that is right for one is wrong for the other. -
An automatic response to a cost shock. The 09-06 spike was handled well and handled by a human. The minimum honest version: a period spend ceiling that narrows the candidate set as it is approached, then refuses with a clear reason. Not "the router is unusable without it" — it demonstrably is not — but it is the difference between a tool that reports a problem and a tool that responds to one.
-
Decide what the classifier is for, given 96.6% cached. Either accept it as session-level and simplify the machinery to match, or make per-turn classification cheap enough to actually run. The current state pays the architecture cost of per-turn classification and gets session-level labels.
-
Retire the eco objective explicitly. Move the energy and attribution findings into the cost model where they are load-bearing; archive the rest — the seed timer, the eco scoring path, the BTU column, open item 3.
-
A premise-expiry mechanism. The thesis. Every setting whose comment says "revisit when X" should have X as a check that fails loudly — a test, a
/metricswarning, a line in the poller. Four have expired silently.
Items 1-5 are independent and small, and 1 and 2 should land first because everything downstream reads what they produce. Items 6, 8, 9 and 10 are decisions rather than tasks, and should be made before more machinery is added to the subsystems they touch.