Files
6krrt/plans/ten-thousand-foot-review.md
adlee-was-taken 0111da9bbf feat(telemetry): router-observed wall-clock and TTFT columns, with plan docs
Adds router_wall_seconds and router_ttft_seconds to energy_observations:

- router_wall_seconds: time.monotonic() from just before the accepted
  candidate's POST/connection-open to the complete response body (buffered)
  or the last byte forwarded (streaming). Re-marked per candidate so
  failover time is excluded — a dead model's 30s stall is not charged to
  the healthy one that replaced it.
- router_ttft_seconds: streaming-only. First delta carrying non-empty
  content or a tool_calls fragment, excluding the role-only opening delta.
  NULL on buffered rows (not applicable) and on streams that produced no
  output token (a broken upstream — not the same as a literal 0).

Separate from duration_seconds (provider's reported serving time) on
purpose: OpenRouter reports duration_seconds on 0 of 3,850 rows, while
the router can always measure its own clock. The two quantities are stored
independently and never written into each other.

log_observation() defaults both new params to None so seed_energy.py and
eval_proficiency.py pass unchanged — a reprise of 6e729ad's bug where new
keyword-only arguments killed the seed timer.

Also:
- docs/data-model.md: document the new columns, span definitions, and
  the distinction from duration_seconds
- config/schema.sql: full CREATE TABLE declaration
- plans/token-waste-waves.md: status update for Wave 1
- plans/ten-thousand-foot-review.md: companion diagnosis
2026-09-13 12:25:39 -04:00

22 KiB

A 10,000ft review: what this router needs to be a more useful tool

Status: reference -- a diagnosis, not a queue item; the executable work it feeds is scoped in plans/token-waste-waves.md.

Measured on the LIVE database at /home/alee/Sources/6krrt/router.db (read-only), 2026-09-13, at repo state 26398d7. An earlier draft of this file measured a stale 11:47 copy of router.db sitting inside the profiles-header-fix worktree and drew a wrong headline from it; that copy is ~10 hours behind and missing the traffic that matters most. If you are measuring from a worktree, check the mtime first.

Audience: the operator, and whoever scopes the next round of work. Deliberately does not repeat the item-level work already scoped in code-analysis-refactoring-opportunities.md.


The headline: the router works, and it just proved it

On 2026-09-06 NeuralWatt billed $19.85 in one day across 4,129 decisions. What followed was the most consequential routing change in the project's history: OpenRouter came back as an allowlist-gated provider and traffic moved.

day decisions billed
2026-09-06 4,129 $19.85
2026-09-09 594 $0.32
2026-09-10 722 $0.25
2026-09-11 755 $0.25
2026-09-12 116 $0.40
2026-09-13 154 $0.21

Volume held at 600-750 decisions a day. Billed spend fell roughly 50x. That is the system doing precisely what it exists to do, and it is the strongest evidence in the database that the thing is useful.

Two qualifications, and they are where the rest of this document lives:

  • The human did the adaptation, not the router. The migration was an operator decision — add a provider, write an allowlist, let ranking follow. The router has no automatic response to a cost shock; plan_kwh_per_period gates nothing and max_energy_per_request is null, so every cost mechanism in the system is advisory. It reported the spike well. It did nothing about it.
  • The migration moved 73% of traffic onto a provider the router cannot measure. That is the single most actionable finding here, and it is new since the migration.

The thesis

This project is unusually good at recording why a decision was made, and has no mechanism for noticing when a recorded reason has expired.

setting the comment's own premise what is true now
quality_tolerance: 0.1 "scores currently rest on 2-3 samples per category, so a 0.05 gap is indistinguishable from sampling variation... Narrow it as samples accumulate" 46 rows per category, all outcome-backed; outcome_blended rows average 51.8 client samples, max 460
cost-as-tiebreak, not an objective "60% of every decision adjudicated fractions of a cent (all real traffic to date totals $0.07)" $118.79 billed. 1,697x the figure the whole argument rests on
assumed_cache_rate: 0.917 measured once, 2026-08-23, over 40.7M tokens cached_prompt_tokens is NULL in all 35,086 rows. The router receives this field and drops it
cost feedback generally catalog estimates "track the real ordering" because billing is capped at 3x list true for NeuralWatt, untestable for OpenRouter: 95.7% of its rows have NULL cost_usd

Four settings, four expiry conditions met, nothing re-fired. That pattern is worth fixing before any individual number. The documentation is a real strength — the reasoning is preserved, measurements are dated, caveats are honest. What is missing is anything that acts when a premise the code depends on stops holding.


1. The provider that now carries the traffic is the one the router cannot see

Recent picks, since 2026-09-09 (2,341 decisions):

provider model n
openrouter xiaomi/mimo-v2.5 744
openrouter deepseek/deepseek-v4-flash 659
neuralwatt deepseek-v4-flash 266
neuralwatt qwen3.6-35b 247
openrouter nvidia/nemotron-3-nano-omni-...:free 178
neuralwatt qwen3.6-35b-fast 119
openrouter z-ai/glm-5.3-flash 95
openrouter xiaomi/mimo-v2.5-pro 26

OpenRouter is 72.8% of current decisions. Telemetry coverage in energy_observations:

provider rows NULL cost_usd NULL duration_seconds
neuralwatt 31,188 66 66
openrouter 3,838 3,673 (95.7%) 3,838 (100%)
ollama-local 60 60 60

So on the majority of current traffic the router has no billed cost to calibrate its estimator against and no completion time at all. Every feedback loop that makes this system smarter than a static config is blind on the provider it now prefers.

This inverts the priority of everything below: recalibrating the cost model (item 3) cannot work for 73% of traffic until this is fixed. OpenRouter returns usage and generation-cost data; it is an ingestion gap, not a provider limitation.

2. feedback.py is the only loop without a timer

proficiency was written in a single batch at 2026-09-10T00:44:55 — every row, one instant, three days ago. Since then:

client outcomes applied 2,224
client outcomes unapplied 95

And the unapplied 95 are exactly the current traffic:

model provider unapplied of which failed
xiaomi/mimo-v2.5 openrouter 41 9
deepseek/deepseek-v4-flash openrouter 36 15
deepseek-v4-flash neuralwatt 17 3
xiaomi/mimo-v2.5-pro openrouter 1 0

Meanwhile every OpenRouter model's proficiency row is outcome_prior with outcome_samples = 0 — an inherited peer-rate prior, not a measurement. xiaomi/mimo-v2.5 is the most-routed model in the catalog right now (744 picks) and scores 0.749 / 0.770 / 0.894 on coding_general / coding_refactor / tool_use_agentic entirely on inheritance. Its 41 real outcomes (78% pass) are sitting in verifications unapplied.

deploy/ ships timers for the poller, the seed sweep, backup and offsite sync, and ~/.config/systemd/user additionally has a baseline-report timer. There is no feedback timer. The loop the docs call "the only ground truth" is the only one that has to be cranked by hand.

This is the cheapest high-value fix in the document: a oneshot service and a timer, matching the four that already exist.

3. The cost estimator is wrong in a model-dependent way

routing.estimated_cost breaks every tie, and the ranking reduces to it whenever candidates land in the same quality band. Per request, on rows with a real billed figure:

model provider est µ$ billed µ$ est/billed
gemma-4-31b neuralwatt 3,822 6,041 0.63x
kimi-k2.7-code-fast neuralwatt 28,623 21,747 1.32x
kimi-k3 neuralwatt 87,910 42,388 2.07x
deepseek-v4-flash neuralwatt 4,154 1,796 2.31x
glm-5.3 neuralwatt 48,232 20,056 2.40x
deepseek/deepseek-v4-flash openrouter 3,830 1,383 2.77x
kimi-k2.7-code neuralwatt 22,253 3,906 5.70x
qwen3.6-35b neuralwatt 4,091 484 8.46x
qwen3.6-35b-fast neuralwatt 3,773 282 13.37x
xiaomi/mimo-v2.5 openrouter 9,689 715 13.55x

Total: $401.99 estimated against $118.79 billed. The scale error does not matter; the 21x spread in the error does, because it reorders candidates. gemma-4-31b is the estimator's cheapest model in the catalog and billing's most expensive per request of the ten. The router's current favourite, xiaomi/mimo-v2.5, is the one it overestimates most — it is winning on quality despite the cost model, not because of it.

The cause is already written down in CLAUDE.md, two sections apart, never reconciled: attribution ratio spans 750x between models and is "most of the real cost difference in the catalog", while estimated_cost prices from catalog token prices, which contain no attribution term at all.

Cost moved off measurement to fix a real problem — a 400-token sweep cannot price a 150k-token workload — and in the move discarded the term the project had already proven was dominant. Third recurrence of one pattern: list price ranks models backwards; the reference sweep ranks them backwards for real traffic; catalog pricing misorders them because it omits pool concurrency.

The fix needs no sweep. 31,119 NeuralWatt rows carry a real billed figure, joinable to decisions by request_id. A per-model correction factor over a trailing window, refreshed by the poller, recalibrates against the bill and keeps the per-request shape sensitivity that motivated the move. For OpenRouter it needs item 1 first.

4. Switching models mid-session costs 2.5x, and the estimator prices it as free

No sticky routing, no concept of the previously-used model. On turns with prompts over 20k tokens and a real billed figure:

turn n avg prompt µ$ / prompt token spend
same model as previous 11,246 100,994 0.037 $47.91
switched model 858 90,309 0.091 $9.68

2.46x, on smaller prompts, so the effect is understated. Switch turns are 7.1% of turns and 16.8% of joined spend.

The mechanism is not mysterious: a switch lands a ~100k-token prompt on a provider that has never seen it, so the prefix cache is cold and ~92% of the prompt is suddenly billed fresh. But rank_candidates prices every candidate at assumed_cache_rate: 0.917, including the one it is about to switch to, where the true rate is 0.

Pricing the incumbent at the assumed rate and every challenger at 0 is a small change that pays for itself, and it makes the router's own decision an input to its own cost model for the first time.

5. Latency is the axis the operator feels and the one nothing optimizes

NeuralWatt, where it is recorded at all:

model avg s max s tok/s
gemma-4-31b 22.49 1,251.1 22.6
glm-5.3 7.45 163.7 143.0
qwen3.6-35b 4.63 172.5 99.4
kimi-k2.7-code 3.65 226.9 93.5
deepseek-v4-flash-flex 1.90 14.8 210.5

An 11.8x spread in the mean and a 21-minute worst case. latency_tolerance exists but only gates -flex rows; duration_seconds is read by nothing in routing. And per item 1 it is not collected at all for OpenRouter or ollama-local — so for 73% of current traffic there is not even a number to ignore.

gemma-4-31b is where every signal disagrees at once: the estimator's cheapest, billing's most expensive per request, and 6x slower than the field.

6. 96.6% of requests are never classified

The classifier is the conceptual center of the system — three backends, a four-step failure cascade, its own circuit breaker, a degraded-share warning in /metrics, an attribution-exclusion rule in report_outcome, and a fine-tuning roadmap. Last five days:

classification_source share
cached 96.6%
classifier 3.2% (avg 1,129 ms)
override / null 0.3%

The consequence is not only that the machinery is oversized. task_category on a given turn is whatever the session was doing at the last cache refresh, not a classification of that turn — one session ran 3,724 turns and cycled through all nine categories. So ~96% of turns carry a borrowed label, and POST /outcome attributes results to (model, task_category). The quality table is largely trained on labels computed for a different turn, and nothing measures that noise.

7. The feedback loop worked, and the tolerance band was never narrowed

CLAUDE.md is stale here in the good direction. The categories it records as "flat at 1.00 awaiting real traffic" are nothing of the kind:

category rows spread outcome-backed
coding_general 46 0.375 - 0.843 46/46
coding_refactor 46 0.423 - 0.846 46/46
debugging 46 0.307 - 0.836 46/46
tool_use_agentic 46 0.298 - 0.918 46/46
reasoning_math 46 0.000 - 0.956 46/46
summarization 46 0.250 - 0.556 46/46
docs_writing 46 0.607 - 0.949 46/46

2,319 client outcomes have arrived. The design bet — a 43-task benchmark is the prior, real traffic is the posterior — paid off comprehensively. Note also that scores came down as evidence thickened (coding_general topped out at 0.947 three days ago, 0.843 now) and that summarization now caps at 0.556: nothing in the catalog is good at it, which no benchmark had revealed.

quality_tolerance is still 0.1, on a comment that says to narrow it as samples accumulate. Two things compound: the top band is 0.1 wide, so models differing by up to 10 points of outcome-backed pass rate are treated as equal; and the tiebreak inside that band is the estimator from item 3, wrong by up to 13.6x and misordered by 21x. The system trades up to 10 points of measured quality for a cost saving it cannot compute correctly.

Caveat in the other direction: 379 of 474 proficiency rows are outcome_prior with zero direct samples, including every OpenRouter row. The band is too wide for the rows that are measured and arguably right for the ones that are not, which is an argument for making the band depend on sample depth rather than picking a new constant.

8. Structural verification is a 99% no-op on this workload

kind verdict n
structural unverifiable 28,958
structural malformed 218
structural ok 64
structural truncated 23
local_llm ok 480
local_llm malformed 47
client_outcome succeeded 1,840
client_outcome failed 479

98.96% of structural verdicts are unverifiable, and that is correct behavior: agent turns end in tool calls, and has_tool_calls short-circuits both checkers, which is the documented fix to a real false-failure incident. The conclusion not drawn is that after that fix the free structural checker checks nothing on the only workload this router serves. 305 of 29,263 verdicts were substantive. All of the subsystem's value is in client_outcome, the channel added last — and 527 local_llm verdicts have never been applied to anything.

9. The router has no concept of what a model is, only numbers on an id

On 2026-09-06, before the allowlist existed, the ranker selected google/lyria-3-clip-preview — a music generation model — 38 times for coding, refactoring, debugging and summarization, 17 of them as the ranked winner rather than an exploration gamble. No energy_observations rows exist for any: all 38 failed at dispatch.

The allowlist fixed the symptom the same day and is the right pragmatic gate. But it is manual curation standing in for a structural absence: every filter in the system is a threshold on a metric, and none asks whether a candidate is a chat model. A new catalog row arrives with no proficiency — deliberately treated as unproven rather than bad — and is immediately eligible to win.


Paradigms being shoehorned

A. "Classify, then dispatch" is a single-request frame on a conversation

The unit of work is a session: hundreds to thousands of turns sharing a monotonically growing prefix, ~92% of which is cached. The router models each turn as an independent classify-and-route problem, then bolts on a session cache to make that affordable — and the cache is now 96.6% of the answer.

Inverting the frame dissolves several items at once. If the session is the unit and a turn is a delta: classification becomes "has the task changed?", cheap and usually no, instead of a ~1.1s full-taxonomy inference; the cache stops being an optimization and becomes the model, making its staleness window a real parameter rather than a cost hack; switching models becomes a visibly expensive act (item 4); and min_tool_proficiency becomes expressible, because "can this model be trusted with tools" is a session-level property — which is how the docs already describe it.

B. Quality as a per-(model, category) scalar is a leaderboard paradigm

One number per model per category, inherited from benchmarks. The project has already found twice that this cannot hold what it learns.

deepseek-v4-flash scored 1.00 on all three coding categories and 0.33 on tool_use_agentic, and the documented reason is a failure mode: given both times in "it is 1:20pm and my meeting is at 3pm", it calls two tools instead of subtracting. That is not a lower skill level on a category; it is a disposition that is hazardous whenever tools are on the table, whatever the task is.

Flattening a mode into a category score forced a second mechanism to express the real rule — routing.min_tool_proficiency, a hard filter — and that filter is null because the flattening made it unusable: opencode sends tools on nearly every request, so switching it on excludes the cheap model from all traffic. The paradigm mismatch is what left a known hazard ungated.

The inheritance structure shows the same strain from the other side: 379 rows are peer-rate priors carrying a number with zero evidence behind it, and nothing downstream distinguishes a 0.749 that was measured from a 0.749 that was copied — except source, which the ranker does not read.

C. Energy and carbon are structurally central and functionally vestigial

eco is "not an objective." max_energy_per_request is null. plan_kwh_per_period gates nothing. The project's own conclusion is that billed energy "ranks how busy the provider was, not how efficient the model is."

Around that sit: two columns and a comment block in energy_observations, a BTU column kept "purely for comedic dashboard value", seed_energy.py, a 6-hourly timer whose accumulation is the stated prerequisite for trusting the eco ordering, seed_local_dispatch_energy.py, grid-intensity and carbon_source logging, open follow-up item 3, and a large share of a 71KB CLAUDE.md.

The measurement work was worth doing — the per-kWh billing discovery and the attribution decomposition are the most valuable findings in the project, and item 3 above is an argument for using them more. But they belong to the cost model now, not to an eco objective nothing optimizes. And the provider now carrying 73% of traffic reports no energy at all, which makes the framing actively misleading about what the router can see.


Scale check

Python 55,711 lines; dispatcher.py 4,974, admin.py 2,498, metrics.py 2,474
tests 1,155
config 129 leaf keys across 22 sections
CLAUDE.md 71 KB
plans/ 62 documents
purpose route one operator's coding agent

The engineering quality is high in ways that are rare and worth keeping: the schema-drift and warnings tripwires, the inherited_from migration that fails safe on NULL, the refusal to fabricate leaderboard priors, the four recorded harness bugs that scored the rig rather than the model. None of that is the problem. The problem is the absence of any forcing function that retires machinery, or that re-fires a premise — 62 plan documents and no mechanism that closes one out.


What it needs, ranked

  1. A feedback.py timer. One oneshot service and one timer, matching the four that already exist. The ground-truth loop is the only one cranked by hand, proficiency is three days stale, and the 95 unapplied outcomes are exactly the models carrying today's traffic. Cheapest item here by a wide margin.

  2. Capture OpenRouter cost and latency telemetry. 95.7% of its rows have NULL cost_usd, 100% have NULL duration_seconds, and it is 73% of current decisions. This is a prerequisite for items 3 and 5, not a parallel task.

  3. Calibrate estimated_cost against the 31,119 billed rows already in the DB. A per-model correction factor over a trailing window. Fixes the 21x misordering without a sweep and without abandoning per-request shape sensitivity.

  4. Price the switch. Charge challengers a cold-cache rate in the tiebreak. ~17% of spend sits on 7% of turns, and it makes the router's own decision an input to its own cost model.

  5. Put latency in the objective. A p50/p95 per model and a latency term or ceiling for INTERACTIVE requests. Needs item 2 for most of the catalog.

  6. Make quality_tolerance depend on sample depth rather than replacing one constant with another. 95 of 474 rows are measured, 379 are inherited priors; a band that is right for one is wrong for the other.

  7. An automatic response to a cost shock. The 09-06 spike was handled well and handled by a human. The minimum honest version: a period spend ceiling that narrows the candidate set as it is approached, then refuses with a clear reason. Not "the router is unusable without it" — it demonstrably is not — but it is the difference between a tool that reports a problem and a tool that responds to one.

  8. Decide what the classifier is for, given 96.6% cached. Either accept it as session-level and simplify the machinery to match, or make per-turn classification cheap enough to actually run. The current state pays the architecture cost of per-turn classification and gets session-level labels.

  9. Retire the eco objective explicitly. Move the energy and attribution findings into the cost model where they are load-bearing; archive the rest — the seed timer, the eco scoring path, the BTU column, open item 3.

  10. A premise-expiry mechanism. The thesis. Every setting whose comment says "revisit when X" should have X as a check that fails loudly — a test, a /metrics warning, a line in the poller. Four have expired silently.

Items 1-5 are independent and small, and 1 and 2 should land first because everything downstream reads what they produce. Items 6, 8, 9 and 10 are decisions rather than tasks, and should be made before more machinery is added to the subsystems they touch.