Files
6krrt/plans/ten-thousand-foot-review.md
adlee-was-taken 0111da9bbf feat(telemetry): router-observed wall-clock and TTFT columns, with plan docs
Adds router_wall_seconds and router_ttft_seconds to energy_observations:

- router_wall_seconds: time.monotonic() from just before the accepted
  candidate's POST/connection-open to the complete response body (buffered)
  or the last byte forwarded (streaming). Re-marked per candidate so
  failover time is excluded — a dead model's 30s stall is not charged to
  the healthy one that replaced it.
- router_ttft_seconds: streaming-only. First delta carrying non-empty
  content or a tool_calls fragment, excluding the role-only opening delta.
  NULL on buffered rows (not applicable) and on streams that produced no
  output token (a broken upstream — not the same as a literal 0).

Separate from duration_seconds (provider's reported serving time) on
purpose: OpenRouter reports duration_seconds on 0 of 3,850 rows, while
the router can always measure its own clock. The two quantities are stored
independently and never written into each other.

log_observation() defaults both new params to None so seed_energy.py and
eval_proficiency.py pass unchanged — a reprise of 6e729ad's bug where new
keyword-only arguments killed the seed timer.

Also:
- docs/data-model.md: document the new columns, span definitions, and
  the distinction from duration_seconds
- config/schema.sql: full CREATE TABLE declaration
- plans/token-waste-waves.md: status update for Wave 1
- plans/ten-thousand-foot-review.md: companion diagnosis
2026-09-13 12:25:39 -04:00

454 lines
22 KiB
Markdown

# A 10,000ft review: what this router needs to be a more useful tool
Status: reference -- a diagnosis, not a queue item; the executable work it
feeds is scoped in `plans/token-waste-waves.md`.
**Measured on the LIVE database** at `/home/alee/Sources/6krrt/router.db`
(read-only), 2026-09-13, at repo state `26398d7`. An earlier draft of this file
measured a stale 11:47 copy of `router.db` sitting inside the
`profiles-header-fix` worktree and drew a wrong headline from it; that copy is
~10 hours behind and missing the traffic that matters most. If you are
measuring from a worktree, check the mtime first.
Audience: the operator, and whoever scopes the next round of work. Deliberately
does not repeat the item-level work already scoped in
`code-analysis-refactoring-opportunities.md`.
---
## The headline: the router works, and it just proved it
On 2026-09-06 NeuralWatt billed **$19.85 in one day** across 4,129 decisions.
What followed was the most consequential routing change in the project's
history: OpenRouter came back as an allowlist-gated provider and traffic moved.
| day | decisions | billed |
|---|---|---|
| 2026-09-06 | 4,129 | **$19.85** |
| 2026-09-09 | 594 | $0.32 |
| 2026-09-10 | 722 | $0.25 |
| 2026-09-11 | 755 | $0.25 |
| 2026-09-12 | 116 | $0.40 |
| 2026-09-13 | 154 | $0.21 |
Volume held at 600-750 decisions a day. Billed spend fell roughly **50x**.
That is the system doing precisely what it exists to do, and it is the
strongest evidence in the database that the thing is useful.
Two qualifications, and they are where the rest of this document lives:
- **The human did the adaptation, not the router.** The migration was an
operator decision — add a provider, write an allowlist, let ranking follow.
The router has no automatic response to a cost shock; `plan_kwh_per_period`
gates nothing and `max_energy_per_request` is null, so every cost mechanism
in the system is advisory. It reported the spike well. It did nothing about
it.
- **The migration moved 73% of traffic onto a provider the router cannot
measure.** That is the single most actionable finding here, and it is new
since the migration.
## The thesis
This project is unusually good at recording *why* a decision was made, and has
no mechanism for noticing when a recorded reason has expired.
| setting | the comment's own premise | what is true now |
|---|---|---|
| `quality_tolerance: 0.1` | "scores currently rest on 2-3 samples per category, so a 0.05 gap is indistinguishable from sampling variation... **Narrow it as samples accumulate**" | 46 rows per category, all outcome-backed; `outcome_blended` rows average 51.8 client samples, max 460 |
| cost-as-tiebreak, not an objective | "60% of every decision adjudicated fractions of a cent (**all real traffic to date totals $0.07**)" | **$118.79** billed. 1,697x the figure the whole argument rests on |
| `assumed_cache_rate: 0.917` | measured once, 2026-08-23, over 40.7M tokens | `cached_prompt_tokens` is NULL in **all 35,086 rows**. The router receives this field and drops it |
| cost feedback generally | catalog estimates "track the real ordering" because billing is capped at 3x list | true for NeuralWatt, untestable for OpenRouter: **95.7% of its rows have NULL `cost_usd`** |
Four settings, four expiry conditions met, nothing re-fired. That pattern is
worth fixing before any individual number. The documentation is a real
strength — the reasoning is preserved, measurements are dated, caveats are
honest. What is missing is anything that *acts* when a premise the code depends
on stops holding.
---
## 1. The provider that now carries the traffic is the one the router cannot see
Recent picks, since 2026-09-09 (2,341 decisions):
| provider | model | n |
|---|---|---|
| openrouter | `xiaomi/mimo-v2.5` | 744 |
| openrouter | `deepseek/deepseek-v4-flash` | 659 |
| neuralwatt | `deepseek-v4-flash` | 266 |
| neuralwatt | `qwen3.6-35b` | 247 |
| openrouter | `nvidia/nemotron-3-nano-omni-...:free` | 178 |
| neuralwatt | `qwen3.6-35b-fast` | 119 |
| openrouter | `z-ai/glm-5.3-flash` | 95 |
| openrouter | `xiaomi/mimo-v2.5-pro` | 26 |
**OpenRouter is 72.8% of current decisions.** Telemetry coverage in
`energy_observations`:
| provider | rows | NULL `cost_usd` | NULL `duration_seconds` |
|---|---|---|---|
| neuralwatt | 31,188 | 66 | 66 |
| **openrouter** | **3,838** | **3,673 (95.7%)** | **3,838 (100%)** |
| ollama-local | 60 | 60 | 60 |
So on the majority of current traffic the router has no billed cost to
calibrate its estimator against and no completion time at all. Every feedback
loop that makes this system smarter than a static config is blind on the
provider it now prefers.
This inverts the priority of everything below: recalibrating the cost model
(item 3) cannot work for 73% of traffic until this is fixed. OpenRouter returns
usage and generation-cost data; it is an ingestion gap, not a provider
limitation.
## 2. `feedback.py` is the only loop without a timer
`proficiency` was written in a single batch at **2026-09-10T00:44:55** — every
row, one instant, three days ago. Since then:
| | |
|---|---|
| client outcomes applied | 2,224 |
| client outcomes **unapplied** | **95** |
And the unapplied 95 are exactly the current traffic:
| model | provider | unapplied | of which failed |
|---|---|---|---|
| `xiaomi/mimo-v2.5` | openrouter | 41 | 9 |
| `deepseek/deepseek-v4-flash` | openrouter | 36 | 15 |
| `deepseek-v4-flash` | neuralwatt | 17 | 3 |
| `xiaomi/mimo-v2.5-pro` | openrouter | 1 | 0 |
Meanwhile **every OpenRouter model's proficiency row is `outcome_prior` with
`outcome_samples = 0`** — an inherited peer-rate prior, not a measurement.
`xiaomi/mimo-v2.5` is the most-routed model in the catalog right now (744
picks) and scores 0.749 / 0.770 / 0.894 on coding_general / coding_refactor /
tool_use_agentic entirely on inheritance. Its 41 real outcomes (78% pass) are
sitting in `verifications` unapplied.
`deploy/` ships timers for the poller, the seed sweep, backup and offsite sync,
and `~/.config/systemd/user` additionally has a baseline-report timer. There is
**no feedback timer**. The loop the docs call "the only ground truth" is the
only one that has to be cranked by hand.
This is the cheapest high-value fix in the document: a oneshot service and a
timer, matching the four that already exist.
## 3. The cost estimator is wrong in a model-dependent way
`routing.estimated_cost` breaks every tie, and the ranking reduces to it
whenever candidates land in the same quality band. Per request, on rows with a
real billed figure:
| model | provider | est µ$ | billed µ$ | est/billed |
|---|---|---|---|---|
| gemma-4-31b | neuralwatt | 3,822 | 6,041 | **0.63x** |
| kimi-k2.7-code-fast | neuralwatt | 28,623 | 21,747 | 1.32x |
| kimi-k3 | neuralwatt | 87,910 | 42,388 | 2.07x |
| deepseek-v4-flash | neuralwatt | 4,154 | 1,796 | 2.31x |
| glm-5.3 | neuralwatt | 48,232 | 20,056 | 2.40x |
| `deepseek/deepseek-v4-flash` | openrouter | 3,830 | 1,383 | 2.77x |
| kimi-k2.7-code | neuralwatt | 22,253 | 3,906 | 5.70x |
| qwen3.6-35b | neuralwatt | 4,091 | 484 | 8.46x |
| `qwen3.6-35b-fast` | neuralwatt | 3,773 | 282 | 13.37x |
| `xiaomi/mimo-v2.5` | openrouter | 9,689 | 715 | **13.55x** |
Total: $401.99 estimated against $118.79 billed. The scale error does not
matter; the **21x spread in the error** does, because it reorders candidates.
`gemma-4-31b` is the estimator's cheapest model in the catalog and billing's
most expensive per request of the ten. The router's current favourite,
`xiaomi/mimo-v2.5`, is the one it overestimates most — it is winning on
quality *despite* the cost model, not because of it.
The cause is already written down in `CLAUDE.md`, two sections apart, never
reconciled: attribution ratio spans **750x between models** and is "most of
the real cost difference in the catalog", while `estimated_cost` prices from
catalog token prices, which contain **no attribution term at all**.
Cost moved off measurement to fix a real problem — a 400-token sweep cannot
price a 150k-token workload — and in the move discarded the term the project
had already proven was dominant. Third recurrence of one pattern: list price
ranks models backwards; the reference sweep ranks them backwards for real
traffic; catalog pricing misorders them because it omits pool concurrency.
**The fix needs no sweep.** 31,119 NeuralWatt rows carry a real billed figure,
joinable to decisions by `request_id`. A per-model correction factor over a
trailing window, refreshed by the poller, recalibrates against the bill and
keeps the per-request shape sensitivity that motivated the move. For OpenRouter
it needs item 1 first.
## 4. Switching models mid-session costs 2.5x, and the estimator prices it as free
No sticky routing, no concept of the previously-used model. On turns with
prompts over 20k tokens and a real billed figure:
| turn | n | avg prompt | µ$ / prompt token | spend |
|---|---|---|---|---|
| same model as previous | 11,246 | 100,994 | 0.037 | $47.91 |
| **switched model** | 858 | 90,309 | **0.091** | **$9.68** |
2.46x, on *smaller* prompts, so the effect is understated. Switch turns are
7.1% of turns and **16.8% of joined spend**.
The mechanism is not mysterious: a switch lands a ~100k-token prompt on a
provider that has never seen it, so the prefix cache is cold and ~92% of the
prompt is suddenly billed fresh. But `rank_candidates` prices every candidate
at `assumed_cache_rate: 0.917`, including the one it is about to switch to,
where the true rate is 0.
Pricing the incumbent at the assumed rate and every challenger at 0 is a small
change that pays for itself, and it makes the router's own decision an input to
its own cost model for the first time.
## 5. Latency is the axis the operator feels and the one nothing optimizes
NeuralWatt, where it is recorded at all:
| model | avg s | max s | tok/s |
|---|---|---|---|
| gemma-4-31b | **22.49** | **1,251.1** | 22.6 |
| glm-5.3 | 7.45 | 163.7 | 143.0 |
| qwen3.6-35b | 4.63 | 172.5 | 99.4 |
| kimi-k2.7-code | 3.65 | 226.9 | 93.5 |
| deepseek-v4-flash-flex | 1.90 | 14.8 | 210.5 |
An 11.8x spread in the mean and a 21-minute worst case. `latency_tolerance`
exists but only gates `-flex` rows; `duration_seconds` is read by nothing in
routing. And per item 1 it is **not collected at all** for OpenRouter or
`ollama-local` — so for 73% of current traffic there is not even a number to
ignore.
`gemma-4-31b` is where every signal disagrees at once: the estimator's
cheapest, billing's most expensive per request, and 6x slower than the field.
## 6. 96.6% of requests are never classified
The classifier is the conceptual center of the system — three backends, a
four-step failure cascade, its own circuit breaker, a degraded-share warning in
`/metrics`, an attribution-exclusion rule in `report_outcome`, and a
fine-tuning roadmap. Last five days:
| `classification_source` | share |
|---|---|
| `cached` | **96.6%** |
| `classifier` | 3.2% (avg 1,129 ms) |
| `override` / null | 0.3% |
The consequence is not only that the machinery is oversized. `task_category` on
a given turn is **whatever the session was doing at the last cache refresh**,
not a classification of that turn — one session ran 3,724 turns and cycled
through all nine categories. So ~96% of turns carry a borrowed label, and
`POST /outcome` attributes results to `(model, task_category)`. The quality
table is largely trained on labels computed for a different turn, and nothing
measures that noise.
## 7. The feedback loop worked, and the tolerance band was never narrowed
`CLAUDE.md` is stale here in the good direction. The categories it records as
"flat at 1.00 awaiting real traffic" are nothing of the kind:
| category | rows | spread | outcome-backed |
|---|---|---|---|
| coding_general | 46 | 0.375 - 0.843 | 46/46 |
| coding_refactor | 46 | 0.423 - 0.846 | 46/46 |
| debugging | 46 | 0.307 - 0.836 | 46/46 |
| tool_use_agentic | 46 | 0.298 - 0.918 | 46/46 |
| reasoning_math | 46 | 0.000 - 0.956 | 46/46 |
| summarization | 46 | 0.250 - 0.556 | 46/46 |
| docs_writing | 46 | 0.607 - 0.949 | 46/46 |
2,319 client outcomes have arrived. The design bet — a 43-task benchmark is the
prior, real traffic is the posterior — paid off comprehensively. Note also that
scores came *down* as evidence thickened (`coding_general` topped out at 0.947
three days ago, 0.843 now) and that `summarization` now caps at 0.556: nothing
in the catalog is good at it, which no benchmark had revealed.
`quality_tolerance` is still 0.1, on a comment that says to narrow it as
samples accumulate. Two things compound: the top band is 0.1 wide, so models
differing by up to 10 points of **outcome-backed pass rate** are treated as
equal; and the tiebreak inside that band is the estimator from item 3, wrong by
up to 13.6x and misordered by 21x. The system trades up to 10 points of
measured quality for a cost saving it cannot compute correctly.
Caveat in the other direction: 379 of 474 proficiency rows are `outcome_prior`
with zero direct samples, including every OpenRouter row. The band is too wide
for the rows that *are* measured and arguably right for the ones that are not,
which is an argument for making the band depend on sample depth rather than
picking a new constant.
## 8. Structural verification is a 99% no-op on this workload
| kind | verdict | n |
|---|---|---|
| structural | **unverifiable** | **28,958** |
| structural | malformed | 218 |
| structural | ok | 64 |
| structural | truncated | 23 |
| local_llm | ok | 480 |
| local_llm | malformed | 47 |
| client_outcome | succeeded | 1,840 |
| client_outcome | failed | 479 |
98.96% of structural verdicts are `unverifiable`, and that is *correct*
behavior: agent turns end in tool calls, and `has_tool_calls` short-circuits
both checkers, which is the documented fix to a real false-failure incident.
The conclusion not drawn is that after that fix the free structural checker
checks nothing on the only workload this router serves. 305 of 29,263 verdicts
were substantive. All of the subsystem's value is in `client_outcome`, the
channel added last — and 527 `local_llm` verdicts have never been applied to
anything.
## 9. The router has no concept of what a model is, only numbers on an id
On 2026-09-06, before the allowlist existed, the ranker selected
`google/lyria-3-clip-preview` — a music generation model — 38 times for coding,
refactoring, debugging and summarization, **17 of them as the ranked winner
rather than an exploration gamble**. No `energy_observations` rows exist for
any: all 38 failed at dispatch.
The allowlist fixed the symptom the same day and is the right pragmatic gate.
But it is manual curation standing in for a structural absence: every filter in
the system is a threshold on a metric, and none asks whether a candidate is a
chat model. A new catalog row arrives with no proficiency — deliberately
treated as unproven rather than bad — and is immediately eligible to win.
---
## Paradigms being shoehorned
### A. "Classify, then dispatch" is a single-request frame on a conversation
The unit of work is a session: hundreds to thousands of turns sharing a
monotonically growing prefix, ~92% of which is cached. The router models each
turn as an independent classify-and-route problem, then bolts on a session
cache to make that affordable — and the cache is now 96.6% of the answer.
Inverting the frame dissolves several items at once. If the session is the unit
and a turn is a delta: classification becomes "has the task changed?", cheap
and usually no, instead of a ~1.1s full-taxonomy inference; the cache stops
being an optimization and becomes the model, making its staleness window a real
parameter rather than a cost hack; switching models becomes a visibly expensive
act (item 4); and `min_tool_proficiency` becomes expressible, because "can this
model be trusted with tools" is a session-level property — which is how the
docs already describe it.
### B. Quality as a per-(model, category) scalar is a leaderboard paradigm
One number per model per category, inherited from benchmarks. The project has
already found twice that this cannot hold what it learns.
`deepseek-v4-flash` scored 1.00 on all three coding categories and 0.33 on
`tool_use_agentic`, and the documented reason is a *failure mode*: given both
times in "it is 1:20pm and my meeting is at 3pm", it calls two tools instead of
subtracting. That is not a lower skill level on a category; it is a disposition
that is hazardous whenever tools are on the table, whatever the task is.
Flattening a mode into a category score forced a second mechanism to express
the real rule — `routing.min_tool_proficiency`, a hard filter — and that filter
is `null` because the flattening made it unusable: opencode sends `tools` on
nearly every request, so switching it on excludes the cheap model from all
traffic. The paradigm mismatch is what left a known hazard ungated.
The inheritance structure shows the same strain from the other side: 379 rows
are peer-rate priors carrying a number with zero evidence behind it, and
nothing downstream distinguishes a 0.749 that was measured from a 0.749 that
was copied — except `source`, which the ranker does not read.
### C. Energy and carbon are structurally central and functionally vestigial
`eco` is "not an objective." `max_energy_per_request` is null.
`plan_kwh_per_period` gates nothing. The project's own conclusion is that
billed energy "ranks how busy the provider was, not how efficient the model
is."
Around that sit: two columns and a comment block in `energy_observations`, a
BTU column kept "purely for comedic dashboard value", `seed_energy.py`, a
6-hourly timer whose accumulation is the stated prerequisite for trusting the
eco ordering, `seed_local_dispatch_energy.py`, grid-intensity and
`carbon_source` logging, open follow-up item 3, and a large share of a 71KB
`CLAUDE.md`.
The measurement work was worth doing — the per-kWh billing discovery and the
attribution decomposition are the most valuable findings in the project, and
item 3 above is an argument for using them *more*. But they belong to the cost
model now, not to an eco objective nothing optimizes. And the provider now
carrying 73% of traffic reports no energy at all, which makes the framing
actively misleading about what the router can see.
---
## Scale check
| | |
|---|---|
| Python | 55,711 lines; `dispatcher.py` 4,974, `admin.py` 2,498, `metrics.py` 2,474 |
| tests | 1,155 |
| config | 129 leaf keys across 22 sections |
| `CLAUDE.md` | 71 KB |
| `plans/` | 62 documents |
| purpose | route one operator's coding agent |
The engineering quality is high in ways that are rare and worth keeping: the
schema-drift and warnings tripwires, the `inherited_from` migration that fails
safe on NULL, the refusal to fabricate leaderboard priors, the four recorded
harness bugs that scored the rig rather than the model. None of that is the
problem. The problem is the absence of any forcing function that retires
machinery, or that re-fires a premise — 62 plan documents and no mechanism that
closes one out.
---
## What it needs, ranked
1. **A `feedback.py` timer.** One oneshot service and one timer, matching the
four that already exist. The ground-truth loop is the only one cranked by
hand, `proficiency` is three days stale, and the 95 unapplied outcomes are
exactly the models carrying today's traffic. Cheapest item here by a wide
margin.
2. **Capture OpenRouter cost and latency telemetry.** 95.7% of its rows have
NULL `cost_usd`, 100% have NULL `duration_seconds`, and it is 73% of current
decisions. This is a prerequisite for items 3 and 5, not a parallel task.
3. **Calibrate `estimated_cost` against the 31,119 billed rows already in the
DB.** A per-model correction factor over a trailing window. Fixes the 21x
misordering without a sweep and without abandoning per-request shape
sensitivity.
4. **Price the switch.** Charge challengers a cold-cache rate in the tiebreak.
~17% of spend sits on 7% of turns, and it makes the router's own decision an
input to its own cost model.
5. **Put latency in the objective.** A p50/p95 per model and a latency term or
ceiling for `INTERACTIVE` requests. Needs item 2 for most of the catalog.
6. **Make `quality_tolerance` depend on sample depth** rather than replacing one
constant with another. 95 of 474 rows are measured, 379 are inherited
priors; a band that is right for one is wrong for the other.
7. **An automatic response to a cost shock.** The 09-06 spike was handled well
and handled by a human. The minimum honest version: a period spend ceiling
that narrows the candidate set as it is approached, then refuses with a clear
reason. Not "the router is unusable without it" — it demonstrably is not —
but it is the difference between a tool that reports a problem and a tool
that responds to one.
8. **Decide what the classifier is for, given 96.6% cached.** Either accept it
as session-level and simplify the machinery to match, or make per-turn
classification cheap enough to actually run. The current state pays the
architecture cost of per-turn classification and gets session-level labels.
9. **Retire the eco objective explicitly.** Move the energy and attribution
findings into the cost model where they are load-bearing; archive the rest —
the seed timer, the eco scoring path, the BTU column, open item 3.
10. **A premise-expiry mechanism.** The thesis. Every setting whose comment says
"revisit when X" should have X as a check that fails loudly — a test, a
`/metrics` warning, a line in the poller. Four have expired silently.
Items 1-5 are independent and small, and 1 and 2 should land first because
everything downstream reads what they produce. Items 6, 8, 9 and 10 are
decisions rather than tasks, and should be made before more machinery is added
to the subsystems they touch.