Adds router_wall_seconds and router_ttft_seconds to energy_observations: - router_wall_seconds: time.monotonic() from just before the accepted candidate's POST/connection-open to the complete response body (buffered) or the last byte forwarded (streaming). Re-marked per candidate so failover time is excluded — a dead model's 30s stall is not charged to the healthy one that replaced it. - router_ttft_seconds: streaming-only. First delta carrying non-empty content or a tool_calls fragment, excluding the role-only opening delta. NULL on buffered rows (not applicable) and on streams that produced no output token (a broken upstream — not the same as a literal 0). Separate from duration_seconds (provider's reported serving time) on purpose: OpenRouter reports duration_seconds on 0 of 3,850 rows, while the router can always measure its own clock. The two quantities are stored independently and never written into each other. log_observation() defaults both new params to None so seed_energy.py and eval_proficiency.py pass unchanged — a reprise of 6e729ad's bug where new keyword-only arguments killed the seed timer. Also: - docs/data-model.md: document the new columns, span definitions, and the distinction from duration_seconds - config/schema.sql: full CREATE TABLE declaration - plans/token-waste-waves.md: status update for Wave 1 - plans/ten-thousand-foot-review.md: companion diagnosis
454 lines
22 KiB
Markdown
454 lines
22 KiB
Markdown
# A 10,000ft review: what this router needs to be a more useful tool
|
|
|
|
Status: reference -- a diagnosis, not a queue item; the executable work it
|
|
feeds is scoped in `plans/token-waste-waves.md`.
|
|
|
|
**Measured on the LIVE database** at `/home/alee/Sources/6krrt/router.db`
|
|
(read-only), 2026-09-13, at repo state `26398d7`. An earlier draft of this file
|
|
measured a stale 11:47 copy of `router.db` sitting inside the
|
|
`profiles-header-fix` worktree and drew a wrong headline from it; that copy is
|
|
~10 hours behind and missing the traffic that matters most. If you are
|
|
measuring from a worktree, check the mtime first.
|
|
|
|
Audience: the operator, and whoever scopes the next round of work. Deliberately
|
|
does not repeat the item-level work already scoped in
|
|
`code-analysis-refactoring-opportunities.md`.
|
|
|
|
---
|
|
|
|
## The headline: the router works, and it just proved it
|
|
|
|
On 2026-09-06 NeuralWatt billed **$19.85 in one day** across 4,129 decisions.
|
|
What followed was the most consequential routing change in the project's
|
|
history: OpenRouter came back as an allowlist-gated provider and traffic moved.
|
|
|
|
| day | decisions | billed |
|
|
|---|---|---|
|
|
| 2026-09-06 | 4,129 | **$19.85** |
|
|
| 2026-09-09 | 594 | $0.32 |
|
|
| 2026-09-10 | 722 | $0.25 |
|
|
| 2026-09-11 | 755 | $0.25 |
|
|
| 2026-09-12 | 116 | $0.40 |
|
|
| 2026-09-13 | 154 | $0.21 |
|
|
|
|
Volume held at 600-750 decisions a day. Billed spend fell roughly **50x**.
|
|
That is the system doing precisely what it exists to do, and it is the
|
|
strongest evidence in the database that the thing is useful.
|
|
|
|
Two qualifications, and they are where the rest of this document lives:
|
|
|
|
- **The human did the adaptation, not the router.** The migration was an
|
|
operator decision — add a provider, write an allowlist, let ranking follow.
|
|
The router has no automatic response to a cost shock; `plan_kwh_per_period`
|
|
gates nothing and `max_energy_per_request` is null, so every cost mechanism
|
|
in the system is advisory. It reported the spike well. It did nothing about
|
|
it.
|
|
- **The migration moved 73% of traffic onto a provider the router cannot
|
|
measure.** That is the single most actionable finding here, and it is new
|
|
since the migration.
|
|
|
|
## The thesis
|
|
|
|
This project is unusually good at recording *why* a decision was made, and has
|
|
no mechanism for noticing when a recorded reason has expired.
|
|
|
|
| setting | the comment's own premise | what is true now |
|
|
|---|---|---|
|
|
| `quality_tolerance: 0.1` | "scores currently rest on 2-3 samples per category, so a 0.05 gap is indistinguishable from sampling variation... **Narrow it as samples accumulate**" | 46 rows per category, all outcome-backed; `outcome_blended` rows average 51.8 client samples, max 460 |
|
|
| cost-as-tiebreak, not an objective | "60% of every decision adjudicated fractions of a cent (**all real traffic to date totals $0.07**)" | **$118.79** billed. 1,697x the figure the whole argument rests on |
|
|
| `assumed_cache_rate: 0.917` | measured once, 2026-08-23, over 40.7M tokens | `cached_prompt_tokens` is NULL in **all 35,086 rows**. The router receives this field and drops it |
|
|
| cost feedback generally | catalog estimates "track the real ordering" because billing is capped at 3x list | true for NeuralWatt, untestable for OpenRouter: **95.7% of its rows have NULL `cost_usd`** |
|
|
|
|
Four settings, four expiry conditions met, nothing re-fired. That pattern is
|
|
worth fixing before any individual number. The documentation is a real
|
|
strength — the reasoning is preserved, measurements are dated, caveats are
|
|
honest. What is missing is anything that *acts* when a premise the code depends
|
|
on stops holding.
|
|
|
|
---
|
|
|
|
## 1. The provider that now carries the traffic is the one the router cannot see
|
|
|
|
Recent picks, since 2026-09-09 (2,341 decisions):
|
|
|
|
| provider | model | n |
|
|
|---|---|---|
|
|
| openrouter | `xiaomi/mimo-v2.5` | 744 |
|
|
| openrouter | `deepseek/deepseek-v4-flash` | 659 |
|
|
| neuralwatt | `deepseek-v4-flash` | 266 |
|
|
| neuralwatt | `qwen3.6-35b` | 247 |
|
|
| openrouter | `nvidia/nemotron-3-nano-omni-...:free` | 178 |
|
|
| neuralwatt | `qwen3.6-35b-fast` | 119 |
|
|
| openrouter | `z-ai/glm-5.3-flash` | 95 |
|
|
| openrouter | `xiaomi/mimo-v2.5-pro` | 26 |
|
|
|
|
**OpenRouter is 72.8% of current decisions.** Telemetry coverage in
|
|
`energy_observations`:
|
|
|
|
| provider | rows | NULL `cost_usd` | NULL `duration_seconds` |
|
|
|---|---|---|---|
|
|
| neuralwatt | 31,188 | 66 | 66 |
|
|
| **openrouter** | **3,838** | **3,673 (95.7%)** | **3,838 (100%)** |
|
|
| ollama-local | 60 | 60 | 60 |
|
|
|
|
So on the majority of current traffic the router has no billed cost to
|
|
calibrate its estimator against and no completion time at all. Every feedback
|
|
loop that makes this system smarter than a static config is blind on the
|
|
provider it now prefers.
|
|
|
|
This inverts the priority of everything below: recalibrating the cost model
|
|
(item 3) cannot work for 73% of traffic until this is fixed. OpenRouter returns
|
|
usage and generation-cost data; it is an ingestion gap, not a provider
|
|
limitation.
|
|
|
|
## 2. `feedback.py` is the only loop without a timer
|
|
|
|
`proficiency` was written in a single batch at **2026-09-10T00:44:55** — every
|
|
row, one instant, three days ago. Since then:
|
|
|
|
| | |
|
|
|---|---|
|
|
| client outcomes applied | 2,224 |
|
|
| client outcomes **unapplied** | **95** |
|
|
|
|
And the unapplied 95 are exactly the current traffic:
|
|
|
|
| model | provider | unapplied | of which failed |
|
|
|---|---|---|---|
|
|
| `xiaomi/mimo-v2.5` | openrouter | 41 | 9 |
|
|
| `deepseek/deepseek-v4-flash` | openrouter | 36 | 15 |
|
|
| `deepseek-v4-flash` | neuralwatt | 17 | 3 |
|
|
| `xiaomi/mimo-v2.5-pro` | openrouter | 1 | 0 |
|
|
|
|
Meanwhile **every OpenRouter model's proficiency row is `outcome_prior` with
|
|
`outcome_samples = 0`** — an inherited peer-rate prior, not a measurement.
|
|
`xiaomi/mimo-v2.5` is the most-routed model in the catalog right now (744
|
|
picks) and scores 0.749 / 0.770 / 0.894 on coding_general / coding_refactor /
|
|
tool_use_agentic entirely on inheritance. Its 41 real outcomes (78% pass) are
|
|
sitting in `verifications` unapplied.
|
|
|
|
`deploy/` ships timers for the poller, the seed sweep, backup and offsite sync,
|
|
and `~/.config/systemd/user` additionally has a baseline-report timer. There is
|
|
**no feedback timer**. The loop the docs call "the only ground truth" is the
|
|
only one that has to be cranked by hand.
|
|
|
|
This is the cheapest high-value fix in the document: a oneshot service and a
|
|
timer, matching the four that already exist.
|
|
|
|
## 3. The cost estimator is wrong in a model-dependent way
|
|
|
|
`routing.estimated_cost` breaks every tie, and the ranking reduces to it
|
|
whenever candidates land in the same quality band. Per request, on rows with a
|
|
real billed figure:
|
|
|
|
| model | provider | est µ$ | billed µ$ | est/billed |
|
|
|---|---|---|---|---|
|
|
| gemma-4-31b | neuralwatt | 3,822 | 6,041 | **0.63x** |
|
|
| kimi-k2.7-code-fast | neuralwatt | 28,623 | 21,747 | 1.32x |
|
|
| kimi-k3 | neuralwatt | 87,910 | 42,388 | 2.07x |
|
|
| deepseek-v4-flash | neuralwatt | 4,154 | 1,796 | 2.31x |
|
|
| glm-5.3 | neuralwatt | 48,232 | 20,056 | 2.40x |
|
|
| `deepseek/deepseek-v4-flash` | openrouter | 3,830 | 1,383 | 2.77x |
|
|
| kimi-k2.7-code | neuralwatt | 22,253 | 3,906 | 5.70x |
|
|
| qwen3.6-35b | neuralwatt | 4,091 | 484 | 8.46x |
|
|
| `qwen3.6-35b-fast` | neuralwatt | 3,773 | 282 | 13.37x |
|
|
| `xiaomi/mimo-v2.5` | openrouter | 9,689 | 715 | **13.55x** |
|
|
|
|
Total: $401.99 estimated against $118.79 billed. The scale error does not
|
|
matter; the **21x spread in the error** does, because it reorders candidates.
|
|
`gemma-4-31b` is the estimator's cheapest model in the catalog and billing's
|
|
most expensive per request of the ten. The router's current favourite,
|
|
`xiaomi/mimo-v2.5`, is the one it overestimates most — it is winning on
|
|
quality *despite* the cost model, not because of it.
|
|
|
|
The cause is already written down in `CLAUDE.md`, two sections apart, never
|
|
reconciled: attribution ratio spans **750x between models** and is "most of
|
|
the real cost difference in the catalog", while `estimated_cost` prices from
|
|
catalog token prices, which contain **no attribution term at all**.
|
|
|
|
Cost moved off measurement to fix a real problem — a 400-token sweep cannot
|
|
price a 150k-token workload — and in the move discarded the term the project
|
|
had already proven was dominant. Third recurrence of one pattern: list price
|
|
ranks models backwards; the reference sweep ranks them backwards for real
|
|
traffic; catalog pricing misorders them because it omits pool concurrency.
|
|
|
|
**The fix needs no sweep.** 31,119 NeuralWatt rows carry a real billed figure,
|
|
joinable to decisions by `request_id`. A per-model correction factor over a
|
|
trailing window, refreshed by the poller, recalibrates against the bill and
|
|
keeps the per-request shape sensitivity that motivated the move. For OpenRouter
|
|
it needs item 1 first.
|
|
|
|
## 4. Switching models mid-session costs 2.5x, and the estimator prices it as free
|
|
|
|
No sticky routing, no concept of the previously-used model. On turns with
|
|
prompts over 20k tokens and a real billed figure:
|
|
|
|
| turn | n | avg prompt | µ$ / prompt token | spend |
|
|
|---|---|---|---|---|
|
|
| same model as previous | 11,246 | 100,994 | 0.037 | $47.91 |
|
|
| **switched model** | 858 | 90,309 | **0.091** | **$9.68** |
|
|
|
|
2.46x, on *smaller* prompts, so the effect is understated. Switch turns are
|
|
7.1% of turns and **16.8% of joined spend**.
|
|
|
|
The mechanism is not mysterious: a switch lands a ~100k-token prompt on a
|
|
provider that has never seen it, so the prefix cache is cold and ~92% of the
|
|
prompt is suddenly billed fresh. But `rank_candidates` prices every candidate
|
|
at `assumed_cache_rate: 0.917`, including the one it is about to switch to,
|
|
where the true rate is 0.
|
|
|
|
Pricing the incumbent at the assumed rate and every challenger at 0 is a small
|
|
change that pays for itself, and it makes the router's own decision an input to
|
|
its own cost model for the first time.
|
|
|
|
## 5. Latency is the axis the operator feels and the one nothing optimizes
|
|
|
|
NeuralWatt, where it is recorded at all:
|
|
|
|
| model | avg s | max s | tok/s |
|
|
|---|---|---|---|
|
|
| gemma-4-31b | **22.49** | **1,251.1** | 22.6 |
|
|
| glm-5.3 | 7.45 | 163.7 | 143.0 |
|
|
| qwen3.6-35b | 4.63 | 172.5 | 99.4 |
|
|
| kimi-k2.7-code | 3.65 | 226.9 | 93.5 |
|
|
| deepseek-v4-flash-flex | 1.90 | 14.8 | 210.5 |
|
|
|
|
An 11.8x spread in the mean and a 21-minute worst case. `latency_tolerance`
|
|
exists but only gates `-flex` rows; `duration_seconds` is read by nothing in
|
|
routing. And per item 1 it is **not collected at all** for OpenRouter or
|
|
`ollama-local` — so for 73% of current traffic there is not even a number to
|
|
ignore.
|
|
|
|
`gemma-4-31b` is where every signal disagrees at once: the estimator's
|
|
cheapest, billing's most expensive per request, and 6x slower than the field.
|
|
|
|
## 6. 96.6% of requests are never classified
|
|
|
|
The classifier is the conceptual center of the system — three backends, a
|
|
four-step failure cascade, its own circuit breaker, a degraded-share warning in
|
|
`/metrics`, an attribution-exclusion rule in `report_outcome`, and a
|
|
fine-tuning roadmap. Last five days:
|
|
|
|
| `classification_source` | share |
|
|
|---|---|
|
|
| `cached` | **96.6%** |
|
|
| `classifier` | 3.2% (avg 1,129 ms) |
|
|
| `override` / null | 0.3% |
|
|
|
|
The consequence is not only that the machinery is oversized. `task_category` on
|
|
a given turn is **whatever the session was doing at the last cache refresh**,
|
|
not a classification of that turn — one session ran 3,724 turns and cycled
|
|
through all nine categories. So ~96% of turns carry a borrowed label, and
|
|
`POST /outcome` attributes results to `(model, task_category)`. The quality
|
|
table is largely trained on labels computed for a different turn, and nothing
|
|
measures that noise.
|
|
|
|
## 7. The feedback loop worked, and the tolerance band was never narrowed
|
|
|
|
`CLAUDE.md` is stale here in the good direction. The categories it records as
|
|
"flat at 1.00 awaiting real traffic" are nothing of the kind:
|
|
|
|
| category | rows | spread | outcome-backed |
|
|
|---|---|---|---|
|
|
| coding_general | 46 | 0.375 - 0.843 | 46/46 |
|
|
| coding_refactor | 46 | 0.423 - 0.846 | 46/46 |
|
|
| debugging | 46 | 0.307 - 0.836 | 46/46 |
|
|
| tool_use_agentic | 46 | 0.298 - 0.918 | 46/46 |
|
|
| reasoning_math | 46 | 0.000 - 0.956 | 46/46 |
|
|
| summarization | 46 | 0.250 - 0.556 | 46/46 |
|
|
| docs_writing | 46 | 0.607 - 0.949 | 46/46 |
|
|
|
|
2,319 client outcomes have arrived. The design bet — a 43-task benchmark is the
|
|
prior, real traffic is the posterior — paid off comprehensively. Note also that
|
|
scores came *down* as evidence thickened (`coding_general` topped out at 0.947
|
|
three days ago, 0.843 now) and that `summarization` now caps at 0.556: nothing
|
|
in the catalog is good at it, which no benchmark had revealed.
|
|
|
|
`quality_tolerance` is still 0.1, on a comment that says to narrow it as
|
|
samples accumulate. Two things compound: the top band is 0.1 wide, so models
|
|
differing by up to 10 points of **outcome-backed pass rate** are treated as
|
|
equal; and the tiebreak inside that band is the estimator from item 3, wrong by
|
|
up to 13.6x and misordered by 21x. The system trades up to 10 points of
|
|
measured quality for a cost saving it cannot compute correctly.
|
|
|
|
Caveat in the other direction: 379 of 474 proficiency rows are `outcome_prior`
|
|
with zero direct samples, including every OpenRouter row. The band is too wide
|
|
for the rows that *are* measured and arguably right for the ones that are not,
|
|
which is an argument for making the band depend on sample depth rather than
|
|
picking a new constant.
|
|
|
|
## 8. Structural verification is a 99% no-op on this workload
|
|
|
|
| kind | verdict | n |
|
|
|---|---|---|
|
|
| structural | **unverifiable** | **28,958** |
|
|
| structural | malformed | 218 |
|
|
| structural | ok | 64 |
|
|
| structural | truncated | 23 |
|
|
| local_llm | ok | 480 |
|
|
| local_llm | malformed | 47 |
|
|
| client_outcome | succeeded | 1,840 |
|
|
| client_outcome | failed | 479 |
|
|
|
|
98.96% of structural verdicts are `unverifiable`, and that is *correct*
|
|
behavior: agent turns end in tool calls, and `has_tool_calls` short-circuits
|
|
both checkers, which is the documented fix to a real false-failure incident.
|
|
The conclusion not drawn is that after that fix the free structural checker
|
|
checks nothing on the only workload this router serves. 305 of 29,263 verdicts
|
|
were substantive. All of the subsystem's value is in `client_outcome`, the
|
|
channel added last — and 527 `local_llm` verdicts have never been applied to
|
|
anything.
|
|
|
|
## 9. The router has no concept of what a model is, only numbers on an id
|
|
|
|
On 2026-09-06, before the allowlist existed, the ranker selected
|
|
`google/lyria-3-clip-preview` — a music generation model — 38 times for coding,
|
|
refactoring, debugging and summarization, **17 of them as the ranked winner
|
|
rather than an exploration gamble**. No `energy_observations` rows exist for
|
|
any: all 38 failed at dispatch.
|
|
|
|
The allowlist fixed the symptom the same day and is the right pragmatic gate.
|
|
But it is manual curation standing in for a structural absence: every filter in
|
|
the system is a threshold on a metric, and none asks whether a candidate is a
|
|
chat model. A new catalog row arrives with no proficiency — deliberately
|
|
treated as unproven rather than bad — and is immediately eligible to win.
|
|
|
|
---
|
|
|
|
## Paradigms being shoehorned
|
|
|
|
### A. "Classify, then dispatch" is a single-request frame on a conversation
|
|
|
|
The unit of work is a session: hundreds to thousands of turns sharing a
|
|
monotonically growing prefix, ~92% of which is cached. The router models each
|
|
turn as an independent classify-and-route problem, then bolts on a session
|
|
cache to make that affordable — and the cache is now 96.6% of the answer.
|
|
|
|
Inverting the frame dissolves several items at once. If the session is the unit
|
|
and a turn is a delta: classification becomes "has the task changed?", cheap
|
|
and usually no, instead of a ~1.1s full-taxonomy inference; the cache stops
|
|
being an optimization and becomes the model, making its staleness window a real
|
|
parameter rather than a cost hack; switching models becomes a visibly expensive
|
|
act (item 4); and `min_tool_proficiency` becomes expressible, because "can this
|
|
model be trusted with tools" is a session-level property — which is how the
|
|
docs already describe it.
|
|
|
|
### B. Quality as a per-(model, category) scalar is a leaderboard paradigm
|
|
|
|
One number per model per category, inherited from benchmarks. The project has
|
|
already found twice that this cannot hold what it learns.
|
|
|
|
`deepseek-v4-flash` scored 1.00 on all three coding categories and 0.33 on
|
|
`tool_use_agentic`, and the documented reason is a *failure mode*: given both
|
|
times in "it is 1:20pm and my meeting is at 3pm", it calls two tools instead of
|
|
subtracting. That is not a lower skill level on a category; it is a disposition
|
|
that is hazardous whenever tools are on the table, whatever the task is.
|
|
|
|
Flattening a mode into a category score forced a second mechanism to express
|
|
the real rule — `routing.min_tool_proficiency`, a hard filter — and that filter
|
|
is `null` because the flattening made it unusable: opencode sends `tools` on
|
|
nearly every request, so switching it on excludes the cheap model from all
|
|
traffic. The paradigm mismatch is what left a known hazard ungated.
|
|
|
|
The inheritance structure shows the same strain from the other side: 379 rows
|
|
are peer-rate priors carrying a number with zero evidence behind it, and
|
|
nothing downstream distinguishes a 0.749 that was measured from a 0.749 that
|
|
was copied — except `source`, which the ranker does not read.
|
|
|
|
### C. Energy and carbon are structurally central and functionally vestigial
|
|
|
|
`eco` is "not an objective." `max_energy_per_request` is null.
|
|
`plan_kwh_per_period` gates nothing. The project's own conclusion is that
|
|
billed energy "ranks how busy the provider was, not how efficient the model
|
|
is."
|
|
|
|
Around that sit: two columns and a comment block in `energy_observations`, a
|
|
BTU column kept "purely for comedic dashboard value", `seed_energy.py`, a
|
|
6-hourly timer whose accumulation is the stated prerequisite for trusting the
|
|
eco ordering, `seed_local_dispatch_energy.py`, grid-intensity and
|
|
`carbon_source` logging, open follow-up item 3, and a large share of a 71KB
|
|
`CLAUDE.md`.
|
|
|
|
The measurement work was worth doing — the per-kWh billing discovery and the
|
|
attribution decomposition are the most valuable findings in the project, and
|
|
item 3 above is an argument for using them *more*. But they belong to the cost
|
|
model now, not to an eco objective nothing optimizes. And the provider now
|
|
carrying 73% of traffic reports no energy at all, which makes the framing
|
|
actively misleading about what the router can see.
|
|
|
|
---
|
|
|
|
## Scale check
|
|
|
|
| | |
|
|
|---|---|
|
|
| Python | 55,711 lines; `dispatcher.py` 4,974, `admin.py` 2,498, `metrics.py` 2,474 |
|
|
| tests | 1,155 |
|
|
| config | 129 leaf keys across 22 sections |
|
|
| `CLAUDE.md` | 71 KB |
|
|
| `plans/` | 62 documents |
|
|
| purpose | route one operator's coding agent |
|
|
|
|
The engineering quality is high in ways that are rare and worth keeping: the
|
|
schema-drift and warnings tripwires, the `inherited_from` migration that fails
|
|
safe on NULL, the refusal to fabricate leaderboard priors, the four recorded
|
|
harness bugs that scored the rig rather than the model. None of that is the
|
|
problem. The problem is the absence of any forcing function that retires
|
|
machinery, or that re-fires a premise — 62 plan documents and no mechanism that
|
|
closes one out.
|
|
|
|
---
|
|
|
|
## What it needs, ranked
|
|
|
|
1. **A `feedback.py` timer.** One oneshot service and one timer, matching the
|
|
four that already exist. The ground-truth loop is the only one cranked by
|
|
hand, `proficiency` is three days stale, and the 95 unapplied outcomes are
|
|
exactly the models carrying today's traffic. Cheapest item here by a wide
|
|
margin.
|
|
|
|
2. **Capture OpenRouter cost and latency telemetry.** 95.7% of its rows have
|
|
NULL `cost_usd`, 100% have NULL `duration_seconds`, and it is 73% of current
|
|
decisions. This is a prerequisite for items 3 and 5, not a parallel task.
|
|
|
|
3. **Calibrate `estimated_cost` against the 31,119 billed rows already in the
|
|
DB.** A per-model correction factor over a trailing window. Fixes the 21x
|
|
misordering without a sweep and without abandoning per-request shape
|
|
sensitivity.
|
|
|
|
4. **Price the switch.** Charge challengers a cold-cache rate in the tiebreak.
|
|
~17% of spend sits on 7% of turns, and it makes the router's own decision an
|
|
input to its own cost model.
|
|
|
|
5. **Put latency in the objective.** A p50/p95 per model and a latency term or
|
|
ceiling for `INTERACTIVE` requests. Needs item 2 for most of the catalog.
|
|
|
|
6. **Make `quality_tolerance` depend on sample depth** rather than replacing one
|
|
constant with another. 95 of 474 rows are measured, 379 are inherited
|
|
priors; a band that is right for one is wrong for the other.
|
|
|
|
7. **An automatic response to a cost shock.** The 09-06 spike was handled well
|
|
and handled by a human. The minimum honest version: a period spend ceiling
|
|
that narrows the candidate set as it is approached, then refuses with a clear
|
|
reason. Not "the router is unusable without it" — it demonstrably is not —
|
|
but it is the difference between a tool that reports a problem and a tool
|
|
that responds to one.
|
|
|
|
8. **Decide what the classifier is for, given 96.6% cached.** Either accept it
|
|
as session-level and simplify the machinery to match, or make per-turn
|
|
classification cheap enough to actually run. The current state pays the
|
|
architecture cost of per-turn classification and gets session-level labels.
|
|
|
|
9. **Retire the eco objective explicitly.** Move the energy and attribution
|
|
findings into the cost model where they are load-bearing; archive the rest —
|
|
the seed timer, the eco scoring path, the BTU column, open item 3.
|
|
|
|
10. **A premise-expiry mechanism.** The thesis. Every setting whose comment says
|
|
"revisit when X" should have X as a check that fails loudly — a test, a
|
|
`/metrics` warning, a line in the poller. Four have expired silently.
|
|
|
|
Items 1-5 are independent and small, and 1 and 2 should land first because
|
|
everything downstream reads what they produce. Items 6, 8, 9 and 10 are
|
|
decisions rather than tasks, and should be made before more machinery is added
|
|
to the subsystems they touch.
|