Files
6krrt/plans/token-waste-waves.md
adlee-was-taken 25d95d4090 test(admin): a new config knob reaches the portal, or says why not
Wave 2 shipped objective.incumbent_cache_pricing and
objective.incumbent_challenger_cache_rate with no admin control. Nobody
decided they should not have one; it never came up, and the plan had no
step that would have made it come up. That matters most for a dial whose
rationale is tuning from neutral to full penalty WITHOUT reverting code:
if turning it means hand-editing a tracked file and restarting, the
tuning loop is too slow to walk.

Two halves, both small.

plans/token-waste-waves.md gains a standing per-wave acceptance gate,
placed ahead of Wave 1 so it is read before executing and checked before
a wave is called done. The escape clause is load-bearing, not hedging --
objective.credit_attenuation.enabled is deliberately off the allowlist
and off provider edits, because enabling it must be a config edit plus a
restart. The rule is that the ABSENCE of a control is a decision someone
made, not an oversight nobody noticed.

tests/test_admin_knob_coverage.py enforces it, in the shape
test_tui_schema_drift and test_tui_warnings already set here: covered is
DERIVED from admin._CONFIG_ALLOWLIST and admin._BOOL_KNOBS rather than
hand-copied, DELIBERATELY_NOT_IN_ADMIN carries reason strings rather than
bare names, and every failure names the knob. Exactness is asserted in
both directions, so the excuse list cannot rot into a rubber stamp as
knobs quietly gain controls.

Scope is the judgement call. RouterConfig has ~130 scalar leaves, and
demanding a decision on all of them produces a baseline nobody reads --
which is the rubber stamp being guarded against. Two clauses cut it to
58: a section is in scope iff the portal already reaches it (where it
reaches, it must reach completely), and deployment wiring -- endpoints,
model ids, credentials, paths, devices -- is out, being configuration of
where the router points rather than of how it behaves. 15 are covered
today, 43 excused with reasons. The docstring draws the line and
justifies it, including the classifier section, which is out because it
has its own dedicated admin card rather than a generic allowlist entry.

Two entries are marked PENDING feat/admin-incumbent-knobs: that branch is
adding controls for exactly those two knobs, and this branch is based on
origin/main where they do not exist yet. When it merges, the exactness
test FAILS on both until the entries are deleted. That is deliberate --
the test announcing its own cleanup beats a stale excuse sitting here
silently.

No config knob added, removed or changed; src/admin.py untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-14 20:01:49 -04:00

515 lines
24 KiB
Markdown

# Token waste: a five-wave plan
Status: in progress -- Wave 1 shipped, deployed and EVALUATED (PR #88; gate run
2026-09-14 on 1,288 decisions at full coverage); Wave 2 built and gated-passed
(PR #89, ships off, awaiting the enable decision); Wave 3 DEMOTED to correctness
-- its savings case was refuted by the same gate.
Measured read-only against the live `/home/alee/Sources/6krrt/router.db` on
2026-09-13 at repo state `26398d7`. Companion diagnosis:
`plans/ten-thousand-foot-review.md`.
**Intended executor: opencode (Atlas), directly.** Every item lands as its own
commit. Waves are ordered by dependency, not by size.
---
## The one-paragraph thesis
The provider's cache is worth ~92% of every prompt on this workload (measured
2026-09-14 at full coverage: **0.919** on same-model turns, n=1,227). The
router's only real action, choosing a model, is also the act that throws that
cache away — a switch drops the rate to **0.348** (n=55) and doubles the billed
cost per prompt token — and it takes that action without knowing the cache
exists. Everything below follows from that.
**Narrowed 2026-09-14.** This paragraph previously read "token waste here is
almost entirely cache destruction, not volume", which was too broad. Cache
destruction by **model switching** is real and is the leak. Cache destruction by
**editing the payload** is not: these providers tolerate a mid-prompt change
without invalidating the remainder, so the router's pruning does not cost cache
(see Wave 3, demoted on that evidence). The distinction matters because it moves
one whole wave out of the savings column.
## Read this before executing: what the data supports, and what it does not
Two claims were wrong during analysis and are corrected here so nothing gets
built on them. This repo has a history of scoring the rig instead of the model;
the same discipline applies to this plan.
**CONFIRMED — model switching destroys the cache.** Re-measured 2026-09-14 on
the post-Wave-1 window at ~100% cached-token coverage, which supersedes the
thin 12-22%-coverage figures this section first carried:
| turn | n | cache rate | µ$ / prompt token |
|---|---|---|---|
| same model as previous | 1,227 | **0.919** | 0.0369 |
| switched model | 55 | **0.348** | 0.0734 |
| first in session | 5 | 0.000 | 0.0403 |
The gap **widened** on the honest sample (0.571 against the thin sample's
0.474) while the cost ratio settled to **1.99x** from 2.46x. Switch turns are
4.3% of traffic. Two independent measurements — cache rate and billed cost per
prompt token — agree in direction and magnitude. This is the real leak, and it
is the only one that survived re-measurement.
**REFUTED 2026-09-14 — payload rewriting does NOT cost cache.** The mechanism
is real: a direct test of `prune_context` showed the relevance path diverging at
message index 19 of 82 with a literally identical relevance order, where the
uniform path diverges nowhere. A growing `target_save` walks a relevance-ordered
list, so the newly-compressed candidate lands at an arbitrary *position*.
But the inference drawn from it was wrong. At full coverage, turns with **half
the prompt sitting after the divergence point still cache at 0.910** — which a
strict prefix cache cannot do. These providers tolerate a mid-prompt edit
without invalidating the remainder. Wave 3 is demoted on that evidence; see it
for the full tables.
**This entry has now been wrong in both directions, which is the useful part.**
It first read NOT CONFIRMED on a 0.924 aggregate — one number averaged across
two cohorts at 12-22% coverage, too coarse to separate them. It was then flipped
to CONFIRMED on a cohort split from a single hour (n=46, 0.705 vs 0.940). The
full window (n=410 rewritten) put the gap at 0.013 and the mechanism test
removed it entirely.
Three readings, two reversals, one underlying question. What distinguished the
final answer was not more data but a **mechanism test** — asking whether cache
rate falls as the share of the prompt after the divergence rises, rather than
comparing two buckets and trusting the difference. A cohort split can only tell
you two groups differ; it cannot tell you why, and it will happily report noise
as signal at small n.
The lesson generalises past this entry: an aggregate that mixes two populations
will report the majority's number and hide the minority's, and "the data does
not show it biting" is a claim about resolution as much as about reality.
---
## Standing acceptance gate — applies to every wave
In addition to each wave's own criteria below:
> **Every config knob a wave introduces ships with an admin control — runtime,
> persisted, or both as its mechanism warrants — or a recorded decision saying
> why it must not.**
The escape clause is load-bearing, not hedging.
`objective.credit_attenuation.enabled` is deliberately absent from the admin
allowlist and from provider edits, because turning it on must be a config edit
plus a restart (CLAUDE.md, "Routing notes"). So the rule is not "everything must
have a control" — it is that the ABSENCE of one is a decision someone made,
rather than an oversight nobody noticed.
Wave 2 is why this is written down: `objective.incumbent_cache_pricing` and
`objective.incumbent_challenger_cache_rate` shipped with no control and no
decision, and a dial whose whole point is tuning from neutral to full penalty
without reverting code is worth little if turning it means hand-editing a
tracked file and restarting the service.
`tests/test_admin_knob_coverage.py` enforces this for the sections the portal
already reaches, and fails naming the knob.
---
## Wave 1 — See the leak
No behavior change. Every later wave's acceptance criteria read what this
produces, so nothing else should start first.
Commit `1bd8c6e` (2026-09-10) already landed streamed cost capture and
`cached_prompt_tokens` parsing at `dispatcher.py:2056`, which is why data
begins on 09-11. The instrument exists; it is under-covered.
### 1.1 OpenRouter usage opt-in
Coverage since the instrument landed:
| provider | obs | cost | cached tokens | duration |
|---|---|---|---|---|
| neuralwatt | 805 | 804 (99.9%) | 103 (12.8%) | 804 (99.9%) |
| openrouter | 734 | 165 (22.5%) | 164 (22.3%) | **0 (0%)** |
OpenRouter's cost coverage (22.5%) and cached coverage (22.3%) match almost
exactly. That is **one** gap, not two: the usage block is simply absent on ~78%
of requests. `dispatcher.py:4254` sets
`stream_options: {"include_usage": True}`, which is the OpenAI spelling;
OpenRouter requires its own `usage: {"include": true}` in the request body to
return accounting. Add it per-provider — the `reports_cost_in_usage` flag from
`1bd8c6e` is the right place to hang it.
This single change is the highest-leverage item in the plan: it takes OpenRouter
from ~22% to ~100% on cost, cached tokens **and** latency at once, on the
provider carrying ~73% of decisions.
**Verify**: coverage for all three fields above 95% on OpenRouter rows written
after the change.
### 1.2 Decide what an absent `cached_tokens` means
NeuralWatt reports cost on 99.9% of rows but `prompt_tokens_details` on only
12.8%. Determine whether the field is omitted on a full cache miss or is
model-dependent. If omitted on a miss, record an explicit `0` rather than NULL,
because every cache-rate reading in this plan is otherwise conditioned on a hit
having occurred. Nine rows currently store `0`, so zeros do reach the DB
sometimes — that needs explaining before the 0.924 figure can be trusted.
**Verify**: a test pinning the parse for a usage block with `cached_tokens: 0`,
one with the key absent, and one with `prompt_tokens_details` absent entirely,
asserting the three map to distinct stored values.
### 1.3 Prefix-stability probe
The decisive instrument for Wave 3, and about twenty lines. After
`prune_context` returns, hash the pruned payload cumulatively by message and
store the hash of the longest stable prefix (or simply a per-turn list of
message hashes) on `route_decisions`. Within a session, compare consecutive
turns to get the real first-divergence index and the tokens after it.
That converts Wave 3 from an argument into a measurement, and it will either
justify Wave 3 or retire it.
**Verify**: on a replayed synthetic session, the probe reports divergence at the
index the direct test predicts.
### 1.4 Cache rate in `/metrics`
A cache-rate series per `(provider, model)` over a trailing window, plus a
warning when a session's rate falls below a configurable floor. Follow the
existing detector conventions in `metrics.py` — novelty-or-rate, not bare
presence, per the reactive rejection detector.
This is also the premise-expiry check for `assumed_cache_rate: 0.917`: the
constant was measured 2026-08-23 and nothing has re-measured it since.
---
## Wave 2 — Stop the confirmed leak
The router has no incumbent. `rank_candidates` does not know what ran last
turn, so there is no hysteresis and no switching cost. Three sub-items, and the
first is most of the work.
### 2.1 Thread the incumbent into ranking
The session's last `selected_model` is already on disk in `route_decisions` and
usually in memory. Pass it into `rank_candidates` as the incumbent.
Keep the existing shape: quality band first, cost as tiebreak. The incumbent
changes only the cost key.
### 2.2 Price the cache loss
In the tiebreak, price the incumbent at the measured cache rate and every
challenger at a cold rate. Concretely, `estimated_cost` already takes
`cache_rate`; pass `cfg.objective.assumed_cache_rate` for the incumbent and
`0.0` for challengers. A challenger then has to beat the incumbent by more than
the cache it is about to discard, which on a 100k prompt is most of the prompt.
This is the first time the router's own decision becomes an input to its own
cost model, and it is why it belongs before the deeper reframe in Wave 5.
Two guards so this cannot become stickiness-at-any-cost:
- the quality band is computed **before** the cost key, exactly as today, so a
genuine quality gap still wins outright and the incumbent gets no quality
advantage;
- a hard-filter failure (context ceiling, capability gate, circuit breaker,
profile allowlist) still removes the incumbent unconditionally.
### 2.3 Exploration becomes session-scoped
Epsilon-greedy prices a suboptimal pull as the arm's cost difference. Here
exploring *means* switching, so the real cost is that difference plus a cold
prompt — and `exploration.py` cannot see it. Move the coin flip to session
start rather than per turn: one exploratory session costs one cold prompt
instead of one per turn, and it produces a cleaner outcome signal because the
whole session is attributable to the explored model.
`exploration.py` takes an injected RNG and holds no mutable state, so this is a
call-site change, not a rewrite.
**Wave 2 acceptance**, read off Wave 1's instruments:
- switch rate per session falls;
- the same-model / switched cache-rate gap (0.952 vs 0.478) narrows, or the
remaining switches are deliberate — driven by a quality band or a hard
filter, not by a cost re-rank;
- median billed µ$ per prompt token moves toward the same-model figure.
---
## Wave 3 — Make prefix stability guaranteed rather than incidental
**DEMOTED 2026-09-14, back to correctness rather than savings.** This section
was promoted on 2026-09-13 on the probe's first hour of data, which showed
rewritten turns at 0.705 cache rate against 0.940 for turns that merely grew.
That promotion carried an explicit caveat — "n=46 from roughly one hour, do not
treat 0.705 as settled, re-measure on a fuller sample." The re-measurement is
in, and the caveat is the part that held.
### What the full window actually measured
1,288 decisions over ~5 hours, at ~100% cached-token coverage, joined on
`cached_tokens_source = 'reported'`:
| prefix state | pruning | n | avg tokens after divergence | cache rate |
|---|---|---|---|---|
| grew only | pruned | 854 | 702 | 0.904 |
| **rewritten** | pruned | 387 | 32,901 | **0.891** |
| grew only | under budget | 14 | 6,322 | 0.765 |
| rewritten | under budget | 23 | 9,075 | 0.548 |
Among pruned turns the gap is **0.013**, not the 0.235 the first hour showed.
### The mechanism test, which is what settles it
If a mid-prompt edit invalidated everything after it, cache rate would fall as
the share of the prompt sitting after the divergence point rises. It does not:
| tokens after divergence | n | share of prompt | cache rate |
|---|---|---|---|
| <5k | 24 | 7% | 0.813 |
| 5-20k | 89 | 28% | 0.888 |
| 20-50k | 227 | **50%** | **0.910** |
| >50k | 66 | 50% | 0.835 |
No monotonic relationship. Turns with **half the prompt after the divergence
point** still cache at 0.910 — arithmetically impossible under a strict prefix
cache, which would cap that bucket near 0.50.
**Conclusion: these providers do not use a strict prefix cache.** A mid-prompt
edit does not invalidate the remainder. That assumption was load-bearing under
this wave's savings case and under the reading of the `prune_context`
simulation, and it is wrong.
The simulation itself was never wrong about what it measured — the relevance
path really does rewrite the payload at an arbitrary position, and the uniform
path really does not. What was wrong was the inference that a rewritten payload
costs cache. It measurably does not.
**Gate 3 is untouched by this.** Switching models is not a mid-prompt edit; it
is a different cache namespace entirely, and it still measures 0.919 against
0.348. Nothing here weakens Wave 2.
### What this changes for the sub-items
- **3.1 (monotone region) is now optional and probably not worth it.** It buys
no measurable cache. It would buy determinism in a working path, at the cost
of editing that path. Do not do it for savings; there are none.
- **3.2 (tripwire test)** as originally specced asserts prefix stability, which
the relevance path violates by design. Without 3.1 it would ship red. Either
it follows 3.1 or it is dropped — it is not independently landable.
- **3.3 (config comment reconciliation) is still worth doing**, and is now a
clean decision. With cache out of the picture the only difference between the
paths is token volume against context quality: the direct test measured the
relevance path shipping 109,606 tokens against uniform's 94,716 on the same
input, because it stops as soon as the deficit is covered. So relevance costs
~15k more tokens per turn and buys keeping the most relevant content verbatim.
That is a real tradeoff, just not a cache one — decide it on its own terms.
- **3.4 (non-pruning rewrite source) drops to low priority.** The under-budget
rewritten cohort (n=23, 0.548) is small and its low rate is better explained
by session-opening turns than by a distinct leak.
### An opportunity this created
If a mid-prompt edit costs no cache, pruning is **cheaper than assumed** and
could be more aggressive without a cache penalty. Pruned turns currently run
132k -> 89k (a 33% cut); `pinch.budget_tokens` could go lower for real prompt-
side savings. The binding constraint is answer quality, not cache. That is a
new, evidence-created item and it belongs in Wave 4 or 5, not here.
### A second rewrite source exists, outside pruning
Cross-referencing the same rows against whether pruning actually fired:
| prefix state | pruning fired | under budget |
|---|---|---|
| grew only | 193 | 6 |
| **rewritten** | **39** | **7** |
Pruning explains 39 of 46 rewrites. **Seven turns rewrote the prefix with
pruning never engaged at all** — under budget, so `prune_context` returned the
messages untouched. On those turns the payload is the client's own message list,
which means something upstream of the router edited its own history (opencode
compaction, a changed tool-definition array, or a mutated system prompt are the
candidates).
Two consequences, and both matter for how this wave is judged:
- **3.1 cannot fix all of it.** Fixing the relevance path addresses at most 39
of 46. Crediting Wave 3 with the whole gap would overstate it.
- **The residual is worth identifying before it is assumed benign.** Among
under-budget turns the rewrite rate is 7 of 13 — proportionally *higher* than
the pruned cohort's 39 of 232, though on a sample far too small to lean on.
If the client rewrites its own history routinely, that is a larger cache
leak than pruning and the router cannot fix it by changing `prune_context`.
### 3.1 Constrain relevance compression to a monotone region
Keep the relevance *ranking* and change what it decides. Instead of compressing
a scattered relevance-ordered subset, let relevance pick **where the compressed
boundary sits** on first crossing, then only ever extend that boundary forward.
Once a message is compressed it is never un-compressed.
That preserves the feature's intent — relevance still decides what is worth
keeping verbatim — while restoring the property the uniform path has for free:
appends cannot change a byte before the boundary.
### 3.2 A prefix-stability tripwire test
This repo already writes exactly this kind of test
(`test_tui_schema_drift.py`, `test_tui_warnings.py`). Assert that across a
simulated multi-turn session, `prune_context`'s output for turn N+1 is
byte-identical to turn N's up to the appended messages. Fail naming the
divergent message index.
The uniform path passes this today; the relevance path does not. That asymmetry
is the whole point of the test.
### 3.3 Reconcile the config with reality
`pinch.relevance.enabled: true`, while the comment directly above it still
reads *"Off by default… ship it, watch route_decisions / pinch stats on real
traffic, then decide the default."* It was switched on and the comment never
followed. Same shape as `session_cache` below it. Decide the default on Wave 1
data and rewrite both comments to say what is actually true.
The Wave 1 data now exists, so this is decidable rather than deferred. Note the
decision is not automatically "turn it off": the earlier direct test showed the
relevance path also ships **more** tokens than uniform (109,606 against 94,716
on the same input, because it stops as soon as the deficit is covered rather
than compressing every candidate). If 3.1 makes it prefix-stable, relevance
becomes strictly better than uniform. If 3.1 proves harder than expected,
disabling it is the cheap fallback that wins on both axes today.
### 3.4 Identify the non-pruning rewrite source
Scoped by the measurement above, not speculative. Seven of 46 rewrites occurred
with pruning disengaged, so something upstream of `prune_context` is editing
conversation history between turns.
This is an investigation, not a fix: determine whether the client is compacting
its own context, whether the tool-definition array changes between turns, or
whether the system prompt is mutated. The probe already stores what is needed
to find the turns; the question is what differs across them.
Sequence it **after 3.1** so the pruning-caused rewrites are removed from the
population first, leaving a clean residual to study. Attempting it now means
diagnosing two overlapping causes at once.
### Acceptance for Wave 3
The probe that promoted this wave is also its acceptance instrument, which is
the point of having built it first:
1. The **rewritten share** among pruned turns falls toward zero (39 of 232
today).
2. The **cache-rate gap between the two cohorts closes** — 0.705 against 0.940
today. Re-measure both on a fuller sample before and after, per the
re-measure discipline in Wave 2's acceptance; do not hardcode today's
figures as the target.
3. The tripwire test in 3.2 stays green on both compression paths.
4. Any residual rewriting is attributed to a named cause by 3.4, not left as
unexplained variance.
---
## Wave 4 — Retries, and the half nobody guards
### 4.1 Denominate the iteration budget in prompt re-bills
`iteration.py` counts attempts. A retry on a 100k-token conversation is a 100k
prompt re-bill to redo a ~400-token answer, and `malformed` escalates to the
*next-ranked candidate*, which makes it a cold one. Tier 3 allows two.
Budget in re-bills instead: gate a retry on prompt size, and prefer
same-model-with-more-tokens (which keeps the cache) over escalation (which does
not) wherever the failure permits it. The existing failure taxonomy already
supports this — `truncated` retries the same model by design; it is `malformed`
that escalates.
### 4.2 Accept that structural verification is inert here
98.96% of structural verdicts are `unverifiable` (28,958 of 29,263), because
agent turns end in tool calls and `has_tool_calls` short-circuits both
checkers. That is correct behavior and the documented fix to a real
false-failure incident. The conclusion not yet drawn is that the free checker
now checks nothing on the only workload this router serves: 305 substantive
verdicts out of 29,263.
Meanwhile a completion token costs ~201x a prompt token, so the expensive half
of the ledger has no guard and the cheap half carries all the machinery.
Two actions, both small:
- **Give `feedback.py` a timer.** It is the only loop without one — `deploy/`
ships timers for the poller, seed sweep, backup and offsite sync, and a
baseline-report timer is installed. `proficiency` was written in one batch at
2026-09-10T00:44:55 and the 95 unapplied outcomes are exactly the models
carrying today's traffic. Client outcomes are the only guard on wasted
completions, and they are applied by hand.
- Stop reporting the structural checker as a safety net in the docs, and either
narrow it to the paths where it still fires or retire it.
---
## Wave 5 — The reframe, and the expiry checks
Decisions, not tasks. Make them before adding machinery to the subsystems they
touch.
### 5.1 The session is the routing unit
Waves 2 and 3 patch a per-request frame. The unit is wrong: the workload is a
session of hundreds to thousands of turns sharing a monotonically growing
prefix, and `estimated_cost` is a function of `prompt_tokens`, so the ranking's
key input changes every turn even when nothing else does. `CLAUDE.md` documents
the consequence as a feature — the winner at 50k differs from the winner at
120k — which inside a session is a cache dump.
If the session is the unit and a turn a delta, then: routing decides once and
re-decides only on a threshold that includes the cache loss; classification
becomes "has the task changed?" rather than a full taxonomy inference (96.6% of
turns already reuse a cached label); exploration is naturally session-scoped;
and `routing.min_tool_proficiency` becomes expressible, because "can this model
be trusted with tools" is a session-level property, which is how the docs
already describe it.
### 5.2 Recalibrate `estimated_cost` against billed rows
Depends on 1.1. Per-request estimate against bill spans **0.63x to 13.55x**, a
21x spread in the error, which reorders candidates. 31,119 NeuralWatt rows carry
a real billed figure joinable by `request_id`. A per-model correction factor over
a trailing window, refreshed by the poller, fixes the ordering without a sweep.
The cause is already in `CLAUDE.md` two sections apart and never reconciled:
attribution ratio spans 750x between models and is "most of the real cost
difference in the catalog", while `estimated_cost` prices from catalog token
prices, which carry no attribution term.
### 5.3 Premise-expiry checks
Four settings carry an explicit revisit condition in their own comment, all four
conditions are met, and none re-fired: `quality_tolerance: 0.1` ("narrow it as
samples accumulate" — now 46 outcome-backed rows per category averaging 51.8
samples); cost-as-tiebreak ("all real traffic to date totals $0.07" — now
$118.79); `assumed_cache_rate: 0.917` (measured 2026-08-23, never re-measured);
`pinch.relevance` ("off by default… then decide").
Every such comment should have its condition as a check that fails loudly — a
test, a `/metrics` warning, a line in the poller. 1.4 is the first instance;
generalize it.
---
## Order of execution
| wave | items | gate to proceed |
|---|---|---|
| 1 | 1.1, 1.2, 1.3, 1.4 | coverage >95% on OpenRouter cost/cached/duration |
| 2 | 2.1, 2.2, 2.3 | switch cache-rate gap narrows, or switches are deliberate |
| 3 | **demoted** — 3.3 only (a token-volume vs context-quality decision); 3.1/3.2 optional, 3.4 low | no savings gate; payload rewriting was measured not to cost cache |
| 4 | 4.1, 4.2 | `proficiency` refreshing on a timer, unapplied backlog ~0 |
| 5 | decisions | — |
1.1 first, alone, and measure for a day before anything else lands. It is small,
it is reversible, and until it is in place every acceptance criterion below it is
reading a 22%-covered sample.