Files
6krrt/plans/token-waste-waves.md
adlee-was-taken 25d95d4090 test(admin): a new config knob reaches the portal, or says why not
Wave 2 shipped objective.incumbent_cache_pricing and
objective.incumbent_challenger_cache_rate with no admin control. Nobody
decided they should not have one; it never came up, and the plan had no
step that would have made it come up. That matters most for a dial whose
rationale is tuning from neutral to full penalty WITHOUT reverting code:
if turning it means hand-editing a tracked file and restarting, the
tuning loop is too slow to walk.

Two halves, both small.

plans/token-waste-waves.md gains a standing per-wave acceptance gate,
placed ahead of Wave 1 so it is read before executing and checked before
a wave is called done. The escape clause is load-bearing, not hedging --
objective.credit_attenuation.enabled is deliberately off the allowlist
and off provider edits, because enabling it must be a config edit plus a
restart. The rule is that the ABSENCE of a control is a decision someone
made, not an oversight nobody noticed.

tests/test_admin_knob_coverage.py enforces it, in the shape
test_tui_schema_drift and test_tui_warnings already set here: covered is
DERIVED from admin._CONFIG_ALLOWLIST and admin._BOOL_KNOBS rather than
hand-copied, DELIBERATELY_NOT_IN_ADMIN carries reason strings rather than
bare names, and every failure names the knob. Exactness is asserted in
both directions, so the excuse list cannot rot into a rubber stamp as
knobs quietly gain controls.

Scope is the judgement call. RouterConfig has ~130 scalar leaves, and
demanding a decision on all of them produces a baseline nobody reads --
which is the rubber stamp being guarded against. Two clauses cut it to
58: a section is in scope iff the portal already reaches it (where it
reaches, it must reach completely), and deployment wiring -- endpoints,
model ids, credentials, paths, devices -- is out, being configuration of
where the router points rather than of how it behaves. 15 are covered
today, 43 excused with reasons. The docstring draws the line and
justifies it, including the classifier section, which is out because it
has its own dedicated admin card rather than a generic allowlist entry.

Two entries are marked PENDING feat/admin-incumbent-knobs: that branch is
adding controls for exactly those two knobs, and this branch is based on
origin/main where they do not exist yet. When it merges, the exactness
test FAILS on both until the entries are deleted. That is deliberate --
the test announcing its own cleanup beats a stale excuse sitting here
silently.

No config knob added, removed or changed; src/admin.py untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-14 20:01:49 -04:00

24 KiB

Token waste: a five-wave plan

Status: in progress -- Wave 1 shipped, deployed and EVALUATED (PR #88; gate run 2026-09-14 on 1,288 decisions at full coverage); Wave 2 built and gated-passed (PR #89, ships off, awaiting the enable decision); Wave 3 DEMOTED to correctness -- its savings case was refuted by the same gate.

Measured read-only against the live /home/alee/Sources/6krrt/router.db on 2026-09-13 at repo state 26398d7. Companion diagnosis: plans/ten-thousand-foot-review.md.

Intended executor: opencode (Atlas), directly. Every item lands as its own commit. Waves are ordered by dependency, not by size.


The one-paragraph thesis

The provider's cache is worth ~92% of every prompt on this workload (measured 2026-09-14 at full coverage: 0.919 on same-model turns, n=1,227). The router's only real action, choosing a model, is also the act that throws that cache away — a switch drops the rate to 0.348 (n=55) and doubles the billed cost per prompt token — and it takes that action without knowing the cache exists. Everything below follows from that.

Narrowed 2026-09-14. This paragraph previously read "token waste here is almost entirely cache destruction, not volume", which was too broad. Cache destruction by model switching is real and is the leak. Cache destruction by editing the payload is not: these providers tolerate a mid-prompt change without invalidating the remainder, so the router's pruning does not cost cache (see Wave 3, demoted on that evidence). The distinction matters because it moves one whole wave out of the savings column.

Read this before executing: what the data supports, and what it does not

Two claims were wrong during analysis and are corrected here so nothing gets built on them. This repo has a history of scoring the rig instead of the model; the same discipline applies to this plan.

CONFIRMED — model switching destroys the cache. Re-measured 2026-09-14 on the post-Wave-1 window at ~100% cached-token coverage, which supersedes the thin 12-22%-coverage figures this section first carried:

turn n cache rate µ$ / prompt token
same model as previous 1,227 0.919 0.0369
switched model 55 0.348 0.0734
first in session 5 0.000 0.0403

The gap widened on the honest sample (0.571 against the thin sample's 0.474) while the cost ratio settled to 1.99x from 2.46x. Switch turns are 4.3% of traffic. Two independent measurements — cache rate and billed cost per prompt token — agree in direction and magnitude. This is the real leak, and it is the only one that survived re-measurement.

REFUTED 2026-09-14 — payload rewriting does NOT cost cache. The mechanism is real: a direct test of prune_context showed the relevance path diverging at message index 19 of 82 with a literally identical relevance order, where the uniform path diverges nowhere. A growing target_save walks a relevance-ordered list, so the newly-compressed candidate lands at an arbitrary position.

But the inference drawn from it was wrong. At full coverage, turns with half the prompt sitting after the divergence point still cache at 0.910 — which a strict prefix cache cannot do. These providers tolerate a mid-prompt edit without invalidating the remainder. Wave 3 is demoted on that evidence; see it for the full tables.

This entry has now been wrong in both directions, which is the useful part. It first read NOT CONFIRMED on a 0.924 aggregate — one number averaged across two cohorts at 12-22% coverage, too coarse to separate them. It was then flipped to CONFIRMED on a cohort split from a single hour (n=46, 0.705 vs 0.940). The full window (n=410 rewritten) put the gap at 0.013 and the mechanism test removed it entirely.

Three readings, two reversals, one underlying question. What distinguished the final answer was not more data but a mechanism test — asking whether cache rate falls as the share of the prompt after the divergence rises, rather than comparing two buckets and trusting the difference. A cohort split can only tell you two groups differ; it cannot tell you why, and it will happily report noise as signal at small n.

The lesson generalises past this entry: an aggregate that mixes two populations will report the majority's number and hide the minority's, and "the data does not show it biting" is a claim about resolution as much as about reality.


Standing acceptance gate — applies to every wave

In addition to each wave's own criteria below:

Every config knob a wave introduces ships with an admin control — runtime, persisted, or both as its mechanism warrants — or a recorded decision saying why it must not.

The escape clause is load-bearing, not hedging. objective.credit_attenuation.enabled is deliberately absent from the admin allowlist and from provider edits, because turning it on must be a config edit plus a restart (CLAUDE.md, "Routing notes"). So the rule is not "everything must have a control" — it is that the ABSENCE of one is a decision someone made, rather than an oversight nobody noticed.

Wave 2 is why this is written down: objective.incumbent_cache_pricing and objective.incumbent_challenger_cache_rate shipped with no control and no decision, and a dial whose whole point is tuning from neutral to full penalty without reverting code is worth little if turning it means hand-editing a tracked file and restarting the service.

tests/test_admin_knob_coverage.py enforces this for the sections the portal already reaches, and fails naming the knob.


Wave 1 — See the leak

No behavior change. Every later wave's acceptance criteria read what this produces, so nothing else should start first.

Commit 1bd8c6e (2026-09-10) already landed streamed cost capture and cached_prompt_tokens parsing at dispatcher.py:2056, which is why data begins on 09-11. The instrument exists; it is under-covered.

1.1 OpenRouter usage opt-in

Coverage since the instrument landed:

provider obs cost cached tokens duration
neuralwatt 805 804 (99.9%) 103 (12.8%) 804 (99.9%)
openrouter 734 165 (22.5%) 164 (22.3%) 0 (0%)

OpenRouter's cost coverage (22.5%) and cached coverage (22.3%) match almost exactly. That is one gap, not two: the usage block is simply absent on ~78% of requests. dispatcher.py:4254 sets stream_options: {"include_usage": True}, which is the OpenAI spelling; OpenRouter requires its own usage: {"include": true} in the request body to return accounting. Add it per-provider — the reports_cost_in_usage flag from 1bd8c6e is the right place to hang it.

This single change is the highest-leverage item in the plan: it takes OpenRouter from ~22% to ~100% on cost, cached tokens and latency at once, on the provider carrying ~73% of decisions.

Verify: coverage for all three fields above 95% on OpenRouter rows written after the change.

1.2 Decide what an absent cached_tokens means

NeuralWatt reports cost on 99.9% of rows but prompt_tokens_details on only 12.8%. Determine whether the field is omitted on a full cache miss or is model-dependent. If omitted on a miss, record an explicit 0 rather than NULL, because every cache-rate reading in this plan is otherwise conditioned on a hit having occurred. Nine rows currently store 0, so zeros do reach the DB sometimes — that needs explaining before the 0.924 figure can be trusted.

Verify: a test pinning the parse for a usage block with cached_tokens: 0, one with the key absent, and one with prompt_tokens_details absent entirely, asserting the three map to distinct stored values.

1.3 Prefix-stability probe

The decisive instrument for Wave 3, and about twenty lines. After prune_context returns, hash the pruned payload cumulatively by message and store the hash of the longest stable prefix (or simply a per-turn list of message hashes) on route_decisions. Within a session, compare consecutive turns to get the real first-divergence index and the tokens after it.

That converts Wave 3 from an argument into a measurement, and it will either justify Wave 3 or retire it.

Verify: on a replayed synthetic session, the probe reports divergence at the index the direct test predicts.

1.4 Cache rate in /metrics

A cache-rate series per (provider, model) over a trailing window, plus a warning when a session's rate falls below a configurable floor. Follow the existing detector conventions in metrics.py — novelty-or-rate, not bare presence, per the reactive rejection detector.

This is also the premise-expiry check for assumed_cache_rate: 0.917: the constant was measured 2026-08-23 and nothing has re-measured it since.


Wave 2 — Stop the confirmed leak

The router has no incumbent. rank_candidates does not know what ran last turn, so there is no hysteresis and no switching cost. Three sub-items, and the first is most of the work.

2.1 Thread the incumbent into ranking

The session's last selected_model is already on disk in route_decisions and usually in memory. Pass it into rank_candidates as the incumbent.

Keep the existing shape: quality band first, cost as tiebreak. The incumbent changes only the cost key.

2.2 Price the cache loss

In the tiebreak, price the incumbent at the measured cache rate and every challenger at a cold rate. Concretely, estimated_cost already takes cache_rate; pass cfg.objective.assumed_cache_rate for the incumbent and 0.0 for challengers. A challenger then has to beat the incumbent by more than the cache it is about to discard, which on a 100k prompt is most of the prompt.

This is the first time the router's own decision becomes an input to its own cost model, and it is why it belongs before the deeper reframe in Wave 5.

Two guards so this cannot become stickiness-at-any-cost:

  • the quality band is computed before the cost key, exactly as today, so a genuine quality gap still wins outright and the incumbent gets no quality advantage;
  • a hard-filter failure (context ceiling, capability gate, circuit breaker, profile allowlist) still removes the incumbent unconditionally.

2.3 Exploration becomes session-scoped

Epsilon-greedy prices a suboptimal pull as the arm's cost difference. Here exploring means switching, so the real cost is that difference plus a cold prompt — and exploration.py cannot see it. Move the coin flip to session start rather than per turn: one exploratory session costs one cold prompt instead of one per turn, and it produces a cleaner outcome signal because the whole session is attributable to the explored model.

exploration.py takes an injected RNG and holds no mutable state, so this is a call-site change, not a rewrite.

Wave 2 acceptance, read off Wave 1's instruments:

  • switch rate per session falls;
  • the same-model / switched cache-rate gap (0.952 vs 0.478) narrows, or the remaining switches are deliberate — driven by a quality band or a hard filter, not by a cost re-rank;
  • median billed µ$ per prompt token moves toward the same-model figure.

Wave 3 — Make prefix stability guaranteed rather than incidental

DEMOTED 2026-09-14, back to correctness rather than savings. This section was promoted on 2026-09-13 on the probe's first hour of data, which showed rewritten turns at 0.705 cache rate against 0.940 for turns that merely grew. That promotion carried an explicit caveat — "n=46 from roughly one hour, do not treat 0.705 as settled, re-measure on a fuller sample." The re-measurement is in, and the caveat is the part that held.

What the full window actually measured

1,288 decisions over ~5 hours, at ~100% cached-token coverage, joined on cached_tokens_source = 'reported':

prefix state pruning n avg tokens after divergence cache rate
grew only pruned 854 702 0.904
rewritten pruned 387 32,901 0.891
grew only under budget 14 6,322 0.765
rewritten under budget 23 9,075 0.548

Among pruned turns the gap is 0.013, not the 0.235 the first hour showed.

The mechanism test, which is what settles it

If a mid-prompt edit invalidated everything after it, cache rate would fall as the share of the prompt sitting after the divergence point rises. It does not:

tokens after divergence n share of prompt cache rate
<5k 24 7% 0.813
5-20k 89 28% 0.888
20-50k 227 50% 0.910
>50k 66 50% 0.835

No monotonic relationship. Turns with half the prompt after the divergence point still cache at 0.910 — arithmetically impossible under a strict prefix cache, which would cap that bucket near 0.50.

Conclusion: these providers do not use a strict prefix cache. A mid-prompt edit does not invalidate the remainder. That assumption was load-bearing under this wave's savings case and under the reading of the prune_context simulation, and it is wrong.

The simulation itself was never wrong about what it measured — the relevance path really does rewrite the payload at an arbitrary position, and the uniform path really does not. What was wrong was the inference that a rewritten payload costs cache. It measurably does not.

Gate 3 is untouched by this. Switching models is not a mid-prompt edit; it is a different cache namespace entirely, and it still measures 0.919 against 0.348. Nothing here weakens Wave 2.

What this changes for the sub-items

  • 3.1 (monotone region) is now optional and probably not worth it. It buys no measurable cache. It would buy determinism in a working path, at the cost of editing that path. Do not do it for savings; there are none.
  • 3.2 (tripwire test) as originally specced asserts prefix stability, which the relevance path violates by design. Without 3.1 it would ship red. Either it follows 3.1 or it is dropped — it is not independently landable.
  • 3.3 (config comment reconciliation) is still worth doing, and is now a clean decision. With cache out of the picture the only difference between the paths is token volume against context quality: the direct test measured the relevance path shipping 109,606 tokens against uniform's 94,716 on the same input, because it stops as soon as the deficit is covered. So relevance costs ~15k more tokens per turn and buys keeping the most relevant content verbatim. That is a real tradeoff, just not a cache one — decide it on its own terms.
  • 3.4 (non-pruning rewrite source) drops to low priority. The under-budget rewritten cohort (n=23, 0.548) is small and its low rate is better explained by session-opening turns than by a distinct leak.

An opportunity this created

If a mid-prompt edit costs no cache, pruning is cheaper than assumed and could be more aggressive without a cache penalty. Pruned turns currently run 132k -> 89k (a 33% cut); pinch.budget_tokens could go lower for real prompt- side savings. The binding constraint is answer quality, not cache. That is a new, evidence-created item and it belongs in Wave 4 or 5, not here.

A second rewrite source exists, outside pruning

Cross-referencing the same rows against whether pruning actually fired:

prefix state pruning fired under budget
grew only 193 6
rewritten 39 7

Pruning explains 39 of 46 rewrites. Seven turns rewrote the prefix with pruning never engaged at all — under budget, so prune_context returned the messages untouched. On those turns the payload is the client's own message list, which means something upstream of the router edited its own history (opencode compaction, a changed tool-definition array, or a mutated system prompt are the candidates).

Two consequences, and both matter for how this wave is judged:

  • 3.1 cannot fix all of it. Fixing the relevance path addresses at most 39 of 46. Crediting Wave 3 with the whole gap would overstate it.
  • The residual is worth identifying before it is assumed benign. Among under-budget turns the rewrite rate is 7 of 13 — proportionally higher than the pruned cohort's 39 of 232, though on a sample far too small to lean on. If the client rewrites its own history routinely, that is a larger cache leak than pruning and the router cannot fix it by changing prune_context.

3.1 Constrain relevance compression to a monotone region

Keep the relevance ranking and change what it decides. Instead of compressing a scattered relevance-ordered subset, let relevance pick where the compressed boundary sits on first crossing, then only ever extend that boundary forward. Once a message is compressed it is never un-compressed.

That preserves the feature's intent — relevance still decides what is worth keeping verbatim — while restoring the property the uniform path has for free: appends cannot change a byte before the boundary.

3.2 A prefix-stability tripwire test

This repo already writes exactly this kind of test (test_tui_schema_drift.py, test_tui_warnings.py). Assert that across a simulated multi-turn session, prune_context's output for turn N+1 is byte-identical to turn N's up to the appended messages. Fail naming the divergent message index.

The uniform path passes this today; the relevance path does not. That asymmetry is the whole point of the test.

3.3 Reconcile the config with reality

pinch.relevance.enabled: true, while the comment directly above it still reads "Off by default… ship it, watch route_decisions / pinch stats on real traffic, then decide the default." It was switched on and the comment never followed. Same shape as session_cache below it. Decide the default on Wave 1 data and rewrite both comments to say what is actually true.

The Wave 1 data now exists, so this is decidable rather than deferred. Note the decision is not automatically "turn it off": the earlier direct test showed the relevance path also ships more tokens than uniform (109,606 against 94,716 on the same input, because it stops as soon as the deficit is covered rather than compressing every candidate). If 3.1 makes it prefix-stable, relevance becomes strictly better than uniform. If 3.1 proves harder than expected, disabling it is the cheap fallback that wins on both axes today.

3.4 Identify the non-pruning rewrite source

Scoped by the measurement above, not speculative. Seven of 46 rewrites occurred with pruning disengaged, so something upstream of prune_context is editing conversation history between turns.

This is an investigation, not a fix: determine whether the client is compacting its own context, whether the tool-definition array changes between turns, or whether the system prompt is mutated. The probe already stores what is needed to find the turns; the question is what differs across them.

Sequence it after 3.1 so the pruning-caused rewrites are removed from the population first, leaving a clean residual to study. Attempting it now means diagnosing two overlapping causes at once.

Acceptance for Wave 3

The probe that promoted this wave is also its acceptance instrument, which is the point of having built it first:

  1. The rewritten share among pruned turns falls toward zero (39 of 232 today).
  2. The cache-rate gap between the two cohorts closes — 0.705 against 0.940 today. Re-measure both on a fuller sample before and after, per the re-measure discipline in Wave 2's acceptance; do not hardcode today's figures as the target.
  3. The tripwire test in 3.2 stays green on both compression paths.
  4. Any residual rewriting is attributed to a named cause by 3.4, not left as unexplained variance.

Wave 4 — Retries, and the half nobody guards

4.1 Denominate the iteration budget in prompt re-bills

iteration.py counts attempts. A retry on a 100k-token conversation is a 100k prompt re-bill to redo a ~400-token answer, and malformed escalates to the next-ranked candidate, which makes it a cold one. Tier 3 allows two.

Budget in re-bills instead: gate a retry on prompt size, and prefer same-model-with-more-tokens (which keeps the cache) over escalation (which does not) wherever the failure permits it. The existing failure taxonomy already supports this — truncated retries the same model by design; it is malformed that escalates.

4.2 Accept that structural verification is inert here

98.96% of structural verdicts are unverifiable (28,958 of 29,263), because agent turns end in tool calls and has_tool_calls short-circuits both checkers. That is correct behavior and the documented fix to a real false-failure incident. The conclusion not yet drawn is that the free checker now checks nothing on the only workload this router serves: 305 substantive verdicts out of 29,263.

Meanwhile a completion token costs ~201x a prompt token, so the expensive half of the ledger has no guard and the cheap half carries all the machinery.

Two actions, both small:

  • Give feedback.py a timer. It is the only loop without one — deploy/ ships timers for the poller, seed sweep, backup and offsite sync, and a baseline-report timer is installed. proficiency was written in one batch at 2026-09-10T00:44:55 and the 95 unapplied outcomes are exactly the models carrying today's traffic. Client outcomes are the only guard on wasted completions, and they are applied by hand.
  • Stop reporting the structural checker as a safety net in the docs, and either narrow it to the paths where it still fires or retire it.

Wave 5 — The reframe, and the expiry checks

Decisions, not tasks. Make them before adding machinery to the subsystems they touch.

5.1 The session is the routing unit

Waves 2 and 3 patch a per-request frame. The unit is wrong: the workload is a session of hundreds to thousands of turns sharing a monotonically growing prefix, and estimated_cost is a function of prompt_tokens, so the ranking's key input changes every turn even when nothing else does. CLAUDE.md documents the consequence as a feature — the winner at 50k differs from the winner at 120k — which inside a session is a cache dump.

If the session is the unit and a turn a delta, then: routing decides once and re-decides only on a threshold that includes the cache loss; classification becomes "has the task changed?" rather than a full taxonomy inference (96.6% of turns already reuse a cached label); exploration is naturally session-scoped; and routing.min_tool_proficiency becomes expressible, because "can this model be trusted with tools" is a session-level property, which is how the docs already describe it.

5.2 Recalibrate estimated_cost against billed rows

Depends on 1.1. Per-request estimate against bill spans 0.63x to 13.55x, a 21x spread in the error, which reorders candidates. 31,119 NeuralWatt rows carry a real billed figure joinable by request_id. A per-model correction factor over a trailing window, refreshed by the poller, fixes the ordering without a sweep.

The cause is already in CLAUDE.md two sections apart and never reconciled: attribution ratio spans 750x between models and is "most of the real cost difference in the catalog", while estimated_cost prices from catalog token prices, which carry no attribution term.

5.3 Premise-expiry checks

Four settings carry an explicit revisit condition in their own comment, all four conditions are met, and none re-fired: quality_tolerance: 0.1 ("narrow it as samples accumulate" — now 46 outcome-backed rows per category averaging 51.8 samples); cost-as-tiebreak ("all real traffic to date totals $0.07" — now $118.79); assumed_cache_rate: 0.917 (measured 2026-08-23, never re-measured); pinch.relevance ("off by default… then decide").

Every such comment should have its condition as a check that fails loudly — a test, a /metrics warning, a line in the poller. 1.4 is the first instance; generalize it.


Order of execution

wave items gate to proceed
1 1.1, 1.2, 1.3, 1.4 coverage >95% on OpenRouter cost/cached/duration
2 2.1, 2.2, 2.3 switch cache-rate gap narrows, or switches are deliberate
3 demoted — 3.3 only (a token-volume vs context-quality decision); 3.1/3.2 optional, 3.4 low no savings gate; payload rewriting was measured not to cost cache
4 4.1, 4.2 proficiency refreshing on a timer, unapplied backlog ~0
5 decisions —

1.1 first, alone, and measure for a day before anything else lands. It is small, it is reversible, and until it is in place every acceptance criterion below it is reading a 22%-covered sample.