Wave 2 shipped objective.incumbent_cache_pricing and objective.incumbent_challenger_cache_rate with no admin control. Nobody decided they should not have one; it never came up, and the plan had no step that would have made it come up. That matters most for a dial whose rationale is tuning from neutral to full penalty WITHOUT reverting code: if turning it means hand-editing a tracked file and restarting, the tuning loop is too slow to walk. Two halves, both small. plans/token-waste-waves.md gains a standing per-wave acceptance gate, placed ahead of Wave 1 so it is read before executing and checked before a wave is called done. The escape clause is load-bearing, not hedging -- objective.credit_attenuation.enabled is deliberately off the allowlist and off provider edits, because enabling it must be a config edit plus a restart. The rule is that the ABSENCE of a control is a decision someone made, not an oversight nobody noticed. tests/test_admin_knob_coverage.py enforces it, in the shape test_tui_schema_drift and test_tui_warnings already set here: covered is DERIVED from admin._CONFIG_ALLOWLIST and admin._BOOL_KNOBS rather than hand-copied, DELIBERATELY_NOT_IN_ADMIN carries reason strings rather than bare names, and every failure names the knob. Exactness is asserted in both directions, so the excuse list cannot rot into a rubber stamp as knobs quietly gain controls. Scope is the judgement call. RouterConfig has ~130 scalar leaves, and demanding a decision on all of them produces a baseline nobody reads -- which is the rubber stamp being guarded against. Two clauses cut it to 58: a section is in scope iff the portal already reaches it (where it reaches, it must reach completely), and deployment wiring -- endpoints, model ids, credentials, paths, devices -- is out, being configuration of where the router points rather than of how it behaves. 15 are covered today, 43 excused with reasons. The docstring draws the line and justifies it, including the classifier section, which is out because it has its own dedicated admin card rather than a generic allowlist entry. Two entries are marked PENDING feat/admin-incumbent-knobs: that branch is adding controls for exactly those two knobs, and this branch is based on origin/main where they do not exist yet. When it merges, the exactness test FAILS on both until the entries are deleted. That is deliberate -- the test announcing its own cleanup beats a stale excuse sitting here silently. No config knob added, removed or changed; src/admin.py untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
515 lines
24 KiB
Markdown
515 lines
24 KiB
Markdown
# Token waste: a five-wave plan
|
|
|
|
Status: in progress -- Wave 1 shipped, deployed and EVALUATED (PR #88; gate run
|
|
2026-09-14 on 1,288 decisions at full coverage); Wave 2 built and gated-passed
|
|
(PR #89, ships off, awaiting the enable decision); Wave 3 DEMOTED to correctness
|
|
-- its savings case was refuted by the same gate.
|
|
|
|
Measured read-only against the live `/home/alee/Sources/6krrt/router.db` on
|
|
2026-09-13 at repo state `26398d7`. Companion diagnosis:
|
|
`plans/ten-thousand-foot-review.md`.
|
|
|
|
**Intended executor: opencode (Atlas), directly.** Every item lands as its own
|
|
commit. Waves are ordered by dependency, not by size.
|
|
|
|
---
|
|
|
|
## The one-paragraph thesis
|
|
|
|
The provider's cache is worth ~92% of every prompt on this workload (measured
|
|
2026-09-14 at full coverage: **0.919** on same-model turns, n=1,227). The
|
|
router's only real action, choosing a model, is also the act that throws that
|
|
cache away — a switch drops the rate to **0.348** (n=55) and doubles the billed
|
|
cost per prompt token — and it takes that action without knowing the cache
|
|
exists. Everything below follows from that.
|
|
|
|
**Narrowed 2026-09-14.** This paragraph previously read "token waste here is
|
|
almost entirely cache destruction, not volume", which was too broad. Cache
|
|
destruction by **model switching** is real and is the leak. Cache destruction by
|
|
**editing the payload** is not: these providers tolerate a mid-prompt change
|
|
without invalidating the remainder, so the router's pruning does not cost cache
|
|
(see Wave 3, demoted on that evidence). The distinction matters because it moves
|
|
one whole wave out of the savings column.
|
|
|
|
## Read this before executing: what the data supports, and what it does not
|
|
|
|
Two claims were wrong during analysis and are corrected here so nothing gets
|
|
built on them. This repo has a history of scoring the rig instead of the model;
|
|
the same discipline applies to this plan.
|
|
|
|
**CONFIRMED — model switching destroys the cache.** Re-measured 2026-09-14 on
|
|
the post-Wave-1 window at ~100% cached-token coverage, which supersedes the
|
|
thin 12-22%-coverage figures this section first carried:
|
|
|
|
| turn | n | cache rate | µ$ / prompt token |
|
|
|---|---|---|---|
|
|
| same model as previous | 1,227 | **0.919** | 0.0369 |
|
|
| switched model | 55 | **0.348** | 0.0734 |
|
|
| first in session | 5 | 0.000 | 0.0403 |
|
|
|
|
The gap **widened** on the honest sample (0.571 against the thin sample's
|
|
0.474) while the cost ratio settled to **1.99x** from 2.46x. Switch turns are
|
|
4.3% of traffic. Two independent measurements — cache rate and billed cost per
|
|
prompt token — agree in direction and magnitude. This is the real leak, and it
|
|
is the only one that survived re-measurement.
|
|
|
|
**REFUTED 2026-09-14 — payload rewriting does NOT cost cache.** The mechanism
|
|
is real: a direct test of `prune_context` showed the relevance path diverging at
|
|
message index 19 of 82 with a literally identical relevance order, where the
|
|
uniform path diverges nowhere. A growing `target_save` walks a relevance-ordered
|
|
list, so the newly-compressed candidate lands at an arbitrary *position*.
|
|
|
|
But the inference drawn from it was wrong. At full coverage, turns with **half
|
|
the prompt sitting after the divergence point still cache at 0.910** — which a
|
|
strict prefix cache cannot do. These providers tolerate a mid-prompt edit
|
|
without invalidating the remainder. Wave 3 is demoted on that evidence; see it
|
|
for the full tables.
|
|
|
|
**This entry has now been wrong in both directions, which is the useful part.**
|
|
It first read NOT CONFIRMED on a 0.924 aggregate — one number averaged across
|
|
two cohorts at 12-22% coverage, too coarse to separate them. It was then flipped
|
|
to CONFIRMED on a cohort split from a single hour (n=46, 0.705 vs 0.940). The
|
|
full window (n=410 rewritten) put the gap at 0.013 and the mechanism test
|
|
removed it entirely.
|
|
|
|
Three readings, two reversals, one underlying question. What distinguished the
|
|
final answer was not more data but a **mechanism test** — asking whether cache
|
|
rate falls as the share of the prompt after the divergence rises, rather than
|
|
comparing two buckets and trusting the difference. A cohort split can only tell
|
|
you two groups differ; it cannot tell you why, and it will happily report noise
|
|
as signal at small n.
|
|
|
|
The lesson generalises past this entry: an aggregate that mixes two populations
|
|
will report the majority's number and hide the minority's, and "the data does
|
|
not show it biting" is a claim about resolution as much as about reality.
|
|
|
|
---
|
|
|
|
## Standing acceptance gate — applies to every wave
|
|
|
|
In addition to each wave's own criteria below:
|
|
|
|
> **Every config knob a wave introduces ships with an admin control — runtime,
|
|
> persisted, or both as its mechanism warrants — or a recorded decision saying
|
|
> why it must not.**
|
|
|
|
The escape clause is load-bearing, not hedging.
|
|
`objective.credit_attenuation.enabled` is deliberately absent from the admin
|
|
allowlist and from provider edits, because turning it on must be a config edit
|
|
plus a restart (CLAUDE.md, "Routing notes"). So the rule is not "everything must
|
|
have a control" — it is that the ABSENCE of one is a decision someone made,
|
|
rather than an oversight nobody noticed.
|
|
|
|
Wave 2 is why this is written down: `objective.incumbent_cache_pricing` and
|
|
`objective.incumbent_challenger_cache_rate` shipped with no control and no
|
|
decision, and a dial whose whole point is tuning from neutral to full penalty
|
|
without reverting code is worth little if turning it means hand-editing a
|
|
tracked file and restarting the service.
|
|
|
|
`tests/test_admin_knob_coverage.py` enforces this for the sections the portal
|
|
already reaches, and fails naming the knob.
|
|
|
|
---
|
|
|
|
## Wave 1 — See the leak
|
|
|
|
No behavior change. Every later wave's acceptance criteria read what this
|
|
produces, so nothing else should start first.
|
|
|
|
Commit `1bd8c6e` (2026-09-10) already landed streamed cost capture and
|
|
`cached_prompt_tokens` parsing at `dispatcher.py:2056`, which is why data
|
|
begins on 09-11. The instrument exists; it is under-covered.
|
|
|
|
### 1.1 OpenRouter usage opt-in
|
|
|
|
Coverage since the instrument landed:
|
|
|
|
| provider | obs | cost | cached tokens | duration |
|
|
|---|---|---|---|---|
|
|
| neuralwatt | 805 | 804 (99.9%) | 103 (12.8%) | 804 (99.9%) |
|
|
| openrouter | 734 | 165 (22.5%) | 164 (22.3%) | **0 (0%)** |
|
|
|
|
OpenRouter's cost coverage (22.5%) and cached coverage (22.3%) match almost
|
|
exactly. That is **one** gap, not two: the usage block is simply absent on ~78%
|
|
of requests. `dispatcher.py:4254` sets
|
|
`stream_options: {"include_usage": True}`, which is the OpenAI spelling;
|
|
OpenRouter requires its own `usage: {"include": true}` in the request body to
|
|
return accounting. Add it per-provider — the `reports_cost_in_usage` flag from
|
|
`1bd8c6e` is the right place to hang it.
|
|
|
|
This single change is the highest-leverage item in the plan: it takes OpenRouter
|
|
from ~22% to ~100% on cost, cached tokens **and** latency at once, on the
|
|
provider carrying ~73% of decisions.
|
|
|
|
**Verify**: coverage for all three fields above 95% on OpenRouter rows written
|
|
after the change.
|
|
|
|
### 1.2 Decide what an absent `cached_tokens` means
|
|
|
|
NeuralWatt reports cost on 99.9% of rows but `prompt_tokens_details` on only
|
|
12.8%. Determine whether the field is omitted on a full cache miss or is
|
|
model-dependent. If omitted on a miss, record an explicit `0` rather than NULL,
|
|
because every cache-rate reading in this plan is otherwise conditioned on a hit
|
|
having occurred. Nine rows currently store `0`, so zeros do reach the DB
|
|
sometimes — that needs explaining before the 0.924 figure can be trusted.
|
|
|
|
**Verify**: a test pinning the parse for a usage block with `cached_tokens: 0`,
|
|
one with the key absent, and one with `prompt_tokens_details` absent entirely,
|
|
asserting the three map to distinct stored values.
|
|
|
|
### 1.3 Prefix-stability probe
|
|
|
|
The decisive instrument for Wave 3, and about twenty lines. After
|
|
`prune_context` returns, hash the pruned payload cumulatively by message and
|
|
store the hash of the longest stable prefix (or simply a per-turn list of
|
|
message hashes) on `route_decisions`. Within a session, compare consecutive
|
|
turns to get the real first-divergence index and the tokens after it.
|
|
|
|
That converts Wave 3 from an argument into a measurement, and it will either
|
|
justify Wave 3 or retire it.
|
|
|
|
**Verify**: on a replayed synthetic session, the probe reports divergence at the
|
|
index the direct test predicts.
|
|
|
|
### 1.4 Cache rate in `/metrics`
|
|
|
|
A cache-rate series per `(provider, model)` over a trailing window, plus a
|
|
warning when a session's rate falls below a configurable floor. Follow the
|
|
existing detector conventions in `metrics.py` — novelty-or-rate, not bare
|
|
presence, per the reactive rejection detector.
|
|
|
|
This is also the premise-expiry check for `assumed_cache_rate: 0.917`: the
|
|
constant was measured 2026-08-23 and nothing has re-measured it since.
|
|
|
|
---
|
|
|
|
## Wave 2 — Stop the confirmed leak
|
|
|
|
The router has no incumbent. `rank_candidates` does not know what ran last
|
|
turn, so there is no hysteresis and no switching cost. Three sub-items, and the
|
|
first is most of the work.
|
|
|
|
### 2.1 Thread the incumbent into ranking
|
|
|
|
The session's last `selected_model` is already on disk in `route_decisions` and
|
|
usually in memory. Pass it into `rank_candidates` as the incumbent.
|
|
|
|
Keep the existing shape: quality band first, cost as tiebreak. The incumbent
|
|
changes only the cost key.
|
|
|
|
### 2.2 Price the cache loss
|
|
|
|
In the tiebreak, price the incumbent at the measured cache rate and every
|
|
challenger at a cold rate. Concretely, `estimated_cost` already takes
|
|
`cache_rate`; pass `cfg.objective.assumed_cache_rate` for the incumbent and
|
|
`0.0` for challengers. A challenger then has to beat the incumbent by more than
|
|
the cache it is about to discard, which on a 100k prompt is most of the prompt.
|
|
|
|
This is the first time the router's own decision becomes an input to its own
|
|
cost model, and it is why it belongs before the deeper reframe in Wave 5.
|
|
|
|
Two guards so this cannot become stickiness-at-any-cost:
|
|
- the quality band is computed **before** the cost key, exactly as today, so a
|
|
genuine quality gap still wins outright and the incumbent gets no quality
|
|
advantage;
|
|
- a hard-filter failure (context ceiling, capability gate, circuit breaker,
|
|
profile allowlist) still removes the incumbent unconditionally.
|
|
|
|
### 2.3 Exploration becomes session-scoped
|
|
|
|
Epsilon-greedy prices a suboptimal pull as the arm's cost difference. Here
|
|
exploring *means* switching, so the real cost is that difference plus a cold
|
|
prompt — and `exploration.py` cannot see it. Move the coin flip to session
|
|
start rather than per turn: one exploratory session costs one cold prompt
|
|
instead of one per turn, and it produces a cleaner outcome signal because the
|
|
whole session is attributable to the explored model.
|
|
|
|
`exploration.py` takes an injected RNG and holds no mutable state, so this is a
|
|
call-site change, not a rewrite.
|
|
|
|
**Wave 2 acceptance**, read off Wave 1's instruments:
|
|
- switch rate per session falls;
|
|
- the same-model / switched cache-rate gap (0.952 vs 0.478) narrows, or the
|
|
remaining switches are deliberate — driven by a quality band or a hard
|
|
filter, not by a cost re-rank;
|
|
- median billed µ$ per prompt token moves toward the same-model figure.
|
|
|
|
---
|
|
|
|
## Wave 3 — Make prefix stability guaranteed rather than incidental
|
|
|
|
**DEMOTED 2026-09-14, back to correctness rather than savings.** This section
|
|
was promoted on 2026-09-13 on the probe's first hour of data, which showed
|
|
rewritten turns at 0.705 cache rate against 0.940 for turns that merely grew.
|
|
That promotion carried an explicit caveat — "n=46 from roughly one hour, do not
|
|
treat 0.705 as settled, re-measure on a fuller sample." The re-measurement is
|
|
in, and the caveat is the part that held.
|
|
|
|
### What the full window actually measured
|
|
|
|
1,288 decisions over ~5 hours, at ~100% cached-token coverage, joined on
|
|
`cached_tokens_source = 'reported'`:
|
|
|
|
| prefix state | pruning | n | avg tokens after divergence | cache rate |
|
|
|---|---|---|---|---|
|
|
| grew only | pruned | 854 | 702 | 0.904 |
|
|
| **rewritten** | pruned | 387 | 32,901 | **0.891** |
|
|
| grew only | under budget | 14 | 6,322 | 0.765 |
|
|
| rewritten | under budget | 23 | 9,075 | 0.548 |
|
|
|
|
Among pruned turns the gap is **0.013**, not the 0.235 the first hour showed.
|
|
|
|
### The mechanism test, which is what settles it
|
|
|
|
If a mid-prompt edit invalidated everything after it, cache rate would fall as
|
|
the share of the prompt sitting after the divergence point rises. It does not:
|
|
|
|
| tokens after divergence | n | share of prompt | cache rate |
|
|
|---|---|---|---|
|
|
| <5k | 24 | 7% | 0.813 |
|
|
| 5-20k | 89 | 28% | 0.888 |
|
|
| 20-50k | 227 | **50%** | **0.910** |
|
|
| >50k | 66 | 50% | 0.835 |
|
|
|
|
No monotonic relationship. Turns with **half the prompt after the divergence
|
|
point** still cache at 0.910 — arithmetically impossible under a strict prefix
|
|
cache, which would cap that bucket near 0.50.
|
|
|
|
**Conclusion: these providers do not use a strict prefix cache.** A mid-prompt
|
|
edit does not invalidate the remainder. That assumption was load-bearing under
|
|
this wave's savings case and under the reading of the `prune_context`
|
|
simulation, and it is wrong.
|
|
|
|
The simulation itself was never wrong about what it measured — the relevance
|
|
path really does rewrite the payload at an arbitrary position, and the uniform
|
|
path really does not. What was wrong was the inference that a rewritten payload
|
|
costs cache. It measurably does not.
|
|
|
|
**Gate 3 is untouched by this.** Switching models is not a mid-prompt edit; it
|
|
is a different cache namespace entirely, and it still measures 0.919 against
|
|
0.348. Nothing here weakens Wave 2.
|
|
|
|
### What this changes for the sub-items
|
|
|
|
- **3.1 (monotone region) is now optional and probably not worth it.** It buys
|
|
no measurable cache. It would buy determinism in a working path, at the cost
|
|
of editing that path. Do not do it for savings; there are none.
|
|
- **3.2 (tripwire test)** as originally specced asserts prefix stability, which
|
|
the relevance path violates by design. Without 3.1 it would ship red. Either
|
|
it follows 3.1 or it is dropped — it is not independently landable.
|
|
- **3.3 (config comment reconciliation) is still worth doing**, and is now a
|
|
clean decision. With cache out of the picture the only difference between the
|
|
paths is token volume against context quality: the direct test measured the
|
|
relevance path shipping 109,606 tokens against uniform's 94,716 on the same
|
|
input, because it stops as soon as the deficit is covered. So relevance costs
|
|
~15k more tokens per turn and buys keeping the most relevant content verbatim.
|
|
That is a real tradeoff, just not a cache one — decide it on its own terms.
|
|
- **3.4 (non-pruning rewrite source) drops to low priority.** The under-budget
|
|
rewritten cohort (n=23, 0.548) is small and its low rate is better explained
|
|
by session-opening turns than by a distinct leak.
|
|
|
|
### An opportunity this created
|
|
|
|
If a mid-prompt edit costs no cache, pruning is **cheaper than assumed** and
|
|
could be more aggressive without a cache penalty. Pruned turns currently run
|
|
132k -> 89k (a 33% cut); `pinch.budget_tokens` could go lower for real prompt-
|
|
side savings. The binding constraint is answer quality, not cache. That is a
|
|
new, evidence-created item and it belongs in Wave 4 or 5, not here.
|
|
|
|
### A second rewrite source exists, outside pruning
|
|
|
|
Cross-referencing the same rows against whether pruning actually fired:
|
|
|
|
| prefix state | pruning fired | under budget |
|
|
|---|---|---|
|
|
| grew only | 193 | 6 |
|
|
| **rewritten** | **39** | **7** |
|
|
|
|
Pruning explains 39 of 46 rewrites. **Seven turns rewrote the prefix with
|
|
pruning never engaged at all** — under budget, so `prune_context` returned the
|
|
messages untouched. On those turns the payload is the client's own message list,
|
|
which means something upstream of the router edited its own history (opencode
|
|
compaction, a changed tool-definition array, or a mutated system prompt are the
|
|
candidates).
|
|
|
|
Two consequences, and both matter for how this wave is judged:
|
|
|
|
- **3.1 cannot fix all of it.** Fixing the relevance path addresses at most 39
|
|
of 46. Crediting Wave 3 with the whole gap would overstate it.
|
|
- **The residual is worth identifying before it is assumed benign.** Among
|
|
under-budget turns the rewrite rate is 7 of 13 — proportionally *higher* than
|
|
the pruned cohort's 39 of 232, though on a sample far too small to lean on.
|
|
If the client rewrites its own history routinely, that is a larger cache
|
|
leak than pruning and the router cannot fix it by changing `prune_context`.
|
|
|
|
### 3.1 Constrain relevance compression to a monotone region
|
|
|
|
Keep the relevance *ranking* and change what it decides. Instead of compressing
|
|
a scattered relevance-ordered subset, let relevance pick **where the compressed
|
|
boundary sits** on first crossing, then only ever extend that boundary forward.
|
|
Once a message is compressed it is never un-compressed.
|
|
|
|
That preserves the feature's intent — relevance still decides what is worth
|
|
keeping verbatim — while restoring the property the uniform path has for free:
|
|
appends cannot change a byte before the boundary.
|
|
|
|
### 3.2 A prefix-stability tripwire test
|
|
|
|
This repo already writes exactly this kind of test
|
|
(`test_tui_schema_drift.py`, `test_tui_warnings.py`). Assert that across a
|
|
simulated multi-turn session, `prune_context`'s output for turn N+1 is
|
|
byte-identical to turn N's up to the appended messages. Fail naming the
|
|
divergent message index.
|
|
|
|
The uniform path passes this today; the relevance path does not. That asymmetry
|
|
is the whole point of the test.
|
|
|
|
### 3.3 Reconcile the config with reality
|
|
|
|
`pinch.relevance.enabled: true`, while the comment directly above it still
|
|
reads *"Off by default… ship it, watch route_decisions / pinch stats on real
|
|
traffic, then decide the default."* It was switched on and the comment never
|
|
followed. Same shape as `session_cache` below it. Decide the default on Wave 1
|
|
data and rewrite both comments to say what is actually true.
|
|
|
|
The Wave 1 data now exists, so this is decidable rather than deferred. Note the
|
|
decision is not automatically "turn it off": the earlier direct test showed the
|
|
relevance path also ships **more** tokens than uniform (109,606 against 94,716
|
|
on the same input, because it stops as soon as the deficit is covered rather
|
|
than compressing every candidate). If 3.1 makes it prefix-stable, relevance
|
|
becomes strictly better than uniform. If 3.1 proves harder than expected,
|
|
disabling it is the cheap fallback that wins on both axes today.
|
|
|
|
### 3.4 Identify the non-pruning rewrite source
|
|
|
|
Scoped by the measurement above, not speculative. Seven of 46 rewrites occurred
|
|
with pruning disengaged, so something upstream of `prune_context` is editing
|
|
conversation history between turns.
|
|
|
|
This is an investigation, not a fix: determine whether the client is compacting
|
|
its own context, whether the tool-definition array changes between turns, or
|
|
whether the system prompt is mutated. The probe already stores what is needed
|
|
to find the turns; the question is what differs across them.
|
|
|
|
Sequence it **after 3.1** so the pruning-caused rewrites are removed from the
|
|
population first, leaving a clean residual to study. Attempting it now means
|
|
diagnosing two overlapping causes at once.
|
|
|
|
### Acceptance for Wave 3
|
|
|
|
The probe that promoted this wave is also its acceptance instrument, which is
|
|
the point of having built it first:
|
|
|
|
1. The **rewritten share** among pruned turns falls toward zero (39 of 232
|
|
today).
|
|
2. The **cache-rate gap between the two cohorts closes** — 0.705 against 0.940
|
|
today. Re-measure both on a fuller sample before and after, per the
|
|
re-measure discipline in Wave 2's acceptance; do not hardcode today's
|
|
figures as the target.
|
|
3. The tripwire test in 3.2 stays green on both compression paths.
|
|
4. Any residual rewriting is attributed to a named cause by 3.4, not left as
|
|
unexplained variance.
|
|
|
|
---
|
|
|
|
## Wave 4 — Retries, and the half nobody guards
|
|
|
|
### 4.1 Denominate the iteration budget in prompt re-bills
|
|
|
|
`iteration.py` counts attempts. A retry on a 100k-token conversation is a 100k
|
|
prompt re-bill to redo a ~400-token answer, and `malformed` escalates to the
|
|
*next-ranked candidate*, which makes it a cold one. Tier 3 allows two.
|
|
|
|
Budget in re-bills instead: gate a retry on prompt size, and prefer
|
|
same-model-with-more-tokens (which keeps the cache) over escalation (which does
|
|
not) wherever the failure permits it. The existing failure taxonomy already
|
|
supports this — `truncated` retries the same model by design; it is `malformed`
|
|
that escalates.
|
|
|
|
### 4.2 Accept that structural verification is inert here
|
|
|
|
98.96% of structural verdicts are `unverifiable` (28,958 of 29,263), because
|
|
agent turns end in tool calls and `has_tool_calls` short-circuits both
|
|
checkers. That is correct behavior and the documented fix to a real
|
|
false-failure incident. The conclusion not yet drawn is that the free checker
|
|
now checks nothing on the only workload this router serves: 305 substantive
|
|
verdicts out of 29,263.
|
|
|
|
Meanwhile a completion token costs ~201x a prompt token, so the expensive half
|
|
of the ledger has no guard and the cheap half carries all the machinery.
|
|
|
|
Two actions, both small:
|
|
- **Give `feedback.py` a timer.** It is the only loop without one — `deploy/`
|
|
ships timers for the poller, seed sweep, backup and offsite sync, and a
|
|
baseline-report timer is installed. `proficiency` was written in one batch at
|
|
2026-09-10T00:44:55 and the 95 unapplied outcomes are exactly the models
|
|
carrying today's traffic. Client outcomes are the only guard on wasted
|
|
completions, and they are applied by hand.
|
|
- Stop reporting the structural checker as a safety net in the docs, and either
|
|
narrow it to the paths where it still fires or retire it.
|
|
|
|
---
|
|
|
|
## Wave 5 — The reframe, and the expiry checks
|
|
|
|
Decisions, not tasks. Make them before adding machinery to the subsystems they
|
|
touch.
|
|
|
|
### 5.1 The session is the routing unit
|
|
|
|
Waves 2 and 3 patch a per-request frame. The unit is wrong: the workload is a
|
|
session of hundreds to thousands of turns sharing a monotonically growing
|
|
prefix, and `estimated_cost` is a function of `prompt_tokens`, so the ranking's
|
|
key input changes every turn even when nothing else does. `CLAUDE.md` documents
|
|
the consequence as a feature — the winner at 50k differs from the winner at
|
|
120k — which inside a session is a cache dump.
|
|
|
|
If the session is the unit and a turn a delta, then: routing decides once and
|
|
re-decides only on a threshold that includes the cache loss; classification
|
|
becomes "has the task changed?" rather than a full taxonomy inference (96.6% of
|
|
turns already reuse a cached label); exploration is naturally session-scoped;
|
|
and `routing.min_tool_proficiency` becomes expressible, because "can this model
|
|
be trusted with tools" is a session-level property, which is how the docs
|
|
already describe it.
|
|
|
|
### 5.2 Recalibrate `estimated_cost` against billed rows
|
|
|
|
Depends on 1.1. Per-request estimate against bill spans **0.63x to 13.55x**, a
|
|
21x spread in the error, which reorders candidates. 31,119 NeuralWatt rows carry
|
|
a real billed figure joinable by `request_id`. A per-model correction factor over
|
|
a trailing window, refreshed by the poller, fixes the ordering without a sweep.
|
|
|
|
The cause is already in `CLAUDE.md` two sections apart and never reconciled:
|
|
attribution ratio spans 750x between models and is "most of the real cost
|
|
difference in the catalog", while `estimated_cost` prices from catalog token
|
|
prices, which carry no attribution term.
|
|
|
|
### 5.3 Premise-expiry checks
|
|
|
|
Four settings carry an explicit revisit condition in their own comment, all four
|
|
conditions are met, and none re-fired: `quality_tolerance: 0.1` ("narrow it as
|
|
samples accumulate" — now 46 outcome-backed rows per category averaging 51.8
|
|
samples); cost-as-tiebreak ("all real traffic to date totals $0.07" — now
|
|
$118.79); `assumed_cache_rate: 0.917` (measured 2026-08-23, never re-measured);
|
|
`pinch.relevance` ("off by default… then decide").
|
|
|
|
Every such comment should have its condition as a check that fails loudly — a
|
|
test, a `/metrics` warning, a line in the poller. 1.4 is the first instance;
|
|
generalize it.
|
|
|
|
---
|
|
|
|
## Order of execution
|
|
|
|
| wave | items | gate to proceed |
|
|
|---|---|---|
|
|
| 1 | 1.1, 1.2, 1.3, 1.4 | coverage >95% on OpenRouter cost/cached/duration |
|
|
| 2 | 2.1, 2.2, 2.3 | switch cache-rate gap narrows, or switches are deliberate |
|
|
| 3 | **demoted** — 3.3 only (a token-volume vs context-quality decision); 3.1/3.2 optional, 3.4 low | no savings gate; payload rewriting was measured not to cost cache |
|
|
| 4 | 4.1, 4.2 | `proficiency` refreshing on a timer, unapplied backlog ~0 |
|
|
| 5 | decisions | — |
|
|
|
|
1.1 first, alone, and measure for a day before anything else lands. It is small,
|
|
it is reversible, and until it is in place every acceptance criterion below it is
|
|
reading a 22%-covered sample.
|