plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
655 lines
30 KiB
Markdown
655 lines
30 KiB
Markdown
# Conceptual review: is the premise sound, and is execution converging on it?
|
|
|
|
Status: done -- review of shipped work
|
|
|
|
**Date:** 2026-09-01
|
|
**Scope:** premise and architecture, not bugs. Every number below is measured
|
|
from the live `router.db` and the checked-in code on this date; the scripts are
|
|
inlined so they can be re-run.
|
|
**Verdict:** the premise is sound and the hard part is built. Three structural
|
|
problems sit between the current state and the stated goal, and one of them
|
|
will actively move routing the wrong way the moment the next obvious step is
|
|
taken.
|
|
|
|
---
|
|
|
|
## 0. What the evidence base is
|
|
|
|
| | |
|
|
|---|---|
|
|
| `route_decisions` | 9,025 rows (9,007 `chat`), 2026-08-24 → 2026-09-01 |
|
|
| `energy_observations` | 13,966 rows, **$49.94** billed, **6.51 kWh**, 1,122 gCO₂eq, from 2026-08-17 |
|
|
| `verifications` | 13,501 rows; 987 of them client outcomes (816 succeeded / 171 failed) |
|
|
| `proficiency` | 141 (model, category) rows |
|
|
| tests | **843 pass in 36.8s**, all offline |
|
|
|
|
This is a real deployment with real traffic, not a demo. The review is worth
|
|
doing precisely because there is now enough data to check the design against
|
|
itself.
|
|
|
|
---
|
|
|
|
## 1. The premise is sound, and the hardest claim in it is now proven
|
|
|
|
Design doc §1: use a local model to classify the task, then route to the
|
|
cheapest/best-fit open-weight model instead of defaulting everything to one
|
|
expensive model.
|
|
|
|
The load-bearing, non-obvious claim in that sentence is that **task category
|
|
should change the model**. For a long stretch it demonstrably did not — the
|
|
classifier computed a category, the router paid ~10s for it, and
|
|
`proficiency` was empty so it changed nothing. That is fixed, and it is
|
|
measurable now. Replaying all 9,001 real decisions with per-category
|
|
proficiency against a category-agnostic mean:
|
|
|
|
```
|
|
category-aware == category-agnostic: 5,917/9,001 (65.7%)
|
|
category changed the winner: 3,084/9,001 (34.3%)
|
|
```
|
|
|
|
**One decision in three turns on the category.** That is the design's central
|
|
bet paying off, and it took populating `proficiency`, fixing four harness bugs,
|
|
and rebuilding the cost axis to get there. It should be stated as a result.
|
|
|
|
Two more things are working and deserve to be said plainly before the
|
|
criticism:
|
|
|
|
- **The session cache removed the latency tax.** It landed 2026-08-30 and is
|
|
running at 98%+ hit rate (2026-09-01: 1,402 cached / 28 fresh). The
|
|
"~10s of local overhead on every message" problem in CLAUDE.md is solved for
|
|
the sessions that matter. Fresh classifications average 2,297ms (max 66,319ms
|
|
— one outlier worth a look, but not a design issue).
|
|
- **The pure-module discipline is the best thing in the codebase.**
|
|
`routing.py`, `scoring.py`, `tiering.py`, `proficiency.py`, `context_prune.py`
|
|
are I/O-free, take thresholds as arguments, and are exhaustively tested. That
|
|
is why 843 tests run offline in 37 seconds, and it is why every one of the
|
|
measurement-driven reversals in CLAUDE.md was cheap to make. Keep it.
|
|
|
|
---
|
|
|
|
## 2. The premise quietly changed, and the docs did not follow
|
|
|
|
Design §7 names the metric: **"$ saved vs. always-frontier baseline"**, with
|
|
the warning that "a router that's cheaper but quietly worse isn't a win."
|
|
Nobody has checked the other side of that: a router that's *better but quietly
|
|
more expensive* isn't obviously a win either, and that is where this landed.
|
|
|
|
`baseline_report.py` already tells the story, and it appears not to have been
|
|
read recently:
|
|
|
|
```
|
|
$ PYTHONPATH=src python -m baseline_report --since 2026-08-24
|
|
decisions: 9018
|
|
total cost: actual $113.89 cheapest $51.39 best-prof $166.59
|
|
dominance (selected == cheapest): 4,972/9,018 (55.1%)
|
|
mean proficiency: actual 0.991 cheapest 0.788 best-prof 0.998
|
|
```
|
|
|
|
The router sits at **0.991 of a maximum 0.998** on quality and **2.2x the
|
|
cheapest** on cost. It is not a cost/quality tradeoff engine; it is a
|
|
quality-maximizer that takes a discount when quality is exactly tied. That may
|
|
be what you want — but it is not what §1 says, and it is not what the README
|
|
says either ("dispatches to the cheapest/best-fit model").
|
|
|
|
The comparison that matters more is against the alternative you would actually
|
|
have used. Replaying all 9,007 chat decisions against every single-model policy
|
|
(same request shapes, same catalog prices, `assumed_cache_rate` 0.917;
|
|
`infeasible` = requests exceeding that model's effective window):
|
|
|
|
| single-model policy | est. cost | vs router | infeasible |
|
|
|---|---|---|---|
|
|
| `gemma-4-31b` | $18.91 | 0.17x | 1,170 |
|
|
| `qwen3.6-35b` | $19.54 | 0.17x | 4,032 |
|
|
| **`deepseek-v4-flash`** | **$36.59** | **0.32x** | **0** |
|
|
| `kimi-k2.7-code` | $136.41 | 1.20x | 963 |
|
|
| `glm-5.2-fast` / `glm-5.3` | $260.20 | 2.28x | 0 |
|
|
| `kimi-k3` | $563.97 | 4.95x | 0 |
|
|
| *(actual routed)* | *$113.89* | *1.00x* | *0* |
|
|
|
|
Against always-frontier (`kimi-k3`) the router saves **4.95x** — the design's
|
|
own metric, comfortably met. Against "just always use `deepseek-v4-flash`" it
|
|
costs **3.1x more**, and deepseek is the only other policy that can serve 100%
|
|
of the traffic without a context failure.
|
|
|
|
Which baseline is honest depends entirely on what you would otherwise have
|
|
done. For opencode agent traffic against this catalog, the realistic
|
|
counterfactual is not `kimi-k3` — it is deepseek. **The router should report
|
|
both baselines, and `baseline_report.py` should grow a single-model column.**
|
|
|
|
The operational consequence is live right now: **6.51 kWh burned against a
|
|
`plan_kwh_per_period` of 6.25.** At roughly $7.67/kWh realized, an
|
|
always-deepseek policy would have used about a third of that and stayed inside
|
|
the allowance.
|
|
|
|
---
|
|
|
|
## 3. Finding #1 — the feedback loop rewards models for not being used
|
|
|
|
**This is the most important item in the review, and it is a reason not to
|
|
take the next obvious step until it is fixed.**
|
|
|
|
987 client outcomes are sitting in `verifications`; 961 of them unapplied.
|
|
`POST /outcome` is correctly identified in CLAUDE.md as "the only ground
|
|
truth," so folding them in looks like the highest-value pending action. It is
|
|
not, as currently built.
|
|
|
|
Simulated on a copy of the DB — run `feedback.py` for real, then replay all
|
|
9,007 decisions through `select_candidates` + `rank_candidates`:
|
|
|
|
```
|
|
BEFORE est cost $137.76 AFTER est cost $179.28 (1.30x)
|
|
kimi-k2.7-code 2,722 glm-5.2-fast 2,243
|
|
qwen3.6-35b 2,646 qwen3.6-35b-fast 1,385
|
|
qwen3.6-35b-fast 1,236 glm-5.3 1,318
|
|
glm-5.3 745 kimi-k2.7-code 1,239
|
|
gemma-4-31b 587 qwen3.6-35b 1,150
|
|
deepseek-v4-flash 500 kimi-k2.7-code-fast 947
|
|
glm-5.2-fast 417 deepseek-v4-flash 596
|
|
```
|
|
|
|
Ingesting ground truth makes routing **30% more expensive** and hands 2,243
|
|
decisions to `glm-5.2-fast`, one of the priciest common rows. The mechanism is
|
|
visible in the scores. `coding_refactor`, before → after:
|
|
|
|
| model | before | n | after | n | real outcomes folded in |
|
|
|---|---|---|---|---|---|
|
|
| `glm-5.2-fast` | 0.994 | 36 | **0.994** | 36 | **zero** |
|
|
| `kimi-k2.7-code` | 1.000 | 18 | 0.813 | 80 | 62 |
|
|
| `deepseek-v4-flash` | 0.826 | 36 | 0.791 | 135 | 99 |
|
|
| `qwen3.6-35b` | 0.900 | 18 | 0.641 | 44 | 26 |
|
|
|
|
**`glm-5.2-fast` wins the category outright by never having been measured on
|
|
real traffic.**
|
|
|
|
The cause is structural, not a bug: `feedback.py` and `eval_proficiency.py`
|
|
both write through `proficiency_store.add_self_eval`, into the *same*
|
|
`self_eval_score` running mean, with equal weight per sample. Benchmark tasks
|
|
score ~1.00. Real agent turns score ~0.80. So every real sample drags a score
|
|
down, and the drag is proportional to how much the model has been used. The
|
|
router then routes away from the used model toward the unused one, gathers
|
|
outcomes on *that* one, penalizes it, and moves on.
|
|
|
|
That is a rotation, not a convergence. And because exposure has been
|
|
concentrated on the cheap models — they win the cost tiebreak — the rotation is
|
|
systematically toward expensive ones.
|
|
|
|
**Directions (pick one; all are cheap):**
|
|
|
|
1. **Keep the two sources separate.** Real-traffic outcomes are a different
|
|
measurement from benchmark tasks and should not share a mean. Add
|
|
`outcome_score` / `outcome_samples` columns and blend them explicitly, the
|
|
way `leaderboard` and `self_eval` already are. This also makes the
|
|
~0.80-vs-1.00 scale difference visible instead of silently mixed.
|
|
2. **Shrink toward the prior by sample size.** A score with n=36 and no real
|
|
exposure should not outrank one with n=135 that includes 99 real outcomes.
|
|
Today low sample count is an *advantage*; `self_eval_min_samples` gates the
|
|
leaderboard blend but nothing penalizes thin evidence in the ranking itself.
|
|
3. **Compare like with like.** Score a model against the *per-category mean*
|
|
outcome rate rather than absolute, so an 82% pass rate in a category where
|
|
everyone scores 80% is neutral, not a 0.18 penalty.
|
|
|
|
Until one of these lands, **do not run `feedback.py` on the pending backlog.**
|
|
The 961 rows are the most valuable data this project has; spending them through
|
|
the current path converts them into a more expensive router.
|
|
|
|
---
|
|
|
|
## 4. Finding #2 — exclusions are self-sealing; there is no exploration
|
|
|
|
`deepseek-v4-flash` scored **0.333 on `tool_use_agentic`** across **3**
|
|
benchmark tasks. Blended, it now reads 0.5. With `quality_tolerance: 0.1` that
|
|
puts it in band 5 while eleven models sit in band 0, so it loses every
|
|
`tool_use_agentic` ranking outright.
|
|
|
|
Consequence, measured: **deepseek won 0 of 2,512 `tool_use_agentic` decisions,
|
|
and has 0 client outcomes in that category.** A 3-sample estimate has
|
|
permanently removed the cheapest capable model from the largest category of
|
|
traffic, and the design contains no path by which that estimate can ever be
|
|
revised.
|
|
|
|
`config/config.yaml` proposes the experiment that would settle it:
|
|
|
|
> run with it off, let `POST /outcome` report real pass/fail, and compare
|
|
> `tool_use_agentic` proficiency for deepseek before and after
|
|
|
|
**That experiment cannot run.** `min_tool_proficiency` is already `null` — the
|
|
filter is off — and it changes nothing, because the *quality band* excludes
|
|
deepseek from the category regardless. The config comment says "with it off,
|
|
that advantage applies." It does not.
|
|
|
|
The general shape: this is a contextual bandit running pure exploitation. Once
|
|
a model wins a (category, context-size) cell it wins it forever; alternatives
|
|
never accumulate the evidence that would overturn the ranking. The consequence
|
|
is worse than a missed saving — it means every proficiency number is
|
|
conditioned on the routing that produced it, which is exactly the confound that
|
|
makes the outcome data below hard to read.
|
|
|
|
**Direction:** a small explicit exploration budget. An ε of 2-3% of requests
|
|
routed to the highest-cost-advantage *excluded* candidate per category would
|
|
have produced ~75 deepseek `tool_use_agentic` samples over this window — enough
|
|
to confirm or kill the 0.333 outright — at an estimated cost of under $2.
|
|
Alternatively a "probation" rule: any candidate excluded solely by a
|
|
proficiency score with `self_eval_samples < 10` gets N requests per week
|
|
regardless.
|
|
|
|
### What the ground truth actually says (with the confound stated)
|
|
|
|
Client outcomes, Wilson 95% CIs, ≥20 samples:
|
|
|
|
| model | n | pass | |
|
|
|---|---|---|---|
|
|
| `deepseek-v4-flash` | 390 | **82.8%** | [79%, 86%] |
|
|
| `kimi-k2.7-code` | 269 | 80.3% | [75%, 85%] |
|
|
| `qwen3.6-35b` | 198 | 85.9% | [80%, 90%] |
|
|
| `kimi-k2.7-code-fast` | 52 | 80.8% | [68%, 89%] |
|
|
| `glm-5.2-fast` | 44 | 88.6% | [76%, 95%] |
|
|
|
|
Within `coding_refactor`, where three models have real samples:
|
|
|
|
| model | n | pass | rel. price |
|
|
|---|---|---|---|
|
|
| `deepseek-v4-flash` | 95 | **81.1%** [72, 88] | 1x |
|
|
| `kimi-k2.7-code` | 62 | 75.8% [64, 85] | ~6.8x |
|
|
| `qwen3.6-35b` | 25 | 48.0% [30, 67] | ~2.1x |
|
|
|
|
The cheapest model is at the top, and `qwen3.6-35b`'s CI does not overlap it.
|
|
|
|
**Read this carefully, not triumphantly.** The comparison is confounded, and
|
|
the confound runs *against* the expensive models: context window determines
|
|
eligibility, so `qwen3.6-35b` (94k effective) only ever sees short requests
|
|
while `kimi-k2.7-code` (192k) takes the 94-192k band and `glm-5.2-fast` (782k)
|
|
takes the largest. Bigger context correlates with longer, harder sessions. So
|
|
the expensive models are being handed the harder work, and the table cannot
|
|
separate "cheaper model is as good" from "cheaper model got easier requests."
|
|
|
|
That is the point. **The router has ~1,000 samples of its highest-value signal
|
|
and cannot draw a conclusion from them, because it never randomizes.** Adding
|
|
exploration is what makes this data interpretable, not just what makes deepseek
|
|
eligible.
|
|
|
|
---
|
|
|
|
## 5. Finding #3 — cost is 93% prompt tokens, and the heuristics gate the other 7%
|
|
|
|
Decomposing the estimated cost of all 9,007 real decisions at
|
|
`assumed_cache_rate: 0.917`:
|
|
|
|
```
|
|
prompt share of est cost: $105.48 (92.6%)
|
|
completion share: $8.40 (7.4%)
|
|
|
|
real billed tokens: 1,190,607,380 prompt / 5,880,453 completion = 202:1
|
|
```
|
|
|
|
Completion price spans 54x across the catalog ($0.28 → $15.00) and decides
|
|
**7.4%** of the bill. Cached-prompt price spans 21x ($0.0144 → $0.30) and
|
|
decides **92.6%**.
|
|
|
|
Two consequences:
|
|
|
|
- **`tiering.cheap_completion_max: 1.00` gates tier 1 on the wrong axis.** Tier
|
|
is the single most consequential filter in the system — it decides the
|
|
eligible set before ranking runs — and it is resolved from completion price
|
|
plus reasoning mode. This is the same substitution the project has now caught
|
|
twice ("list price ranks models backwards"; "cheapness is not a capability
|
|
ceiling"), appearing a third time. For this workload the honest cheapness
|
|
signal is `cost_per_1m_prompt_cached`.
|
|
|
|
- **Pinch is the highest-leverage lever in the codebase and the least
|
|
instrumented.** It is the only thing that touches the 93%. Post-pinch
|
|
`required_context_tokens` (line 2450 records `measured`, computed on
|
|
`send_messages`) distributes as:
|
|
|
|
```
|
|
p10 49,817 p25 66,097 p50 88,069 p75 127,517 p90 195,683 p99 279,995
|
|
```
|
|
|
|
against `pinch.budget_tokens: 50000`. **The median request ships at 1.8x the
|
|
pinch budget and p90 at 3.9x** — pinch is running and not reaching its own
|
|
target, which makes sense given it may only trim tool results outside the
|
|
last 4 turns while opencode's ~32k system prompt and tool definitions are
|
|
untouchable. Its stats log at `logs.debug` while `logging.level: info`, so
|
|
zero pinch lines exist in seven days of journal. Nothing in `/metrics`, the
|
|
TUI, or the admin dashboard reports pinch effectiveness either.
|
|
|
|
A 20% reduction in prompt tokens is worth more than every routing decision in
|
|
this window combined, and right now there is no way to tell whether pinch
|
|
delivers 2% or 40%. **Promote the pinch line to info, and add
|
|
`tokens_saved` / `original_tokens` to `/metrics`.** That is the cheapest
|
|
high-value change on this list.
|
|
|
|
### Related: the tier ladder does not produce a capability gradient
|
|
|
|
| tier | decisions | avg est cost | outcome pass rate (approx. join, n) |
|
|
|---|---|---|---|
|
|
| 1 | 896 | $0.00907 | 76.9% (13) |
|
|
| 2 | 6,663 | $0.01337 | 83.6% (1,201) |
|
|
| 3 | 1,448 | $0.01155 | **70.4%** (226) |
|
|
|
|
Tier 3 costs *less* per decision than tier 2 and fails *more*. The failure rate
|
|
is the good news — it means the classifier's tier call carries real signal about
|
|
task difficulty. The cost figure is the problem: "frontier / high-stakes" is
|
|
not buying a more capable model, only a reasoning-enabled subset, because tier 3
|
|
is resolved from `reasoning_default_enabled` while deepseek (1M window, 1.00 on
|
|
all three coding categories) is pinned at tier 2 and excluded from all 1,448 of
|
|
them. That is the tier-1 lesson from CLAUDE.md recurring one rung up the ladder.
|
|
*(Outcome join is `session_key` + 5s window, so treat as indicative.)*
|
|
|
|
---
|
|
|
|
## 6. Smaller structural notes
|
|
|
|
**Sample-depth trap, recurring.** CLAUDE.md documents this precisely once
|
|
already — `docs_writing` read 0.70-1.00 at n=2 with a model at the ceiling, and
|
|
six more passes spread it 0.66-0.97 without touching a task. The same shape is
|
|
sitting in four more categories right now:
|
|
|
|
| category | rows | at exactly 1.00 | mean samples |
|
|
|---|---|---|---|
|
|
| `reasoning_math` | 16 | 13 | 5.4 |
|
|
| `translation` | 16 | 11 | 3.2 |
|
|
| `general_chat` | 13 | 11 | 3.5 |
|
|
| `summarization` | 16 | 3 | 3.4 |
|
|
|
|
The lesson was learned and written down; it has not yet been applied to the
|
|
categories that still show the symptom. "Try samples before hardening" is
|
|
already the project's own guidance.
|
|
|
|
**"Quality is the objective" is operationally inverted.** 72 of 141 proficiency
|
|
rows sit at exactly 1.00, and 98 of 141 (70%) fall inside the top
|
|
`quality_tolerance` band. For most requests quality is a tie and **cost is the
|
|
sole decider**, with quality acting as a coarse veto on the worst ~30%. That is
|
|
a defensible design — arguably the right one — but CLAUDE.md and
|
|
`docs/routing.md` read the other way ("there is no weight to tune — quality is
|
|
the objective and cost is the tiebreak"), which will mislead anyone tuning it.
|
|
|
|
**Structural verification is inert for this traffic.** 12,047 of 13,501
|
|
verification rows (89.2%) are `unverifiable` — correctly, since agent turns end
|
|
in tool calls. It costs nothing (pure Python, post-response), so there is no
|
|
harm, but it occupies a verdict-mix panel on the dashboard and a section in the
|
|
docs that together imply coverage that does not exist. `feedback.py`'s own
|
|
`coverage()` already prints the warning; the dashboard should too.
|
|
|
|
**`dispatcher.py` is where the discipline stops.** 3,104 lines, with
|
|
`chat_completions` at **662 lines** in a single function. Every other module in
|
|
this project is small, pure, and injectable; this one holds routing, pinch,
|
|
classification, capability checks, local vision, streaming proxy, telemetry
|
|
sniffing, and verification scheduling in one call frame. It is the one place
|
|
where the next change is expensive, and the one place a reader cannot hold the
|
|
whole path in their head. Extracting the pre-dispatch pipeline (prune →
|
|
measure → classify → gate → rank) into a pure function taking messages + config
|
|
and returning a decision would bring it in line with the rest and would be
|
|
directly testable.
|
|
|
|
**Minor doc drift.** `design/local-llm-model-router.md` §4 still presents the
|
|
weighted composite as the scoring model and §9 item 6 still calls proficiency
|
|
"the last inert axis" — both superseded. `scoring.composite_score` and
|
|
`eco_score` are dead but tested. `README.md:8` and `docs/routing.md`'s
|
|
"## Weighted Scoring" heading carry the same stale framing. CLAUDE.md already
|
|
declares itself the source of truth, so this is low-stakes, but the design doc
|
|
is the file a new reader opens first.
|
|
|
|
---
|
|
|
|
## 7. Recommended order
|
|
|
|
1. **Do not run `feedback.py` on the 961-row backlog** until the exposure bias
|
|
in §3 is fixed. Separate outcome samples from benchmark samples, or weight
|
|
by sample size. This is the one item that will actively degrade routing if
|
|
the obvious next step is taken.
|
|
2. **Promote the pinch log to `info` and surface `tokens_saved` in
|
|
`/metrics`.** Cheapest change here, on the axis that carries 93% of the bill.
|
|
3. **Add an exploration budget** (ε ≈ 2-3%, or probation for candidates
|
|
excluded on `self_eval_samples < 10`). Without it, no proficiency number in
|
|
the table is interpretable and the experiment `config.yaml` already proposes
|
|
cannot run.
|
|
4. **Add a single-model column to `baseline_report.py`.** "vs always-cheapest"
|
|
and "vs always-frontier" bracket the answer; "vs always-deepseek" is the
|
|
decision the operator actually faces.
|
|
5. **Re-tier on `cost_per_1m_prompt_cached`**, or drop the price term from
|
|
tiering entirely and keep the context/reasoning signals.
|
|
6. Run the eval harness for `reasoning_math`, `translation`, `general_chat`,
|
|
`summarization` — the same six-pass treatment that spread `docs_writing`.
|
|
7. Extract the pre-dispatch pipeline out of `chat_completions`.
|
|
|
|
---
|
|
|
|
## 8. Bottom line
|
|
|
|
The premise holds and the machine works. The measurement culture is the real
|
|
asset here — this project has overturned its own conclusions on evidence
|
|
repeatedly (list price ranks backwards, tier-from-price, benchmark-vs-real cost,
|
|
four harness bugs) and each reversal is written down with the reasoning intact.
|
|
That is rarer than the router.
|
|
|
|
The gap is that the loop is not yet closed. Every measurement so far has been
|
|
*taken*, then reasoned about, then encoded by hand. The two mechanisms that
|
|
were supposed to close the loop automatically — folding client outcomes into
|
|
proficiency, and letting the tool-proficiency experiment settle itself — are
|
|
both currently pointed slightly the wrong way: one rewards models for being
|
|
unmeasured, the other cannot gather the measurement it needs. Neither is a
|
|
large fix. Both are the difference between a router that learns from your
|
|
traffic and one that is very carefully hand-tuned to it.
|
|
|
|
---
|
|
|
|
## Appendix — reproduction scripts
|
|
|
|
Read-only. Each was run against `router.db` on 2026-09-01 to produce the
|
|
numbers above. Save to a scratch dir and run with `PYTHONPATH=src` from the
|
|
repo root.
|
|
|
|
### A. Single-model cost baselines (§2)
|
|
|
|
```python
|
|
"""What would a single-model policy have cost, on the same 9k real requests?"""
|
|
import sqlite3, sys
|
|
sys.path.insert(0, "src")
|
|
from routing import estimated_cost
|
|
|
|
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
|
|
models = {r["model_id"]: dict(r) for r in conn.execute("select * from models")}
|
|
rows = conn.execute("""
|
|
select required_context_tokens rc, est_cost_usd, selected_model, task_tier
|
|
from route_decisions where kind='chat' and required_context_tokens is not null
|
|
""").fetchall()
|
|
|
|
CACHE, COMP = 0.917, 500
|
|
actual = sum(r["est_cost_usd"] or 0 for r in rows)
|
|
print(f"decisions: {len(rows)} actual router est: ${actual:,.2f}\n")
|
|
print(f"{'single-model policy':<26} {'cost':>10} {'vs router':>10} {'infeasible':>11}")
|
|
print("-"*62)
|
|
res=[]
|
|
for mid, m in models.items():
|
|
if m["availability"] != "active" or m["access_level"] != "public" or m["latency_class"]=="flex":
|
|
continue
|
|
eff = m["effective_context_window"] or 0
|
|
tot = 0.0; bad = 0
|
|
for r in rows:
|
|
if r["rc"] > eff:
|
|
bad += 1; continue
|
|
tot += estimated_cost(m, r["rc"], COMP, CACHE) or 0
|
|
res.append((tot, mid, bad))
|
|
for tot, mid, bad in sorted(res):
|
|
print(f"{mid:<26} ${tot:>9,.2f} {tot/actual:>9.2f}x {bad:>10,}")
|
|
```
|
|
|
|
### B. Prompt-vs-completion cost anatomy (§5)
|
|
|
|
```python
|
|
import sqlite3, sys, statistics
|
|
sys.path.insert(0,"src")
|
|
conn=sqlite3.connect("router.db"); conn.row_factory=sqlite3.Row
|
|
m={r["model_id"]:dict(r) for r in conn.execute("select * from models")}
|
|
rows=conn.execute("""select selected_model sm, required_context_tokens rc, est_cost_usd c
|
|
from route_decisions where kind='chat' and selected_model is not null
|
|
and required_context_tokens is not null""").fetchall()
|
|
CACHE,COMP=0.917,500
|
|
pt=ct=0.0
|
|
for r in rows:
|
|
mm=m.get(r["sm"]);
|
|
if not mm: continue
|
|
p=r["rc"]*((1-CACHE)*mm["cost_per_1m_prompt"]+CACHE*(mm["cost_per_1m_prompt_cached"] or mm["cost_per_1m_prompt"]))/1e6
|
|
c=COMP*mm["cost_per_1m_completion"]/1e6
|
|
pt+=p; ct+=c
|
|
print(f"prompt share of est cost: ${pt:.2f} ({pt/(pt+ct)*100:.1f}%)")
|
|
print(f"completion share: ${ct:.2f} ({ct/(pt+ct)*100:.1f}%)")
|
|
print()
|
|
rc=[r["rc"] for r in rows]
|
|
rc.sort()
|
|
print("required_context_tokens percentiles:")
|
|
for q in (0.1,0.25,0.5,0.75,0.9,0.99):
|
|
print(f" p{int(q*100):>2}: {rc[int(q*len(rc))]:>9,}")
|
|
print(f" max: {rc[-1]:,} mean: {statistics.mean(rc):,.0f}")
|
|
print()
|
|
# actual billed vs prompt size
|
|
print("actual billed energy_observations, prompt vs completion tokens:")
|
|
r=conn.execute("select sum(prompt_tokens) p, sum(completion_tokens) c, count(*) n from energy_observations where task_category!='seed_reference'").fetchone()
|
|
print(f" prompt {r['p']:,} completion {r['c']:,} ratio {r['p']/r['c']:.0f}:1 n={r['n']:,}")
|
|
```
|
|
|
|
### C. Does category change the winner? (§1)
|
|
|
|
```python
|
|
"""Does the classifier's CATEGORY output change the routing decision?
|
|
|
|
Replays every real chat decision twice: once with per-category proficiency
|
|
(what ships), once with a category-agnostic mean proficiency per model.
|
|
"""
|
|
import sqlite3, sys, statistics
|
|
sys.path.insert(0,"src")
|
|
from config import load_config
|
|
from routing import select_candidates, rank_candidates
|
|
|
|
cfg = load_config("config/config.yaml")
|
|
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
|
|
models = [dict(r) for r in conn.execute("select * from models")]
|
|
prof = {(r["model_id"], r["category"]): r["blended_score"]
|
|
for r in conn.execute("select model_id, category, blended_score from proficiency")}
|
|
# category-agnostic: mean of a model's blended scores across all categories
|
|
avg = {}
|
|
for (m,c),v in prof.items():
|
|
if v is not None: avg.setdefault(m,[]).append(v)
|
|
avg = {m: statistics.fmean(v) for m,v in avg.items()}
|
|
|
|
decisions = conn.execute("""select task_category c, task_tier t, required_context_tokens rc,
|
|
latency_tolerance lt, selected_model sm from route_decisions
|
|
where kind='chat' and selected_model is not null and task_category is not null
|
|
and required_context_tokens is not null and task_tier is not null""").fetchall()
|
|
|
|
def pick(d, use_category):
|
|
rows=[]
|
|
for m in models:
|
|
r=dict(m)
|
|
r["proficiency"] = prof.get((m["model_id"], d["c"])) if use_category else avg.get(m["model_id"])
|
|
rows.append(r)
|
|
cand = select_candidates(rows,
|
|
required_context_tokens=d["rc"], required_tier=d["t"],
|
|
latency_tolerance=d["lt"] or "interactive",
|
|
allowed_access_levels=cfg.routing.allowed_access_levels,
|
|
exclude_stale=cfg.freshness.exclude_stale,
|
|
exclude_deprecated=cfg.freshness.exclude_deprecated)
|
|
rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
|
|
prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
|
|
cache_rate=cfg.objective.assumed_cache_rate)
|
|
return (rk[0]["model_id"], rk[0]["cost"]) if rk else (None,0)
|
|
|
|
same=diff=0; cost_cat=cost_flat=0.0
|
|
for d in decisions:
|
|
a,ca = pick(d, True); b,cb = pick(d, False)
|
|
cost_cat+=ca; cost_flat+=cb
|
|
if a==b: same+=1
|
|
else: diff+=1
|
|
n=len(decisions)
|
|
print(f"replayed {n} real decisions")
|
|
print(f" category-aware == category-agnostic: {same}/{n} ({same/n*100:.1f}%)")
|
|
print(f" category changed the winner: {diff}/{n} ({diff/n*100:.1f}%)")
|
|
print(f" cost with category: ${cost_cat:,.2f}")
|
|
print(f" cost without category: ${cost_flat:,.2f} ({cost_flat/cost_cat:.2f}x)")
|
|
```
|
|
|
|
### D. Simulated feedback fold-in (§3)
|
|
|
|
Copies the DB first; never mutates `router.db`.
|
|
|
|
```python
|
|
"""Simulate: apply the 1,039 pending client outcomes, then replay routing."""
|
|
import sqlite3, sys, os, shutil
|
|
sys.path.insert(0,"src")
|
|
from config import load_config
|
|
import feedback
|
|
from routing import select_candidates, rank_candidates
|
|
|
|
cfg = load_config("config/config.yaml")
|
|
db = os.environ["CLAUDE_JOB_DIR"] + "/tmp/sim.db"
|
|
|
|
def replay(dbpath, label):
|
|
conn = sqlite3.connect(dbpath); conn.row_factory = sqlite3.Row
|
|
models=[dict(r) for r in conn.execute("select * from models")]
|
|
prof={(r["model_id"],r["category"]):r["blended_score"]
|
|
for r in conn.execute("select model_id,category,blended_score from proficiency")}
|
|
ds=conn.execute("""select task_category c,task_tier t,required_context_tokens rc,
|
|
latency_tolerance lt from route_decisions where kind='chat'
|
|
and task_category is not null and required_context_tokens is not null
|
|
and task_tier is not null""").fetchall()
|
|
tot=0.0; mix={}
|
|
for d in ds:
|
|
rows=[{**m,"proficiency":prof.get((m["model_id"],d["c"]))} for m in models]
|
|
cand=select_candidates(rows,required_context_tokens=d["rc"],required_tier=d["t"],
|
|
latency_tolerance=d["lt"] or "interactive",
|
|
allowed_access_levels=cfg.routing.allowed_access_levels,
|
|
exclude_stale=cfg.freshness.exclude_stale,
|
|
exclude_deprecated=cfg.freshness.exclude_deprecated)
|
|
rk=rank_candidates(cand,quality_tolerance=cfg.objective.quality_tolerance,
|
|
prompt_tokens=d["rc"],completion_tokens=cfg.objective.assumed_completion_tokens,
|
|
cache_rate=cfg.objective.assumed_cache_rate)
|
|
if rk:
|
|
tot+=rk[0]["cost"] or 0; mix[rk[0]["model_id"]]=mix.get(rk[0]["model_id"],0)+1
|
|
conn.close()
|
|
print(f"\n=== {label} === replayed {len(ds)} decisions est cost ${tot:,.2f}")
|
|
for m,n in sorted(mix.items(),key=lambda kv:-kv[1]):
|
|
print(f" {m:<24}{n:>6}")
|
|
return tot
|
|
|
|
before = replay(db, "BEFORE (current proficiency)")
|
|
conn=sqlite3.connect(db)
|
|
rows=feedback.unapplied_failures(conn)
|
|
feedback.apply_failures(conn, cfg, feedback.summarize(rows), dry_run=False)
|
|
conn.close()
|
|
after = replay(db, "AFTER folding in 1,039 client outcomes")
|
|
print(f"\ncost change: ${before:,.2f} -> ${after:,.2f} ({after/before:.2f}x)")
|
|
```
|
|
|
|
### E. Client-outcome pass rates with Wilson CIs (§4)
|
|
|
|
```python
|
|
import sqlite3, math
|
|
conn=sqlite3.connect("router.db"); conn.row_factory=sqlite3.Row
|
|
q="""select model_id m, task_category c,
|
|
sum(verdict='succeeded') s, sum(verdict='failed') f
|
|
from verifications where kind='client_outcome' group by 1,2"""
|
|
rows=[dict(r) for r in conn.execute(q)]
|
|
def wilson(s,n):
|
|
if n==0: return (0,0)
|
|
z=1.96; p=s/n; d=1+z*z/n
|
|
c=(p+z*z/(2*n))/d; h=z*math.sqrt(p*(1-p)/n+z*z/(4*n*n))/d
|
|
return (c-h,c+h)
|
|
print(f"{'model':<22}{'category':<18}{'n':>5}{'pass':>7} 95% CI")
|
|
print("-"*66)
|
|
for r in sorted(rows,key=lambda r:(-(r['s']+r['f']))):
|
|
n=r['s']+r['f']
|
|
if n<15: continue
|
|
lo,hi=wilson(r['s'],n)
|
|
print(f"{r['m']:<22}{r['c']:<18}{n:>5}{r['s']/n*100:>6.1f}% [{lo*100:.0f}%, {hi*100:.0f}%]")
|
|
print()
|
|
print("--- pooled per model (all categories) ---")
|
|
agg={}
|
|
for r in rows:
|
|
a=agg.setdefault(r['m'],[0,0]); a[0]+=r['s']; a[1]+=r['f']
|
|
for m,(s,f) in sorted(agg.items(),key=lambda kv:-(kv[1][0]+kv[1][1])):
|
|
n=s+f
|
|
if n<20: continue
|
|
lo,hi=wilson(s,n)
|
|
print(f"{m:<22}{n:>5}{s/n*100:>6.1f}% [{lo*100:.0f}%, {hi*100:.0f}%]")
|
|
```
|