plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
30 KiB
Conceptual review: is the premise sound, and is execution converging on it?
Status: done -- review of shipped work
Date: 2026-09-01
Scope: premise and architecture, not bugs. Every number below is measured
from the live router.db and the checked-in code on this date; the scripts are
inlined so they can be re-run.
Verdict: the premise is sound and the hard part is built. Three structural
problems sit between the current state and the stated goal, and one of them
will actively move routing the wrong way the moment the next obvious step is
taken.
0. What the evidence base is
route_decisions |
9,025 rows (9,007 chat), 2026-08-24 → 2026-09-01 |
energy_observations |
13,966 rows, $49.94 billed, 6.51 kWh, 1,122 gCO₂eq, from 2026-08-17 |
verifications |
13,501 rows; 987 of them client outcomes (816 succeeded / 171 failed) |
proficiency |
141 (model, category) rows |
| tests | 843 pass in 36.8s, all offline |
This is a real deployment with real traffic, not a demo. The review is worth doing precisely because there is now enough data to check the design against itself.
1. The premise is sound, and the hardest claim in it is now proven
Design doc §1: use a local model to classify the task, then route to the cheapest/best-fit open-weight model instead of defaulting everything to one expensive model.
The load-bearing, non-obvious claim in that sentence is that task category
should change the model. For a long stretch it demonstrably did not — the
classifier computed a category, the router paid ~10s for it, and
proficiency was empty so it changed nothing. That is fixed, and it is
measurable now. Replaying all 9,001 real decisions with per-category
proficiency against a category-agnostic mean:
category-aware == category-agnostic: 5,917/9,001 (65.7%)
category changed the winner: 3,084/9,001 (34.3%)
One decision in three turns on the category. That is the design's central
bet paying off, and it took populating proficiency, fixing four harness bugs,
and rebuilding the cost axis to get there. It should be stated as a result.
Two more things are working and deserve to be said plainly before the criticism:
- The session cache removed the latency tax. It landed 2026-08-30 and is running at 98%+ hit rate (2026-09-01: 1,402 cached / 28 fresh). The "~10s of local overhead on every message" problem in CLAUDE.md is solved for the sessions that matter. Fresh classifications average 2,297ms (max 66,319ms — one outlier worth a look, but not a design issue).
- The pure-module discipline is the best thing in the codebase.
routing.py,scoring.py,tiering.py,proficiency.py,context_prune.pyare I/O-free, take thresholds as arguments, and are exhaustively tested. That is why 843 tests run offline in 37 seconds, and it is why every one of the measurement-driven reversals in CLAUDE.md was cheap to make. Keep it.
2. The premise quietly changed, and the docs did not follow
Design §7 names the metric: "$ saved vs. always-frontier baseline", with the warning that "a router that's cheaper but quietly worse isn't a win." Nobody has checked the other side of that: a router that's better but quietly more expensive isn't obviously a win either, and that is where this landed.
baseline_report.py already tells the story, and it appears not to have been
read recently:
$ PYTHONPATH=src python -m baseline_report --since 2026-08-24
decisions: 9018
total cost: actual $113.89 cheapest $51.39 best-prof $166.59
dominance (selected == cheapest): 4,972/9,018 (55.1%)
mean proficiency: actual 0.991 cheapest 0.788 best-prof 0.998
The router sits at 0.991 of a maximum 0.998 on quality and 2.2x the cheapest on cost. It is not a cost/quality tradeoff engine; it is a quality-maximizer that takes a discount when quality is exactly tied. That may be what you want — but it is not what §1 says, and it is not what the README says either ("dispatches to the cheapest/best-fit model").
The comparison that matters more is against the alternative you would actually
have used. Replaying all 9,007 chat decisions against every single-model policy
(same request shapes, same catalog prices, assumed_cache_rate 0.917;
infeasible = requests exceeding that model's effective window):
| single-model policy | est. cost | vs router | infeasible |
|---|---|---|---|
gemma-4-31b |
$18.91 | 0.17x | 1,170 |
qwen3.6-35b |
$19.54 | 0.17x | 4,032 |
deepseek-v4-flash |
$36.59 | 0.32x | 0 |
kimi-k2.7-code |
$136.41 | 1.20x | 963 |
glm-5.2-fast / glm-5.3 |
$260.20 | 2.28x | 0 |
kimi-k3 |
$563.97 | 4.95x | 0 |
| (actual routed) | $113.89 | 1.00x | 0 |
Against always-frontier (kimi-k3) the router saves 4.95x — the design's
own metric, comfortably met. Against "just always use deepseek-v4-flash" it
costs 3.1x more, and deepseek is the only other policy that can serve 100%
of the traffic without a context failure.
Which baseline is honest depends entirely on what you would otherwise have
done. For opencode agent traffic against this catalog, the realistic
counterfactual is not kimi-k3 — it is deepseek. The router should report
both baselines, and baseline_report.py should grow a single-model column.
The operational consequence is live right now: 6.51 kWh burned against a
plan_kwh_per_period of 6.25. At roughly $7.67/kWh realized, an
always-deepseek policy would have used about a third of that and stayed inside
the allowance.
3. Finding #1 — the feedback loop rewards models for not being used
This is the most important item in the review, and it is a reason not to take the next obvious step until it is fixed.
987 client outcomes are sitting in verifications; 961 of them unapplied.
POST /outcome is correctly identified in CLAUDE.md as "the only ground
truth," so folding them in looks like the highest-value pending action. It is
not, as currently built.
Simulated on a copy of the DB — run feedback.py for real, then replay all
9,007 decisions through select_candidates + rank_candidates:
BEFORE est cost $137.76 AFTER est cost $179.28 (1.30x)
kimi-k2.7-code 2,722 glm-5.2-fast 2,243
qwen3.6-35b 2,646 qwen3.6-35b-fast 1,385
qwen3.6-35b-fast 1,236 glm-5.3 1,318
glm-5.3 745 kimi-k2.7-code 1,239
gemma-4-31b 587 qwen3.6-35b 1,150
deepseek-v4-flash 500 kimi-k2.7-code-fast 947
glm-5.2-fast 417 deepseek-v4-flash 596
Ingesting ground truth makes routing 30% more expensive and hands 2,243
decisions to glm-5.2-fast, one of the priciest common rows. The mechanism is
visible in the scores. coding_refactor, before → after:
| model | before | n | after | n | real outcomes folded in |
|---|---|---|---|---|---|
glm-5.2-fast |
0.994 | 36 | 0.994 | 36 | zero |
kimi-k2.7-code |
1.000 | 18 | 0.813 | 80 | 62 |
deepseek-v4-flash |
0.826 | 36 | 0.791 | 135 | 99 |
qwen3.6-35b |
0.900 | 18 | 0.641 | 44 | 26 |
glm-5.2-fast wins the category outright by never having been measured on
real traffic.
The cause is structural, not a bug: feedback.py and eval_proficiency.py
both write through proficiency_store.add_self_eval, into the same
self_eval_score running mean, with equal weight per sample. Benchmark tasks
score ~1.00. Real agent turns score ~0.80. So every real sample drags a score
down, and the drag is proportional to how much the model has been used. The
router then routes away from the used model toward the unused one, gathers
outcomes on that one, penalizes it, and moves on.
That is a rotation, not a convergence. And because exposure has been concentrated on the cheap models — they win the cost tiebreak — the rotation is systematically toward expensive ones.
Directions (pick one; all are cheap):
- Keep the two sources separate. Real-traffic outcomes are a different
measurement from benchmark tasks and should not share a mean. Add
outcome_score/outcome_samplescolumns and blend them explicitly, the wayleaderboardandself_evalalready are. This also makes the ~0.80-vs-1.00 scale difference visible instead of silently mixed. - Shrink toward the prior by sample size. A score with n=36 and no real
exposure should not outrank one with n=135 that includes 99 real outcomes.
Today low sample count is an advantage;
self_eval_min_samplesgates the leaderboard blend but nothing penalizes thin evidence in the ranking itself. - Compare like with like. Score a model against the per-category mean outcome rate rather than absolute, so an 82% pass rate in a category where everyone scores 80% is neutral, not a 0.18 penalty.
Until one of these lands, do not run feedback.py on the pending backlog.
The 961 rows are the most valuable data this project has; spending them through
the current path converts them into a more expensive router.
4. Finding #2 — exclusions are self-sealing; there is no exploration
deepseek-v4-flash scored 0.333 on tool_use_agentic across 3
benchmark tasks. Blended, it now reads 0.5. With quality_tolerance: 0.1 that
puts it in band 5 while eleven models sit in band 0, so it loses every
tool_use_agentic ranking outright.
Consequence, measured: deepseek won 0 of 2,512 tool_use_agentic decisions,
and has 0 client outcomes in that category. A 3-sample estimate has
permanently removed the cheapest capable model from the largest category of
traffic, and the design contains no path by which that estimate can ever be
revised.
config/config.yaml proposes the experiment that would settle it:
run with it off, let
POST /outcomereport real pass/fail, and comparetool_use_agenticproficiency for deepseek before and after
That experiment cannot run. min_tool_proficiency is already null — the
filter is off — and it changes nothing, because the quality band excludes
deepseek from the category regardless. The config comment says "with it off,
that advantage applies." It does not.
The general shape: this is a contextual bandit running pure exploitation. Once a model wins a (category, context-size) cell it wins it forever; alternatives never accumulate the evidence that would overturn the ranking. The consequence is worse than a missed saving — it means every proficiency number is conditioned on the routing that produced it, which is exactly the confound that makes the outcome data below hard to read.
Direction: a small explicit exploration budget. An ε of 2-3% of requests
routed to the highest-cost-advantage excluded candidate per category would
have produced ~75 deepseek tool_use_agentic samples over this window — enough
to confirm or kill the 0.333 outright — at an estimated cost of under $2.
Alternatively a "probation" rule: any candidate excluded solely by a
proficiency score with self_eval_samples < 10 gets N requests per week
regardless.
What the ground truth actually says (with the confound stated)
Client outcomes, Wilson 95% CIs, ≥20 samples:
| model | n | pass | |
|---|---|---|---|
deepseek-v4-flash |
390 | 82.8% | [79%, 86%] |
kimi-k2.7-code |
269 | 80.3% | [75%, 85%] |
qwen3.6-35b |
198 | 85.9% | [80%, 90%] |
kimi-k2.7-code-fast |
52 | 80.8% | [68%, 89%] |
glm-5.2-fast |
44 | 88.6% | [76%, 95%] |
Within coding_refactor, where three models have real samples:
| model | n | pass | rel. price |
|---|---|---|---|
deepseek-v4-flash |
95 | 81.1% [72, 88] | 1x |
kimi-k2.7-code |
62 | 75.8% [64, 85] | ~6.8x |
qwen3.6-35b |
25 | 48.0% [30, 67] | ~2.1x |
The cheapest model is at the top, and qwen3.6-35b's CI does not overlap it.
Read this carefully, not triumphantly. The comparison is confounded, and
the confound runs against the expensive models: context window determines
eligibility, so qwen3.6-35b (94k effective) only ever sees short requests
while kimi-k2.7-code (192k) takes the 94-192k band and glm-5.2-fast (782k)
takes the largest. Bigger context correlates with longer, harder sessions. So
the expensive models are being handed the harder work, and the table cannot
separate "cheaper model is as good" from "cheaper model got easier requests."
That is the point. The router has ~1,000 samples of its highest-value signal and cannot draw a conclusion from them, because it never randomizes. Adding exploration is what makes this data interpretable, not just what makes deepseek eligible.
5. Finding #3 — cost is 93% prompt tokens, and the heuristics gate the other 7%
Decomposing the estimated cost of all 9,007 real decisions at
assumed_cache_rate: 0.917:
prompt share of est cost: $105.48 (92.6%)
completion share: $8.40 (7.4%)
real billed tokens: 1,190,607,380 prompt / 5,880,453 completion = 202:1
Completion price spans 54x across the catalog ($0.28 → $15.00) and decides 7.4% of the bill. Cached-prompt price spans 21x ($0.0144 → $0.30) and decides 92.6%.
Two consequences:
-
tiering.cheap_completion_max: 1.00gates tier 1 on the wrong axis. Tier is the single most consequential filter in the system — it decides the eligible set before ranking runs — and it is resolved from completion price plus reasoning mode. This is the same substitution the project has now caught twice ("list price ranks models backwards"; "cheapness is not a capability ceiling"), appearing a third time. For this workload the honest cheapness signal iscost_per_1m_prompt_cached. -
Pinch is the highest-leverage lever in the codebase and the least instrumented. It is the only thing that touches the 93%. Post-pinch
required_context_tokens(line 2450 recordsmeasured, computed onsend_messages) distributes as:p10 49,817 p25 66,097 p50 88,069 p75 127,517 p90 195,683 p99 279,995against
pinch.budget_tokens: 50000. The median request ships at 1.8x the pinch budget and p90 at 3.9x — pinch is running and not reaching its own target, which makes sense given it may only trim tool results outside the last 4 turns while opencode's ~32k system prompt and tool definitions are untouchable. Its stats log atlogs.debugwhilelogging.level: info, so zero pinch lines exist in seven days of journal. Nothing in/metrics, the TUI, or the admin dashboard reports pinch effectiveness either.A 20% reduction in prompt tokens is worth more than every routing decision in this window combined, and right now there is no way to tell whether pinch delivers 2% or 40%. Promote the pinch line to info, and add
tokens_saved/original_tokensto/metrics. That is the cheapest high-value change on this list.
Related: the tier ladder does not produce a capability gradient
| tier | decisions | avg est cost | outcome pass rate (approx. join, n) |
|---|---|---|---|
| 1 | 896 | $0.00907 | 76.9% (13) |
| 2 | 6,663 | $0.01337 | 83.6% (1,201) |
| 3 | 1,448 | $0.01155 | 70.4% (226) |
Tier 3 costs less per decision than tier 2 and fails more. The failure rate
is the good news — it means the classifier's tier call carries real signal about
task difficulty. The cost figure is the problem: "frontier / high-stakes" is
not buying a more capable model, only a reasoning-enabled subset, because tier 3
is resolved from reasoning_default_enabled while deepseek (1M window, 1.00 on
all three coding categories) is pinned at tier 2 and excluded from all 1,448 of
them. That is the tier-1 lesson from CLAUDE.md recurring one rung up the ladder.
(Outcome join is session_key + 5s window, so treat as indicative.)
6. Smaller structural notes
Sample-depth trap, recurring. CLAUDE.md documents this precisely once
already — docs_writing read 0.70-1.00 at n=2 with a model at the ceiling, and
six more passes spread it 0.66-0.97 without touching a task. The same shape is
sitting in four more categories right now:
| category | rows | at exactly 1.00 | mean samples |
|---|---|---|---|
reasoning_math |
16 | 13 | 5.4 |
translation |
16 | 11 | 3.2 |
general_chat |
13 | 11 | 3.5 |
summarization |
16 | 3 | 3.4 |
The lesson was learned and written down; it has not yet been applied to the categories that still show the symptom. "Try samples before hardening" is already the project's own guidance.
"Quality is the objective" is operationally inverted. 72 of 141 proficiency
rows sit at exactly 1.00, and 98 of 141 (70%) fall inside the top
quality_tolerance band. For most requests quality is a tie and cost is the
sole decider, with quality acting as a coarse veto on the worst ~30%. That is
a defensible design — arguably the right one — but CLAUDE.md and
docs/routing.md read the other way ("there is no weight to tune — quality is
the objective and cost is the tiebreak"), which will mislead anyone tuning it.
Structural verification is inert for this traffic. 12,047 of 13,501
verification rows (89.2%) are unverifiable — correctly, since agent turns end
in tool calls. It costs nothing (pure Python, post-response), so there is no
harm, but it occupies a verdict-mix panel on the dashboard and a section in the
docs that together imply coverage that does not exist. feedback.py's own
coverage() already prints the warning; the dashboard should too.
dispatcher.py is where the discipline stops. 3,104 lines, with
chat_completions at 662 lines in a single function. Every other module in
this project is small, pure, and injectable; this one holds routing, pinch,
classification, capability checks, local vision, streaming proxy, telemetry
sniffing, and verification scheduling in one call frame. It is the one place
where the next change is expensive, and the one place a reader cannot hold the
whole path in their head. Extracting the pre-dispatch pipeline (prune →
measure → classify → gate → rank) into a pure function taking messages + config
and returning a decision would bring it in line with the rest and would be
directly testable.
Minor doc drift. design/local-llm-model-router.md §4 still presents the
weighted composite as the scoring model and §9 item 6 still calls proficiency
"the last inert axis" — both superseded. scoring.composite_score and
eco_score are dead but tested. README.md:8 and docs/routing.md's
"## Weighted Scoring" heading carry the same stale framing. CLAUDE.md already
declares itself the source of truth, so this is low-stakes, but the design doc
is the file a new reader opens first.
7. Recommended order
- Do not run
feedback.pyon the 961-row backlog until the exposure bias in §3 is fixed. Separate outcome samples from benchmark samples, or weight by sample size. This is the one item that will actively degrade routing if the obvious next step is taken. - Promote the pinch log to
infoand surfacetokens_savedin/metrics. Cheapest change here, on the axis that carries 93% of the bill. - Add an exploration budget (ε ≈ 2-3%, or probation for candidates
excluded on
self_eval_samples < 10). Without it, no proficiency number in the table is interpretable and the experimentconfig.yamlalready proposes cannot run. - Add a single-model column to
baseline_report.py. "vs always-cheapest" and "vs always-frontier" bracket the answer; "vs always-deepseek" is the decision the operator actually faces. - Re-tier on
cost_per_1m_prompt_cached, or drop the price term from tiering entirely and keep the context/reasoning signals. - Run the eval harness for
reasoning_math,translation,general_chat,summarization— the same six-pass treatment that spreaddocs_writing. - Extract the pre-dispatch pipeline out of
chat_completions.
8. Bottom line
The premise holds and the machine works. The measurement culture is the real asset here — this project has overturned its own conclusions on evidence repeatedly (list price ranks backwards, tier-from-price, benchmark-vs-real cost, four harness bugs) and each reversal is written down with the reasoning intact. That is rarer than the router.
The gap is that the loop is not yet closed. Every measurement so far has been taken, then reasoned about, then encoded by hand. The two mechanisms that were supposed to close the loop automatically — folding client outcomes into proficiency, and letting the tool-proficiency experiment settle itself — are both currently pointed slightly the wrong way: one rewards models for being unmeasured, the other cannot gather the measurement it needs. Neither is a large fix. Both are the difference between a router that learns from your traffic and one that is very carefully hand-tuned to it.
Appendix — reproduction scripts
Read-only. Each was run against router.db on 2026-09-01 to produce the
numbers above. Save to a scratch dir and run with PYTHONPATH=src from the
repo root.
A. Single-model cost baselines (§2)
"""What would a single-model policy have cost, on the same 9k real requests?"""
import sqlite3, sys
sys.path.insert(0, "src")
from routing import estimated_cost
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = {r["model_id"]: dict(r) for r in conn.execute("select * from models")}
rows = conn.execute("""
select required_context_tokens rc, est_cost_usd, selected_model, task_tier
from route_decisions where kind='chat' and required_context_tokens is not null
""").fetchall()
CACHE, COMP = 0.917, 500
actual = sum(r["est_cost_usd"] or 0 for r in rows)
print(f"decisions: {len(rows)} actual router est: ${actual:,.2f}\n")
print(f"{'single-model policy':<26} {'cost':>10} {'vs router':>10} {'infeasible':>11}")
print("-"*62)
res=[]
for mid, m in models.items():
if m["availability"] != "active" or m["access_level"] != "public" or m["latency_class"]=="flex":
continue
eff = m["effective_context_window"] or 0
tot = 0.0; bad = 0
for r in rows:
if r["rc"] > eff:
bad += 1; continue
tot += estimated_cost(m, r["rc"], COMP, CACHE) or 0
res.append((tot, mid, bad))
for tot, mid, bad in sorted(res):
print(f"{mid:<26} ${tot:>9,.2f} {tot/actual:>9.2f}x {bad:>10,}")
B. Prompt-vs-completion cost anatomy (§5)
import sqlite3, sys, statistics
sys.path.insert(0,"src")
conn=sqlite3.connect("router.db"); conn.row_factory=sqlite3.Row
m={r["model_id"]:dict(r) for r in conn.execute("select * from models")}
rows=conn.execute("""select selected_model sm, required_context_tokens rc, est_cost_usd c
from route_decisions where kind='chat' and selected_model is not null
and required_context_tokens is not null""").fetchall()
CACHE,COMP=0.917,500
pt=ct=0.0
for r in rows:
mm=m.get(r["sm"]);
if not mm: continue
p=r["rc"]*((1-CACHE)*mm["cost_per_1m_prompt"]+CACHE*(mm["cost_per_1m_prompt_cached"] or mm["cost_per_1m_prompt"]))/1e6
c=COMP*mm["cost_per_1m_completion"]/1e6
pt+=p; ct+=c
print(f"prompt share of est cost: ${pt:.2f} ({pt/(pt+ct)*100:.1f}%)")
print(f"completion share: ${ct:.2f} ({ct/(pt+ct)*100:.1f}%)")
print()
rc=[r["rc"] for r in rows]
rc.sort()
print("required_context_tokens percentiles:")
for q in (0.1,0.25,0.5,0.75,0.9,0.99):
print(f" p{int(q*100):>2}: {rc[int(q*len(rc))]:>9,}")
print(f" max: {rc[-1]:,} mean: {statistics.mean(rc):,.0f}")
print()
# actual billed vs prompt size
print("actual billed energy_observations, prompt vs completion tokens:")
r=conn.execute("select sum(prompt_tokens) p, sum(completion_tokens) c, count(*) n from energy_observations where task_category!='seed_reference'").fetchone()
print(f" prompt {r['p']:,} completion {r['c']:,} ratio {r['p']/r['c']:.0f}:1 n={r['n']:,}")
C. Does category change the winner? (§1)
"""Does the classifier's CATEGORY output change the routing decision?
Replays every real chat decision twice: once with per-category proficiency
(what ships), once with a category-agnostic mean proficiency per model.
"""
import sqlite3, sys, statistics
sys.path.insert(0,"src")
from config import load_config
from routing import select_candidates, rank_candidates
cfg = load_config("config/config.yaml")
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = [dict(r) for r in conn.execute("select * from models")]
prof = {(r["model_id"], r["category"]): r["blended_score"]
for r in conn.execute("select model_id, category, blended_score from proficiency")}
# category-agnostic: mean of a model's blended scores across all categories
avg = {}
for (m,c),v in prof.items():
if v is not None: avg.setdefault(m,[]).append(v)
avg = {m: statistics.fmean(v) for m,v in avg.items()}
decisions = conn.execute("""select task_category c, task_tier t, required_context_tokens rc,
latency_tolerance lt, selected_model sm from route_decisions
where kind='chat' and selected_model is not null and task_category is not null
and required_context_tokens is not null and task_tier is not null""").fetchall()
def pick(d, use_category):
rows=[]
for m in models:
r=dict(m)
r["proficiency"] = prof.get((m["model_id"], d["c"])) if use_category else avg.get(m["model_id"])
rows.append(r)
cand = select_candidates(rows,
required_context_tokens=d["rc"], required_tier=d["t"],
latency_tolerance=d["lt"] or "interactive",
allowed_access_levels=cfg.routing.allowed_access_levels,
exclude_stale=cfg.freshness.exclude_stale,
exclude_deprecated=cfg.freshness.exclude_deprecated)
rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
cache_rate=cfg.objective.assumed_cache_rate)
return (rk[0]["model_id"], rk[0]["cost"]) if rk else (None,0)
same=diff=0; cost_cat=cost_flat=0.0
for d in decisions:
a,ca = pick(d, True); b,cb = pick(d, False)
cost_cat+=ca; cost_flat+=cb
if a==b: same+=1
else: diff+=1
n=len(decisions)
print(f"replayed {n} real decisions")
print(f" category-aware == category-agnostic: {same}/{n} ({same/n*100:.1f}%)")
print(f" category changed the winner: {diff}/{n} ({diff/n*100:.1f}%)")
print(f" cost with category: ${cost_cat:,.2f}")
print(f" cost without category: ${cost_flat:,.2f} ({cost_flat/cost_cat:.2f}x)")
D. Simulated feedback fold-in (§3)
Copies the DB first; never mutates router.db.
"""Simulate: apply the 1,039 pending client outcomes, then replay routing."""
import sqlite3, sys, os, shutil
sys.path.insert(0,"src")
from config import load_config
import feedback
from routing import select_candidates, rank_candidates
cfg = load_config("config/config.yaml")
db = os.environ["CLAUDE_JOB_DIR"] + "/tmp/sim.db"
def replay(dbpath, label):
conn = sqlite3.connect(dbpath); conn.row_factory = sqlite3.Row
models=[dict(r) for r in conn.execute("select * from models")]
prof={(r["model_id"],r["category"]):r["blended_score"]
for r in conn.execute("select model_id,category,blended_score from proficiency")}
ds=conn.execute("""select task_category c,task_tier t,required_context_tokens rc,
latency_tolerance lt from route_decisions where kind='chat'
and task_category is not null and required_context_tokens is not null
and task_tier is not null""").fetchall()
tot=0.0; mix={}
for d in ds:
rows=[{**m,"proficiency":prof.get((m["model_id"],d["c"]))} for m in models]
cand=select_candidates(rows,required_context_tokens=d["rc"],required_tier=d["t"],
latency_tolerance=d["lt"] or "interactive",
allowed_access_levels=cfg.routing.allowed_access_levels,
exclude_stale=cfg.freshness.exclude_stale,
exclude_deprecated=cfg.freshness.exclude_deprecated)
rk=rank_candidates(cand,quality_tolerance=cfg.objective.quality_tolerance,
prompt_tokens=d["rc"],completion_tokens=cfg.objective.assumed_completion_tokens,
cache_rate=cfg.objective.assumed_cache_rate)
if rk:
tot+=rk[0]["cost"] or 0; mix[rk[0]["model_id"]]=mix.get(rk[0]["model_id"],0)+1
conn.close()
print(f"\n=== {label} === replayed {len(ds)} decisions est cost ${tot:,.2f}")
for m,n in sorted(mix.items(),key=lambda kv:-kv[1]):
print(f" {m:<24}{n:>6}")
return tot
before = replay(db, "BEFORE (current proficiency)")
conn=sqlite3.connect(db)
rows=feedback.unapplied_failures(conn)
feedback.apply_failures(conn, cfg, feedback.summarize(rows), dry_run=False)
conn.close()
after = replay(db, "AFTER folding in 1,039 client outcomes")
print(f"\ncost change: ${before:,.2f} -> ${after:,.2f} ({after/before:.2f}x)")
E. Client-outcome pass rates with Wilson CIs (§4)
import sqlite3, math
conn=sqlite3.connect("router.db"); conn.row_factory=sqlite3.Row
q="""select model_id m, task_category c,
sum(verdict='succeeded') s, sum(verdict='failed') f
from verifications where kind='client_outcome' group by 1,2"""
rows=[dict(r) for r in conn.execute(q)]
def wilson(s,n):
if n==0: return (0,0)
z=1.96; p=s/n; d=1+z*z/n
c=(p+z*z/(2*n))/d; h=z*math.sqrt(p*(1-p)/n+z*z/(4*n*n))/d
return (c-h,c+h)
print(f"{'model':<22}{'category':<18}{'n':>5}{'pass':>7} 95% CI")
print("-"*66)
for r in sorted(rows,key=lambda r:(-(r['s']+r['f']))):
n=r['s']+r['f']
if n<15: continue
lo,hi=wilson(r['s'],n)
print(f"{r['m']:<22}{r['c']:<18}{n:>5}{r['s']/n*100:>6.1f}% [{lo*100:.0f}%, {hi*100:.0f}%]")
print()
print("--- pooled per model (all categories) ---")
agg={}
for r in rows:
a=agg.setdefault(r['m'],[0,0]); a[0]+=r['s']; a[1]+=r['f']
for m,(s,f) in sorted(agg.items(),key=lambda kv:-(kv[1][0]+kv[1][1])):
n=s+f
if n<20: continue
lo,hi=wilson(s,n)
print(f"{m:<22}{n:>5}{s/n*100:>6.1f}% [{lo*100:.0f}%, {hi*100:.0f}%]")