Files
6krrt/plans/conceptual-review-premise-and-execution.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

30 KiB

Conceptual review: is the premise sound, and is execution converging on it?

Status: done -- review of shipped work

Date: 2026-09-01 Scope: premise and architecture, not bugs. Every number below is measured from the live router.db and the checked-in code on this date; the scripts are inlined so they can be re-run. Verdict: the premise is sound and the hard part is built. Three structural problems sit between the current state and the stated goal, and one of them will actively move routing the wrong way the moment the next obvious step is taken.


0. What the evidence base is

route_decisions 9,025 rows (9,007 chat), 2026-08-24 → 2026-09-01
energy_observations 13,966 rows, $49.94 billed, 6.51 kWh, 1,122 gCO₂eq, from 2026-08-17
verifications 13,501 rows; 987 of them client outcomes (816 succeeded / 171 failed)
proficiency 141 (model, category) rows
tests 843 pass in 36.8s, all offline

This is a real deployment with real traffic, not a demo. The review is worth doing precisely because there is now enough data to check the design against itself.


1. The premise is sound, and the hardest claim in it is now proven

Design doc §1: use a local model to classify the task, then route to the cheapest/best-fit open-weight model instead of defaulting everything to one expensive model.

The load-bearing, non-obvious claim in that sentence is that task category should change the model. For a long stretch it demonstrably did not — the classifier computed a category, the router paid ~10s for it, and proficiency was empty so it changed nothing. That is fixed, and it is measurable now. Replaying all 9,001 real decisions with per-category proficiency against a category-agnostic mean:

category-aware == category-agnostic: 5,917/9,001 (65.7%)
category changed the winner:         3,084/9,001 (34.3%)

One decision in three turns on the category. That is the design's central bet paying off, and it took populating proficiency, fixing four harness bugs, and rebuilding the cost axis to get there. It should be stated as a result.

Two more things are working and deserve to be said plainly before the criticism:

  • The session cache removed the latency tax. It landed 2026-08-30 and is running at 98%+ hit rate (2026-09-01: 1,402 cached / 28 fresh). The "~10s of local overhead on every message" problem in CLAUDE.md is solved for the sessions that matter. Fresh classifications average 2,297ms (max 66,319ms — one outlier worth a look, but not a design issue).
  • The pure-module discipline is the best thing in the codebase. routing.py, scoring.py, tiering.py, proficiency.py, context_prune.py are I/O-free, take thresholds as arguments, and are exhaustively tested. That is why 843 tests run offline in 37 seconds, and it is why every one of the measurement-driven reversals in CLAUDE.md was cheap to make. Keep it.

2. The premise quietly changed, and the docs did not follow

Design §7 names the metric: "$ saved vs. always-frontier baseline", with the warning that "a router that's cheaper but quietly worse isn't a win." Nobody has checked the other side of that: a router that's better but quietly more expensive isn't obviously a win either, and that is where this landed.

baseline_report.py already tells the story, and it appears not to have been read recently:

$ PYTHONPATH=src python -m baseline_report --since 2026-08-24
  decisions: 9018
  total cost:   actual $113.89   cheapest $51.39   best-prof $166.59
  dominance (selected == cheapest): 4,972/9,018 (55.1%)
  mean proficiency: actual 0.991   cheapest 0.788   best-prof 0.998

The router sits at 0.991 of a maximum 0.998 on quality and 2.2x the cheapest on cost. It is not a cost/quality tradeoff engine; it is a quality-maximizer that takes a discount when quality is exactly tied. That may be what you want — but it is not what §1 says, and it is not what the README says either ("dispatches to the cheapest/best-fit model").

The comparison that matters more is against the alternative you would actually have used. Replaying all 9,007 chat decisions against every single-model policy (same request shapes, same catalog prices, assumed_cache_rate 0.917; infeasible = requests exceeding that model's effective window):

single-model policy est. cost vs router infeasible
gemma-4-31b $18.91 0.17x 1,170
qwen3.6-35b $19.54 0.17x 4,032
deepseek-v4-flash $36.59 0.32x 0
kimi-k2.7-code $136.41 1.20x 963
glm-5.2-fast / glm-5.3 $260.20 2.28x 0
kimi-k3 $563.97 4.95x 0
(actual routed) $113.89 1.00x 0

Against always-frontier (kimi-k3) the router saves 4.95x — the design's own metric, comfortably met. Against "just always use deepseek-v4-flash" it costs 3.1x more, and deepseek is the only other policy that can serve 100% of the traffic without a context failure.

Which baseline is honest depends entirely on what you would otherwise have done. For opencode agent traffic against this catalog, the realistic counterfactual is not kimi-k3 — it is deepseek. The router should report both baselines, and baseline_report.py should grow a single-model column.

The operational consequence is live right now: 6.51 kWh burned against a plan_kwh_per_period of 6.25. At roughly $7.67/kWh realized, an always-deepseek policy would have used about a third of that and stayed inside the allowance.


3. Finding #1 — the feedback loop rewards models for not being used

This is the most important item in the review, and it is a reason not to take the next obvious step until it is fixed.

987 client outcomes are sitting in verifications; 961 of them unapplied. POST /outcome is correctly identified in CLAUDE.md as "the only ground truth," so folding them in looks like the highest-value pending action. It is not, as currently built.

Simulated on a copy of the DB — run feedback.py for real, then replay all 9,007 decisions through select_candidates + rank_candidates:

BEFORE   est cost $137.76      AFTER   est cost $179.28   (1.30x)
  kimi-k2.7-code      2,722      glm-5.2-fast        2,243
  qwen3.6-35b         2,646      qwen3.6-35b-fast    1,385
  qwen3.6-35b-fast    1,236      glm-5.3             1,318
  glm-5.3               745      kimi-k2.7-code      1,239
  gemma-4-31b           587      qwen3.6-35b         1,150
  deepseek-v4-flash     500      kimi-k2.7-code-fast   947
  glm-5.2-fast          417      deepseek-v4-flash     596

Ingesting ground truth makes routing 30% more expensive and hands 2,243 decisions to glm-5.2-fast, one of the priciest common rows. The mechanism is visible in the scores. coding_refactor, before → after:

model before n after n real outcomes folded in
glm-5.2-fast 0.994 36 0.994 36 zero
kimi-k2.7-code 1.000 18 0.813 80 62
deepseek-v4-flash 0.826 36 0.791 135 99
qwen3.6-35b 0.900 18 0.641 44 26

glm-5.2-fast wins the category outright by never having been measured on real traffic.

The cause is structural, not a bug: feedback.py and eval_proficiency.py both write through proficiency_store.add_self_eval, into the same self_eval_score running mean, with equal weight per sample. Benchmark tasks score ~1.00. Real agent turns score ~0.80. So every real sample drags a score down, and the drag is proportional to how much the model has been used. The router then routes away from the used model toward the unused one, gathers outcomes on that one, penalizes it, and moves on.

That is a rotation, not a convergence. And because exposure has been concentrated on the cheap models — they win the cost tiebreak — the rotation is systematically toward expensive ones.

Directions (pick one; all are cheap):

  1. Keep the two sources separate. Real-traffic outcomes are a different measurement from benchmark tasks and should not share a mean. Add outcome_score / outcome_samples columns and blend them explicitly, the way leaderboard and self_eval already are. This also makes the ~0.80-vs-1.00 scale difference visible instead of silently mixed.
  2. Shrink toward the prior by sample size. A score with n=36 and no real exposure should not outrank one with n=135 that includes 99 real outcomes. Today low sample count is an advantage; self_eval_min_samples gates the leaderboard blend but nothing penalizes thin evidence in the ranking itself.
  3. Compare like with like. Score a model against the per-category mean outcome rate rather than absolute, so an 82% pass rate in a category where everyone scores 80% is neutral, not a 0.18 penalty.

Until one of these lands, do not run feedback.py on the pending backlog. The 961 rows are the most valuable data this project has; spending them through the current path converts them into a more expensive router.


4. Finding #2 — exclusions are self-sealing; there is no exploration

deepseek-v4-flash scored 0.333 on tool_use_agentic across 3 benchmark tasks. Blended, it now reads 0.5. With quality_tolerance: 0.1 that puts it in band 5 while eleven models sit in band 0, so it loses every tool_use_agentic ranking outright.

Consequence, measured: deepseek won 0 of 2,512 tool_use_agentic decisions, and has 0 client outcomes in that category. A 3-sample estimate has permanently removed the cheapest capable model from the largest category of traffic, and the design contains no path by which that estimate can ever be revised.

config/config.yaml proposes the experiment that would settle it:

run with it off, let POST /outcome report real pass/fail, and compare tool_use_agentic proficiency for deepseek before and after

That experiment cannot run. min_tool_proficiency is already null — the filter is off — and it changes nothing, because the quality band excludes deepseek from the category regardless. The config comment says "with it off, that advantage applies." It does not.

The general shape: this is a contextual bandit running pure exploitation. Once a model wins a (category, context-size) cell it wins it forever; alternatives never accumulate the evidence that would overturn the ranking. The consequence is worse than a missed saving — it means every proficiency number is conditioned on the routing that produced it, which is exactly the confound that makes the outcome data below hard to read.

Direction: a small explicit exploration budget. An ε of 2-3% of requests routed to the highest-cost-advantage excluded candidate per category would have produced ~75 deepseek tool_use_agentic samples over this window — enough to confirm or kill the 0.333 outright — at an estimated cost of under $2. Alternatively a "probation" rule: any candidate excluded solely by a proficiency score with self_eval_samples < 10 gets N requests per week regardless.

What the ground truth actually says (with the confound stated)

Client outcomes, Wilson 95% CIs, ≥20 samples:

model n pass
deepseek-v4-flash 390 82.8% [79%, 86%]
kimi-k2.7-code 269 80.3% [75%, 85%]
qwen3.6-35b 198 85.9% [80%, 90%]
kimi-k2.7-code-fast 52 80.8% [68%, 89%]
glm-5.2-fast 44 88.6% [76%, 95%]

Within coding_refactor, where three models have real samples:

model n pass rel. price
deepseek-v4-flash 95 81.1% [72, 88] 1x
kimi-k2.7-code 62 75.8% [64, 85] ~6.8x
qwen3.6-35b 25 48.0% [30, 67] ~2.1x

The cheapest model is at the top, and qwen3.6-35b's CI does not overlap it.

Read this carefully, not triumphantly. The comparison is confounded, and the confound runs against the expensive models: context window determines eligibility, so qwen3.6-35b (94k effective) only ever sees short requests while kimi-k2.7-code (192k) takes the 94-192k band and glm-5.2-fast (782k) takes the largest. Bigger context correlates with longer, harder sessions. So the expensive models are being handed the harder work, and the table cannot separate "cheaper model is as good" from "cheaper model got easier requests."

That is the point. The router has ~1,000 samples of its highest-value signal and cannot draw a conclusion from them, because it never randomizes. Adding exploration is what makes this data interpretable, not just what makes deepseek eligible.


5. Finding #3 — cost is 93% prompt tokens, and the heuristics gate the other 7%

Decomposing the estimated cost of all 9,007 real decisions at assumed_cache_rate: 0.917:

prompt share of est cost:  $105.48  (92.6%)
completion share:            $8.40   (7.4%)

real billed tokens: 1,190,607,380 prompt / 5,880,453 completion  = 202:1

Completion price spans 54x across the catalog ($0.28 → $15.00) and decides 7.4% of the bill. Cached-prompt price spans 21x ($0.0144 → $0.30) and decides 92.6%.

Two consequences:

  • tiering.cheap_completion_max: 1.00 gates tier 1 on the wrong axis. Tier is the single most consequential filter in the system — it decides the eligible set before ranking runs — and it is resolved from completion price plus reasoning mode. This is the same substitution the project has now caught twice ("list price ranks models backwards"; "cheapness is not a capability ceiling"), appearing a third time. For this workload the honest cheapness signal is cost_per_1m_prompt_cached.

  • Pinch is the highest-leverage lever in the codebase and the least instrumented. It is the only thing that touches the 93%. Post-pinch required_context_tokens (line 2450 records measured, computed on send_messages) distributes as:

    p10 49,817   p25 66,097   p50 88,069   p75 127,517   p90 195,683   p99 279,995
    

    against pinch.budget_tokens: 50000. The median request ships at 1.8x the pinch budget and p90 at 3.9x — pinch is running and not reaching its own target, which makes sense given it may only trim tool results outside the last 4 turns while opencode's ~32k system prompt and tool definitions are untouchable. Its stats log at logs.debug while logging.level: info, so zero pinch lines exist in seven days of journal. Nothing in /metrics, the TUI, or the admin dashboard reports pinch effectiveness either.

    A 20% reduction in prompt tokens is worth more than every routing decision in this window combined, and right now there is no way to tell whether pinch delivers 2% or 40%. Promote the pinch line to info, and add tokens_saved / original_tokens to /metrics. That is the cheapest high-value change on this list.

tier decisions avg est cost outcome pass rate (approx. join, n)
1 896 $0.00907 76.9% (13)
2 6,663 $0.01337 83.6% (1,201)
3 1,448 $0.01155 70.4% (226)

Tier 3 costs less per decision than tier 2 and fails more. The failure rate is the good news — it means the classifier's tier call carries real signal about task difficulty. The cost figure is the problem: "frontier / high-stakes" is not buying a more capable model, only a reasoning-enabled subset, because tier 3 is resolved from reasoning_default_enabled while deepseek (1M window, 1.00 on all three coding categories) is pinned at tier 2 and excluded from all 1,448 of them. That is the tier-1 lesson from CLAUDE.md recurring one rung up the ladder. (Outcome join is session_key + 5s window, so treat as indicative.)


6. Smaller structural notes

Sample-depth trap, recurring. CLAUDE.md documents this precisely once already — docs_writing read 0.70-1.00 at n=2 with a model at the ceiling, and six more passes spread it 0.66-0.97 without touching a task. The same shape is sitting in four more categories right now:

category rows at exactly 1.00 mean samples
reasoning_math 16 13 5.4
translation 16 11 3.2
general_chat 13 11 3.5
summarization 16 3 3.4

The lesson was learned and written down; it has not yet been applied to the categories that still show the symptom. "Try samples before hardening" is already the project's own guidance.

"Quality is the objective" is operationally inverted. 72 of 141 proficiency rows sit at exactly 1.00, and 98 of 141 (70%) fall inside the top quality_tolerance band. For most requests quality is a tie and cost is the sole decider, with quality acting as a coarse veto on the worst ~30%. That is a defensible design — arguably the right one — but CLAUDE.md and docs/routing.md read the other way ("there is no weight to tune — quality is the objective and cost is the tiebreak"), which will mislead anyone tuning it.

Structural verification is inert for this traffic. 12,047 of 13,501 verification rows (89.2%) are unverifiable — correctly, since agent turns end in tool calls. It costs nothing (pure Python, post-response), so there is no harm, but it occupies a verdict-mix panel on the dashboard and a section in the docs that together imply coverage that does not exist. feedback.py's own coverage() already prints the warning; the dashboard should too.

dispatcher.py is where the discipline stops. 3,104 lines, with chat_completions at 662 lines in a single function. Every other module in this project is small, pure, and injectable; this one holds routing, pinch, classification, capability checks, local vision, streaming proxy, telemetry sniffing, and verification scheduling in one call frame. It is the one place where the next change is expensive, and the one place a reader cannot hold the whole path in their head. Extracting the pre-dispatch pipeline (prune → measure → classify → gate → rank) into a pure function taking messages + config and returning a decision would bring it in line with the rest and would be directly testable.

Minor doc drift. design/local-llm-model-router.md §4 still presents the weighted composite as the scoring model and §9 item 6 still calls proficiency "the last inert axis" — both superseded. scoring.composite_score and eco_score are dead but tested. README.md:8 and docs/routing.md's "## Weighted Scoring" heading carry the same stale framing. CLAUDE.md already declares itself the source of truth, so this is low-stakes, but the design doc is the file a new reader opens first.


  1. Do not run feedback.py on the 961-row backlog until the exposure bias in §3 is fixed. Separate outcome samples from benchmark samples, or weight by sample size. This is the one item that will actively degrade routing if the obvious next step is taken.
  2. Promote the pinch log to info and surface tokens_saved in /metrics. Cheapest change here, on the axis that carries 93% of the bill.
  3. Add an exploration budget (ε ≈ 2-3%, or probation for candidates excluded on self_eval_samples < 10). Without it, no proficiency number in the table is interpretable and the experiment config.yaml already proposes cannot run.
  4. Add a single-model column to baseline_report.py. "vs always-cheapest" and "vs always-frontier" bracket the answer; "vs always-deepseek" is the decision the operator actually faces.
  5. Re-tier on cost_per_1m_prompt_cached, or drop the price term from tiering entirely and keep the context/reasoning signals.
  6. Run the eval harness for reasoning_math, translation, general_chat, summarization — the same six-pass treatment that spread docs_writing.
  7. Extract the pre-dispatch pipeline out of chat_completions.

8. Bottom line

The premise holds and the machine works. The measurement culture is the real asset here — this project has overturned its own conclusions on evidence repeatedly (list price ranks backwards, tier-from-price, benchmark-vs-real cost, four harness bugs) and each reversal is written down with the reasoning intact. That is rarer than the router.

The gap is that the loop is not yet closed. Every measurement so far has been taken, then reasoned about, then encoded by hand. The two mechanisms that were supposed to close the loop automatically — folding client outcomes into proficiency, and letting the tool-proficiency experiment settle itself — are both currently pointed slightly the wrong way: one rewards models for being unmeasured, the other cannot gather the measurement it needs. Neither is a large fix. Both are the difference between a router that learns from your traffic and one that is very carefully hand-tuned to it.


Appendix — reproduction scripts

Read-only. Each was run against router.db on 2026-09-01 to produce the numbers above. Save to a scratch dir and run with PYTHONPATH=src from the repo root.

A. Single-model cost baselines (§2)

"""What would a single-model policy have cost, on the same 9k real requests?"""
import sqlite3, sys
sys.path.insert(0, "src")
from routing import estimated_cost

conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = {r["model_id"]: dict(r) for r in conn.execute("select * from models")}
rows = conn.execute("""
    select required_context_tokens rc, est_cost_usd, selected_model, task_tier
      from route_decisions where kind='chat' and required_context_tokens is not null
""").fetchall()

CACHE, COMP = 0.917, 500
actual = sum(r["est_cost_usd"] or 0 for r in rows)
print(f"decisions: {len(rows)}   actual router est: ${actual:,.2f}\n")
print(f"{'single-model policy':<26} {'cost':>10} {'vs router':>10} {'infeasible':>11}")
print("-"*62)
res=[]
for mid, m in models.items():
    if m["availability"] != "active" or m["access_level"] != "public" or m["latency_class"]=="flex":
        continue
    eff = m["effective_context_window"] or 0
    tot = 0.0; bad = 0
    for r in rows:
        if r["rc"] > eff:
            bad += 1; continue
        tot += estimated_cost(m, r["rc"], COMP, CACHE) or 0
    res.append((tot, mid, bad))
for tot, mid, bad in sorted(res):
    print(f"{mid:<26} ${tot:>9,.2f} {tot/actual:>9.2f}x {bad:>10,}")

B. Prompt-vs-completion cost anatomy (§5)

import sqlite3, sys, statistics
sys.path.insert(0,"src")
conn=sqlite3.connect("router.db"); conn.row_factory=sqlite3.Row
m={r["model_id"]:dict(r) for r in conn.execute("select * from models")}
rows=conn.execute("""select selected_model sm, required_context_tokens rc, est_cost_usd c
   from route_decisions where kind='chat' and selected_model is not null
   and required_context_tokens is not null""").fetchall()
CACHE,COMP=0.917,500
pt=ct=0.0
for r in rows:
    mm=m.get(r["sm"]);
    if not mm: continue
    p=r["rc"]*((1-CACHE)*mm["cost_per_1m_prompt"]+CACHE*(mm["cost_per_1m_prompt_cached"] or mm["cost_per_1m_prompt"]))/1e6
    c=COMP*mm["cost_per_1m_completion"]/1e6
    pt+=p; ct+=c
print(f"prompt share of est cost: ${pt:.2f} ({pt/(pt+ct)*100:.1f}%)")
print(f"completion share:         ${ct:.2f} ({ct/(pt+ct)*100:.1f}%)")
print()
rc=[r["rc"] for r in rows]
rc.sort()
print("required_context_tokens percentiles:")
for q in (0.1,0.25,0.5,0.75,0.9,0.99):
    print(f"  p{int(q*100):>2}: {rc[int(q*len(rc))]:>9,}")
print(f"  max: {rc[-1]:,}   mean: {statistics.mean(rc):,.0f}")
print()
# actual billed vs prompt size
print("actual billed energy_observations, prompt vs completion tokens:")
r=conn.execute("select sum(prompt_tokens) p, sum(completion_tokens) c, count(*) n from energy_observations where task_category!='seed_reference'").fetchone()
print(f"  prompt {r['p']:,}   completion {r['c']:,}   ratio {r['p']/r['c']:.0f}:1   n={r['n']:,}")

C. Does category change the winner? (§1)

"""Does the classifier's CATEGORY output change the routing decision?

Replays every real chat decision twice: once with per-category proficiency
(what ships), once with a category-agnostic mean proficiency per model.
"""
import sqlite3, sys, statistics
sys.path.insert(0,"src")
from config import load_config
from routing import select_candidates, rank_candidates

cfg = load_config("config/config.yaml")
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = [dict(r) for r in conn.execute("select * from models")]
prof = {(r["model_id"], r["category"]): r["blended_score"]
        for r in conn.execute("select model_id, category, blended_score from proficiency")}
# category-agnostic: mean of a model's blended scores across all categories
avg = {}
for (m,c),v in prof.items():
    if v is not None: avg.setdefault(m,[]).append(v)
avg = {m: statistics.fmean(v) for m,v in avg.items()}

decisions = conn.execute("""select task_category c, task_tier t, required_context_tokens rc,
       latency_tolerance lt, selected_model sm from route_decisions
       where kind='chat' and selected_model is not null and task_category is not null
       and required_context_tokens is not null and task_tier is not null""").fetchall()

def pick(d, use_category):
    rows=[]
    for m in models:
        r=dict(m)
        r["proficiency"] = prof.get((m["model_id"], d["c"])) if use_category else avg.get(m["model_id"])
        rows.append(r)
    cand = select_candidates(rows,
        required_context_tokens=d["rc"], required_tier=d["t"],
        latency_tolerance=d["lt"] or "interactive",
        allowed_access_levels=cfg.routing.allowed_access_levels,
        exclude_stale=cfg.freshness.exclude_stale,
        exclude_deprecated=cfg.freshness.exclude_deprecated)
    rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
        prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
        cache_rate=cfg.objective.assumed_cache_rate)
    return (rk[0]["model_id"], rk[0]["cost"]) if rk else (None,0)

same=diff=0; cost_cat=cost_flat=0.0
for d in decisions:
    a,ca = pick(d, True); b,cb = pick(d, False)
    cost_cat+=ca; cost_flat+=cb
    if a==b: same+=1
    else: diff+=1
n=len(decisions)
print(f"replayed {n} real decisions")
print(f"  category-aware == category-agnostic: {same}/{n} ({same/n*100:.1f}%)")
print(f"  category changed the winner:         {diff}/{n} ({diff/n*100:.1f}%)")
print(f"  cost with category:    ${cost_cat:,.2f}")
print(f"  cost without category: ${cost_flat:,.2f}  ({cost_flat/cost_cat:.2f}x)")

D. Simulated feedback fold-in (§3)

Copies the DB first; never mutates router.db.

"""Simulate: apply the 1,039 pending client outcomes, then replay routing."""
import sqlite3, sys, os, shutil
sys.path.insert(0,"src")
from config import load_config
import feedback
from routing import select_candidates, rank_candidates

cfg = load_config("config/config.yaml")
db = os.environ["CLAUDE_JOB_DIR"] + "/tmp/sim.db"

def replay(dbpath, label):
    conn = sqlite3.connect(dbpath); conn.row_factory = sqlite3.Row
    models=[dict(r) for r in conn.execute("select * from models")]
    prof={(r["model_id"],r["category"]):r["blended_score"]
          for r in conn.execute("select model_id,category,blended_score from proficiency")}
    ds=conn.execute("""select task_category c,task_tier t,required_context_tokens rc,
        latency_tolerance lt from route_decisions where kind='chat'
        and task_category is not null and required_context_tokens is not null
        and task_tier is not null""").fetchall()
    tot=0.0; mix={}
    for d in ds:
        rows=[{**m,"proficiency":prof.get((m["model_id"],d["c"]))} for m in models]
        cand=select_candidates(rows,required_context_tokens=d["rc"],required_tier=d["t"],
            latency_tolerance=d["lt"] or "interactive",
            allowed_access_levels=cfg.routing.allowed_access_levels,
            exclude_stale=cfg.freshness.exclude_stale,
            exclude_deprecated=cfg.freshness.exclude_deprecated)
        rk=rank_candidates(cand,quality_tolerance=cfg.objective.quality_tolerance,
            prompt_tokens=d["rc"],completion_tokens=cfg.objective.assumed_completion_tokens,
            cache_rate=cfg.objective.assumed_cache_rate)
        if rk:
            tot+=rk[0]["cost"] or 0; mix[rk[0]["model_id"]]=mix.get(rk[0]["model_id"],0)+1
    conn.close()
    print(f"\n=== {label} ===  replayed {len(ds)} decisions   est cost ${tot:,.2f}")
    for m,n in sorted(mix.items(),key=lambda kv:-kv[1]):
        print(f"   {m:<24}{n:>6}")
    return tot

before = replay(db, "BEFORE (current proficiency)")
conn=sqlite3.connect(db)
rows=feedback.unapplied_failures(conn)
feedback.apply_failures(conn, cfg, feedback.summarize(rows), dry_run=False)
conn.close()
after = replay(db, "AFTER folding in 1,039 client outcomes")
print(f"\ncost change: ${before:,.2f} -> ${after:,.2f}  ({after/before:.2f}x)")

E. Client-outcome pass rates with Wilson CIs (§4)

import sqlite3, math
conn=sqlite3.connect("router.db"); conn.row_factory=sqlite3.Row
q="""select model_id m, task_category c,
      sum(verdict='succeeded') s, sum(verdict='failed') f
      from verifications where kind='client_outcome' group by 1,2"""
rows=[dict(r) for r in conn.execute(q)]
def wilson(s,n):
    if n==0: return (0,0)
    z=1.96; p=s/n; d=1+z*z/n
    c=(p+z*z/(2*n))/d; h=z*math.sqrt(p*(1-p)/n+z*z/(4*n*n))/d
    return (c-h,c+h)
print(f"{'model':<22}{'category':<18}{'n':>5}{'pass':>7}   95% CI")
print("-"*66)
for r in sorted(rows,key=lambda r:(-(r['s']+r['f']))):
    n=r['s']+r['f']
    if n<15: continue
    lo,hi=wilson(r['s'],n)
    print(f"{r['m']:<22}{r['c']:<18}{n:>5}{r['s']/n*100:>6.1f}%   [{lo*100:.0f}%, {hi*100:.0f}%]")
print()
print("--- pooled per model (all categories) ---")
agg={}
for r in rows:
    a=agg.setdefault(r['m'],[0,0]); a[0]+=r['s']; a[1]+=r['f']
for m,(s,f) in sorted(agg.items(),key=lambda kv:-(kv[1][0]+kv[1][1])):
    n=s+f
    if n<20: continue
    lo,hi=wilson(s,n)
    print(f"{m:<22}{n:>5}{s/n*100:>6.1f}%   [{lo*100:.0f}%, {hi*100:.0f}%]")