# Conceptual review: is the premise sound, and is execution converging on it? Status: done -- review of shipped work **Date:** 2026-09-01 **Scope:** premise and architecture, not bugs. Every number below is measured from the live `router.db` and the checked-in code on this date; the scripts are inlined so they can be re-run. **Verdict:** the premise is sound and the hard part is built. Three structural problems sit between the current state and the stated goal, and one of them will actively move routing the wrong way the moment the next obvious step is taken. --- ## 0. What the evidence base is | | | |---|---| | `route_decisions` | 9,025 rows (9,007 `chat`), 2026-08-24 → 2026-09-01 | | `energy_observations` | 13,966 rows, **$49.94** billed, **6.51 kWh**, 1,122 gCO₂eq, from 2026-08-17 | | `verifications` | 13,501 rows; 987 of them client outcomes (816 succeeded / 171 failed) | | `proficiency` | 141 (model, category) rows | | tests | **843 pass in 36.8s**, all offline | This is a real deployment with real traffic, not a demo. The review is worth doing precisely because there is now enough data to check the design against itself. --- ## 1. The premise is sound, and the hardest claim in it is now proven Design doc §1: use a local model to classify the task, then route to the cheapest/best-fit open-weight model instead of defaulting everything to one expensive model. The load-bearing, non-obvious claim in that sentence is that **task category should change the model**. For a long stretch it demonstrably did not — the classifier computed a category, the router paid ~10s for it, and `proficiency` was empty so it changed nothing. That is fixed, and it is measurable now. Replaying all 9,001 real decisions with per-category proficiency against a category-agnostic mean: ``` category-aware == category-agnostic: 5,917/9,001 (65.7%) category changed the winner: 3,084/9,001 (34.3%) ``` **One decision in three turns on the category.** That is the design's central bet paying off, and it took populating `proficiency`, fixing four harness bugs, and rebuilding the cost axis to get there. It should be stated as a result. Two more things are working and deserve to be said plainly before the criticism: - **The session cache removed the latency tax.** It landed 2026-08-30 and is running at 98%+ hit rate (2026-09-01: 1,402 cached / 28 fresh). The "~10s of local overhead on every message" problem in CLAUDE.md is solved for the sessions that matter. Fresh classifications average 2,297ms (max 66,319ms — one outlier worth a look, but not a design issue). - **The pure-module discipline is the best thing in the codebase.** `routing.py`, `scoring.py`, `tiering.py`, `proficiency.py`, `context_prune.py` are I/O-free, take thresholds as arguments, and are exhaustively tested. That is why 843 tests run offline in 37 seconds, and it is why every one of the measurement-driven reversals in CLAUDE.md was cheap to make. Keep it. --- ## 2. The premise quietly changed, and the docs did not follow Design §7 names the metric: **"$ saved vs. always-frontier baseline"**, with the warning that "a router that's cheaper but quietly worse isn't a win." Nobody has checked the other side of that: a router that's *better but quietly more expensive* isn't obviously a win either, and that is where this landed. `baseline_report.py` already tells the story, and it appears not to have been read recently: ``` $ PYTHONPATH=src python -m baseline_report --since 2026-08-24 decisions: 9018 total cost: actual $113.89 cheapest $51.39 best-prof $166.59 dominance (selected == cheapest): 4,972/9,018 (55.1%) mean proficiency: actual 0.991 cheapest 0.788 best-prof 0.998 ``` The router sits at **0.991 of a maximum 0.998** on quality and **2.2x the cheapest** on cost. It is not a cost/quality tradeoff engine; it is a quality-maximizer that takes a discount when quality is exactly tied. That may be what you want — but it is not what §1 says, and it is not what the README says either ("dispatches to the cheapest/best-fit model"). The comparison that matters more is against the alternative you would actually have used. Replaying all 9,007 chat decisions against every single-model policy (same request shapes, same catalog prices, `assumed_cache_rate` 0.917; `infeasible` = requests exceeding that model's effective window): | single-model policy | est. cost | vs router | infeasible | |---|---|---|---| | `gemma-4-31b` | $18.91 | 0.17x | 1,170 | | `qwen3.6-35b` | $19.54 | 0.17x | 4,032 | | **`deepseek-v4-flash`** | **$36.59** | **0.32x** | **0** | | `kimi-k2.7-code` | $136.41 | 1.20x | 963 | | `glm-5.2-fast` / `glm-5.3` | $260.20 | 2.28x | 0 | | `kimi-k3` | $563.97 | 4.95x | 0 | | *(actual routed)* | *$113.89* | *1.00x* | *0* | Against always-frontier (`kimi-k3`) the router saves **4.95x** — the design's own metric, comfortably met. Against "just always use `deepseek-v4-flash`" it costs **3.1x more**, and deepseek is the only other policy that can serve 100% of the traffic without a context failure. Which baseline is honest depends entirely on what you would otherwise have done. For opencode agent traffic against this catalog, the realistic counterfactual is not `kimi-k3` — it is deepseek. **The router should report both baselines, and `baseline_report.py` should grow a single-model column.** The operational consequence is live right now: **6.51 kWh burned against a `plan_kwh_per_period` of 6.25.** At roughly $7.67/kWh realized, an always-deepseek policy would have used about a third of that and stayed inside the allowance. --- ## 3. Finding #1 — the feedback loop rewards models for not being used **This is the most important item in the review, and it is a reason not to take the next obvious step until it is fixed.** 987 client outcomes are sitting in `verifications`; 961 of them unapplied. `POST /outcome` is correctly identified in CLAUDE.md as "the only ground truth," so folding them in looks like the highest-value pending action. It is not, as currently built. Simulated on a copy of the DB — run `feedback.py` for real, then replay all 9,007 decisions through `select_candidates` + `rank_candidates`: ``` BEFORE est cost $137.76 AFTER est cost $179.28 (1.30x) kimi-k2.7-code 2,722 glm-5.2-fast 2,243 qwen3.6-35b 2,646 qwen3.6-35b-fast 1,385 qwen3.6-35b-fast 1,236 glm-5.3 1,318 glm-5.3 745 kimi-k2.7-code 1,239 gemma-4-31b 587 qwen3.6-35b 1,150 deepseek-v4-flash 500 kimi-k2.7-code-fast 947 glm-5.2-fast 417 deepseek-v4-flash 596 ``` Ingesting ground truth makes routing **30% more expensive** and hands 2,243 decisions to `glm-5.2-fast`, one of the priciest common rows. The mechanism is visible in the scores. `coding_refactor`, before → after: | model | before | n | after | n | real outcomes folded in | |---|---|---|---|---|---| | `glm-5.2-fast` | 0.994 | 36 | **0.994** | 36 | **zero** | | `kimi-k2.7-code` | 1.000 | 18 | 0.813 | 80 | 62 | | `deepseek-v4-flash` | 0.826 | 36 | 0.791 | 135 | 99 | | `qwen3.6-35b` | 0.900 | 18 | 0.641 | 44 | 26 | **`glm-5.2-fast` wins the category outright by never having been measured on real traffic.** The cause is structural, not a bug: `feedback.py` and `eval_proficiency.py` both write through `proficiency_store.add_self_eval`, into the *same* `self_eval_score` running mean, with equal weight per sample. Benchmark tasks score ~1.00. Real agent turns score ~0.80. So every real sample drags a score down, and the drag is proportional to how much the model has been used. The router then routes away from the used model toward the unused one, gathers outcomes on *that* one, penalizes it, and moves on. That is a rotation, not a convergence. And because exposure has been concentrated on the cheap models — they win the cost tiebreak — the rotation is systematically toward expensive ones. **Directions (pick one; all are cheap):** 1. **Keep the two sources separate.** Real-traffic outcomes are a different measurement from benchmark tasks and should not share a mean. Add `outcome_score` / `outcome_samples` columns and blend them explicitly, the way `leaderboard` and `self_eval` already are. This also makes the ~0.80-vs-1.00 scale difference visible instead of silently mixed. 2. **Shrink toward the prior by sample size.** A score with n=36 and no real exposure should not outrank one with n=135 that includes 99 real outcomes. Today low sample count is an *advantage*; `self_eval_min_samples` gates the leaderboard blend but nothing penalizes thin evidence in the ranking itself. 3. **Compare like with like.** Score a model against the *per-category mean* outcome rate rather than absolute, so an 82% pass rate in a category where everyone scores 80% is neutral, not a 0.18 penalty. Until one of these lands, **do not run `feedback.py` on the pending backlog.** The 961 rows are the most valuable data this project has; spending them through the current path converts them into a more expensive router. --- ## 4. Finding #2 — exclusions are self-sealing; there is no exploration `deepseek-v4-flash` scored **0.333 on `tool_use_agentic`** across **3** benchmark tasks. Blended, it now reads 0.5. With `quality_tolerance: 0.1` that puts it in band 5 while eleven models sit in band 0, so it loses every `tool_use_agentic` ranking outright. Consequence, measured: **deepseek won 0 of 2,512 `tool_use_agentic` decisions, and has 0 client outcomes in that category.** A 3-sample estimate has permanently removed the cheapest capable model from the largest category of traffic, and the design contains no path by which that estimate can ever be revised. `config/config.yaml` proposes the experiment that would settle it: > run with it off, let `POST /outcome` report real pass/fail, and compare > `tool_use_agentic` proficiency for deepseek before and after **That experiment cannot run.** `min_tool_proficiency` is already `null` — the filter is off — and it changes nothing, because the *quality band* excludes deepseek from the category regardless. The config comment says "with it off, that advantage applies." It does not. The general shape: this is a contextual bandit running pure exploitation. Once a model wins a (category, context-size) cell it wins it forever; alternatives never accumulate the evidence that would overturn the ranking. The consequence is worse than a missed saving — it means every proficiency number is conditioned on the routing that produced it, which is exactly the confound that makes the outcome data below hard to read. **Direction:** a small explicit exploration budget. An ε of 2-3% of requests routed to the highest-cost-advantage *excluded* candidate per category would have produced ~75 deepseek `tool_use_agentic` samples over this window — enough to confirm or kill the 0.333 outright — at an estimated cost of under $2. Alternatively a "probation" rule: any candidate excluded solely by a proficiency score with `self_eval_samples < 10` gets N requests per week regardless. ### What the ground truth actually says (with the confound stated) Client outcomes, Wilson 95% CIs, ≥20 samples: | model | n | pass | | |---|---|---|---| | `deepseek-v4-flash` | 390 | **82.8%** | [79%, 86%] | | `kimi-k2.7-code` | 269 | 80.3% | [75%, 85%] | | `qwen3.6-35b` | 198 | 85.9% | [80%, 90%] | | `kimi-k2.7-code-fast` | 52 | 80.8% | [68%, 89%] | | `glm-5.2-fast` | 44 | 88.6% | [76%, 95%] | Within `coding_refactor`, where three models have real samples: | model | n | pass | rel. price | |---|---|---|---| | `deepseek-v4-flash` | 95 | **81.1%** [72, 88] | 1x | | `kimi-k2.7-code` | 62 | 75.8% [64, 85] | ~6.8x | | `qwen3.6-35b` | 25 | 48.0% [30, 67] | ~2.1x | The cheapest model is at the top, and `qwen3.6-35b`'s CI does not overlap it. **Read this carefully, not triumphantly.** The comparison is confounded, and the confound runs *against* the expensive models: context window determines eligibility, so `qwen3.6-35b` (94k effective) only ever sees short requests while `kimi-k2.7-code` (192k) takes the 94-192k band and `glm-5.2-fast` (782k) takes the largest. Bigger context correlates with longer, harder sessions. So the expensive models are being handed the harder work, and the table cannot separate "cheaper model is as good" from "cheaper model got easier requests." That is the point. **The router has ~1,000 samples of its highest-value signal and cannot draw a conclusion from them, because it never randomizes.** Adding exploration is what makes this data interpretable, not just what makes deepseek eligible. --- ## 5. Finding #3 — cost is 93% prompt tokens, and the heuristics gate the other 7% Decomposing the estimated cost of all 9,007 real decisions at `assumed_cache_rate: 0.917`: ``` prompt share of est cost: $105.48 (92.6%) completion share: $8.40 (7.4%) real billed tokens: 1,190,607,380 prompt / 5,880,453 completion = 202:1 ``` Completion price spans 54x across the catalog ($0.28 → $15.00) and decides **7.4%** of the bill. Cached-prompt price spans 21x ($0.0144 → $0.30) and decides **92.6%**. Two consequences: - **`tiering.cheap_completion_max: 1.00` gates tier 1 on the wrong axis.** Tier is the single most consequential filter in the system — it decides the eligible set before ranking runs — and it is resolved from completion price plus reasoning mode. This is the same substitution the project has now caught twice ("list price ranks models backwards"; "cheapness is not a capability ceiling"), appearing a third time. For this workload the honest cheapness signal is `cost_per_1m_prompt_cached`. - **Pinch is the highest-leverage lever in the codebase and the least instrumented.** It is the only thing that touches the 93%. Post-pinch `required_context_tokens` (line 2450 records `measured`, computed on `send_messages`) distributes as: ``` p10 49,817 p25 66,097 p50 88,069 p75 127,517 p90 195,683 p99 279,995 ``` against `pinch.budget_tokens: 50000`. **The median request ships at 1.8x the pinch budget and p90 at 3.9x** — pinch is running and not reaching its own target, which makes sense given it may only trim tool results outside the last 4 turns while opencode's ~32k system prompt and tool definitions are untouchable. Its stats log at `logs.debug` while `logging.level: info`, so zero pinch lines exist in seven days of journal. Nothing in `/metrics`, the TUI, or the admin dashboard reports pinch effectiveness either. A 20% reduction in prompt tokens is worth more than every routing decision in this window combined, and right now there is no way to tell whether pinch delivers 2% or 40%. **Promote the pinch line to info, and add `tokens_saved` / `original_tokens` to `/metrics`.** That is the cheapest high-value change on this list. ### Related: the tier ladder does not produce a capability gradient | tier | decisions | avg est cost | outcome pass rate (approx. join, n) | |---|---|---|---| | 1 | 896 | $0.00907 | 76.9% (13) | | 2 | 6,663 | $0.01337 | 83.6% (1,201) | | 3 | 1,448 | $0.01155 | **70.4%** (226) | Tier 3 costs *less* per decision than tier 2 and fails *more*. The failure rate is the good news — it means the classifier's tier call carries real signal about task difficulty. The cost figure is the problem: "frontier / high-stakes" is not buying a more capable model, only a reasoning-enabled subset, because tier 3 is resolved from `reasoning_default_enabled` while deepseek (1M window, 1.00 on all three coding categories) is pinned at tier 2 and excluded from all 1,448 of them. That is the tier-1 lesson from CLAUDE.md recurring one rung up the ladder. *(Outcome join is `session_key` + 5s window, so treat as indicative.)* --- ## 6. Smaller structural notes **Sample-depth trap, recurring.** CLAUDE.md documents this precisely once already — `docs_writing` read 0.70-1.00 at n=2 with a model at the ceiling, and six more passes spread it 0.66-0.97 without touching a task. The same shape is sitting in four more categories right now: | category | rows | at exactly 1.00 | mean samples | |---|---|---|---| | `reasoning_math` | 16 | 13 | 5.4 | | `translation` | 16 | 11 | 3.2 | | `general_chat` | 13 | 11 | 3.5 | | `summarization` | 16 | 3 | 3.4 | The lesson was learned and written down; it has not yet been applied to the categories that still show the symptom. "Try samples before hardening" is already the project's own guidance. **"Quality is the objective" is operationally inverted.** 72 of 141 proficiency rows sit at exactly 1.00, and 98 of 141 (70%) fall inside the top `quality_tolerance` band. For most requests quality is a tie and **cost is the sole decider**, with quality acting as a coarse veto on the worst ~30%. That is a defensible design — arguably the right one — but CLAUDE.md and `docs/routing.md` read the other way ("there is no weight to tune — quality is the objective and cost is the tiebreak"), which will mislead anyone tuning it. **Structural verification is inert for this traffic.** 12,047 of 13,501 verification rows (89.2%) are `unverifiable` — correctly, since agent turns end in tool calls. It costs nothing (pure Python, post-response), so there is no harm, but it occupies a verdict-mix panel on the dashboard and a section in the docs that together imply coverage that does not exist. `feedback.py`'s own `coverage()` already prints the warning; the dashboard should too. **`dispatcher.py` is where the discipline stops.** 3,104 lines, with `chat_completions` at **662 lines** in a single function. Every other module in this project is small, pure, and injectable; this one holds routing, pinch, classification, capability checks, local vision, streaming proxy, telemetry sniffing, and verification scheduling in one call frame. It is the one place where the next change is expensive, and the one place a reader cannot hold the whole path in their head. Extracting the pre-dispatch pipeline (prune → measure → classify → gate → rank) into a pure function taking messages + config and returning a decision would bring it in line with the rest and would be directly testable. **Minor doc drift.** `design/local-llm-model-router.md` §4 still presents the weighted composite as the scoring model and §9 item 6 still calls proficiency "the last inert axis" — both superseded. `scoring.composite_score` and `eco_score` are dead but tested. `README.md:8` and `docs/routing.md`'s "## Weighted Scoring" heading carry the same stale framing. CLAUDE.md already declares itself the source of truth, so this is low-stakes, but the design doc is the file a new reader opens first. --- ## 7. Recommended order 1. **Do not run `feedback.py` on the 961-row backlog** until the exposure bias in §3 is fixed. Separate outcome samples from benchmark samples, or weight by sample size. This is the one item that will actively degrade routing if the obvious next step is taken. 2. **Promote the pinch log to `info` and surface `tokens_saved` in `/metrics`.** Cheapest change here, on the axis that carries 93% of the bill. 3. **Add an exploration budget** (ε ≈ 2-3%, or probation for candidates excluded on `self_eval_samples < 10`). Without it, no proficiency number in the table is interpretable and the experiment `config.yaml` already proposes cannot run. 4. **Add a single-model column to `baseline_report.py`.** "vs always-cheapest" and "vs always-frontier" bracket the answer; "vs always-deepseek" is the decision the operator actually faces. 5. **Re-tier on `cost_per_1m_prompt_cached`**, or drop the price term from tiering entirely and keep the context/reasoning signals. 6. Run the eval harness for `reasoning_math`, `translation`, `general_chat`, `summarization` — the same six-pass treatment that spread `docs_writing`. 7. Extract the pre-dispatch pipeline out of `chat_completions`. --- ## 8. Bottom line The premise holds and the machine works. The measurement culture is the real asset here — this project has overturned its own conclusions on evidence repeatedly (list price ranks backwards, tier-from-price, benchmark-vs-real cost, four harness bugs) and each reversal is written down with the reasoning intact. That is rarer than the router. The gap is that the loop is not yet closed. Every measurement so far has been *taken*, then reasoned about, then encoded by hand. The two mechanisms that were supposed to close the loop automatically — folding client outcomes into proficiency, and letting the tool-proficiency experiment settle itself — are both currently pointed slightly the wrong way: one rewards models for being unmeasured, the other cannot gather the measurement it needs. Neither is a large fix. Both are the difference between a router that learns from your traffic and one that is very carefully hand-tuned to it. --- ## Appendix — reproduction scripts Read-only. Each was run against `router.db` on 2026-09-01 to produce the numbers above. Save to a scratch dir and run with `PYTHONPATH=src` from the repo root. ### A. Single-model cost baselines (§2) ```python """What would a single-model policy have cost, on the same 9k real requests?""" import sqlite3, sys sys.path.insert(0, "src") from routing import estimated_cost conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row models = {r["model_id"]: dict(r) for r in conn.execute("select * from models")} rows = conn.execute(""" select required_context_tokens rc, est_cost_usd, selected_model, task_tier from route_decisions where kind='chat' and required_context_tokens is not null """).fetchall() CACHE, COMP = 0.917, 500 actual = sum(r["est_cost_usd"] or 0 for r in rows) print(f"decisions: {len(rows)} actual router est: ${actual:,.2f}\n") print(f"{'single-model policy':<26} {'cost':>10} {'vs router':>10} {'infeasible':>11}") print("-"*62) res=[] for mid, m in models.items(): if m["availability"] != "active" or m["access_level"] != "public" or m["latency_class"]=="flex": continue eff = m["effective_context_window"] or 0 tot = 0.0; bad = 0 for r in rows: if r["rc"] > eff: bad += 1; continue tot += estimated_cost(m, r["rc"], COMP, CACHE) or 0 res.append((tot, mid, bad)) for tot, mid, bad in sorted(res): print(f"{mid:<26} ${tot:>9,.2f} {tot/actual:>9.2f}x {bad:>10,}") ``` ### B. Prompt-vs-completion cost anatomy (§5) ```python import sqlite3, sys, statistics sys.path.insert(0,"src") conn=sqlite3.connect("router.db"); conn.row_factory=sqlite3.Row m={r["model_id"]:dict(r) for r in conn.execute("select * from models")} rows=conn.execute("""select selected_model sm, required_context_tokens rc, est_cost_usd c from route_decisions where kind='chat' and selected_model is not null and required_context_tokens is not null""").fetchall() CACHE,COMP=0.917,500 pt=ct=0.0 for r in rows: mm=m.get(r["sm"]); if not mm: continue p=r["rc"]*((1-CACHE)*mm["cost_per_1m_prompt"]+CACHE*(mm["cost_per_1m_prompt_cached"] or mm["cost_per_1m_prompt"]))/1e6 c=COMP*mm["cost_per_1m_completion"]/1e6 pt+=p; ct+=c print(f"prompt share of est cost: ${pt:.2f} ({pt/(pt+ct)*100:.1f}%)") print(f"completion share: ${ct:.2f} ({ct/(pt+ct)*100:.1f}%)") print() rc=[r["rc"] for r in rows] rc.sort() print("required_context_tokens percentiles:") for q in (0.1,0.25,0.5,0.75,0.9,0.99): print(f" p{int(q*100):>2}: {rc[int(q*len(rc))]:>9,}") print(f" max: {rc[-1]:,} mean: {statistics.mean(rc):,.0f}") print() # actual billed vs prompt size print("actual billed energy_observations, prompt vs completion tokens:") r=conn.execute("select sum(prompt_tokens) p, sum(completion_tokens) c, count(*) n from energy_observations where task_category!='seed_reference'").fetchone() print(f" prompt {r['p']:,} completion {r['c']:,} ratio {r['p']/r['c']:.0f}:1 n={r['n']:,}") ``` ### C. Does category change the winner? (§1) ```python """Does the classifier's CATEGORY output change the routing decision? Replays every real chat decision twice: once with per-category proficiency (what ships), once with a category-agnostic mean proficiency per model. """ import sqlite3, sys, statistics sys.path.insert(0,"src") from config import load_config from routing import select_candidates, rank_candidates cfg = load_config("config/config.yaml") conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row models = [dict(r) for r in conn.execute("select * from models")] prof = {(r["model_id"], r["category"]): r["blended_score"] for r in conn.execute("select model_id, category, blended_score from proficiency")} # category-agnostic: mean of a model's blended scores across all categories avg = {} for (m,c),v in prof.items(): if v is not None: avg.setdefault(m,[]).append(v) avg = {m: statistics.fmean(v) for m,v in avg.items()} decisions = conn.execute("""select task_category c, task_tier t, required_context_tokens rc, latency_tolerance lt, selected_model sm from route_decisions where kind='chat' and selected_model is not null and task_category is not null and required_context_tokens is not null and task_tier is not null""").fetchall() def pick(d, use_category): rows=[] for m in models: r=dict(m) r["proficiency"] = prof.get((m["model_id"], d["c"])) if use_category else avg.get(m["model_id"]) rows.append(r) cand = select_candidates(rows, required_context_tokens=d["rc"], required_tier=d["t"], latency_tolerance=d["lt"] or "interactive", allowed_access_levels=cfg.routing.allowed_access_levels, exclude_stale=cfg.freshness.exclude_stale, exclude_deprecated=cfg.freshness.exclude_deprecated) rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance, prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens, cache_rate=cfg.objective.assumed_cache_rate) return (rk[0]["model_id"], rk[0]["cost"]) if rk else (None,0) same=diff=0; cost_cat=cost_flat=0.0 for d in decisions: a,ca = pick(d, True); b,cb = pick(d, False) cost_cat+=ca; cost_flat+=cb if a==b: same+=1 else: diff+=1 n=len(decisions) print(f"replayed {n} real decisions") print(f" category-aware == category-agnostic: {same}/{n} ({same/n*100:.1f}%)") print(f" category changed the winner: {diff}/{n} ({diff/n*100:.1f}%)") print(f" cost with category: ${cost_cat:,.2f}") print(f" cost without category: ${cost_flat:,.2f} ({cost_flat/cost_cat:.2f}x)") ``` ### D. Simulated feedback fold-in (§3) Copies the DB first; never mutates `router.db`. ```python """Simulate: apply the 1,039 pending client outcomes, then replay routing.""" import sqlite3, sys, os, shutil sys.path.insert(0,"src") from config import load_config import feedback from routing import select_candidates, rank_candidates cfg = load_config("config/config.yaml") db = os.environ["CLAUDE_JOB_DIR"] + "/tmp/sim.db" def replay(dbpath, label): conn = sqlite3.connect(dbpath); conn.row_factory = sqlite3.Row models=[dict(r) for r in conn.execute("select * from models")] prof={(r["model_id"],r["category"]):r["blended_score"] for r in conn.execute("select model_id,category,blended_score from proficiency")} ds=conn.execute("""select task_category c,task_tier t,required_context_tokens rc, latency_tolerance lt from route_decisions where kind='chat' and task_category is not null and required_context_tokens is not null and task_tier is not null""").fetchall() tot=0.0; mix={} for d in ds: rows=[{**m,"proficiency":prof.get((m["model_id"],d["c"]))} for m in models] cand=select_candidates(rows,required_context_tokens=d["rc"],required_tier=d["t"], latency_tolerance=d["lt"] or "interactive", allowed_access_levels=cfg.routing.allowed_access_levels, exclude_stale=cfg.freshness.exclude_stale, exclude_deprecated=cfg.freshness.exclude_deprecated) rk=rank_candidates(cand,quality_tolerance=cfg.objective.quality_tolerance, prompt_tokens=d["rc"],completion_tokens=cfg.objective.assumed_completion_tokens, cache_rate=cfg.objective.assumed_cache_rate) if rk: tot+=rk[0]["cost"] or 0; mix[rk[0]["model_id"]]=mix.get(rk[0]["model_id"],0)+1 conn.close() print(f"\n=== {label} === replayed {len(ds)} decisions est cost ${tot:,.2f}") for m,n in sorted(mix.items(),key=lambda kv:-kv[1]): print(f" {m:<24}{n:>6}") return tot before = replay(db, "BEFORE (current proficiency)") conn=sqlite3.connect(db) rows=feedback.unapplied_failures(conn) feedback.apply_failures(conn, cfg, feedback.summarize(rows), dry_run=False) conn.close() after = replay(db, "AFTER folding in 1,039 client outcomes") print(f"\ncost change: ${before:,.2f} -> ${after:,.2f} ({after/before:.2f}x)") ``` ### E. Client-outcome pass rates with Wilson CIs (§4) ```python import sqlite3, math conn=sqlite3.connect("router.db"); conn.row_factory=sqlite3.Row q="""select model_id m, task_category c, sum(verdict='succeeded') s, sum(verdict='failed') f from verifications where kind='client_outcome' group by 1,2""" rows=[dict(r) for r in conn.execute(q)] def wilson(s,n): if n==0: return (0,0) z=1.96; p=s/n; d=1+z*z/n c=(p+z*z/(2*n))/d; h=z*math.sqrt(p*(1-p)/n+z*z/(4*n*n))/d return (c-h,c+h) print(f"{'model':<22}{'category':<18}{'n':>5}{'pass':>7} 95% CI") print("-"*66) for r in sorted(rows,key=lambda r:(-(r['s']+r['f']))): n=r['s']+r['f'] if n<15: continue lo,hi=wilson(r['s'],n) print(f"{r['m']:<22}{r['c']:<18}{n:>5}{r['s']/n*100:>6.1f}% [{lo*100:.0f}%, {hi*100:.0f}%]") print() print("--- pooled per model (all categories) ---") agg={} for r in rows: a=agg.setdefault(r['m'],[0,0]); a[0]+=r['s']; a[1]+=r['f'] for m,(s,f) in sorted(agg.items(),key=lambda kv:-(kv[1][0]+kv[1][1])): n=s+f if n<20: continue lo,hi=wilson(s,n) print(f"{m:<22}{n:>5}{s/n*100:>6.1f}% [{lo*100:.0f}%, {hi*100:.0f}%]") ```