Files
6krrt/plans/conceptual-review-premise-and-execution.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

655 lines
30 KiB
Markdown

# Conceptual review: is the premise sound, and is execution converging on it?
Status: done -- review of shipped work
**Date:** 2026-09-01
**Scope:** premise and architecture, not bugs. Every number below is measured
from the live `router.db` and the checked-in code on this date; the scripts are
inlined so they can be re-run.
**Verdict:** the premise is sound and the hard part is built. Three structural
problems sit between the current state and the stated goal, and one of them
will actively move routing the wrong way the moment the next obvious step is
taken.
---
## 0. What the evidence base is
| | |
|---|---|
| `route_decisions` | 9,025 rows (9,007 `chat`), 2026-08-24 → 2026-09-01 |
| `energy_observations` | 13,966 rows, **$49.94** billed, **6.51 kWh**, 1,122 gCO₂eq, from 2026-08-17 |
| `verifications` | 13,501 rows; 987 of them client outcomes (816 succeeded / 171 failed) |
| `proficiency` | 141 (model, category) rows |
| tests | **843 pass in 36.8s**, all offline |
This is a real deployment with real traffic, not a demo. The review is worth
doing precisely because there is now enough data to check the design against
itself.
---
## 1. The premise is sound, and the hardest claim in it is now proven
Design doc §1: use a local model to classify the task, then route to the
cheapest/best-fit open-weight model instead of defaulting everything to one
expensive model.
The load-bearing, non-obvious claim in that sentence is that **task category
should change the model**. For a long stretch it demonstrably did not — the
classifier computed a category, the router paid ~10s for it, and
`proficiency` was empty so it changed nothing. That is fixed, and it is
measurable now. Replaying all 9,001 real decisions with per-category
proficiency against a category-agnostic mean:
```
category-aware == category-agnostic: 5,917/9,001 (65.7%)
category changed the winner: 3,084/9,001 (34.3%)
```
**One decision in three turns on the category.** That is the design's central
bet paying off, and it took populating `proficiency`, fixing four harness bugs,
and rebuilding the cost axis to get there. It should be stated as a result.
Two more things are working and deserve to be said plainly before the
criticism:
- **The session cache removed the latency tax.** It landed 2026-08-30 and is
running at 98%+ hit rate (2026-09-01: 1,402 cached / 28 fresh). The
"~10s of local overhead on every message" problem in CLAUDE.md is solved for
the sessions that matter. Fresh classifications average 2,297ms (max 66,319ms
— one outlier worth a look, but not a design issue).
- **The pure-module discipline is the best thing in the codebase.**
`routing.py`, `scoring.py`, `tiering.py`, `proficiency.py`, `context_prune.py`
are I/O-free, take thresholds as arguments, and are exhaustively tested. That
is why 843 tests run offline in 37 seconds, and it is why every one of the
measurement-driven reversals in CLAUDE.md was cheap to make. Keep it.
---
## 2. The premise quietly changed, and the docs did not follow
Design §7 names the metric: **"$ saved vs. always-frontier baseline"**, with
the warning that "a router that's cheaper but quietly worse isn't a win."
Nobody has checked the other side of that: a router that's *better but quietly
more expensive* isn't obviously a win either, and that is where this landed.
`baseline_report.py` already tells the story, and it appears not to have been
read recently:
```
$ PYTHONPATH=src python -m baseline_report --since 2026-08-24
decisions: 9018
total cost: actual $113.89 cheapest $51.39 best-prof $166.59
dominance (selected == cheapest): 4,972/9,018 (55.1%)
mean proficiency: actual 0.991 cheapest 0.788 best-prof 0.998
```
The router sits at **0.991 of a maximum 0.998** on quality and **2.2x the
cheapest** on cost. It is not a cost/quality tradeoff engine; it is a
quality-maximizer that takes a discount when quality is exactly tied. That may
be what you want — but it is not what §1 says, and it is not what the README
says either ("dispatches to the cheapest/best-fit model").
The comparison that matters more is against the alternative you would actually
have used. Replaying all 9,007 chat decisions against every single-model policy
(same request shapes, same catalog prices, `assumed_cache_rate` 0.917;
`infeasible` = requests exceeding that model's effective window):
| single-model policy | est. cost | vs router | infeasible |
|---|---|---|---|
| `gemma-4-31b` | $18.91 | 0.17x | 1,170 |
| `qwen3.6-35b` | $19.54 | 0.17x | 4,032 |
| **`deepseek-v4-flash`** | **$36.59** | **0.32x** | **0** |
| `kimi-k2.7-code` | $136.41 | 1.20x | 963 |
| `glm-5.2-fast` / `glm-5.3` | $260.20 | 2.28x | 0 |
| `kimi-k3` | $563.97 | 4.95x | 0 |
| *(actual routed)* | *$113.89* | *1.00x* | *0* |
Against always-frontier (`kimi-k3`) the router saves **4.95x** — the design's
own metric, comfortably met. Against "just always use `deepseek-v4-flash`" it
costs **3.1x more**, and deepseek is the only other policy that can serve 100%
of the traffic without a context failure.
Which baseline is honest depends entirely on what you would otherwise have
done. For opencode agent traffic against this catalog, the realistic
counterfactual is not `kimi-k3` — it is deepseek. **The router should report
both baselines, and `baseline_report.py` should grow a single-model column.**
The operational consequence is live right now: **6.51 kWh burned against a
`plan_kwh_per_period` of 6.25.** At roughly $7.67/kWh realized, an
always-deepseek policy would have used about a third of that and stayed inside
the allowance.
---
## 3. Finding #1 — the feedback loop rewards models for not being used
**This is the most important item in the review, and it is a reason not to
take the next obvious step until it is fixed.**
987 client outcomes are sitting in `verifications`; 961 of them unapplied.
`POST /outcome` is correctly identified in CLAUDE.md as "the only ground
truth," so folding them in looks like the highest-value pending action. It is
not, as currently built.
Simulated on a copy of the DB — run `feedback.py` for real, then replay all
9,007 decisions through `select_candidates` + `rank_candidates`:
```
BEFORE est cost $137.76 AFTER est cost $179.28 (1.30x)
kimi-k2.7-code 2,722 glm-5.2-fast 2,243
qwen3.6-35b 2,646 qwen3.6-35b-fast 1,385
qwen3.6-35b-fast 1,236 glm-5.3 1,318
glm-5.3 745 kimi-k2.7-code 1,239
gemma-4-31b 587 qwen3.6-35b 1,150
deepseek-v4-flash 500 kimi-k2.7-code-fast 947
glm-5.2-fast 417 deepseek-v4-flash 596
```
Ingesting ground truth makes routing **30% more expensive** and hands 2,243
decisions to `glm-5.2-fast`, one of the priciest common rows. The mechanism is
visible in the scores. `coding_refactor`, before → after:
| model | before | n | after | n | real outcomes folded in |
|---|---|---|---|---|---|
| `glm-5.2-fast` | 0.994 | 36 | **0.994** | 36 | **zero** |
| `kimi-k2.7-code` | 1.000 | 18 | 0.813 | 80 | 62 |
| `deepseek-v4-flash` | 0.826 | 36 | 0.791 | 135 | 99 |
| `qwen3.6-35b` | 0.900 | 18 | 0.641 | 44 | 26 |
**`glm-5.2-fast` wins the category outright by never having been measured on
real traffic.**
The cause is structural, not a bug: `feedback.py` and `eval_proficiency.py`
both write through `proficiency_store.add_self_eval`, into the *same*
`self_eval_score` running mean, with equal weight per sample. Benchmark tasks
score ~1.00. Real agent turns score ~0.80. So every real sample drags a score
down, and the drag is proportional to how much the model has been used. The
router then routes away from the used model toward the unused one, gathers
outcomes on *that* one, penalizes it, and moves on.
That is a rotation, not a convergence. And because exposure has been
concentrated on the cheap models — they win the cost tiebreak — the rotation is
systematically toward expensive ones.
**Directions (pick one; all are cheap):**
1. **Keep the two sources separate.** Real-traffic outcomes are a different
measurement from benchmark tasks and should not share a mean. Add
`outcome_score` / `outcome_samples` columns and blend them explicitly, the
way `leaderboard` and `self_eval` already are. This also makes the
~0.80-vs-1.00 scale difference visible instead of silently mixed.
2. **Shrink toward the prior by sample size.** A score with n=36 and no real
exposure should not outrank one with n=135 that includes 99 real outcomes.
Today low sample count is an *advantage*; `self_eval_min_samples` gates the
leaderboard blend but nothing penalizes thin evidence in the ranking itself.
3. **Compare like with like.** Score a model against the *per-category mean*
outcome rate rather than absolute, so an 82% pass rate in a category where
everyone scores 80% is neutral, not a 0.18 penalty.
Until one of these lands, **do not run `feedback.py` on the pending backlog.**
The 961 rows are the most valuable data this project has; spending them through
the current path converts them into a more expensive router.
---
## 4. Finding #2 — exclusions are self-sealing; there is no exploration
`deepseek-v4-flash` scored **0.333 on `tool_use_agentic`** across **3**
benchmark tasks. Blended, it now reads 0.5. With `quality_tolerance: 0.1` that
puts it in band 5 while eleven models sit in band 0, so it loses every
`tool_use_agentic` ranking outright.
Consequence, measured: **deepseek won 0 of 2,512 `tool_use_agentic` decisions,
and has 0 client outcomes in that category.** A 3-sample estimate has
permanently removed the cheapest capable model from the largest category of
traffic, and the design contains no path by which that estimate can ever be
revised.
`config/config.yaml` proposes the experiment that would settle it:
> run with it off, let `POST /outcome` report real pass/fail, and compare
> `tool_use_agentic` proficiency for deepseek before and after
**That experiment cannot run.** `min_tool_proficiency` is already `null` — the
filter is off — and it changes nothing, because the *quality band* excludes
deepseek from the category regardless. The config comment says "with it off,
that advantage applies." It does not.
The general shape: this is a contextual bandit running pure exploitation. Once
a model wins a (category, context-size) cell it wins it forever; alternatives
never accumulate the evidence that would overturn the ranking. The consequence
is worse than a missed saving — it means every proficiency number is
conditioned on the routing that produced it, which is exactly the confound that
makes the outcome data below hard to read.
**Direction:** a small explicit exploration budget. An ε of 2-3% of requests
routed to the highest-cost-advantage *excluded* candidate per category would
have produced ~75 deepseek `tool_use_agentic` samples over this window — enough
to confirm or kill the 0.333 outright — at an estimated cost of under $2.
Alternatively a "probation" rule: any candidate excluded solely by a
proficiency score with `self_eval_samples < 10` gets N requests per week
regardless.
### What the ground truth actually says (with the confound stated)
Client outcomes, Wilson 95% CIs, ≥20 samples:
| model | n | pass | |
|---|---|---|---|
| `deepseek-v4-flash` | 390 | **82.8%** | [79%, 86%] |
| `kimi-k2.7-code` | 269 | 80.3% | [75%, 85%] |
| `qwen3.6-35b` | 198 | 85.9% | [80%, 90%] |
| `kimi-k2.7-code-fast` | 52 | 80.8% | [68%, 89%] |
| `glm-5.2-fast` | 44 | 88.6% | [76%, 95%] |
Within `coding_refactor`, where three models have real samples:
| model | n | pass | rel. price |
|---|---|---|---|
| `deepseek-v4-flash` | 95 | **81.1%** [72, 88] | 1x |
| `kimi-k2.7-code` | 62 | 75.8% [64, 85] | ~6.8x |
| `qwen3.6-35b` | 25 | 48.0% [30, 67] | ~2.1x |
The cheapest model is at the top, and `qwen3.6-35b`'s CI does not overlap it.
**Read this carefully, not triumphantly.** The comparison is confounded, and
the confound runs *against* the expensive models: context window determines
eligibility, so `qwen3.6-35b` (94k effective) only ever sees short requests
while `kimi-k2.7-code` (192k) takes the 94-192k band and `glm-5.2-fast` (782k)
takes the largest. Bigger context correlates with longer, harder sessions. So
the expensive models are being handed the harder work, and the table cannot
separate "cheaper model is as good" from "cheaper model got easier requests."
That is the point. **The router has ~1,000 samples of its highest-value signal
and cannot draw a conclusion from them, because it never randomizes.** Adding
exploration is what makes this data interpretable, not just what makes deepseek
eligible.
---
## 5. Finding #3 — cost is 93% prompt tokens, and the heuristics gate the other 7%
Decomposing the estimated cost of all 9,007 real decisions at
`assumed_cache_rate: 0.917`:
```
prompt share of est cost: $105.48 (92.6%)
completion share: $8.40 (7.4%)
real billed tokens: 1,190,607,380 prompt / 5,880,453 completion = 202:1
```
Completion price spans 54x across the catalog ($0.28 → $15.00) and decides
**7.4%** of the bill. Cached-prompt price spans 21x ($0.0144 → $0.30) and
decides **92.6%**.
Two consequences:
- **`tiering.cheap_completion_max: 1.00` gates tier 1 on the wrong axis.** Tier
is the single most consequential filter in the system — it decides the
eligible set before ranking runs — and it is resolved from completion price
plus reasoning mode. This is the same substitution the project has now caught
twice ("list price ranks models backwards"; "cheapness is not a capability
ceiling"), appearing a third time. For this workload the honest cheapness
signal is `cost_per_1m_prompt_cached`.
- **Pinch is the highest-leverage lever in the codebase and the least
instrumented.** It is the only thing that touches the 93%. Post-pinch
`required_context_tokens` (line 2450 records `measured`, computed on
`send_messages`) distributes as:
```
p10 49,817 p25 66,097 p50 88,069 p75 127,517 p90 195,683 p99 279,995
```
against `pinch.budget_tokens: 50000`. **The median request ships at 1.8x the
pinch budget and p90 at 3.9x** — pinch is running and not reaching its own
target, which makes sense given it may only trim tool results outside the
last 4 turns while opencode's ~32k system prompt and tool definitions are
untouchable. Its stats log at `logs.debug` while `logging.level: info`, so
zero pinch lines exist in seven days of journal. Nothing in `/metrics`, the
TUI, or the admin dashboard reports pinch effectiveness either.
A 20% reduction in prompt tokens is worth more than every routing decision in
this window combined, and right now there is no way to tell whether pinch
delivers 2% or 40%. **Promote the pinch line to info, and add
`tokens_saved` / `original_tokens` to `/metrics`.** That is the cheapest
high-value change on this list.
### Related: the tier ladder does not produce a capability gradient
| tier | decisions | avg est cost | outcome pass rate (approx. join, n) |
|---|---|---|---|
| 1 | 896 | $0.00907 | 76.9% (13) |
| 2 | 6,663 | $0.01337 | 83.6% (1,201) |
| 3 | 1,448 | $0.01155 | **70.4%** (226) |
Tier 3 costs *less* per decision than tier 2 and fails *more*. The failure rate
is the good news — it means the classifier's tier call carries real signal about
task difficulty. The cost figure is the problem: "frontier / high-stakes" is
not buying a more capable model, only a reasoning-enabled subset, because tier 3
is resolved from `reasoning_default_enabled` while deepseek (1M window, 1.00 on
all three coding categories) is pinned at tier 2 and excluded from all 1,448 of
them. That is the tier-1 lesson from CLAUDE.md recurring one rung up the ladder.
*(Outcome join is `session_key` + 5s window, so treat as indicative.)*
---
## 6. Smaller structural notes
**Sample-depth trap, recurring.** CLAUDE.md documents this precisely once
already — `docs_writing` read 0.70-1.00 at n=2 with a model at the ceiling, and
six more passes spread it 0.66-0.97 without touching a task. The same shape is
sitting in four more categories right now:
| category | rows | at exactly 1.00 | mean samples |
|---|---|---|---|
| `reasoning_math` | 16 | 13 | 5.4 |
| `translation` | 16 | 11 | 3.2 |
| `general_chat` | 13 | 11 | 3.5 |
| `summarization` | 16 | 3 | 3.4 |
The lesson was learned and written down; it has not yet been applied to the
categories that still show the symptom. "Try samples before hardening" is
already the project's own guidance.
**"Quality is the objective" is operationally inverted.** 72 of 141 proficiency
rows sit at exactly 1.00, and 98 of 141 (70%) fall inside the top
`quality_tolerance` band. For most requests quality is a tie and **cost is the
sole decider**, with quality acting as a coarse veto on the worst ~30%. That is
a defensible design — arguably the right one — but CLAUDE.md and
`docs/routing.md` read the other way ("there is no weight to tune — quality is
the objective and cost is the tiebreak"), which will mislead anyone tuning it.
**Structural verification is inert for this traffic.** 12,047 of 13,501
verification rows (89.2%) are `unverifiable` — correctly, since agent turns end
in tool calls. It costs nothing (pure Python, post-response), so there is no
harm, but it occupies a verdict-mix panel on the dashboard and a section in the
docs that together imply coverage that does not exist. `feedback.py`'s own
`coverage()` already prints the warning; the dashboard should too.
**`dispatcher.py` is where the discipline stops.** 3,104 lines, with
`chat_completions` at **662 lines** in a single function. Every other module in
this project is small, pure, and injectable; this one holds routing, pinch,
classification, capability checks, local vision, streaming proxy, telemetry
sniffing, and verification scheduling in one call frame. It is the one place
where the next change is expensive, and the one place a reader cannot hold the
whole path in their head. Extracting the pre-dispatch pipeline (prune →
measure → classify → gate → rank) into a pure function taking messages + config
and returning a decision would bring it in line with the rest and would be
directly testable.
**Minor doc drift.** `design/local-llm-model-router.md` §4 still presents the
weighted composite as the scoring model and §9 item 6 still calls proficiency
"the last inert axis" — both superseded. `scoring.composite_score` and
`eco_score` are dead but tested. `README.md:8` and `docs/routing.md`'s
"## Weighted Scoring" heading carry the same stale framing. CLAUDE.md already
declares itself the source of truth, so this is low-stakes, but the design doc
is the file a new reader opens first.
---
## 7. Recommended order
1. **Do not run `feedback.py` on the 961-row backlog** until the exposure bias
in §3 is fixed. Separate outcome samples from benchmark samples, or weight
by sample size. This is the one item that will actively degrade routing if
the obvious next step is taken.
2. **Promote the pinch log to `info` and surface `tokens_saved` in
`/metrics`.** Cheapest change here, on the axis that carries 93% of the bill.
3. **Add an exploration budget** (ε ≈ 2-3%, or probation for candidates
excluded on `self_eval_samples < 10`). Without it, no proficiency number in
the table is interpretable and the experiment `config.yaml` already proposes
cannot run.
4. **Add a single-model column to `baseline_report.py`.** "vs always-cheapest"
and "vs always-frontier" bracket the answer; "vs always-deepseek" is the
decision the operator actually faces.
5. **Re-tier on `cost_per_1m_prompt_cached`**, or drop the price term from
tiering entirely and keep the context/reasoning signals.
6. Run the eval harness for `reasoning_math`, `translation`, `general_chat`,
`summarization` — the same six-pass treatment that spread `docs_writing`.
7. Extract the pre-dispatch pipeline out of `chat_completions`.
---
## 8. Bottom line
The premise holds and the machine works. The measurement culture is the real
asset here — this project has overturned its own conclusions on evidence
repeatedly (list price ranks backwards, tier-from-price, benchmark-vs-real cost,
four harness bugs) and each reversal is written down with the reasoning intact.
That is rarer than the router.
The gap is that the loop is not yet closed. Every measurement so far has been
*taken*, then reasoned about, then encoded by hand. The two mechanisms that
were supposed to close the loop automatically — folding client outcomes into
proficiency, and letting the tool-proficiency experiment settle itself — are
both currently pointed slightly the wrong way: one rewards models for being
unmeasured, the other cannot gather the measurement it needs. Neither is a
large fix. Both are the difference between a router that learns from your
traffic and one that is very carefully hand-tuned to it.
---
## Appendix — reproduction scripts
Read-only. Each was run against `router.db` on 2026-09-01 to produce the
numbers above. Save to a scratch dir and run with `PYTHONPATH=src` from the
repo root.
### A. Single-model cost baselines (§2)
```python
"""What would a single-model policy have cost, on the same 9k real requests?"""
import sqlite3, sys
sys.path.insert(0, "src")
from routing import estimated_cost
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = {r["model_id"]: dict(r) for r in conn.execute("select * from models")}
rows = conn.execute("""
select required_context_tokens rc, est_cost_usd, selected_model, task_tier
from route_decisions where kind='chat' and required_context_tokens is not null
""").fetchall()
CACHE, COMP = 0.917, 500
actual = sum(r["est_cost_usd"] or 0 for r in rows)
print(f"decisions: {len(rows)} actual router est: ${actual:,.2f}\n")
print(f"{'single-model policy':<26} {'cost':>10} {'vs router':>10} {'infeasible':>11}")
print("-"*62)
res=[]
for mid, m in models.items():
if m["availability"] != "active" or m["access_level"] != "public" or m["latency_class"]=="flex":
continue
eff = m["effective_context_window"] or 0
tot = 0.0; bad = 0
for r in rows:
if r["rc"] > eff:
bad += 1; continue
tot += estimated_cost(m, r["rc"], COMP, CACHE) or 0
res.append((tot, mid, bad))
for tot, mid, bad in sorted(res):
print(f"{mid:<26} ${tot:>9,.2f} {tot/actual:>9.2f}x {bad:>10,}")
```
### B. Prompt-vs-completion cost anatomy (§5)
```python
import sqlite3, sys, statistics
sys.path.insert(0,"src")
conn=sqlite3.connect("router.db"); conn.row_factory=sqlite3.Row
m={r["model_id"]:dict(r) for r in conn.execute("select * from models")}
rows=conn.execute("""select selected_model sm, required_context_tokens rc, est_cost_usd c
from route_decisions where kind='chat' and selected_model is not null
and required_context_tokens is not null""").fetchall()
CACHE,COMP=0.917,500
pt=ct=0.0
for r in rows:
mm=m.get(r["sm"]);
if not mm: continue
p=r["rc"]*((1-CACHE)*mm["cost_per_1m_prompt"]+CACHE*(mm["cost_per_1m_prompt_cached"] or mm["cost_per_1m_prompt"]))/1e6
c=COMP*mm["cost_per_1m_completion"]/1e6
pt+=p; ct+=c
print(f"prompt share of est cost: ${pt:.2f} ({pt/(pt+ct)*100:.1f}%)")
print(f"completion share: ${ct:.2f} ({ct/(pt+ct)*100:.1f}%)")
print()
rc=[r["rc"] for r in rows]
rc.sort()
print("required_context_tokens percentiles:")
for q in (0.1,0.25,0.5,0.75,0.9,0.99):
print(f" p{int(q*100):>2}: {rc[int(q*len(rc))]:>9,}")
print(f" max: {rc[-1]:,} mean: {statistics.mean(rc):,.0f}")
print()
# actual billed vs prompt size
print("actual billed energy_observations, prompt vs completion tokens:")
r=conn.execute("select sum(prompt_tokens) p, sum(completion_tokens) c, count(*) n from energy_observations where task_category!='seed_reference'").fetchone()
print(f" prompt {r['p']:,} completion {r['c']:,} ratio {r['p']/r['c']:.0f}:1 n={r['n']:,}")
```
### C. Does category change the winner? (§1)
```python
"""Does the classifier's CATEGORY output change the routing decision?
Replays every real chat decision twice: once with per-category proficiency
(what ships), once with a category-agnostic mean proficiency per model.
"""
import sqlite3, sys, statistics
sys.path.insert(0,"src")
from config import load_config
from routing import select_candidates, rank_candidates
cfg = load_config("config/config.yaml")
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = [dict(r) for r in conn.execute("select * from models")]
prof = {(r["model_id"], r["category"]): r["blended_score"]
for r in conn.execute("select model_id, category, blended_score from proficiency")}
# category-agnostic: mean of a model's blended scores across all categories
avg = {}
for (m,c),v in prof.items():
if v is not None: avg.setdefault(m,[]).append(v)
avg = {m: statistics.fmean(v) for m,v in avg.items()}
decisions = conn.execute("""select task_category c, task_tier t, required_context_tokens rc,
latency_tolerance lt, selected_model sm from route_decisions
where kind='chat' and selected_model is not null and task_category is not null
and required_context_tokens is not null and task_tier is not null""").fetchall()
def pick(d, use_category):
rows=[]
for m in models:
r=dict(m)
r["proficiency"] = prof.get((m["model_id"], d["c"])) if use_category else avg.get(m["model_id"])
rows.append(r)
cand = select_candidates(rows,
required_context_tokens=d["rc"], required_tier=d["t"],
latency_tolerance=d["lt"] or "interactive",
allowed_access_levels=cfg.routing.allowed_access_levels,
exclude_stale=cfg.freshness.exclude_stale,
exclude_deprecated=cfg.freshness.exclude_deprecated)
rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
cache_rate=cfg.objective.assumed_cache_rate)
return (rk[0]["model_id"], rk[0]["cost"]) if rk else (None,0)
same=diff=0; cost_cat=cost_flat=0.0
for d in decisions:
a,ca = pick(d, True); b,cb = pick(d, False)
cost_cat+=ca; cost_flat+=cb
if a==b: same+=1
else: diff+=1
n=len(decisions)
print(f"replayed {n} real decisions")
print(f" category-aware == category-agnostic: {same}/{n} ({same/n*100:.1f}%)")
print(f" category changed the winner: {diff}/{n} ({diff/n*100:.1f}%)")
print(f" cost with category: ${cost_cat:,.2f}")
print(f" cost without category: ${cost_flat:,.2f} ({cost_flat/cost_cat:.2f}x)")
```
### D. Simulated feedback fold-in (§3)
Copies the DB first; never mutates `router.db`.
```python
"""Simulate: apply the 1,039 pending client outcomes, then replay routing."""
import sqlite3, sys, os, shutil
sys.path.insert(0,"src")
from config import load_config
import feedback
from routing import select_candidates, rank_candidates
cfg = load_config("config/config.yaml")
db = os.environ["CLAUDE_JOB_DIR"] + "/tmp/sim.db"
def replay(dbpath, label):
conn = sqlite3.connect(dbpath); conn.row_factory = sqlite3.Row
models=[dict(r) for r in conn.execute("select * from models")]
prof={(r["model_id"],r["category"]):r["blended_score"]
for r in conn.execute("select model_id,category,blended_score from proficiency")}
ds=conn.execute("""select task_category c,task_tier t,required_context_tokens rc,
latency_tolerance lt from route_decisions where kind='chat'
and task_category is not null and required_context_tokens is not null
and task_tier is not null""").fetchall()
tot=0.0; mix={}
for d in ds:
rows=[{**m,"proficiency":prof.get((m["model_id"],d["c"]))} for m in models]
cand=select_candidates(rows,required_context_tokens=d["rc"],required_tier=d["t"],
latency_tolerance=d["lt"] or "interactive",
allowed_access_levels=cfg.routing.allowed_access_levels,
exclude_stale=cfg.freshness.exclude_stale,
exclude_deprecated=cfg.freshness.exclude_deprecated)
rk=rank_candidates(cand,quality_tolerance=cfg.objective.quality_tolerance,
prompt_tokens=d["rc"],completion_tokens=cfg.objective.assumed_completion_tokens,
cache_rate=cfg.objective.assumed_cache_rate)
if rk:
tot+=rk[0]["cost"] or 0; mix[rk[0]["model_id"]]=mix.get(rk[0]["model_id"],0)+1
conn.close()
print(f"\n=== {label} === replayed {len(ds)} decisions est cost ${tot:,.2f}")
for m,n in sorted(mix.items(),key=lambda kv:-kv[1]):
print(f" {m:<24}{n:>6}")
return tot
before = replay(db, "BEFORE (current proficiency)")
conn=sqlite3.connect(db)
rows=feedback.unapplied_failures(conn)
feedback.apply_failures(conn, cfg, feedback.summarize(rows), dry_run=False)
conn.close()
after = replay(db, "AFTER folding in 1,039 client outcomes")
print(f"\ncost change: ${before:,.2f} -> ${after:,.2f} ({after/before:.2f}x)")
```
### E. Client-outcome pass rates with Wilson CIs (§4)
```python
import sqlite3, math
conn=sqlite3.connect("router.db"); conn.row_factory=sqlite3.Row
q="""select model_id m, task_category c,
sum(verdict='succeeded') s, sum(verdict='failed') f
from verifications where kind='client_outcome' group by 1,2"""
rows=[dict(r) for r in conn.execute(q)]
def wilson(s,n):
if n==0: return (0,0)
z=1.96; p=s/n; d=1+z*z/n
c=(p+z*z/(2*n))/d; h=z*math.sqrt(p*(1-p)/n+z*z/(4*n*n))/d
return (c-h,c+h)
print(f"{'model':<22}{'category':<18}{'n':>5}{'pass':>7} 95% CI")
print("-"*66)
for r in sorted(rows,key=lambda r:(-(r['s']+r['f']))):
n=r['s']+r['f']
if n<15: continue
lo,hi=wilson(r['s'],n)
print(f"{r['m']:<22}{r['c']:<18}{n:>5}{r['s']/n*100:>6.1f}% [{lo*100:.0f}%, {hi*100:.0f}%]")
print()
print("--- pooled per model (all categories) ---")
agg={}
for r in rows:
a=agg.setdefault(r['m'],[0,0]); a[0]+=r['s']; a[1]+=r['f']
for m,(s,f) in sorted(agg.items(),key=lambda kv:-(kv[1][0]+kv[1][1])):
n=s+f
if n<20: continue
lo,hi=wilson(s,n)
print(f"{m:<22}{n:>5}{s/n*100:>6.1f}% [{lo*100:.0f}%, {hi*100:.0f}%]")
```