# Fixing exposure bias (#1) and self-sealing exclusions (#2) Status: done -- code shipped; only the operator-gated backlog fold remains (see below) > **Correction, 2026-09-08.** This doc was marked `planned` in the status > backfill on the strength of its own "ready to implement" framing and a > memory note warning not to fold the backlog yet. Both were read as "the > code is not written". It is. Verified against the live tree: > > ``` > exploration: enabled=True epsilon=0.03 max_cost_ratio=4.0 max_tier=2 > proficiency.outcome_prior_strength: 20 > feedback.FAILURE_VERDICTS = ("failed",) # structural verdicts excluded > proficiency.expected_success_rate() # present > ``` > > And the three-branch taxonomy is live in the table, not just in code -- > `outcome_prior` 264, `self_eval_thin` 147, `outcome_blended` 38, > `self_eval` 2. > > What actually remains is not code. **1,136 of 2,221 client outcomes are > unfolded** (all attributable). Spending them mutates `router.db` > irreversibly, which is why the plan sequences it last and why it needs an > operator decision rather than an implementer. That is the whole of the > outstanding work. > > The lesson is the one this repo already has a memory for: check the code > path before calling something a gap. A plan that says "FINAL -- ready to > implement" describes the state when it was written, not now. **Date:** 2026-09-01 **Follows:** `plans/conceptual-review-premise-and-execution.md` findings #1 and #2 **Status:** **FINAL — ready to implement.** Validated by replay against the 9,007 real chat decisions in `router.db`. Nothing here has been implemented; `router.db` was not mutated. Both open decisions were resolved 2026-09-01 and are now stated as instructions, not options: 1. `route_decisions.request_id` ships as part of this work — it is a prerequisite for the validation in §5, not a follow-up. 2. Structural and `local_llm` verdicts **stop feeding proficiency** and remain diagnostics only. See §3. Section 6 is the file-by-file task list. --- ## 0. They are one problem seen from two sides Both findings reduce to the same sentence: **absence of evidence is currently treated as evidence of quality.** - On the **ranking** side (#1), a model with no traffic evidence keeps its saturated benchmark score and wins the band — `glm-5.2-fast` takes 2,243 decisions after the naive fold-in *because* it has never been measured on real work. - On the **selection** side (#2), a model excluded by a 3-sample benchmark score never accumulates the evidence that would overturn it — `deepseek-v4-flash` won 0 of 2,512 `tool_use_agentic` decisions and therefore has 0 outcome samples there, forever. So the fixes are complementary, not alternatives. Fixing the score without adding exploration leaves cells that can never be filled. Adding exploration without fixing the score feeds good data into a path that misreads it. Ordering matters: **fix the ingestion path first**, then turn on exploration, then spend the 961-row backlog. --- ## 1. What is actually wrong with the score `feedback.py` and `eval_proficiency.py` both write through `proficiency_store.add_self_eval` into a single `self_eval_score` running mean, equal weight per sample. But they are measuring different quantities: | | benchmark (`evals/tasks.yaml`) | client outcomes (`POST /outcome`) | |---|---|---| | scale | ~1.00 (saturated; 51% of rows sit at exactly 1.00) | ~0.80 | | comparability | controlled — every model gets the same task | confounded — models get the requests routing sent them | | discrimination | weak | strong | | cost to add | a benchmark run | free | Averaging a saturated absolute score with an uncontrolled success rate produces a number whose value depends on the **mixing ratio**, and the mixing ratio is set by how much the model has been used. That is the bug, stated exactly. The fix is not to weight the two sources — it is to **stop treating the benchmark as a level and start treating it as a prior**, with everything expressed on the scale that matters: expected probability that a request succeeds. --- ## 2. The scoring rule Per (model, category), with `k` a prior strength in samples (`k = 20` below): ``` peer_rate = Σ successes / Σ samples # over models with traffic in this category peer_bench = mean benchmark score # over that same set, so the ratio is calibrated prior_m = min(1.0, peer_rate × bench_m / peer_bench) score_m = (n_m × rate_m + k × prior_m) / (n_m + k) ``` Read it as: *the benchmark says where this model sits relative to its peers; the peer traffic rate says what that position is worth in practice; a model's own traffic pulls the estimate toward its observed rate in proportion to how much of it there is.* Properties that matter: - **An unproven model lands at the peer average, not at the ceiling.** That is the direct fix for #1. It is still admitted (consistent with the project's "absent evidence does not disqualify" rule) — it just does not get to sit above every measured model for free. - **Scores become interpretable.** `blended_score` now means "expected pass rate on your traffic." `quality_tolerance: 0.1` becomes "10 percentage points of real success rate," which is a knob you can reason about. Today it bands an abstract 0-1 quality index whose units are the benchmark's. - **Cold start is unchanged.** A category with no traffic falls back to the benchmark exactly as today; a model with neither returns `None` → neutral 0.5 downstream. - **Shrinkage replaces the sample-count threshold.** `self_eval_min_samples` gates a step change; `k` is continuous, which is the behaviour that section was reaching for. ### Measured effect Replaying all 9,007 real chat decisions through `select_candidates` + `rank_candidates`, changing only the proficiency values: | policy | est. cost | vs today | |---|---|---| | **A. today** (benchmark only) | $137.76 | 1.00x | | **B. naive fold-in** (`feedback.py` as written) | $179.28 | **1.30x** | | **D. proposed** (empirical Bayes, above) | **$109.35** | **0.79x** | The proposed rule is **39% cheaper than the naive fold-in** and 21% cheaper than today, and the traffic moves toward the models with the best *measured real-world* pass rates rather than away from them: ``` today proposed kimi-k2.7-code 2,722 deepseek-v4-flash 3,260 qwen3.6-35b 2,646 qwen3.6-35b 2,646 qwen3.6-35b-fast 1,236 kimi-k2.7-code 2,210 glm-5.3 745 glm-5.2-fast 558 deepseek-v4-flash 500 qwen3.6-35b-fast 133 ``` `coding_refactor` under the rule — note the unproven rows now sit *below* or level with the proven ones instead of above them: | model | bench | n | rate | score | |---|---|---|---|---| | `glm-5.3` | 1.000 | 1 | 100% | 0.795 | | `glm-5.2-flex` | 1.000 | 0 | — | 0.785 | | `kimi-k3-fast` | 1.000 | 0 | — | 0.785 | | `gemma-4-31b` | 0.997 | 0 | — | 0.783 | | `deepseek-v4-flash` | 0.826 | 95 | 81% | 0.782 | All inside one `quality_tolerance` band, so cost decides — and deepseek, the one with 95 real samples at 81%, is by far the cheapest. That is the outcome the review said was missing. ### One honest tradeoff With this little traffic evidence, scores compress (0.78-0.93 in most categories), so a 0.1 band covers much of the range and **cost decides more often than it does today.** That is correct given the evidence — nothing is yet *proven* better — but it is a real behavioural change, not a free win. Two levers: raise `k` to lean harder on the benchmark while traffic is thin, or narrow `quality_tolerance` now that its units are meaningful. Revisit both once exploration has been running a few weeks. ### Rejected: a "proven ceiling" cap The first formulation tried was: cap any model with `< 30` outcome samples at the best *traffic-proven* score in its category. It was simulated and rejected — it only bites when the best proven model happens to score low, so `coding_general` collapsed to a single flat value while `coding_refactor` was untouched. Inconsistent, and it needed a second threshold. The empirical-Bayes form gets the same effect from one formula with no special case. --- ## 3. Implementation — scoring ### Schema (both via the existing `ensure_columns` migration pattern) ```sql ALTER TABLE proficiency ADD COLUMN outcome_score REAL; ALTER TABLE proficiency ADD COLUMN outcome_samples INTEGER DEFAULT 0; ALTER TABLE route_decisions ADD COLUMN request_id TEXT; -- see note below ALTER TABLE route_decisions ADD COLUMN exploration INTEGER DEFAULT 0; ``` `config/schema.sql` gains the same columns for fresh installs. Mirror `proficiency_store.ensure_columns` / `_ensure_route_decisions_table` so a live `router.db` migrates on load and on write, idempotently. **`route_decisions.request_id` is a prerequisite, not a nice-to-have.** There is currently no exact join from a client outcome back to the decision that produced it — the review had to approximate with `session_key` + a 5-second window, which is why the tier/outcome table in §5 there is marked indicative. `report_outcome` already resolves `request_id` against `energy_observations`; recording it on the decision closes the loop and is what makes the validation in §5 below possible. ### `proficiency.py` — one new pure function Stays I/O-free; the per-category aggregates are passed in by the caller. ```python def expected_success_rate( benchmark_score: float | None, outcome_score: float | None, outcome_samples: int, *, peer_rate: float | None, # None when the category has no traffic yet peer_benchmark: float | None, prior_strength: int, ) -> tuple[float | None, Source | None]: ... ``` Returns `(None, None)` when there is neither benchmark nor outcome data, so the neutral-0.5 path downstream is unchanged. Add `"outcome_blended"` to `Source` so provenance stays inspectable the way `self_eval_thin` already is. ### `proficiency_store.py` — a second writer, and a category recompute - New `add_outcome(conn, cfg, model_id, provider, category, scores)` writing `outcome_score` / `outcome_samples` through the existing `accumulate`. Keep the 0-1 clamp — the comment about a harness bug pushing a score to 1.50 and raising `best` for every candidate applies verbatim here. - **`blended_score` becomes a category-level computation**, because `peer_rate` and `peer_benchmark` are aggregates over the category. `_write` cannot produce the final value from one row any more. Split it: - `_write` keeps writing `leaderboard_score` / `self_eval_score` / `self_eval_samples` (and now `outcome_*`), and leaves `blended_score` at the benchmark value from `blend()` so a single write is never internally inconsistent. - New `recompute_category(conn, cfg, category)` reads every row in the category, re-derives each row's benchmark score by calling the **existing** `blend()` on its stored components, computes `peer_rate` / `peer_benchmark` from `outcome_score` / `outcome_samples`, then writes the final `blended_score` + `source` per row. Deriving the benchmark half from the stored components rather than caching it means **no fourth score column and no drift** — the invariant this module exists to hold (it is the only writer, so `blended_score` and `source` can never disagree with their inputs) survives intact, and `dispatcher.load_candidates` stays untouched on the hot path. - Every writer calls `recompute_category` at the end of its transaction: `add_self_eval`, `add_outcome`, `propagate_to_variants`, and the `leaderboard.py` importer. - `propagate_to_variants` must copy `outcome_score` / `outcome_samples` too, and the `inherited_from` guard applies unchanged. - `ensure_columns` gains the two new columns. No backfill is needed or wanted: `outcome_samples` defaults to 0, which is exactly true of every existing row. (Contrast `inherited_from`, where a NULL default was actively wrong and needed `_backfill_inherited`.) ### `config.py` / `config/config.yaml` — the prior strength is a knob `ProficiencyConfig` gains `outcome_prior_strength: int` (the `k` in §2), and it **must be stated in `config/config.yaml`**, not left to a Pydantic default. This project has already been bitten once by a setting that existed only as a default and was therefore invisible to anyone tuning it (`classifier.max_input_chars`, shipped into the wrong section, accepted and discarded because it happened to match the code default). Every knob belongs in the file. ```yaml proficiency: # Prior strength, in samples, for folding real client outcomes into a score. # A model's own traffic outweighs the benchmark-derived prior once it has # more than this many outcome samples. 20 was chosen because the categories # with real evidence carry 25-163 samples, so it lets a well-measured model # move while still holding a 5-sample cell near its peers. outcome_prior_strength: 20 ``` `self_eval_min_samples`, `leaderboard_weight` and `self_eval_weight` are unchanged — they still govern `blend()`, which now produces the *benchmark half* that feeds the prior rather than the final score. ### `feedback.py` — client outcomes only `SUCCESS_VERDICTS` stays. The changes: - Client outcomes route to `add_outcome` instead of `add_self_eval`. `applied_at` idempotency is unchanged. - **`FAILURE_VERDICTS` drops `truncated` and `malformed`** — structural and `local_llm` verdicts stop feeding proficiency entirely and remain diagnostics. They are *failure-only* contributors (a passing structural check is deliberately not recorded), and you cannot form a rate from failures alone: folding them in would bias `outcome_score` downward by exactly however often the checker happened to fire, which is a property of the checker, not the model. They are 134 rows against 987 client outcomes, and the structural checker's one documented encounter with real agent traffic produced ~29 false `malformed`s before the `has_tool_calls` fix. `failed` (the client-outcome failure verdict) stays. - `coverage()` keeps reporting all verdicts, so the diagnostics stay visible where they belong. Its existing 80%-unverifiable warning is unaffected. --- ## 4. Implementation — exploration ### The rule On an ε-share of requests, instead of the rank winner, dispatch to the **hard-filter-eligible candidate with the fewest outcome samples in this category**, tie-broken by lowest cost, and skipped entirely if its estimated cost exceeds `max_cost_ratio ×` the winner's. Every hard filter still applies — context window, tier floor, access level, latency class, vision, JSON mode. Exploration only ever reorders *within the eligible set*, so it can never produce a request the model cannot serve. ### Cost, measured Replayed at ε = 0.03, `max_cost_ratio` 4.0, over the same 9,007 decisions: ``` explored 263 of 9,007 (2.9%) exploit-only $137.76 with explore $139.12 (+$1.36, +1.0%) ``` **$1.36 over nine days.** What it buys: | cell | new samples | had | |---|---|---| | `deepseek-v4-flash` / `tool_use_agentic` | **+66** | **0** | | `kimi-k2.7-code-fast` / `coding_refactor` | +29 | 6 | | `kimi-k3` / `coding_general` | +23 | 0 | | `glm-5.2-fast` / `coding_refactor` | +22 | 0 | | `kimi-k3` / `coding_refactor` | +19 | 0 | | `qwen3.6-35b-fast` / `coding_general` | +18 | 0 | | `qwen3.6-35b` / `docs_writing` | +16 | 0 | | …plus 5 more cells currently at zero | | | The "fewest samples first" rule targets empty cells without being told to, and 66 samples is enough to confirm or kill the 3-task 0.333 that currently bans the cheapest capable model from the largest category of traffic. ### Module shape New `src/exploration.py`, following `circuit_breaker.py` / `session_cache.py`: pure, module-level state only if needed, **injected RNG** the way `circuit_breaker` injects time, and it never imports `dispatcher` or `config`. ```python def choose( ranked: Sequence[dict], sample_counts: dict[str, int], *, epsilon: float, max_cost_ratio: float, rng: random.Random, ) -> tuple[dict, bool]: # (row, was_exploration) ``` `dispatcher.load_candidates` already LEFT JOINs `proficiency`; add `p.outcome_samples` to the SELECT and pass the counts in. The call site is immediately after `rank_candidates`, before `apply_flex_preference` — a flex swap is a serving-class decision and should apply to whatever was chosen. ### Config ```yaml exploration: # Routes a small share of requests to the least-evidenced eligible candidate # so proficiency scores can be corrected by evidence rather than frozen by # the first 3-sample benchmark that touched them. Measured on 9,007 real # decisions: 2.9% of traffic, +$1.36 (+1.0%), and it fills 11 (model, # category) cells that currently hold ZERO outcome samples -- including # deepseek-v4-flash / tool_use_agentic, which the router cannot otherwise # ever measure because its own ranking excludes it. enabled: true epsilon: 0.03 # Never explore into something more than this multiple of the winner's cost. max_cost_ratio: 4.0 # Exploration is for gathering evidence, not for gambling on high-stakes # work. Tier 3 is excluded; 2,034 of 2,512 tool_use_agentic decisions are # tier 2, so this costs almost no coverage. max_tier: 2 ``` `ExplorationConfig(StrictModel)` in `config.py`, registered on `RouterConfig`, mirroring `CircuitBreakerConfig`'s shape (defaults on the model *and* stated in the file). Validate `0.0 <= epsilon <= 1.0` and `max_cost_ratio >= 1.0` — an epsilon above 1 or a ratio below 1 are both silently self-defeating rather than loud. There is deliberately no `min_samples_target` knob. "Fewest outcome samples first" needs no threshold: it targets empty cells on its own, and once a cell fills, the next-emptiest becomes the target automatically. **Shipping this `enabled: true` breaks the project's usual "new knob ships off" convention, and does so deliberately.** Off, it changes nothing and `config.yaml`'s own stated experiment ("run with it off, let `POST /outcome` report real pass/fail, and compare `tool_use_agentic` proficiency for deepseek before and after") stays unrunnable — which is precisely the condition the review flagged. The convention exists to stop unproven knobs changing behaviour silently; this one has a measured cost (+1.0%), a measured benefit (11 empty cells), and a hard cost cap (`max_cost_ratio`). It is the exception that earns itself. --- ## 5. Sequencing Four commits, in this order. Steps 1-2 are inert until step 3 runs, so the service can be restarted between any of them. 1. **Schema + `request_id`.** Migrations only; no behaviour change. Verify a live `router.db` migrates and existing rows are intact. 2. **Scoring path.** `expected_success_rate`, `add_outcome`, `recompute_category`, `feedback.py` retargeted. Still inert — no outcome rows have been applied yet, so every `outcome_samples` is 0 and every score reproduces today's value. That is the acceptance test for this step: **after step 2 and before step 3, replaying the decision stream must still produce $137.76.** 3. **Spend the backlog.** `python -m feedback --dry-run`, then `python -m feedback`, on the 961 unapplied rows through the new path. Expected: the §2 table — traffic shifts toward `deepseek-v4-flash` and total estimated cost falls to ~$109. 4. **Exploration on.** Then re-run `baseline_report.py` weekly. ### Validation, at ~2 weeks The check that matters: `deepseek-v4-flash` / `tool_use_agentic` should hold ≳60 real outcome samples. Compare its measured rate against the 0.333 benchmark score that currently bans it from 2,512 decisions. Either the benchmark was right and the exclusion is now *earned*, or it was a 3-sample artifact costing roughly 3x on the largest category of traffic. Both answers are worth $1.36. Then revisit `quality_tolerance` and `outcome_prior_strength`, with scores that finally mean something and enough evidence to set them from. ### The report this unlocks Once `route_decisions.exploration` and `request_id` exist, the confound named in review §4 becomes addressable: pass rates computed **on explored requests only** are unconfounded by the routing policy, because assignment was random within the eligible set. That is the first genuinely causal comparison this project will have been able to make, and it is the thing that settles whether the expensive models are worth their price. --- ## 6. File-by-file task list **Commit 1 — schema** - `config/schema.sql` — add `outcome_score REAL`, `outcome_samples INTEGER DEFAULT 0` to `proficiency`; add `request_id TEXT`, `exploration INTEGER DEFAULT 0` to the `route_decisions` DDL. - `src/proficiency_store.py::ensure_columns` — the two `proficiency` columns. No backfill (0 is correct for every existing row). - `src/dispatcher.py::ensure_route_decisions` + `_ensure_route_decisions_table` — the two `route_decisions` columns, same idempotent pattern. - `src/dispatcher.py::persist_route_decision` — write `request_id` on every decision, and `exploration` (0 for now). `request_id` is already in scope on the chat path; on `/route` (which spends nothing upstream) it stays NULL. **Commit 2 — scoring** - `src/proficiency.py` — add `expected_success_rate(...)` per §2 and `"outcome_blended"` to the `Source` literal. Keep it pure: peer aggregates are arguments. - `src/proficiency_store.py` — split `_write` (benchmark half only, plus the new columns); add `add_outcome`; add `recompute_category`; call it from `add_self_eval`, `add_outcome`, `propagate_to_variants`; extend `propagate_to_variants` to copy `outcome_score` / `outcome_samples`. - `src/leaderboard.py` — importer calls `recompute_category` after its writes. - `src/config.py` — `ProficiencyConfig.outcome_prior_strength: int`. - `config/config.yaml` — the `outcome_prior_strength` block from §3. - `src/feedback.py` — client outcomes → `add_outcome`; drop `truncated` and `malformed` from `FAILURE_VERDICTS`; docstring updated to say why. **Commit 3 — exploration** - `src/exploration.py` — new, pure, injected RNG, no `config`/`dispatcher` import. - `src/config.py` — `ExplorationConfig(StrictModel)` + field on `RouterConfig`. - `config/config.yaml` — the `exploration:` block from §4. - `src/dispatcher.py` — add `p.outcome_samples` to `load_candidates`'s SELECT; call `exploration.choose` after `rank_candidates` and before `apply_flex_preference`; gate on `cfg.exploration.enabled` and `task_tier <= max_tier`; set `exploration=1` on the decision row and add `explore=` to the existing `logs.info` routing line. **Tests** (`tests/test_proficiency.py`, `tests/test_feedback.py`, new `tests/test_exploration.py`) — pin the properties, not the arithmetic: - an unproven model does not outrank a proven one in the same category; - a category with no outcome data reproduces today's benchmark score exactly (this is the step-2 acceptance test in miniature); - a model with neither source still yields `None` → neutral 0.5 downstream; - `recompute_category` is idempotent and leaves untouched categories alone; - structural/`local_llm` verdicts no longer move any score; - exploration never returns a row that fails a hard filter, never exceeds `max_cost_ratio`, and returns the winner unchanged when `epsilon` is 0; - with a seeded RNG, the explore share lands within tolerance of `epsilon`. **Docs** — `docs/data-model.md` (four new columns), `docs/routing.md` (scoring section: the score is now an expected pass rate, and what `quality_tolerance` means in those units), `docs/evaluation.md` (benchmark is now a prior, not the score). `CLAUDE.md`'s "Proficiency: category now changes routing" and "The only ground truth" sections both need the new story once step 3 has run and the numbers are real. **Check while you are in here** — the admin portal's "apply feedback" operational trigger runs `feedback.py`; confirm it still works after the retarget, and that the models page shows the new `source` value rather than blanking on an unrecognized string. --- ## 7. Summary | | fix | cost | effect | |---|---|---|---| | **#1** | benchmark becomes a prior; outcomes are a separate, shrunk, peer-relative signal on the traffic scale | one formula + 2 columns | $179.28 → **$109.35** on replay; unproven models stop winning | | **#2** | ε=3% exploration to the least-evidenced eligible candidate | **+$1.36 / 9 days (+1.0%)** | 11 empty cells filled, incl. +66 `deepseek`/`tool_use_agentic`; makes the table correctable | Neither is a large change. Together they convert the router from something carefully hand-tuned against your traffic into something that learns from it — which is what the design said it was for. --- ## Appendix — reproduction Scripts used to produce every number above are read-only and replay against `router.db` without mutating it. `fix_sim2.py` (empirical Bayes) and `explore_sim.py` (ε-greedy pricing) are the two that matter; both follow the same shape as the appendix scripts in `plans/conceptual-review-premise-and-execution.md` — load `models` + `proficiency` + `verifications`, rebuild each decision's eligible set with `routing.select_candidates`, rank with `routing.rank_candidates`, and sum `estimated_cost`. Constants used: `k = 20`, `ε = 0.03`, `max_cost_ratio = 4.0`, RNG seed 7, and the shipped `quality_tolerance = 0.1` / `assumed_cache_rate = 0.917` / `assumed_completion_tokens = 500`. ### Script D — empirical-Bayes replay (§2) ```python """Empirical-Bayes version: shrink toward a benchmark-informed peer prior, all expressed on the TRAFFIC scale (expected pass rate). prior_m = peer_rate x (bench_m / peer_bench) # benchmark sets relative position score_m = (n_m * rate_m + k * prior_m) / (n_m + k) An unproven model lands at the category's average traffic performance, adjusted by where the benchmark puts it -- not at the benchmark ceiling. """ import sqlite3, sys sys.path.insert(0, "src") from config import load_config from routing import select_candidates, rank_candidates cfg = load_config("config/config.yaml"); K = 20 conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row models = [dict(r) for r in conn.execute("select * from models")] bench = {(r["model_id"], r["category"]): r["blended_score"] for r in conn.execute("select model_id,category,blended_score from proficiency")} out = {(r["m"], r["c"]): (r["s"], r["n"]) for r in conn.execute( """select model_id m, task_category c, sum(verdict='succeeded') s, count(*) n from verifications where kind='client_outcome' and task_category is not null group by 1,2""")} decisions = conn.execute("""select task_category c, task_tier t, required_context_tokens rc, latency_tolerance lt from route_decisions where kind='chat' and task_category is not null and required_context_tokens is not null and task_tier is not null""").fetchall() conn.close() def eb_scores(cat): obs = {m: (s, n) for (m, c), (s, n) in out.items() if c == cat and n > 0} ids = {m["model_id"] for m in models} if not obs: return {m: bench.get((m, cat)) for m in ids} peer_rate = sum(s for s, n in obs.values()) / sum(n for s, n in obs.values()) bl = [bench[(m, cat)] for m in obs if (m, cat) in bench and bench[(m, cat)] is not None] peer_bench = sum(bl) / len(bl) if bl else 1.0 sc = {} for m in ids: b = bench.get((m, cat)) if b is None: sc[m] = None; continue prior = min(1.0, peer_rate * (b / peer_bench)) if peer_bench else peer_rate s, n = obs.get(m, (0, 0)) sc[m] = (n * (s / n) + K * prior) / (n + K) if n else prior return sc def replay(fn, label, show=8): cache, tot, mix = {}, 0.0, {} for d in decisions: sc = cache.setdefault(d["c"], fn(d["c"])) rows = [{**m, "proficiency": sc.get(m["model_id"])} for m in models] cand = select_candidates(rows, required_context_tokens=d["rc"], required_tier=d["t"], latency_tolerance=d["lt"] or "interactive", allowed_access_levels=cfg.routing.allowed_access_levels, exclude_stale=cfg.freshness.exclude_stale, exclude_deprecated=cfg.freshness.exclude_deprecated) rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance, prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens, cache_rate=cfg.objective.assumed_cache_rate) if rk: tot += rk[0]["cost"] or 0 mix[rk[0]["model_id"]] = mix.get(rk[0]["model_id"], 0) + 1 print(f"\n=== {label} === ${tot:,.2f}") for m, n in sorted(mix.items(), key=lambda kv: -kv[1])[:show]: print(f" {m:<26}{n:>6}") return tot a = replay(lambda c: {m["model_id"]: bench.get((m["model_id"], c)) for m in models}, "A. today") d = replay(eb_scores, "D. empirical-Bayes on traffic scale") print(f"\nA ${a:,.2f} B $179.28 (naive fold-in) D ${d:,.2f} D/A {d/a:.2f}x D/B {d/179.28:.2f}x") for cat in ("coding_refactor", "coding_general", "tool_use_agentic"): sc = eb_scores(cat) print(f"\n{cat} (peer traffic rate anchors the prior)") print(f" {'model':<24}{'bench':>7}{'n':>5}{'rate':>7}{'score':>8}") for m in sorted(sc, key=lambda m: -(sc[m] if sc[m] is not None else -1))[:8]: b = bench.get((m, cat)); s, n = out.get((m, cat), (0, 0)) if b is None: continue print(f" {m:<24}{b:>7.3f}{n:>5}{(f'{s/n*100:.0f}%' if n else '-'):>7}{sc[m]:>8.3f}") ``` ### Script E — epsilon-greedy exploration pricing (§4) ```python """Price an epsilon-greedy exploration budget on the real decision stream.""" import sqlite3, sys, random sys.path.insert(0, "src") from config import load_config from routing import select_candidates, rank_candidates, estimated_cost cfg = load_config("config/config.yaml") conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row models = [dict(r) for r in conn.execute("select * from models")] bench = {(r["model_id"], r["category"]): r["blended_score"] for r in conn.execute("select model_id,category,blended_score from proficiency")} out = {(r["m"], r["c"]): r["n"] for r in conn.execute( """select model_id m, task_category c, count(*) n from verifications where kind='client_outcome' and task_category is not null group by 1,2""")} ds = conn.execute("""select task_category c, task_tier t, required_context_tokens rc, latency_tolerance lt from route_decisions where kind='chat' and task_category is not null and required_context_tokens is not null and task_tier is not null""").fetchall() conn.close() EPS, MAX_RATIO = 0.03, 4.0 rng = random.Random(7) exploit_cost = explore_cost = 0.0 n_explore = 0; gained = {} for d in ds: rows = [{**m, "proficiency": bench.get((m["model_id"], d["c"]))} for m in models] cand = select_candidates(rows, required_context_tokens=d["rc"], required_tier=d["t"], latency_tolerance=d["lt"] or "interactive", allowed_access_levels=cfg.routing.allowed_access_levels, exclude_stale=cfg.freshness.exclude_stale, exclude_deprecated=cfg.freshness.exclude_deprecated) rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance, prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens, cache_rate=cfg.objective.assumed_cache_rate) if not rk: continue win = rk[0]; wc = win["cost"] or 0 exploit_cost += wc # explore: fewest outcome samples in this category, cost-capped, cheapest tiebreak if rng.random() < EPS and len(rk) > 1: pool = [r for r in rk if (r["cost"] or 0) <= MAX_RATIO * max(wc, 1e-9)] pool = [r for r in pool if r["model_id"] != win["model_id"]] if pool: pick = min(pool, key=lambda r: (out.get((r["model_id"], d["c"]), 0), r["cost"] or 0)) explore_cost += pick["cost"] or 0 n_explore += 1 gained[(pick["model_id"], d["c"])] = gained.get((pick["model_id"], d["c"]), 0) + 1 continue explore_cost += wc print(f"decisions {len(ds):,} explored {n_explore:,} ({n_explore/len(ds)*100:.1f}%)") print(f"exploit-only cost ${exploit_cost:,.2f}") print(f"with exploration ${explore_cost:,.2f} (+${explore_cost-exploit_cost:,.2f}, " f"{(explore_cost/exploit_cost-1)*100:+.1f}%)") print(f"\nnew outcome samples this window would have bought (top 12):") for (m, c), n in sorted(gained.items(), key=lambda kv: -kv[1])[:12]: have = out.get((m, c), 0) print(f" {m:<24}{c:<18}+{n:>4} (had {have})") ```