Files
6krrt/plans/proficiency-exposure-bias-and-exploration.md
adlee-was-taken b11f4133fd feat(admin): set the routing default from the profiles page; fold the allowlist
Three UI changes and two plan corrections.

Set as default. routing.default_profile decides what every client that does
not name a profile gets -- including all 13 opencode agents, which send bare
llm-router/auto -- and it was reachable only from a dropdown on Controls,
with the profiles page unable to even show which profile was live. The
profiles list now marks the default, the detail pane badges it, and a button
sets it. It posts to the same allowlisted config endpoint the Controls page
uses, so the validation is the one that already exists rather than a second
rule that can drift. is_default is read from the config store rather than
cfg, because cfg binds at import and would report the pre-restart value at
exactly the moment the operator is looking at it. Delete is disabled on the
current default, saying so before the click instead of after the 422.

Allowlist folds. The two lists ran together in one scroll column with
identical row styling, so the only cue for which list a row belonged to was
whether its button was red or blue -- and the allowed scroller cut a row in
half at the boundary, which read as a rendering fault rather than a divider.
They are now separate collapsible sections, each boxed, each with its count
in the header so a folded one still reports what it holds under the filter.
The catalog starts folded: opening the manager should not dump 425 rows
nobody asked for. The allowed scroller is 7 * 38px so it cuts on a row.

Degraded output plan, second trigger. Measured on the live router while
onlycheaps was default and opencode hammered a free model: 38 of 108 calls
to nemotron-3-nano-omni:free came back MALFORMED EMPTY on HTTP 200 -- 35.2%,
against 0% from three other models over the same window. Nothing was logged
as an upstream failure because nothing failed; the circuit breaker trips on
status >= 400 and cannot see this at all, so the router kept dispatching
with no backoff. verify_response caught every one, and structural verdicts
are diagnostics only, so it detected the degradation 38 times and could do
nothing. That is a stronger case for the plan than the mojibake it was
written for, and it flips the scope decision: encoding faults are
provider-shaped, content faults are model-shaped, so the signature decides
the key.

Exposure-bias plan: marked done, not planned. It was labelled planned in the
status backfill on the strength of its own "FINAL -- ready to implement"
header and a memory note saying "until the fix lands". Both describe when
they were written. The code shipped long ago -- exploration enabled at
epsilon 0.03, outcome_prior_strength 20, FAILURE_VERDICTS ("failed",),
expected_success_rate present, and the taxonomy live in the table
(outcome_prior 264, self_eval_thin 147, outcome_blended 38). What remains is
1,136 unfolded outcomes of 2,221, which is an operator decision about an
irreversible DB mutation, not missing code. Cost a wasted dispatch to Atlas;
the correction is recorded in the doc so it cannot cost another.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 20:49:36 -04:00

32 KiB
Raw Permalink Blame History

Fixing exposure bias (#1) and self-sealing exclusions (#2)

Status: done -- code shipped; only the operator-gated backlog fold remains (see below)

Correction, 2026-09-08. This doc was marked planned in the status backfill on the strength of its own "ready to implement" framing and a memory note warning not to fold the backlog yet. Both were read as "the code is not written". It is. Verified against the live tree:

exploration: enabled=True epsilon=0.03 max_cost_ratio=4.0 max_tier=2
proficiency.outcome_prior_strength: 20
feedback.FAILURE_VERDICTS = ("failed",)     # structural verdicts excluded
proficiency.expected_success_rate()          # present

And the three-branch taxonomy is live in the table, not just in code -- outcome_prior 264, self_eval_thin 147, outcome_blended 38, self_eval 2.

What actually remains is not code. 1,136 of 2,221 client outcomes are unfolded (all attributable). Spending them mutates router.db irreversibly, which is why the plan sequences it last and why it needs an operator decision rather than an implementer. That is the whole of the outstanding work.

The lesson is the one this repo already has a memory for: check the code path before calling something a gap. A plan that says "FINAL -- ready to implement" describes the state when it was written, not now.

Date: 2026-09-01 Follows: plans/conceptual-review-premise-and-execution.md findings #1 and #2 Status: FINAL — ready to implement. Validated by replay against the 9,007 real chat decisions in router.db. Nothing here has been implemented; router.db was not mutated. Both open decisions were resolved 2026-09-01 and are now stated as instructions, not options:

  1. route_decisions.request_id ships as part of this work — it is a prerequisite for the validation in §5, not a follow-up.
  2. Structural and local_llm verdicts stop feeding proficiency and remain diagnostics only. See §3.

Section 6 is the file-by-file task list.


0. They are one problem seen from two sides

Both findings reduce to the same sentence: absence of evidence is currently treated as evidence of quality.

  • On the ranking side (#1), a model with no traffic evidence keeps its saturated benchmark score and wins the band — glm-5.2-fast takes 2,243 decisions after the naive fold-in because it has never been measured on real work.
  • On the selection side (#2), a model excluded by a 3-sample benchmark score never accumulates the evidence that would overturn it — deepseek-v4-flash won 0 of 2,512 tool_use_agentic decisions and therefore has 0 outcome samples there, forever.

So the fixes are complementary, not alternatives. Fixing the score without adding exploration leaves cells that can never be filled. Adding exploration without fixing the score feeds good data into a path that misreads it.

Ordering matters: fix the ingestion path first, then turn on exploration, then spend the 961-row backlog.


1. What is actually wrong with the score

feedback.py and eval_proficiency.py both write through proficiency_store.add_self_eval into a single self_eval_score running mean, equal weight per sample. But they are measuring different quantities:

benchmark (evals/tasks.yaml) client outcomes (POST /outcome)
scale ~1.00 (saturated; 51% of rows sit at exactly 1.00) ~0.80
comparability controlled — every model gets the same task confounded — models get the requests routing sent them
discrimination weak strong
cost to add a benchmark run free

Averaging a saturated absolute score with an uncontrolled success rate produces a number whose value depends on the mixing ratio, and the mixing ratio is set by how much the model has been used. That is the bug, stated exactly.

The fix is not to weight the two sources — it is to stop treating the benchmark as a level and start treating it as a prior, with everything expressed on the scale that matters: expected probability that a request succeeds.


2. The scoring rule

Per (model, category), with k a prior strength in samples (k = 20 below):

peer_rate   = Σ successes / Σ samples          # over models with traffic in this category
peer_bench  = mean benchmark score             # over that same set, so the ratio is calibrated
prior_m     = min(1.0, peer_rate × bench_m / peer_bench)
score_m     = (n_m × rate_m + k × prior_m) / (n_m + k)

Read it as: the benchmark says where this model sits relative to its peers; the peer traffic rate says what that position is worth in practice; a model's own traffic pulls the estimate toward its observed rate in proportion to how much of it there is.

Properties that matter:

  • An unproven model lands at the peer average, not at the ceiling. That is the direct fix for #1. It is still admitted (consistent with the project's "absent evidence does not disqualify" rule) — it just does not get to sit above every measured model for free.
  • Scores become interpretable. blended_score now means "expected pass rate on your traffic." quality_tolerance: 0.1 becomes "10 percentage points of real success rate," which is a knob you can reason about. Today it bands an abstract 0-1 quality index whose units are the benchmark's.
  • Cold start is unchanged. A category with no traffic falls back to the benchmark exactly as today; a model with neither returns None → neutral 0.5 downstream.
  • Shrinkage replaces the sample-count threshold. self_eval_min_samples gates a step change; k is continuous, which is the behaviour that section was reaching for.

Measured effect

Replaying all 9,007 real chat decisions through select_candidates + rank_candidates, changing only the proficiency values:

policy est. cost vs today
A. today (benchmark only) $137.76 1.00x
B. naive fold-in (feedback.py as written) $179.28 1.30x
D. proposed (empirical Bayes, above) $109.35 0.79x

The proposed rule is 39% cheaper than the naive fold-in and 21% cheaper than today, and the traffic moves toward the models with the best measured real-world pass rates rather than away from them:

today                          proposed
  kimi-k2.7-code    2,722        deepseek-v4-flash  3,260
  qwen3.6-35b       2,646        qwen3.6-35b        2,646
  qwen3.6-35b-fast  1,236        kimi-k2.7-code     2,210
  glm-5.3             745        glm-5.2-fast         558
  deepseek-v4-flash   500        qwen3.6-35b-fast     133

coding_refactor under the rule — note the unproven rows now sit below or level with the proven ones instead of above them:

model bench n rate score
glm-5.3 1.000 1 100% 0.795
glm-5.2-flex 1.000 0 — 0.785
kimi-k3-fast 1.000 0 — 0.785
gemma-4-31b 0.997 0 — 0.783
deepseek-v4-flash 0.826 95 81% 0.782

All inside one quality_tolerance band, so cost decides — and deepseek, the one with 95 real samples at 81%, is by far the cheapest. That is the outcome the review said was missing.

One honest tradeoff

With this little traffic evidence, scores compress (0.78-0.93 in most categories), so a 0.1 band covers much of the range and cost decides more often than it does today. That is correct given the evidence — nothing is yet proven better — but it is a real behavioural change, not a free win. Two levers: raise k to lean harder on the benchmark while traffic is thin, or narrow quality_tolerance now that its units are meaningful. Revisit both once exploration has been running a few weeks.

Rejected: a "proven ceiling" cap

The first formulation tried was: cap any model with < 30 outcome samples at the best traffic-proven score in its category. It was simulated and rejected — it only bites when the best proven model happens to score low, so coding_general collapsed to a single flat value while coding_refactor was untouched. Inconsistent, and it needed a second threshold. The empirical-Bayes form gets the same effect from one formula with no special case.


3. Implementation — scoring

Schema (both via the existing ensure_columns migration pattern)

ALTER TABLE proficiency ADD COLUMN outcome_score   REAL;
ALTER TABLE proficiency ADD COLUMN outcome_samples INTEGER DEFAULT 0;

ALTER TABLE route_decisions ADD COLUMN request_id  TEXT;      -- see note below
ALTER TABLE route_decisions ADD COLUMN exploration INTEGER DEFAULT 0;

config/schema.sql gains the same columns for fresh installs. Mirror proficiency_store.ensure_columns / _ensure_route_decisions_table so a live router.db migrates on load and on write, idempotently.

route_decisions.request_id is a prerequisite, not a nice-to-have. There is currently no exact join from a client outcome back to the decision that produced it — the review had to approximate with session_key + a 5-second window, which is why the tier/outcome table in §5 there is marked indicative. report_outcome already resolves request_id against energy_observations; recording it on the decision closes the loop and is what makes the validation in §5 below possible.

proficiency.py — one new pure function

Stays I/O-free; the per-category aggregates are passed in by the caller.

def expected_success_rate(
    benchmark_score: float | None,
    outcome_score: float | None,
    outcome_samples: int,
    *,
    peer_rate: float | None,        # None when the category has no traffic yet
    peer_benchmark: float | None,
    prior_strength: int,
) -> tuple[float | None, Source | None]:
    ...

Returns (None, None) when there is neither benchmark nor outcome data, so the neutral-0.5 path downstream is unchanged. Add "outcome_blended" to Source so provenance stays inspectable the way self_eval_thin already is.

proficiency_store.py — a second writer, and a category recompute

  • New add_outcome(conn, cfg, model_id, provider, category, scores) writing outcome_score / outcome_samples through the existing accumulate. Keep the 0-1 clamp — the comment about a harness bug pushing a score to 1.50 and raising best for every candidate applies verbatim here.

  • blended_score becomes a category-level computation, because peer_rate and peer_benchmark are aggregates over the category. _write cannot produce the final value from one row any more. Split it:

    • _write keeps writing leaderboard_score / self_eval_score / self_eval_samples (and now outcome_*), and leaves blended_score at the benchmark value from blend() so a single write is never internally inconsistent.
    • New recompute_category(conn, cfg, category) reads every row in the category, re-derives each row's benchmark score by calling the existing blend() on its stored components, computes peer_rate / peer_benchmark from outcome_score / outcome_samples, then writes the final blended_score + source per row.

    Deriving the benchmark half from the stored components rather than caching it means no fourth score column and no drift — the invariant this module exists to hold (it is the only writer, so blended_score and source can never disagree with their inputs) survives intact, and dispatcher.load_candidates stays untouched on the hot path.

  • Every writer calls recompute_category at the end of its transaction: add_self_eval, add_outcome, propagate_to_variants, and the leaderboard.py importer.

  • propagate_to_variants must copy outcome_score / outcome_samples too, and the inherited_from guard applies unchanged.

  • ensure_columns gains the two new columns. No backfill is needed or wanted: outcome_samples defaults to 0, which is exactly true of every existing row. (Contrast inherited_from, where a NULL default was actively wrong and needed _backfill_inherited.)

config.py / config/config.yaml — the prior strength is a knob

ProficiencyConfig gains outcome_prior_strength: int (the k in §2), and it must be stated in config/config.yaml, not left to a Pydantic default. This project has already been bitten once by a setting that existed only as a default and was therefore invisible to anyone tuning it (classifier.max_input_chars, shipped into the wrong section, accepted and discarded because it happened to match the code default). Every knob belongs in the file.

proficiency:
  # Prior strength, in samples, for folding real client outcomes into a score.
  # A model's own traffic outweighs the benchmark-derived prior once it has
  # more than this many outcome samples. 20 was chosen because the categories
  # with real evidence carry 25-163 samples, so it lets a well-measured model
  # move while still holding a 5-sample cell near its peers.
  outcome_prior_strength: 20

self_eval_min_samples, leaderboard_weight and self_eval_weight are unchanged — they still govern blend(), which now produces the benchmark half that feeds the prior rather than the final score.

feedback.py — client outcomes only

SUCCESS_VERDICTS stays. The changes:

  • Client outcomes route to add_outcome instead of add_self_eval. applied_at idempotency is unchanged.
  • FAILURE_VERDICTS drops truncated and malformed — structural and local_llm verdicts stop feeding proficiency entirely and remain diagnostics. They are failure-only contributors (a passing structural check is deliberately not recorded), and you cannot form a rate from failures alone: folding them in would bias outcome_score downward by exactly however often the checker happened to fire, which is a property of the checker, not the model. They are 134 rows against 987 client outcomes, and the structural checker's one documented encounter with real agent traffic produced ~29 false malformeds before the has_tool_calls fix. failed (the client-outcome failure verdict) stays.
  • coverage() keeps reporting all verdicts, so the diagnostics stay visible where they belong. Its existing 80%-unverifiable warning is unaffected.

4. Implementation — exploration

The rule

On an ε-share of requests, instead of the rank winner, dispatch to the hard-filter-eligible candidate with the fewest outcome samples in this category, tie-broken by lowest cost, and skipped entirely if its estimated cost exceeds max_cost_ratio × the winner's.

Every hard filter still applies — context window, tier floor, access level, latency class, vision, JSON mode. Exploration only ever reorders within the eligible set, so it can never produce a request the model cannot serve.

Cost, measured

Replayed at ε = 0.03, max_cost_ratio 4.0, over the same 9,007 decisions:

explored 263 of 9,007 (2.9%)
exploit-only   $137.76
with explore   $139.12      (+$1.36, +1.0%)

$1.36 over nine days. What it buys:

cell new samples had
deepseek-v4-flash / tool_use_agentic +66 0
kimi-k2.7-code-fast / coding_refactor +29 6
kimi-k3 / coding_general +23 0
glm-5.2-fast / coding_refactor +22 0
kimi-k3 / coding_refactor +19 0
qwen3.6-35b-fast / coding_general +18 0
qwen3.6-35b / docs_writing +16 0
…plus 5 more cells currently at zero

The "fewest samples first" rule targets empty cells without being told to, and 66 samples is enough to confirm or kill the 3-task 0.333 that currently bans the cheapest capable model from the largest category of traffic.

Module shape

New src/exploration.py, following circuit_breaker.py / session_cache.py: pure, module-level state only if needed, injected RNG the way circuit_breaker injects time, and it never imports dispatcher or config.

def choose(
    ranked: Sequence[dict],
    sample_counts: dict[str, int],
    *,
    epsilon: float,
    max_cost_ratio: float,
    rng: random.Random,
) -> tuple[dict, bool]:      # (row, was_exploration)

dispatcher.load_candidates already LEFT JOINs proficiency; add p.outcome_samples to the SELECT and pass the counts in. The call site is immediately after rank_candidates, before apply_flex_preference — a flex swap is a serving-class decision and should apply to whatever was chosen.

Config

exploration:
  # Routes a small share of requests to the least-evidenced eligible candidate
  # so proficiency scores can be corrected by evidence rather than frozen by
  # the first 3-sample benchmark that touched them. Measured on 9,007 real
  # decisions: 2.9% of traffic, +$1.36 (+1.0%), and it fills 11 (model,
  # category) cells that currently hold ZERO outcome samples -- including
  # deepseek-v4-flash / tool_use_agentic, which the router cannot otherwise
  # ever measure because its own ranking excludes it.
  enabled: true
  epsilon: 0.03
  # Never explore into something more than this multiple of the winner's cost.
  max_cost_ratio: 4.0
  # Exploration is for gathering evidence, not for gambling on high-stakes
  # work. Tier 3 is excluded; 2,034 of 2,512 tool_use_agentic decisions are
  # tier 2, so this costs almost no coverage.
  max_tier: 2

ExplorationConfig(StrictModel) in config.py, registered on RouterConfig, mirroring CircuitBreakerConfig's shape (defaults on the model and stated in the file). Validate 0.0 <= epsilon <= 1.0 and max_cost_ratio >= 1.0 — an epsilon above 1 or a ratio below 1 are both silently self-defeating rather than loud.

There is deliberately no min_samples_target knob. "Fewest outcome samples first" needs no threshold: it targets empty cells on its own, and once a cell fills, the next-emptiest becomes the target automatically.

Shipping this enabled: true breaks the project's usual "new knob ships off" convention, and does so deliberately. Off, it changes nothing and config.yaml's own stated experiment ("run with it off, let POST /outcome report real pass/fail, and compare tool_use_agentic proficiency for deepseek before and after") stays unrunnable — which is precisely the condition the review flagged. The convention exists to stop unproven knobs changing behaviour silently; this one has a measured cost (+1.0%), a measured benefit (11 empty cells), and a hard cost cap (max_cost_ratio). It is the exception that earns itself.


5. Sequencing

Four commits, in this order. Steps 1-2 are inert until step 3 runs, so the service can be restarted between any of them.

  1. Schema + request_id. Migrations only; no behaviour change. Verify a live router.db migrates and existing rows are intact.
  2. Scoring path. expected_success_rate, add_outcome, recompute_category, feedback.py retargeted. Still inert — no outcome rows have been applied yet, so every outcome_samples is 0 and every score reproduces today's value. That is the acceptance test for this step: after step 2 and before step 3, replaying the decision stream must still produce $137.76.
  3. Spend the backlog. python -m feedback --dry-run, then python -m feedback, on the 961 unapplied rows through the new path. Expected: the §2 table — traffic shifts toward deepseek-v4-flash and total estimated cost falls to ~$109.
  4. Exploration on. Then re-run baseline_report.py weekly.

Validation, at ~2 weeks

The check that matters: deepseek-v4-flash / tool_use_agentic should hold ≳60 real outcome samples. Compare its measured rate against the 0.333 benchmark score that currently bans it from 2,512 decisions. Either the benchmark was right and the exclusion is now earned, or it was a 3-sample artifact costing roughly 3x on the largest category of traffic. Both answers are worth $1.36.

Then revisit quality_tolerance and outcome_prior_strength, with scores that finally mean something and enough evidence to set them from.

The report this unlocks

Once route_decisions.exploration and request_id exist, the confound named in review §4 becomes addressable: pass rates computed on explored requests only are unconfounded by the routing policy, because assignment was random within the eligible set. That is the first genuinely causal comparison this project will have been able to make, and it is the thing that settles whether the expensive models are worth their price.


6. File-by-file task list

Commit 1 — schema

  • config/schema.sql — add outcome_score REAL, outcome_samples INTEGER DEFAULT 0 to proficiency; add request_id TEXT, exploration INTEGER DEFAULT 0 to the route_decisions DDL.
  • src/proficiency_store.py::ensure_columns — the two proficiency columns. No backfill (0 is correct for every existing row).
  • src/dispatcher.py::ensure_route_decisions + _ensure_route_decisions_table — the two route_decisions columns, same idempotent pattern.
  • src/dispatcher.py::persist_route_decision — write request_id on every decision, and exploration (0 for now). request_id is already in scope on the chat path; on /route (which spends nothing upstream) it stays NULL.

Commit 2 — scoring

  • src/proficiency.py — add expected_success_rate(...) per §2 and "outcome_blended" to the Source literal. Keep it pure: peer aggregates are arguments.
  • src/proficiency_store.py — split _write (benchmark half only, plus the new columns); add add_outcome; add recompute_category; call it from add_self_eval, add_outcome, propagate_to_variants; extend propagate_to_variants to copy outcome_score / outcome_samples.
  • src/leaderboard.py — importer calls recompute_category after its writes.
  • src/config.py — ProficiencyConfig.outcome_prior_strength: int.
  • config/config.yaml — the outcome_prior_strength block from §3.
  • src/feedback.py — client outcomes → add_outcome; drop truncated and malformed from FAILURE_VERDICTS; docstring updated to say why.

Commit 3 — exploration

  • src/exploration.py — new, pure, injected RNG, no config/dispatcher import.
  • src/config.py — ExplorationConfig(StrictModel) + field on RouterConfig.
  • config/config.yaml — the exploration: block from §4.
  • src/dispatcher.py — add p.outcome_samples to load_candidates's SELECT; call exploration.choose after rank_candidates and before apply_flex_preference; gate on cfg.exploration.enabled and task_tier <= max_tier; set exploration=1 on the decision row and add explore= to the existing logs.info routing line.

Tests (tests/test_proficiency.py, tests/test_feedback.py, new tests/test_exploration.py) — pin the properties, not the arithmetic:

  • an unproven model does not outrank a proven one in the same category;
  • a category with no outcome data reproduces today's benchmark score exactly (this is the step-2 acceptance test in miniature);
  • a model with neither source still yields None → neutral 0.5 downstream;
  • recompute_category is idempotent and leaves untouched categories alone;
  • structural/local_llm verdicts no longer move any score;
  • exploration never returns a row that fails a hard filter, never exceeds max_cost_ratio, and returns the winner unchanged when epsilon is 0;
  • with a seeded RNG, the explore share lands within tolerance of epsilon.

Docs — docs/data-model.md (four new columns), docs/routing.md (scoring section: the score is now an expected pass rate, and what quality_tolerance means in those units), docs/evaluation.md (benchmark is now a prior, not the score). CLAUDE.md's "Proficiency: category now changes routing" and "The only ground truth" sections both need the new story once step 3 has run and the numbers are real.

Check while you are in here — the admin portal's "apply feedback" operational trigger runs feedback.py; confirm it still works after the retarget, and that the models page shows the new source value rather than blanking on an unrecognized string.


7. Summary

fix cost effect
#1 benchmark becomes a prior; outcomes are a separate, shrunk, peer-relative signal on the traffic scale one formula + 2 columns $179.28 → $109.35 on replay; unproven models stop winning
#2 ε=3% exploration to the least-evidenced eligible candidate +$1.36 / 9 days (+1.0%) 11 empty cells filled, incl. +66 deepseek/tool_use_agentic; makes the table correctable

Neither is a large change. Together they convert the router from something carefully hand-tuned against your traffic into something that learns from it — which is what the design said it was for.


Appendix — reproduction

Scripts used to produce every number above are read-only and replay against router.db without mutating it. fix_sim2.py (empirical Bayes) and explore_sim.py (ε-greedy pricing) are the two that matter; both follow the same shape as the appendix scripts in plans/conceptual-review-premise-and-execution.md — load models + proficiency + verifications, rebuild each decision's eligible set with routing.select_candidates, rank with routing.rank_candidates, and sum estimated_cost. Constants used: k = 20, ε = 0.03, max_cost_ratio = 4.0, RNG seed 7, and the shipped quality_tolerance = 0.1 / assumed_cache_rate = 0.917 / assumed_completion_tokens = 500.

Script D — empirical-Bayes replay (§2)

"""Empirical-Bayes version: shrink toward a benchmark-informed peer prior,
all expressed on the TRAFFIC scale (expected pass rate).

    prior_m   = peer_rate x (bench_m / peer_bench)      # benchmark sets relative position
    score_m   = (n_m * rate_m + k * prior_m) / (n_m + k)

An unproven model lands at the category's average traffic performance, adjusted
by where the benchmark puts it -- not at the benchmark ceiling.
"""
import sqlite3, sys
sys.path.insert(0, "src")
from config import load_config
from routing import select_candidates, rank_candidates

cfg = load_config("config/config.yaml"); K = 20
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = [dict(r) for r in conn.execute("select * from models")]
bench = {(r["model_id"], r["category"]): r["blended_score"]
         for r in conn.execute("select model_id,category,blended_score from proficiency")}
out = {(r["m"], r["c"]): (r["s"], r["n"]) for r in conn.execute(
    """select model_id m, task_category c, sum(verdict='succeeded') s, count(*) n
       from verifications where kind='client_outcome' and task_category is not null
       group by 1,2""")}
decisions = conn.execute("""select task_category c, task_tier t,
    required_context_tokens rc, latency_tolerance lt from route_decisions
    where kind='chat' and task_category is not null
      and required_context_tokens is not null and task_tier is not null""").fetchall()
conn.close()

def eb_scores(cat):
    obs = {m: (s, n) for (m, c), (s, n) in out.items() if c == cat and n > 0}
    ids = {m["model_id"] for m in models}
    if not obs:
        return {m: bench.get((m, cat)) for m in ids}
    peer_rate = sum(s for s, n in obs.values()) / sum(n for s, n in obs.values())
    bl = [bench[(m, cat)] for m in obs if (m, cat) in bench and bench[(m, cat)] is not None]
    peer_bench = sum(bl) / len(bl) if bl else 1.0
    sc = {}
    for m in ids:
        b = bench.get((m, cat))
        if b is None:
            sc[m] = None; continue
        prior = min(1.0, peer_rate * (b / peer_bench)) if peer_bench else peer_rate
        s, n = obs.get(m, (0, 0))
        sc[m] = (n * (s / n) + K * prior) / (n + K) if n else prior
    return sc

def replay(fn, label, show=8):
    cache, tot, mix = {}, 0.0, {}
    for d in decisions:
        sc = cache.setdefault(d["c"], fn(d["c"]))
        rows = [{**m, "proficiency": sc.get(m["model_id"])} for m in models]
        cand = select_candidates(rows, required_context_tokens=d["rc"], required_tier=d["t"],
            latency_tolerance=d["lt"] or "interactive",
            allowed_access_levels=cfg.routing.allowed_access_levels,
            exclude_stale=cfg.freshness.exclude_stale,
            exclude_deprecated=cfg.freshness.exclude_deprecated)
        rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
            prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
            cache_rate=cfg.objective.assumed_cache_rate)
        if rk:
            tot += rk[0]["cost"] or 0
            mix[rk[0]["model_id"]] = mix.get(rk[0]["model_id"], 0) + 1
    print(f"\n=== {label} ===  ${tot:,.2f}")
    for m, n in sorted(mix.items(), key=lambda kv: -kv[1])[:show]:
        print(f"   {m:<26}{n:>6}")
    return tot

a = replay(lambda c: {m["model_id"]: bench.get((m["model_id"], c)) for m in models}, "A. today")
d = replay(eb_scores, "D. empirical-Bayes on traffic scale")
print(f"\nA ${a:,.2f}   B $179.28 (naive fold-in)   D ${d:,.2f}   D/A {d/a:.2f}x  D/B {d/179.28:.2f}x")

for cat in ("coding_refactor", "coding_general", "tool_use_agentic"):
    sc = eb_scores(cat)
    print(f"\n{cat}   (peer traffic rate anchors the prior)")
    print(f"  {'model':<24}{'bench':>7}{'n':>5}{'rate':>7}{'score':>8}")
    for m in sorted(sc, key=lambda m: -(sc[m] if sc[m] is not None else -1))[:8]:
        b = bench.get((m, cat)); s, n = out.get((m, cat), (0, 0))
        if b is None: continue
        print(f"  {m:<24}{b:>7.3f}{n:>5}{(f'{s/n*100:.0f}%' if n else '-'):>7}{sc[m]:>8.3f}")

Script E — epsilon-greedy exploration pricing (§4)

"""Price an epsilon-greedy exploration budget on the real decision stream."""
import sqlite3, sys, random
sys.path.insert(0, "src")
from config import load_config
from routing import select_candidates, rank_candidates, estimated_cost

cfg = load_config("config/config.yaml")
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = [dict(r) for r in conn.execute("select * from models")]
bench = {(r["model_id"], r["category"]): r["blended_score"]
         for r in conn.execute("select model_id,category,blended_score from proficiency")}
out = {(r["m"], r["c"]): r["n"] for r in conn.execute(
    """select model_id m, task_category c, count(*) n from verifications
       where kind='client_outcome' and task_category is not null group by 1,2""")}
ds = conn.execute("""select task_category c, task_tier t, required_context_tokens rc,
    latency_tolerance lt from route_decisions where kind='chat'
    and task_category is not null and required_context_tokens is not null
    and task_tier is not null""").fetchall()
conn.close()

EPS, MAX_RATIO = 0.03, 4.0
rng = random.Random(7)
exploit_cost = explore_cost = 0.0
n_explore = 0; gained = {}
for d in ds:
    rows = [{**m, "proficiency": bench.get((m["model_id"], d["c"]))} for m in models]
    cand = select_candidates(rows, required_context_tokens=d["rc"], required_tier=d["t"],
        latency_tolerance=d["lt"] or "interactive",
        allowed_access_levels=cfg.routing.allowed_access_levels,
        exclude_stale=cfg.freshness.exclude_stale,
        exclude_deprecated=cfg.freshness.exclude_deprecated)
    rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
        prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
        cache_rate=cfg.objective.assumed_cache_rate)
    if not rk: continue
    win = rk[0]; wc = win["cost"] or 0
    exploit_cost += wc
    # explore: fewest outcome samples in this category, cost-capped, cheapest tiebreak
    if rng.random() < EPS and len(rk) > 1:
        pool = [r for r in rk if (r["cost"] or 0) <= MAX_RATIO * max(wc, 1e-9)]
        pool = [r for r in pool if r["model_id"] != win["model_id"]]
        if pool:
            pick = min(pool, key=lambda r: (out.get((r["model_id"], d["c"]), 0), r["cost"] or 0))
            explore_cost += pick["cost"] or 0
            n_explore += 1
            gained[(pick["model_id"], d["c"])] = gained.get((pick["model_id"], d["c"]), 0) + 1
            continue
    explore_cost += wc

print(f"decisions {len(ds):,}   explored {n_explore:,} ({n_explore/len(ds)*100:.1f}%)")
print(f"exploit-only cost  ${exploit_cost:,.2f}")
print(f"with exploration   ${explore_cost:,.2f}   (+${explore_cost-exploit_cost:,.2f}, "
      f"{(explore_cost/exploit_cost-1)*100:+.1f}%)")
print(f"\nnew outcome samples this window would have bought (top 12):")
for (m, c), n in sorted(gained.items(), key=lambda kv: -kv[1])[:12]:
    have = out.get((m, c), 0)
    print(f"  {m:<24}{c:<18}+{n:>4}   (had {have})")