Three UI changes and two plan corrections.
Set as default. routing.default_profile decides what every client that does
not name a profile gets -- including all 13 opencode agents, which send bare
llm-router/auto -- and it was reachable only from a dropdown on Controls,
with the profiles page unable to even show which profile was live. The
profiles list now marks the default, the detail pane badges it, and a button
sets it. It posts to the same allowlisted config endpoint the Controls page
uses, so the validation is the one that already exists rather than a second
rule that can drift. is_default is read from the config store rather than
cfg, because cfg binds at import and would report the pre-restart value at
exactly the moment the operator is looking at it. Delete is disabled on the
current default, saying so before the click instead of after the 422.
Allowlist folds. The two lists ran together in one scroll column with
identical row styling, so the only cue for which list a row belonged to was
whether its button was red or blue -- and the allowed scroller cut a row in
half at the boundary, which read as a rendering fault rather than a divider.
They are now separate collapsible sections, each boxed, each with its count
in the header so a folded one still reports what it holds under the filter.
The catalog starts folded: opening the manager should not dump 425 rows
nobody asked for. The allowed scroller is 7 * 38px so it cuts on a row.
Degraded output plan, second trigger. Measured on the live router while
onlycheaps was default and opencode hammered a free model: 38 of 108 calls
to nemotron-3-nano-omni:free came back MALFORMED EMPTY on HTTP 200 -- 35.2%,
against 0% from three other models over the same window. Nothing was logged
as an upstream failure because nothing failed; the circuit breaker trips on
status >= 400 and cannot see this at all, so the router kept dispatching
with no backoff. verify_response caught every one, and structural verdicts
are diagnostics only, so it detected the degradation 38 times and could do
nothing. That is a stronger case for the plan than the mojibake it was
written for, and it flips the scope decision: encoding faults are
provider-shaped, content faults are model-shaped, so the signature decides
the key.
Exposure-bias plan: marked done, not planned. It was labelled planned in the
status backfill on the strength of its own "FINAL -- ready to implement"
header and a memory note saying "until the fix lands". Both describe when
they were written. The code shipped long ago -- exploration enabled at
epsilon 0.03, outcome_prior_strength 20, FAILURE_VERDICTS ("failed",),
expected_success_rate present, and the taxonomy live in the table
(outcome_prior 264, self_eval_thin 147, outcome_blended 38). What remains is
1,136 unfolded outcomes of 2,221, which is an operator decision about an
irreversible DB mutation, not missing code. Cost a wasted dispatch to Atlas;
the correction is recorded in the doc so it cannot cost another.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
32 KiB
Fixing exposure bias (#1) and self-sealing exclusions (#2)
Status: done -- code shipped; only the operator-gated backlog fold remains (see below)
Correction, 2026-09-08. This doc was marked
plannedin the status backfill on the strength of its own "ready to implement" framing and a memory note warning not to fold the backlog yet. Both were read as "the code is not written". It is. Verified against the live tree:exploration: enabled=True epsilon=0.03 max_cost_ratio=4.0 max_tier=2 proficiency.outcome_prior_strength: 20 feedback.FAILURE_VERDICTS = ("failed",) # structural verdicts excluded proficiency.expected_success_rate() # presentAnd the three-branch taxonomy is live in the table, not just in code --
outcome_prior264,self_eval_thin147,outcome_blended38,self_eval2.What actually remains is not code. 1,136 of 2,221 client outcomes are unfolded (all attributable). Spending them mutates
router.dbirreversibly, which is why the plan sequences it last and why it needs an operator decision rather than an implementer. That is the whole of the outstanding work.The lesson is the one this repo already has a memory for: check the code path before calling something a gap. A plan that says "FINAL -- ready to implement" describes the state when it was written, not now.
Date: 2026-09-01
Follows: plans/conceptual-review-premise-and-execution.md findings #1 and #2
Status: FINAL — ready to implement. Validated by replay against the
9,007 real chat decisions in router.db. Nothing here has been implemented;
router.db was not mutated. Both open decisions were resolved 2026-09-01 and
are now stated as instructions, not options:
route_decisions.request_idships as part of this work — it is a prerequisite for the validation in §5, not a follow-up.- Structural and
local_llmverdicts stop feeding proficiency and remain diagnostics only. See §3.
Section 6 is the file-by-file task list.
0. They are one problem seen from two sides
Both findings reduce to the same sentence: absence of evidence is currently treated as evidence of quality.
- On the ranking side (#1), a model with no traffic evidence keeps its
saturated benchmark score and wins the band —
glm-5.2-fasttakes 2,243 decisions after the naive fold-in because it has never been measured on real work. - On the selection side (#2), a model excluded by a 3-sample benchmark score
never accumulates the evidence that would overturn it —
deepseek-v4-flashwon 0 of 2,512tool_use_agenticdecisions and therefore has 0 outcome samples there, forever.
So the fixes are complementary, not alternatives. Fixing the score without adding exploration leaves cells that can never be filled. Adding exploration without fixing the score feeds good data into a path that misreads it.
Ordering matters: fix the ingestion path first, then turn on exploration, then spend the 961-row backlog.
1. What is actually wrong with the score
feedback.py and eval_proficiency.py both write through
proficiency_store.add_self_eval into a single self_eval_score running mean,
equal weight per sample. But they are measuring different quantities:
benchmark (evals/tasks.yaml) |
client outcomes (POST /outcome) |
|
|---|---|---|
| scale | ~1.00 (saturated; 51% of rows sit at exactly 1.00) | ~0.80 |
| comparability | controlled — every model gets the same task | confounded — models get the requests routing sent them |
| discrimination | weak | strong |
| cost to add | a benchmark run | free |
Averaging a saturated absolute score with an uncontrolled success rate produces a number whose value depends on the mixing ratio, and the mixing ratio is set by how much the model has been used. That is the bug, stated exactly.
The fix is not to weight the two sources — it is to stop treating the benchmark as a level and start treating it as a prior, with everything expressed on the scale that matters: expected probability that a request succeeds.
2. The scoring rule
Per (model, category), with k a prior strength in samples (k = 20 below):
peer_rate = Σ successes / Σ samples # over models with traffic in this category
peer_bench = mean benchmark score # over that same set, so the ratio is calibrated
prior_m = min(1.0, peer_rate × bench_m / peer_bench)
score_m = (n_m × rate_m + k × prior_m) / (n_m + k)
Read it as: the benchmark says where this model sits relative to its peers; the peer traffic rate says what that position is worth in practice; a model's own traffic pulls the estimate toward its observed rate in proportion to how much of it there is.
Properties that matter:
- An unproven model lands at the peer average, not at the ceiling. That is the direct fix for #1. It is still admitted (consistent with the project's "absent evidence does not disqualify" rule) — it just does not get to sit above every measured model for free.
- Scores become interpretable.
blended_scorenow means "expected pass rate on your traffic."quality_tolerance: 0.1becomes "10 percentage points of real success rate," which is a knob you can reason about. Today it bands an abstract 0-1 quality index whose units are the benchmark's. - Cold start is unchanged. A category with no traffic falls back to the
benchmark exactly as today; a model with neither returns
None→ neutral 0.5 downstream. - Shrinkage replaces the sample-count threshold.
self_eval_min_samplesgates a step change;kis continuous, which is the behaviour that section was reaching for.
Measured effect
Replaying all 9,007 real chat decisions through
select_candidates + rank_candidates, changing only the proficiency values:
| policy | est. cost | vs today |
|---|---|---|
| A. today (benchmark only) | $137.76 | 1.00x |
B. naive fold-in (feedback.py as written) |
$179.28 | 1.30x |
| D. proposed (empirical Bayes, above) | $109.35 | 0.79x |
The proposed rule is 39% cheaper than the naive fold-in and 21% cheaper than today, and the traffic moves toward the models with the best measured real-world pass rates rather than away from them:
today proposed
kimi-k2.7-code 2,722 deepseek-v4-flash 3,260
qwen3.6-35b 2,646 qwen3.6-35b 2,646
qwen3.6-35b-fast 1,236 kimi-k2.7-code 2,210
glm-5.3 745 glm-5.2-fast 558
deepseek-v4-flash 500 qwen3.6-35b-fast 133
coding_refactor under the rule — note the unproven rows now sit below or
level with the proven ones instead of above them:
| model | bench | n | rate | score |
|---|---|---|---|---|
glm-5.3 |
1.000 | 1 | 100% | 0.795 |
glm-5.2-flex |
1.000 | 0 | — | 0.785 |
kimi-k3-fast |
1.000 | 0 | — | 0.785 |
gemma-4-31b |
0.997 | 0 | — | 0.783 |
deepseek-v4-flash |
0.826 | 95 | 81% | 0.782 |
All inside one quality_tolerance band, so cost decides — and deepseek, the
one with 95 real samples at 81%, is by far the cheapest. That is the outcome
the review said was missing.
One honest tradeoff
With this little traffic evidence, scores compress (0.78-0.93 in most
categories), so a 0.1 band covers much of the range and cost decides more
often than it does today. That is correct given the evidence — nothing is yet
proven better — but it is a real behavioural change, not a free win. Two
levers: raise k to lean harder on the benchmark while traffic is thin, or
narrow quality_tolerance now that its units are meaningful. Revisit both once
exploration has been running a few weeks.
Rejected: a "proven ceiling" cap
The first formulation tried was: cap any model with < 30 outcome samples at
the best traffic-proven score in its category. It was simulated and rejected —
it only bites when the best proven model happens to score low, so coding_general
collapsed to a single flat value while coding_refactor was untouched.
Inconsistent, and it needed a second threshold. The empirical-Bayes form gets
the same effect from one formula with no special case.
3. Implementation — scoring
Schema (both via the existing ensure_columns migration pattern)
ALTER TABLE proficiency ADD COLUMN outcome_score REAL;
ALTER TABLE proficiency ADD COLUMN outcome_samples INTEGER DEFAULT 0;
ALTER TABLE route_decisions ADD COLUMN request_id TEXT; -- see note below
ALTER TABLE route_decisions ADD COLUMN exploration INTEGER DEFAULT 0;
config/schema.sql gains the same columns for fresh installs. Mirror
proficiency_store.ensure_columns / _ensure_route_decisions_table so a live
router.db migrates on load and on write, idempotently.
route_decisions.request_id is a prerequisite, not a nice-to-have. There is
currently no exact join from a client outcome back to the decision that produced
it — the review had to approximate with session_key + a 5-second window, which
is why the tier/outcome table in §5 there is marked indicative. report_outcome
already resolves request_id against energy_observations; recording it on the
decision closes the loop and is what makes the validation in §5 below possible.
proficiency.py — one new pure function
Stays I/O-free; the per-category aggregates are passed in by the caller.
def expected_success_rate(
benchmark_score: float | None,
outcome_score: float | None,
outcome_samples: int,
*,
peer_rate: float | None, # None when the category has no traffic yet
peer_benchmark: float | None,
prior_strength: int,
) -> tuple[float | None, Source | None]:
...
Returns (None, None) when there is neither benchmark nor outcome data, so the
neutral-0.5 path downstream is unchanged. Add "outcome_blended" to Source
so provenance stays inspectable the way self_eval_thin already is.
proficiency_store.py — a second writer, and a category recompute
-
New
add_outcome(conn, cfg, model_id, provider, category, scores)writingoutcome_score/outcome_samplesthrough the existingaccumulate. Keep the 0-1 clamp — the comment about a harness bug pushing a score to 1.50 and raisingbestfor every candidate applies verbatim here. -
blended_scorebecomes a category-level computation, becausepeer_rateandpeer_benchmarkare aggregates over the category._writecannot produce the final value from one row any more. Split it:_writekeeps writingleaderboard_score/self_eval_score/self_eval_samples(and nowoutcome_*), and leavesblended_scoreat the benchmark value fromblend()so a single write is never internally inconsistent.- New
recompute_category(conn, cfg, category)reads every row in the category, re-derives each row's benchmark score by calling the existingblend()on its stored components, computespeer_rate/peer_benchmarkfromoutcome_score/outcome_samples, then writes the finalblended_score+sourceper row.
Deriving the benchmark half from the stored components rather than caching it means no fourth score column and no drift — the invariant this module exists to hold (it is the only writer, so
blended_scoreandsourcecan never disagree with their inputs) survives intact, anddispatcher.load_candidatesstays untouched on the hot path. -
Every writer calls
recompute_categoryat the end of its transaction:add_self_eval,add_outcome,propagate_to_variants, and theleaderboard.pyimporter. -
propagate_to_variantsmust copyoutcome_score/outcome_samplestoo, and theinherited_fromguard applies unchanged. -
ensure_columnsgains the two new columns. No backfill is needed or wanted:outcome_samplesdefaults to 0, which is exactly true of every existing row. (Contrastinherited_from, where a NULL default was actively wrong and needed_backfill_inherited.)
config.py / config/config.yaml — the prior strength is a knob
ProficiencyConfig gains outcome_prior_strength: int (the k in §2), and it
must be stated in config/config.yaml, not left to a Pydantic default. This
project has already been bitten once by a setting that existed only as a default
and was therefore invisible to anyone tuning it (classifier.max_input_chars,
shipped into the wrong section, accepted and discarded because it happened to
match the code default). Every knob belongs in the file.
proficiency:
# Prior strength, in samples, for folding real client outcomes into a score.
# A model's own traffic outweighs the benchmark-derived prior once it has
# more than this many outcome samples. 20 was chosen because the categories
# with real evidence carry 25-163 samples, so it lets a well-measured model
# move while still holding a 5-sample cell near its peers.
outcome_prior_strength: 20
self_eval_min_samples, leaderboard_weight and self_eval_weight are
unchanged — they still govern blend(), which now produces the benchmark half
that feeds the prior rather than the final score.
feedback.py — client outcomes only
SUCCESS_VERDICTS stays. The changes:
- Client outcomes route to
add_outcomeinstead ofadd_self_eval.applied_atidempotency is unchanged. FAILURE_VERDICTSdropstruncatedandmalformed— structural andlocal_llmverdicts stop feeding proficiency entirely and remain diagnostics. They are failure-only contributors (a passing structural check is deliberately not recorded), and you cannot form a rate from failures alone: folding them in would biasoutcome_scoredownward by exactly however often the checker happened to fire, which is a property of the checker, not the model. They are 134 rows against 987 client outcomes, and the structural checker's one documented encounter with real agent traffic produced ~29 falsemalformeds before thehas_tool_callsfix.failed(the client-outcome failure verdict) stays.coverage()keeps reporting all verdicts, so the diagnostics stay visible where they belong. Its existing 80%-unverifiable warning is unaffected.
4. Implementation — exploration
The rule
On an ε-share of requests, instead of the rank winner, dispatch to the
hard-filter-eligible candidate with the fewest outcome samples in this
category, tie-broken by lowest cost, and skipped entirely if its estimated
cost exceeds max_cost_ratio × the winner's.
Every hard filter still applies — context window, tier floor, access level, latency class, vision, JSON mode. Exploration only ever reorders within the eligible set, so it can never produce a request the model cannot serve.
Cost, measured
Replayed at ε = 0.03, max_cost_ratio 4.0, over the same 9,007 decisions:
explored 263 of 9,007 (2.9%)
exploit-only $137.76
with explore $139.12 (+$1.36, +1.0%)
$1.36 over nine days. What it buys:
| cell | new samples | had |
|---|---|---|
deepseek-v4-flash / tool_use_agentic |
+66 | 0 |
kimi-k2.7-code-fast / coding_refactor |
+29 | 6 |
kimi-k3 / coding_general |
+23 | 0 |
glm-5.2-fast / coding_refactor |
+22 | 0 |
kimi-k3 / coding_refactor |
+19 | 0 |
qwen3.6-35b-fast / coding_general |
+18 | 0 |
qwen3.6-35b / docs_writing |
+16 | 0 |
| …plus 5 more cells currently at zero |
The "fewest samples first" rule targets empty cells without being told to, and 66 samples is enough to confirm or kill the 3-task 0.333 that currently bans the cheapest capable model from the largest category of traffic.
Module shape
New src/exploration.py, following circuit_breaker.py / session_cache.py:
pure, module-level state only if needed, injected RNG the way
circuit_breaker injects time, and it never imports dispatcher or config.
def choose(
ranked: Sequence[dict],
sample_counts: dict[str, int],
*,
epsilon: float,
max_cost_ratio: float,
rng: random.Random,
) -> tuple[dict, bool]: # (row, was_exploration)
dispatcher.load_candidates already LEFT JOINs proficiency; add
p.outcome_samples to the SELECT and pass the counts in. The call site is
immediately after rank_candidates, before apply_flex_preference — a flex
swap is a serving-class decision and should apply to whatever was chosen.
Config
exploration:
# Routes a small share of requests to the least-evidenced eligible candidate
# so proficiency scores can be corrected by evidence rather than frozen by
# the first 3-sample benchmark that touched them. Measured on 9,007 real
# decisions: 2.9% of traffic, +$1.36 (+1.0%), and it fills 11 (model,
# category) cells that currently hold ZERO outcome samples -- including
# deepseek-v4-flash / tool_use_agentic, which the router cannot otherwise
# ever measure because its own ranking excludes it.
enabled: true
epsilon: 0.03
# Never explore into something more than this multiple of the winner's cost.
max_cost_ratio: 4.0
# Exploration is for gathering evidence, not for gambling on high-stakes
# work. Tier 3 is excluded; 2,034 of 2,512 tool_use_agentic decisions are
# tier 2, so this costs almost no coverage.
max_tier: 2
ExplorationConfig(StrictModel) in config.py, registered on RouterConfig,
mirroring CircuitBreakerConfig's shape (defaults on the model and stated in
the file). Validate 0.0 <= epsilon <= 1.0 and max_cost_ratio >= 1.0 — an
epsilon above 1 or a ratio below 1 are both silently self-defeating rather than
loud.
There is deliberately no min_samples_target knob. "Fewest outcome samples
first" needs no threshold: it targets empty cells on its own, and once a cell
fills, the next-emptiest becomes the target automatically.
Shipping this enabled: true breaks the project's usual "new knob ships off"
convention, and does so deliberately. Off, it changes nothing and
config.yaml's own stated experiment ("run with it off, let POST /outcome
report real pass/fail, and compare tool_use_agentic proficiency for deepseek
before and after") stays unrunnable — which is precisely the condition the
review flagged. The convention exists to stop unproven knobs changing behaviour
silently; this one has a measured cost (+1.0%), a measured benefit (11 empty
cells), and a hard cost cap (max_cost_ratio). It is the exception that earns
itself.
5. Sequencing
Four commits, in this order. Steps 1-2 are inert until step 3 runs, so the service can be restarted between any of them.
- Schema +
request_id. Migrations only; no behaviour change. Verify a liverouter.dbmigrates and existing rows are intact. - Scoring path.
expected_success_rate,add_outcome,recompute_category,feedback.pyretargeted. Still inert — no outcome rows have been applied yet, so everyoutcome_samplesis 0 and every score reproduces today's value. That is the acceptance test for this step: after step 2 and before step 3, replaying the decision stream must still produce $137.76. - Spend the backlog.
python -m feedback --dry-run, thenpython -m feedback, on the 961 unapplied rows through the new path. Expected: the §2 table — traffic shifts towarddeepseek-v4-flashand total estimated cost falls to ~$109. - Exploration on. Then re-run
baseline_report.pyweekly.
Validation, at ~2 weeks
The check that matters: deepseek-v4-flash / tool_use_agentic should hold
≳60 real outcome samples. Compare its measured rate against the 0.333 benchmark
score that currently bans it from 2,512 decisions. Either the benchmark was
right and the exclusion is now earned, or it was a 3-sample artifact costing
roughly 3x on the largest category of traffic. Both answers are worth $1.36.
Then revisit quality_tolerance and outcome_prior_strength, with scores that
finally mean something and enough evidence to set them from.
The report this unlocks
Once route_decisions.exploration and request_id exist, the confound named
in review §4 becomes addressable: pass rates computed on explored requests
only are unconfounded by the routing policy, because assignment was random
within the eligible set. That is the first genuinely causal comparison this
project will have been able to make, and it is the thing that settles whether
the expensive models are worth their price.
6. File-by-file task list
Commit 1 — schema
config/schema.sql— addoutcome_score REAL,outcome_samples INTEGER DEFAULT 0toproficiency; addrequest_id TEXT,exploration INTEGER DEFAULT 0to theroute_decisionsDDL.src/proficiency_store.py::ensure_columns— the twoproficiencycolumns. No backfill (0 is correct for every existing row).src/dispatcher.py::ensure_route_decisions+_ensure_route_decisions_table— the tworoute_decisionscolumns, same idempotent pattern.src/dispatcher.py::persist_route_decision— writerequest_idon every decision, andexploration(0 for now).request_idis already in scope on the chat path; on/route(which spends nothing upstream) it stays NULL.
Commit 2 — scoring
src/proficiency.py— addexpected_success_rate(...)per §2 and"outcome_blended"to theSourceliteral. Keep it pure: peer aggregates are arguments.src/proficiency_store.py— split_write(benchmark half only, plus the new columns); addadd_outcome; addrecompute_category; call it fromadd_self_eval,add_outcome,propagate_to_variants; extendpropagate_to_variantsto copyoutcome_score/outcome_samples.src/leaderboard.py— importer callsrecompute_categoryafter its writes.src/config.py—ProficiencyConfig.outcome_prior_strength: int.config/config.yaml— theoutcome_prior_strengthblock from §3.src/feedback.py— client outcomes →add_outcome; droptruncatedandmalformedfromFAILURE_VERDICTS; docstring updated to say why.
Commit 3 — exploration
src/exploration.py— new, pure, injected RNG, noconfig/dispatcherimport.src/config.py—ExplorationConfig(StrictModel)+ field onRouterConfig.config/config.yaml— theexploration:block from §4.src/dispatcher.py— addp.outcome_samplestoload_candidates's SELECT; callexploration.chooseafterrank_candidatesand beforeapply_flex_preference; gate oncfg.exploration.enabledandtask_tier <= max_tier; setexploration=1on the decision row and addexplore=to the existinglogs.inforouting line.
Tests (tests/test_proficiency.py, tests/test_feedback.py, new
tests/test_exploration.py) — pin the properties, not the arithmetic:
- an unproven model does not outrank a proven one in the same category;
- a category with no outcome data reproduces today's benchmark score exactly (this is the step-2 acceptance test in miniature);
- a model with neither source still yields
None→ neutral 0.5 downstream; recompute_categoryis idempotent and leaves untouched categories alone;- structural/
local_llmverdicts no longer move any score; - exploration never returns a row that fails a hard filter, never exceeds
max_cost_ratio, and returns the winner unchanged whenepsilonis 0; - with a seeded RNG, the explore share lands within tolerance of
epsilon.
Docs — docs/data-model.md (four new columns), docs/routing.md
(scoring section: the score is now an expected pass rate, and what
quality_tolerance means in those units), docs/evaluation.md (benchmark is
now a prior, not the score). CLAUDE.md's "Proficiency: category now changes
routing" and "The only ground truth" sections both need the new story once
step 3 has run and the numbers are real.
Check while you are in here — the admin portal's "apply feedback"
operational trigger runs feedback.py; confirm it still works after the
retarget, and that the models page shows the new source value rather than
blanking on an unrecognized string.
7. Summary
| fix | cost | effect | |
|---|---|---|---|
| #1 | benchmark becomes a prior; outcomes are a separate, shrunk, peer-relative signal on the traffic scale | one formula + 2 columns | $179.28 → $109.35 on replay; unproven models stop winning |
| #2 | ε=3% exploration to the least-evidenced eligible candidate | +$1.36 / 9 days (+1.0%) | 11 empty cells filled, incl. +66 deepseek/tool_use_agentic; makes the table correctable |
Neither is a large change. Together they convert the router from something carefully hand-tuned against your traffic into something that learns from it — which is what the design said it was for.
Appendix — reproduction
Scripts used to produce every number above are read-only and replay against
router.db without mutating it. fix_sim2.py (empirical Bayes) and
explore_sim.py (ε-greedy pricing) are the two that matter; both follow the
same shape as the appendix scripts in
plans/conceptual-review-premise-and-execution.md — load models +
proficiency + verifications, rebuild each decision's eligible set with
routing.select_candidates, rank with routing.rank_candidates, and sum
estimated_cost. Constants used: k = 20, ε = 0.03,
max_cost_ratio = 4.0, RNG seed 7, and the shipped
quality_tolerance = 0.1 / assumed_cache_rate = 0.917 /
assumed_completion_tokens = 500.
Script D — empirical-Bayes replay (§2)
"""Empirical-Bayes version: shrink toward a benchmark-informed peer prior,
all expressed on the TRAFFIC scale (expected pass rate).
prior_m = peer_rate x (bench_m / peer_bench) # benchmark sets relative position
score_m = (n_m * rate_m + k * prior_m) / (n_m + k)
An unproven model lands at the category's average traffic performance, adjusted
by where the benchmark puts it -- not at the benchmark ceiling.
"""
import sqlite3, sys
sys.path.insert(0, "src")
from config import load_config
from routing import select_candidates, rank_candidates
cfg = load_config("config/config.yaml"); K = 20
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = [dict(r) for r in conn.execute("select * from models")]
bench = {(r["model_id"], r["category"]): r["blended_score"]
for r in conn.execute("select model_id,category,blended_score from proficiency")}
out = {(r["m"], r["c"]): (r["s"], r["n"]) for r in conn.execute(
"""select model_id m, task_category c, sum(verdict='succeeded') s, count(*) n
from verifications where kind='client_outcome' and task_category is not null
group by 1,2""")}
decisions = conn.execute("""select task_category c, task_tier t,
required_context_tokens rc, latency_tolerance lt from route_decisions
where kind='chat' and task_category is not null
and required_context_tokens is not null and task_tier is not null""").fetchall()
conn.close()
def eb_scores(cat):
obs = {m: (s, n) for (m, c), (s, n) in out.items() if c == cat and n > 0}
ids = {m["model_id"] for m in models}
if not obs:
return {m: bench.get((m, cat)) for m in ids}
peer_rate = sum(s for s, n in obs.values()) / sum(n for s, n in obs.values())
bl = [bench[(m, cat)] for m in obs if (m, cat) in bench and bench[(m, cat)] is not None]
peer_bench = sum(bl) / len(bl) if bl else 1.0
sc = {}
for m in ids:
b = bench.get((m, cat))
if b is None:
sc[m] = None; continue
prior = min(1.0, peer_rate * (b / peer_bench)) if peer_bench else peer_rate
s, n = obs.get(m, (0, 0))
sc[m] = (n * (s / n) + K * prior) / (n + K) if n else prior
return sc
def replay(fn, label, show=8):
cache, tot, mix = {}, 0.0, {}
for d in decisions:
sc = cache.setdefault(d["c"], fn(d["c"]))
rows = [{**m, "proficiency": sc.get(m["model_id"])} for m in models]
cand = select_candidates(rows, required_context_tokens=d["rc"], required_tier=d["t"],
latency_tolerance=d["lt"] or "interactive",
allowed_access_levels=cfg.routing.allowed_access_levels,
exclude_stale=cfg.freshness.exclude_stale,
exclude_deprecated=cfg.freshness.exclude_deprecated)
rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
cache_rate=cfg.objective.assumed_cache_rate)
if rk:
tot += rk[0]["cost"] or 0
mix[rk[0]["model_id"]] = mix.get(rk[0]["model_id"], 0) + 1
print(f"\n=== {label} === ${tot:,.2f}")
for m, n in sorted(mix.items(), key=lambda kv: -kv[1])[:show]:
print(f" {m:<26}{n:>6}")
return tot
a = replay(lambda c: {m["model_id"]: bench.get((m["model_id"], c)) for m in models}, "A. today")
d = replay(eb_scores, "D. empirical-Bayes on traffic scale")
print(f"\nA ${a:,.2f} B $179.28 (naive fold-in) D ${d:,.2f} D/A {d/a:.2f}x D/B {d/179.28:.2f}x")
for cat in ("coding_refactor", "coding_general", "tool_use_agentic"):
sc = eb_scores(cat)
print(f"\n{cat} (peer traffic rate anchors the prior)")
print(f" {'model':<24}{'bench':>7}{'n':>5}{'rate':>7}{'score':>8}")
for m in sorted(sc, key=lambda m: -(sc[m] if sc[m] is not None else -1))[:8]:
b = bench.get((m, cat)); s, n = out.get((m, cat), (0, 0))
if b is None: continue
print(f" {m:<24}{b:>7.3f}{n:>5}{(f'{s/n*100:.0f}%' if n else '-'):>7}{sc[m]:>8.3f}")
Script E — epsilon-greedy exploration pricing (§4)
"""Price an epsilon-greedy exploration budget on the real decision stream."""
import sqlite3, sys, random
sys.path.insert(0, "src")
from config import load_config
from routing import select_candidates, rank_candidates, estimated_cost
cfg = load_config("config/config.yaml")
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = [dict(r) for r in conn.execute("select * from models")]
bench = {(r["model_id"], r["category"]): r["blended_score"]
for r in conn.execute("select model_id,category,blended_score from proficiency")}
out = {(r["m"], r["c"]): r["n"] for r in conn.execute(
"""select model_id m, task_category c, count(*) n from verifications
where kind='client_outcome' and task_category is not null group by 1,2""")}
ds = conn.execute("""select task_category c, task_tier t, required_context_tokens rc,
latency_tolerance lt from route_decisions where kind='chat'
and task_category is not null and required_context_tokens is not null
and task_tier is not null""").fetchall()
conn.close()
EPS, MAX_RATIO = 0.03, 4.0
rng = random.Random(7)
exploit_cost = explore_cost = 0.0
n_explore = 0; gained = {}
for d in ds:
rows = [{**m, "proficiency": bench.get((m["model_id"], d["c"]))} for m in models]
cand = select_candidates(rows, required_context_tokens=d["rc"], required_tier=d["t"],
latency_tolerance=d["lt"] or "interactive",
allowed_access_levels=cfg.routing.allowed_access_levels,
exclude_stale=cfg.freshness.exclude_stale,
exclude_deprecated=cfg.freshness.exclude_deprecated)
rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
cache_rate=cfg.objective.assumed_cache_rate)
if not rk: continue
win = rk[0]; wc = win["cost"] or 0
exploit_cost += wc
# explore: fewest outcome samples in this category, cost-capped, cheapest tiebreak
if rng.random() < EPS and len(rk) > 1:
pool = [r for r in rk if (r["cost"] or 0) <= MAX_RATIO * max(wc, 1e-9)]
pool = [r for r in pool if r["model_id"] != win["model_id"]]
if pool:
pick = min(pool, key=lambda r: (out.get((r["model_id"], d["c"]), 0), r["cost"] or 0))
explore_cost += pick["cost"] or 0
n_explore += 1
gained[(pick["model_id"], d["c"])] = gained.get((pick["model_id"], d["c"]), 0) + 1
continue
explore_cost += wc
print(f"decisions {len(ds):,} explored {n_explore:,} ({n_explore/len(ds)*100:.1f}%)")
print(f"exploit-only cost ${exploit_cost:,.2f}")
print(f"with exploration ${explore_cost:,.2f} (+${explore_cost-exploit_cost:,.2f}, "
f"{(explore_cost/exploit_cost-1)*100:+.1f}%)")
print(f"\nnew outcome samples this window would have bought (top 12):")
for (m, c), n in sorted(gained.items(), key=lambda kv: -kv[1])[:12]:
have = out.get((m, c), 0)
print(f" {m:<24}{c:<18}+{n:>4} (had {have})")