Files
6krrt/plans/proficiency-exposure-bias-and-exploration.md
adlee-was-taken b11f4133fd feat(admin): set the routing default from the profiles page; fold the allowlist
Three UI changes and two plan corrections.

Set as default. routing.default_profile decides what every client that does
not name a profile gets -- including all 13 opencode agents, which send bare
llm-router/auto -- and it was reachable only from a dropdown on Controls,
with the profiles page unable to even show which profile was live. The
profiles list now marks the default, the detail pane badges it, and a button
sets it. It posts to the same allowlisted config endpoint the Controls page
uses, so the validation is the one that already exists rather than a second
rule that can drift. is_default is read from the config store rather than
cfg, because cfg binds at import and would report the pre-restart value at
exactly the moment the operator is looking at it. Delete is disabled on the
current default, saying so before the click instead of after the 422.

Allowlist folds. The two lists ran together in one scroll column with
identical row styling, so the only cue for which list a row belonged to was
whether its button was red or blue -- and the allowed scroller cut a row in
half at the boundary, which read as a rendering fault rather than a divider.
They are now separate collapsible sections, each boxed, each with its count
in the header so a folded one still reports what it holds under the filter.
The catalog starts folded: opening the manager should not dump 425 rows
nobody asked for. The allowed scroller is 7 * 38px so it cuts on a row.

Degraded output plan, second trigger. Measured on the live router while
onlycheaps was default and opencode hammered a free model: 38 of 108 calls
to nemotron-3-nano-omni:free came back MALFORMED EMPTY on HTTP 200 -- 35.2%,
against 0% from three other models over the same window. Nothing was logged
as an upstream failure because nothing failed; the circuit breaker trips on
status >= 400 and cannot see this at all, so the router kept dispatching
with no backoff. verify_response caught every one, and structural verdicts
are diagnostics only, so it detected the degradation 38 times and could do
nothing. That is a stronger case for the plan than the mojibake it was
written for, and it flips the scope decision: encoding faults are
provider-shaped, content faults are model-shaped, so the signature decides
the key.

Exposure-bias plan: marked done, not planned. It was labelled planned in the
status backfill on the strength of its own "FINAL -- ready to implement"
header and a memory note saying "until the fix lands". Both describe when
they were written. The code shipped long ago -- exploration enabled at
epsilon 0.03, outcome_prior_strength 20, FAILURE_VERDICTS ("failed",),
expected_success_rate present, and the taxonomy live in the table
(outcome_prior 264, self_eval_thin 147, outcome_blended 38). What remains is
1,136 unfolded outcomes of 2,221, which is an operator decision about an
irreversible DB mutation, not missing code. Cost a wasted dispatch to Atlas;
the correction is recorded in the doc so it cannot cost another.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 20:49:36 -04:00

697 lines
32 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Fixing exposure bias (#1) and self-sealing exclusions (#2)
Status: done -- code shipped; only the operator-gated backlog fold remains (see below)
> **Correction, 2026-09-08.** This doc was marked `planned` in the status
> backfill on the strength of its own "ready to implement" framing and a
> memory note warning not to fold the backlog yet. Both were read as "the
> code is not written". It is. Verified against the live tree:
>
> ```
> exploration: enabled=True epsilon=0.03 max_cost_ratio=4.0 max_tier=2
> proficiency.outcome_prior_strength: 20
> feedback.FAILURE_VERDICTS = ("failed",) # structural verdicts excluded
> proficiency.expected_success_rate() # present
> ```
>
> And the three-branch taxonomy is live in the table, not just in code --
> `outcome_prior` 264, `self_eval_thin` 147, `outcome_blended` 38,
> `self_eval` 2.
>
> What actually remains is not code. **1,136 of 2,221 client outcomes are
> unfolded** (all attributable). Spending them mutates `router.db`
> irreversibly, which is why the plan sequences it last and why it needs an
> operator decision rather than an implementer. That is the whole of the
> outstanding work.
>
> The lesson is the one this repo already has a memory for: check the code
> path before calling something a gap. A plan that says "FINAL -- ready to
> implement" describes the state when it was written, not now.
**Date:** 2026-09-01
**Follows:** `plans/conceptual-review-premise-and-execution.md` findings #1 and #2
**Status:** **FINAL — ready to implement.** Validated by replay against the
9,007 real chat decisions in `router.db`. Nothing here has been implemented;
`router.db` was not mutated. Both open decisions were resolved 2026-09-01 and
are now stated as instructions, not options:
1. `route_decisions.request_id` ships as part of this work — it is a
prerequisite for the validation in §5, not a follow-up.
2. Structural and `local_llm` verdicts **stop feeding proficiency** and remain
diagnostics only. See §3.
Section 6 is the file-by-file task list.
---
## 0. They are one problem seen from two sides
Both findings reduce to the same sentence: **absence of evidence is currently
treated as evidence of quality.**
- On the **ranking** side (#1), a model with no traffic evidence keeps its
saturated benchmark score and wins the band — `glm-5.2-fast` takes 2,243
decisions after the naive fold-in *because* it has never been measured on
real work.
- On the **selection** side (#2), a model excluded by a 3-sample benchmark score
never accumulates the evidence that would overturn it — `deepseek-v4-flash`
won 0 of 2,512 `tool_use_agentic` decisions and therefore has 0 outcome
samples there, forever.
So the fixes are complementary, not alternatives. Fixing the score without
adding exploration leaves cells that can never be filled. Adding exploration
without fixing the score feeds good data into a path that misreads it.
Ordering matters: **fix the ingestion path first**, then turn on exploration,
then spend the 961-row backlog.
---
## 1. What is actually wrong with the score
`feedback.py` and `eval_proficiency.py` both write through
`proficiency_store.add_self_eval` into a single `self_eval_score` running mean,
equal weight per sample. But they are measuring different quantities:
| | benchmark (`evals/tasks.yaml`) | client outcomes (`POST /outcome`) |
|---|---|---|
| scale | ~1.00 (saturated; 51% of rows sit at exactly 1.00) | ~0.80 |
| comparability | controlled — every model gets the same task | confounded — models get the requests routing sent them |
| discrimination | weak | strong |
| cost to add | a benchmark run | free |
Averaging a saturated absolute score with an uncontrolled success rate produces
a number whose value depends on the **mixing ratio**, and the mixing ratio is
set by how much the model has been used. That is the bug, stated exactly.
The fix is not to weight the two sources — it is to **stop treating the
benchmark as a level and start treating it as a prior**, with everything
expressed on the scale that matters: expected probability that a request
succeeds.
---
## 2. The scoring rule
Per (model, category), with `k` a prior strength in samples (`k = 20` below):
```
peer_rate = Σ successes / Σ samples # over models with traffic in this category
peer_bench = mean benchmark score # over that same set, so the ratio is calibrated
prior_m = min(1.0, peer_rate × bench_m / peer_bench)
score_m = (n_m × rate_m + k × prior_m) / (n_m + k)
```
Read it as: *the benchmark says where this model sits relative to its peers;
the peer traffic rate says what that position is worth in practice; a model's
own traffic pulls the estimate toward its observed rate in proportion to how
much of it there is.*
Properties that matter:
- **An unproven model lands at the peer average, not at the ceiling.** That is
the direct fix for #1. It is still admitted (consistent with the project's
"absent evidence does not disqualify" rule) — it just does not get to sit
above every measured model for free.
- **Scores become interpretable.** `blended_score` now means "expected pass
rate on your traffic." `quality_tolerance: 0.1` becomes "10 percentage points
of real success rate," which is a knob you can reason about. Today it bands
an abstract 0-1 quality index whose units are the benchmark's.
- **Cold start is unchanged.** A category with no traffic falls back to the
benchmark exactly as today; a model with neither returns `None` → neutral 0.5
downstream.
- **Shrinkage replaces the sample-count threshold.** `self_eval_min_samples`
gates a step change; `k` is continuous, which is the behaviour that section
was reaching for.
### Measured effect
Replaying all 9,007 real chat decisions through
`select_candidates` + `rank_candidates`, changing only the proficiency values:
| policy | est. cost | vs today |
|---|---|---|
| **A. today** (benchmark only) | $137.76 | 1.00x |
| **B. naive fold-in** (`feedback.py` as written) | $179.28 | **1.30x** |
| **D. proposed** (empirical Bayes, above) | **$109.35** | **0.79x** |
The proposed rule is **39% cheaper than the naive fold-in** and 21% cheaper
than today, and the traffic moves toward the models with the best *measured
real-world* pass rates rather than away from them:
```
today proposed
kimi-k2.7-code 2,722 deepseek-v4-flash 3,260
qwen3.6-35b 2,646 qwen3.6-35b 2,646
qwen3.6-35b-fast 1,236 kimi-k2.7-code 2,210
glm-5.3 745 glm-5.2-fast 558
deepseek-v4-flash 500 qwen3.6-35b-fast 133
```
`coding_refactor` under the rule — note the unproven rows now sit *below* or
level with the proven ones instead of above them:
| model | bench | n | rate | score |
|---|---|---|---|---|
| `glm-5.3` | 1.000 | 1 | 100% | 0.795 |
| `glm-5.2-flex` | 1.000 | 0 | — | 0.785 |
| `kimi-k3-fast` | 1.000 | 0 | — | 0.785 |
| `gemma-4-31b` | 0.997 | 0 | — | 0.783 |
| `deepseek-v4-flash` | 0.826 | 95 | 81% | 0.782 |
All inside one `quality_tolerance` band, so cost decides — and deepseek, the
one with 95 real samples at 81%, is by far the cheapest. That is the outcome
the review said was missing.
### One honest tradeoff
With this little traffic evidence, scores compress (0.78-0.93 in most
categories), so a 0.1 band covers much of the range and **cost decides more
often than it does today.** That is correct given the evidence — nothing is yet
*proven* better — but it is a real behavioural change, not a free win. Two
levers: raise `k` to lean harder on the benchmark while traffic is thin, or
narrow `quality_tolerance` now that its units are meaningful. Revisit both once
exploration has been running a few weeks.
### Rejected: a "proven ceiling" cap
The first formulation tried was: cap any model with `< 30` outcome samples at
the best *traffic-proven* score in its category. It was simulated and rejected —
it only bites when the best proven model happens to score low, so `coding_general`
collapsed to a single flat value while `coding_refactor` was untouched.
Inconsistent, and it needed a second threshold. The empirical-Bayes form gets
the same effect from one formula with no special case.
---
## 3. Implementation — scoring
### Schema (both via the existing `ensure_columns` migration pattern)
```sql
ALTER TABLE proficiency ADD COLUMN outcome_score REAL;
ALTER TABLE proficiency ADD COLUMN outcome_samples INTEGER DEFAULT 0;
ALTER TABLE route_decisions ADD COLUMN request_id TEXT; -- see note below
ALTER TABLE route_decisions ADD COLUMN exploration INTEGER DEFAULT 0;
```
`config/schema.sql` gains the same columns for fresh installs. Mirror
`proficiency_store.ensure_columns` / `_ensure_route_decisions_table` so a live
`router.db` migrates on load and on write, idempotently.
**`route_decisions.request_id` is a prerequisite, not a nice-to-have.** There is
currently no exact join from a client outcome back to the decision that produced
it — the review had to approximate with `session_key` + a 5-second window, which
is why the tier/outcome table in §5 there is marked indicative. `report_outcome`
already resolves `request_id` against `energy_observations`; recording it on the
decision closes the loop and is what makes the validation in §5 below possible.
### `proficiency.py` — one new pure function
Stays I/O-free; the per-category aggregates are passed in by the caller.
```python
def expected_success_rate(
benchmark_score: float | None,
outcome_score: float | None,
outcome_samples: int,
*,
peer_rate: float | None, # None when the category has no traffic yet
peer_benchmark: float | None,
prior_strength: int,
) -> tuple[float | None, Source | None]:
...
```
Returns `(None, None)` when there is neither benchmark nor outcome data, so the
neutral-0.5 path downstream is unchanged. Add `"outcome_blended"` to `Source`
so provenance stays inspectable the way `self_eval_thin` already is.
### `proficiency_store.py` — a second writer, and a category recompute
- New `add_outcome(conn, cfg, model_id, provider, category, scores)` writing
`outcome_score` / `outcome_samples` through the existing `accumulate`. Keep
the 0-1 clamp — the comment about a harness bug pushing a score to 1.50 and
raising `best` for every candidate applies verbatim here.
- **`blended_score` becomes a category-level computation**, because `peer_rate`
and `peer_benchmark` are aggregates over the category. `_write` cannot produce
the final value from one row any more. Split it:
- `_write` keeps writing `leaderboard_score` / `self_eval_score` /
`self_eval_samples` (and now `outcome_*`), and leaves `blended_score` at the
benchmark value from `blend()` so a single write is never internally
inconsistent.
- New `recompute_category(conn, cfg, category)` reads every row in the
category, re-derives each row's benchmark score by calling the **existing**
`blend()` on its stored components, computes `peer_rate` / `peer_benchmark`
from `outcome_score` / `outcome_samples`, then writes the final
`blended_score` + `source` per row.
Deriving the benchmark half from the stored components rather than caching it
means **no fourth score column and no drift** — the invariant this module
exists to hold (it is the only writer, so `blended_score` and `source` can
never disagree with their inputs) survives intact, and
`dispatcher.load_candidates` stays untouched on the hot path.
- Every writer calls `recompute_category` at the end of its transaction:
`add_self_eval`, `add_outcome`, `propagate_to_variants`, and the
`leaderboard.py` importer.
- `propagate_to_variants` must copy `outcome_score` / `outcome_samples` too, and
the `inherited_from` guard applies unchanged.
- `ensure_columns` gains the two new columns. No backfill is needed or wanted:
`outcome_samples` defaults to 0, which is exactly true of every existing row.
(Contrast `inherited_from`, where a NULL default was actively wrong and needed
`_backfill_inherited`.)
### `config.py` / `config/config.yaml` — the prior strength is a knob
`ProficiencyConfig` gains `outcome_prior_strength: int` (the `k` in §2), and it
**must be stated in `config/config.yaml`**, not left to a Pydantic default. This
project has already been bitten once by a setting that existed only as a default
and was therefore invisible to anyone tuning it (`classifier.max_input_chars`,
shipped into the wrong section, accepted and discarded because it happened to
match the code default). Every knob belongs in the file.
```yaml
proficiency:
# Prior strength, in samples, for folding real client outcomes into a score.
# A model's own traffic outweighs the benchmark-derived prior once it has
# more than this many outcome samples. 20 was chosen because the categories
# with real evidence carry 25-163 samples, so it lets a well-measured model
# move while still holding a 5-sample cell near its peers.
outcome_prior_strength: 20
```
`self_eval_min_samples`, `leaderboard_weight` and `self_eval_weight` are
unchanged — they still govern `blend()`, which now produces the *benchmark half*
that feeds the prior rather than the final score.
### `feedback.py` — client outcomes only
`SUCCESS_VERDICTS` stays. The changes:
- Client outcomes route to `add_outcome` instead of `add_self_eval`.
`applied_at` idempotency is unchanged.
- **`FAILURE_VERDICTS` drops `truncated` and `malformed`** — structural and
`local_llm` verdicts stop feeding proficiency entirely and remain diagnostics.
They are *failure-only* contributors (a passing structural check is
deliberately not recorded), and you cannot form a rate from failures alone:
folding them in would bias `outcome_score` downward by exactly however often
the checker happened to fire, which is a property of the checker, not the
model. They are 134 rows against 987 client outcomes, and the structural
checker's one documented encounter with real agent traffic produced ~29 false
`malformed`s before the `has_tool_calls` fix. `failed` (the client-outcome
failure verdict) stays.
- `coverage()` keeps reporting all verdicts, so the diagnostics stay visible
where they belong. Its existing 80%-unverifiable warning is unaffected.
---
## 4. Implementation — exploration
### The rule
On an ε-share of requests, instead of the rank winner, dispatch to the
**hard-filter-eligible candidate with the fewest outcome samples in this
category**, tie-broken by lowest cost, and skipped entirely if its estimated
cost exceeds `max_cost_ratio ×` the winner's.
Every hard filter still applies — context window, tier floor, access level,
latency class, vision, JSON mode. Exploration only ever reorders *within the
eligible set*, so it can never produce a request the model cannot serve.
### Cost, measured
Replayed at ε = 0.03, `max_cost_ratio` 4.0, over the same 9,007 decisions:
```
explored 263 of 9,007 (2.9%)
exploit-only $137.76
with explore $139.12 (+$1.36, +1.0%)
```
**$1.36 over nine days.** What it buys:
| cell | new samples | had |
|---|---|---|
| `deepseek-v4-flash` / `tool_use_agentic` | **+66** | **0** |
| `kimi-k2.7-code-fast` / `coding_refactor` | +29 | 6 |
| `kimi-k3` / `coding_general` | +23 | 0 |
| `glm-5.2-fast` / `coding_refactor` | +22 | 0 |
| `kimi-k3` / `coding_refactor` | +19 | 0 |
| `qwen3.6-35b-fast` / `coding_general` | +18 | 0 |
| `qwen3.6-35b` / `docs_writing` | +16 | 0 |
| …plus 5 more cells currently at zero | | |
The "fewest samples first" rule targets empty cells without being told to, and
66 samples is enough to confirm or kill the 3-task 0.333 that currently bans
the cheapest capable model from the largest category of traffic.
### Module shape
New `src/exploration.py`, following `circuit_breaker.py` / `session_cache.py`:
pure, module-level state only if needed, **injected RNG** the way
`circuit_breaker` injects time, and it never imports `dispatcher` or `config`.
```python
def choose(
ranked: Sequence[dict],
sample_counts: dict[str, int],
*,
epsilon: float,
max_cost_ratio: float,
rng: random.Random,
) -> tuple[dict, bool]: # (row, was_exploration)
```
`dispatcher.load_candidates` already LEFT JOINs `proficiency`; add
`p.outcome_samples` to the SELECT and pass the counts in. The call site is
immediately after `rank_candidates`, before `apply_flex_preference` — a flex
swap is a serving-class decision and should apply to whatever was chosen.
### Config
```yaml
exploration:
# Routes a small share of requests to the least-evidenced eligible candidate
# so proficiency scores can be corrected by evidence rather than frozen by
# the first 3-sample benchmark that touched them. Measured on 9,007 real
# decisions: 2.9% of traffic, +$1.36 (+1.0%), and it fills 11 (model,
# category) cells that currently hold ZERO outcome samples -- including
# deepseek-v4-flash / tool_use_agentic, which the router cannot otherwise
# ever measure because its own ranking excludes it.
enabled: true
epsilon: 0.03
# Never explore into something more than this multiple of the winner's cost.
max_cost_ratio: 4.0
# Exploration is for gathering evidence, not for gambling on high-stakes
# work. Tier 3 is excluded; 2,034 of 2,512 tool_use_agentic decisions are
# tier 2, so this costs almost no coverage.
max_tier: 2
```
`ExplorationConfig(StrictModel)` in `config.py`, registered on `RouterConfig`,
mirroring `CircuitBreakerConfig`'s shape (defaults on the model *and* stated in
the file). Validate `0.0 <= epsilon <= 1.0` and `max_cost_ratio >= 1.0` — an
epsilon above 1 or a ratio below 1 are both silently self-defeating rather than
loud.
There is deliberately no `min_samples_target` knob. "Fewest outcome samples
first" needs no threshold: it targets empty cells on its own, and once a cell
fills, the next-emptiest becomes the target automatically.
**Shipping this `enabled: true` breaks the project's usual "new knob ships off"
convention, and does so deliberately.** Off, it changes nothing and
`config.yaml`'s own stated experiment ("run with it off, let `POST /outcome`
report real pass/fail, and compare `tool_use_agentic` proficiency for deepseek
before and after") stays unrunnable — which is precisely the condition the
review flagged. The convention exists to stop unproven knobs changing behaviour
silently; this one has a measured cost (+1.0%), a measured benefit (11 empty
cells), and a hard cost cap (`max_cost_ratio`). It is the exception that earns
itself.
---
## 5. Sequencing
Four commits, in this order. Steps 1-2 are inert until step 3 runs, so the
service can be restarted between any of them.
1. **Schema + `request_id`.** Migrations only; no behaviour change. Verify a
live `router.db` migrates and existing rows are intact.
2. **Scoring path.** `expected_success_rate`, `add_outcome`,
`recompute_category`, `feedback.py` retargeted. Still inert — no outcome rows
have been applied yet, so every `outcome_samples` is 0 and every score
reproduces today's value. That is the acceptance test for this step: **after
step 2 and before step 3, replaying the decision stream must still produce
$137.76.**
3. **Spend the backlog.** `python -m feedback --dry-run`, then
`python -m feedback`, on the 961 unapplied rows through the new path.
Expected: the §2 table — traffic shifts toward `deepseek-v4-flash` and total
estimated cost falls to ~$109.
4. **Exploration on.** Then re-run `baseline_report.py` weekly.
### Validation, at ~2 weeks
The check that matters: `deepseek-v4-flash` / `tool_use_agentic` should hold
≳60 real outcome samples. Compare its measured rate against the 0.333 benchmark
score that currently bans it from 2,512 decisions. Either the benchmark was
right and the exclusion is now *earned*, or it was a 3-sample artifact costing
roughly 3x on the largest category of traffic. Both answers are worth $1.36.
Then revisit `quality_tolerance` and `outcome_prior_strength`, with scores that
finally mean something and enough evidence to set them from.
### The report this unlocks
Once `route_decisions.exploration` and `request_id` exist, the confound named
in review §4 becomes addressable: pass rates computed **on explored requests
only** are unconfounded by the routing policy, because assignment was random
within the eligible set. That is the first genuinely causal comparison this
project will have been able to make, and it is the thing that settles whether
the expensive models are worth their price.
---
## 6. File-by-file task list
**Commit 1 — schema**
- `config/schema.sql` — add `outcome_score REAL`, `outcome_samples INTEGER
DEFAULT 0` to `proficiency`; add `request_id TEXT`, `exploration INTEGER
DEFAULT 0` to the `route_decisions` DDL.
- `src/proficiency_store.py::ensure_columns` — the two `proficiency` columns.
No backfill (0 is correct for every existing row).
- `src/dispatcher.py::ensure_route_decisions` + `_ensure_route_decisions_table`
— the two `route_decisions` columns, same idempotent pattern.
- `src/dispatcher.py::persist_route_decision` — write `request_id` on every
decision, and `exploration` (0 for now). `request_id` is already in scope on
the chat path; on `/route` (which spends nothing upstream) it stays NULL.
**Commit 2 — scoring**
- `src/proficiency.py` — add `expected_success_rate(...)` per §2 and
`"outcome_blended"` to the `Source` literal. Keep it pure: peer aggregates
are arguments.
- `src/proficiency_store.py` — split `_write` (benchmark half only, plus the
new columns); add `add_outcome`; add `recompute_category`; call it from
`add_self_eval`, `add_outcome`, `propagate_to_variants`; extend
`propagate_to_variants` to copy `outcome_score` / `outcome_samples`.
- `src/leaderboard.py` — importer calls `recompute_category` after its writes.
- `src/config.py` — `ProficiencyConfig.outcome_prior_strength: int`.
- `config/config.yaml` — the `outcome_prior_strength` block from §3.
- `src/feedback.py` — client outcomes → `add_outcome`; drop `truncated` and
`malformed` from `FAILURE_VERDICTS`; docstring updated to say why.
**Commit 3 — exploration**
- `src/exploration.py` — new, pure, injected RNG, no `config`/`dispatcher`
import.
- `src/config.py` — `ExplorationConfig(StrictModel)` + field on `RouterConfig`.
- `config/config.yaml` — the `exploration:` block from §4.
- `src/dispatcher.py` — add `p.outcome_samples` to `load_candidates`'s SELECT;
call `exploration.choose` after `rank_candidates` and before
`apply_flex_preference`; gate on `cfg.exploration.enabled` and
`task_tier <= max_tier`; set `exploration=1` on the decision row and add
`explore=` to the existing `logs.info` routing line.
**Tests** (`tests/test_proficiency.py`, `tests/test_feedback.py`, new
`tests/test_exploration.py`) — pin the properties, not the arithmetic:
- an unproven model does not outrank a proven one in the same category;
- a category with no outcome data reproduces today's benchmark score exactly
(this is the step-2 acceptance test in miniature);
- a model with neither source still yields `None` → neutral 0.5 downstream;
- `recompute_category` is idempotent and leaves untouched categories alone;
- structural/`local_llm` verdicts no longer move any score;
- exploration never returns a row that fails a hard filter, never exceeds
`max_cost_ratio`, and returns the winner unchanged when `epsilon` is 0;
- with a seeded RNG, the explore share lands within tolerance of `epsilon`.
**Docs** — `docs/data-model.md` (four new columns), `docs/routing.md`
(scoring section: the score is now an expected pass rate, and what
`quality_tolerance` means in those units), `docs/evaluation.md` (benchmark is
now a prior, not the score). `CLAUDE.md`'s "Proficiency: category now changes
routing" and "The only ground truth" sections both need the new story once
step 3 has run and the numbers are real.
**Check while you are in here** — the admin portal's "apply feedback"
operational trigger runs `feedback.py`; confirm it still works after the
retarget, and that the models page shows the new `source` value rather than
blanking on an unrecognized string.
---
## 7. Summary
| | fix | cost | effect |
|---|---|---|---|
| **#1** | benchmark becomes a prior; outcomes are a separate, shrunk, peer-relative signal on the traffic scale | one formula + 2 columns | $179.28 → **$109.35** on replay; unproven models stop winning |
| **#2** | ε=3% exploration to the least-evidenced eligible candidate | **+$1.36 / 9 days (+1.0%)** | 11 empty cells filled, incl. +66 `deepseek`/`tool_use_agentic`; makes the table correctable |
Neither is a large change. Together they convert the router from something
carefully hand-tuned against your traffic into something that learns from it —
which is what the design said it was for.
---
## Appendix — reproduction
Scripts used to produce every number above are read-only and replay against
`router.db` without mutating it. `fix_sim2.py` (empirical Bayes) and
`explore_sim.py` (ε-greedy pricing) are the two that matter; both follow the
same shape as the appendix scripts in
`plans/conceptual-review-premise-and-execution.md` — load `models` +
`proficiency` + `verifications`, rebuild each decision's eligible set with
`routing.select_candidates`, rank with `routing.rank_candidates`, and sum
`estimated_cost`. Constants used: `k = 20`, `ε = 0.03`,
`max_cost_ratio = 4.0`, RNG seed 7, and the shipped
`quality_tolerance = 0.1` / `assumed_cache_rate = 0.917` /
`assumed_completion_tokens = 500`.
### Script D — empirical-Bayes replay (§2)
```python
"""Empirical-Bayes version: shrink toward a benchmark-informed peer prior,
all expressed on the TRAFFIC scale (expected pass rate).
prior_m = peer_rate x (bench_m / peer_bench) # benchmark sets relative position
score_m = (n_m * rate_m + k * prior_m) / (n_m + k)
An unproven model lands at the category's average traffic performance, adjusted
by where the benchmark puts it -- not at the benchmark ceiling.
"""
import sqlite3, sys
sys.path.insert(0, "src")
from config import load_config
from routing import select_candidates, rank_candidates
cfg = load_config("config/config.yaml"); K = 20
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = [dict(r) for r in conn.execute("select * from models")]
bench = {(r["model_id"], r["category"]): r["blended_score"]
for r in conn.execute("select model_id,category,blended_score from proficiency")}
out = {(r["m"], r["c"]): (r["s"], r["n"]) for r in conn.execute(
"""select model_id m, task_category c, sum(verdict='succeeded') s, count(*) n
from verifications where kind='client_outcome' and task_category is not null
group by 1,2""")}
decisions = conn.execute("""select task_category c, task_tier t,
required_context_tokens rc, latency_tolerance lt from route_decisions
where kind='chat' and task_category is not null
and required_context_tokens is not null and task_tier is not null""").fetchall()
conn.close()
def eb_scores(cat):
obs = {m: (s, n) for (m, c), (s, n) in out.items() if c == cat and n > 0}
ids = {m["model_id"] for m in models}
if not obs:
return {m: bench.get((m, cat)) for m in ids}
peer_rate = sum(s for s, n in obs.values()) / sum(n for s, n in obs.values())
bl = [bench[(m, cat)] for m in obs if (m, cat) in bench and bench[(m, cat)] is not None]
peer_bench = sum(bl) / len(bl) if bl else 1.0
sc = {}
for m in ids:
b = bench.get((m, cat))
if b is None:
sc[m] = None; continue
prior = min(1.0, peer_rate * (b / peer_bench)) if peer_bench else peer_rate
s, n = obs.get(m, (0, 0))
sc[m] = (n * (s / n) + K * prior) / (n + K) if n else prior
return sc
def replay(fn, label, show=8):
cache, tot, mix = {}, 0.0, {}
for d in decisions:
sc = cache.setdefault(d["c"], fn(d["c"]))
rows = [{**m, "proficiency": sc.get(m["model_id"])} for m in models]
cand = select_candidates(rows, required_context_tokens=d["rc"], required_tier=d["t"],
latency_tolerance=d["lt"] or "interactive",
allowed_access_levels=cfg.routing.allowed_access_levels,
exclude_stale=cfg.freshness.exclude_stale,
exclude_deprecated=cfg.freshness.exclude_deprecated)
rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
cache_rate=cfg.objective.assumed_cache_rate)
if rk:
tot += rk[0]["cost"] or 0
mix[rk[0]["model_id"]] = mix.get(rk[0]["model_id"], 0) + 1
print(f"\n=== {label} === ${tot:,.2f}")
for m, n in sorted(mix.items(), key=lambda kv: -kv[1])[:show]:
print(f" {m:<26}{n:>6}")
return tot
a = replay(lambda c: {m["model_id"]: bench.get((m["model_id"], c)) for m in models}, "A. today")
d = replay(eb_scores, "D. empirical-Bayes on traffic scale")
print(f"\nA ${a:,.2f} B $179.28 (naive fold-in) D ${d:,.2f} D/A {d/a:.2f}x D/B {d/179.28:.2f}x")
for cat in ("coding_refactor", "coding_general", "tool_use_agentic"):
sc = eb_scores(cat)
print(f"\n{cat} (peer traffic rate anchors the prior)")
print(f" {'model':<24}{'bench':>7}{'n':>5}{'rate':>7}{'score':>8}")
for m in sorted(sc, key=lambda m: -(sc[m] if sc[m] is not None else -1))[:8]:
b = bench.get((m, cat)); s, n = out.get((m, cat), (0, 0))
if b is None: continue
print(f" {m:<24}{b:>7.3f}{n:>5}{(f'{s/n*100:.0f}%' if n else '-'):>7}{sc[m]:>8.3f}")
```
### Script E — epsilon-greedy exploration pricing (§4)
```python
"""Price an epsilon-greedy exploration budget on the real decision stream."""
import sqlite3, sys, random
sys.path.insert(0, "src")
from config import load_config
from routing import select_candidates, rank_candidates, estimated_cost
cfg = load_config("config/config.yaml")
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
models = [dict(r) for r in conn.execute("select * from models")]
bench = {(r["model_id"], r["category"]): r["blended_score"]
for r in conn.execute("select model_id,category,blended_score from proficiency")}
out = {(r["m"], r["c"]): r["n"] for r in conn.execute(
"""select model_id m, task_category c, count(*) n from verifications
where kind='client_outcome' and task_category is not null group by 1,2""")}
ds = conn.execute("""select task_category c, task_tier t, required_context_tokens rc,
latency_tolerance lt from route_decisions where kind='chat'
and task_category is not null and required_context_tokens is not null
and task_tier is not null""").fetchall()
conn.close()
EPS, MAX_RATIO = 0.03, 4.0
rng = random.Random(7)
exploit_cost = explore_cost = 0.0
n_explore = 0; gained = {}
for d in ds:
rows = [{**m, "proficiency": bench.get((m["model_id"], d["c"]))} for m in models]
cand = select_candidates(rows, required_context_tokens=d["rc"], required_tier=d["t"],
latency_tolerance=d["lt"] or "interactive",
allowed_access_levels=cfg.routing.allowed_access_levels,
exclude_stale=cfg.freshness.exclude_stale,
exclude_deprecated=cfg.freshness.exclude_deprecated)
rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
cache_rate=cfg.objective.assumed_cache_rate)
if not rk: continue
win = rk[0]; wc = win["cost"] or 0
exploit_cost += wc
# explore: fewest outcome samples in this category, cost-capped, cheapest tiebreak
if rng.random() < EPS and len(rk) > 1:
pool = [r for r in rk if (r["cost"] or 0) <= MAX_RATIO * max(wc, 1e-9)]
pool = [r for r in pool if r["model_id"] != win["model_id"]]
if pool:
pick = min(pool, key=lambda r: (out.get((r["model_id"], d["c"]), 0), r["cost"] or 0))
explore_cost += pick["cost"] or 0
n_explore += 1
gained[(pick["model_id"], d["c"])] = gained.get((pick["model_id"], d["c"]), 0) + 1
continue
explore_cost += wc
print(f"decisions {len(ds):,} explored {n_explore:,} ({n_explore/len(ds)*100:.1f}%)")
print(f"exploit-only cost ${exploit_cost:,.2f}")
print(f"with exploration ${explore_cost:,.2f} (+${explore_cost-exploit_cost:,.2f}, "
f"{(explore_cost/exploit_cost-1)*100:+.1f}%)")
print(f"\nnew outcome samples this window would have bought (top 12):")
for (m, c), n in sorted(gained.items(), key=lambda kv: -kv[1])[:12]:
have = out.get((m, c), 0)
print(f" {m:<24}{c:<18}+{n:>4} (had {have})")
```