Three UI changes and two plan corrections.
Set as default. routing.default_profile decides what every client that does
not name a profile gets -- including all 13 opencode agents, which send bare
llm-router/auto -- and it was reachable only from a dropdown on Controls,
with the profiles page unable to even show which profile was live. The
profiles list now marks the default, the detail pane badges it, and a button
sets it. It posts to the same allowlisted config endpoint the Controls page
uses, so the validation is the one that already exists rather than a second
rule that can drift. is_default is read from the config store rather than
cfg, because cfg binds at import and would report the pre-restart value at
exactly the moment the operator is looking at it. Delete is disabled on the
current default, saying so before the click instead of after the 422.
Allowlist folds. The two lists ran together in one scroll column with
identical row styling, so the only cue for which list a row belonged to was
whether its button was red or blue -- and the allowed scroller cut a row in
half at the boundary, which read as a rendering fault rather than a divider.
They are now separate collapsible sections, each boxed, each with its count
in the header so a folded one still reports what it holds under the filter.
The catalog starts folded: opening the manager should not dump 425 rows
nobody asked for. The allowed scroller is 7 * 38px so it cuts on a row.
Degraded output plan, second trigger. Measured on the live router while
onlycheaps was default and opencode hammered a free model: 38 of 108 calls
to nemotron-3-nano-omni:free came back MALFORMED EMPTY on HTTP 200 -- 35.2%,
against 0% from three other models over the same window. Nothing was logged
as an upstream failure because nothing failed; the circuit breaker trips on
status >= 400 and cannot see this at all, so the router kept dispatching
with no backoff. verify_response caught every one, and structural verdicts
are diagnostics only, so it detected the degradation 38 times and could do
nothing. That is a stronger case for the plan than the mojibake it was
written for, and it flips the scope decision: encoding faults are
provider-shaped, content faults are model-shaped, so the signature decides
the key.
Exposure-bias plan: marked done, not planned. It was labelled planned in the
status backfill on the strength of its own "FINAL -- ready to implement"
header and a memory note saying "until the fix lands". Both describe when
they were written. The code shipped long ago -- exploration enabled at
epsilon 0.03, outcome_prior_strength 20, FAILURE_VERDICTS ("failed",),
expected_success_rate present, and the taxonomy live in the table
(outcome_prior 264, self_eval_thin 147, outcome_blended 38). What remains is
1,136 unfolded outcomes of 2,221, which is an operator decision about an
irreversible DB mutation, not missing code. Cost a wasted dispatch to Atlas;
the correction is recorded in the doc so it cannot cost another.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
697 lines
32 KiB
Markdown
697 lines
32 KiB
Markdown
# Fixing exposure bias (#1) and self-sealing exclusions (#2)
|
||
|
||
Status: done -- code shipped; only the operator-gated backlog fold remains (see below)
|
||
|
||
> **Correction, 2026-09-08.** This doc was marked `planned` in the status
|
||
> backfill on the strength of its own "ready to implement" framing and a
|
||
> memory note warning not to fold the backlog yet. Both were read as "the
|
||
> code is not written". It is. Verified against the live tree:
|
||
>
|
||
> ```
|
||
> exploration: enabled=True epsilon=0.03 max_cost_ratio=4.0 max_tier=2
|
||
> proficiency.outcome_prior_strength: 20
|
||
> feedback.FAILURE_VERDICTS = ("failed",) # structural verdicts excluded
|
||
> proficiency.expected_success_rate() # present
|
||
> ```
|
||
>
|
||
> And the three-branch taxonomy is live in the table, not just in code --
|
||
> `outcome_prior` 264, `self_eval_thin` 147, `outcome_blended` 38,
|
||
> `self_eval` 2.
|
||
>
|
||
> What actually remains is not code. **1,136 of 2,221 client outcomes are
|
||
> unfolded** (all attributable). Spending them mutates `router.db`
|
||
> irreversibly, which is why the plan sequences it last and why it needs an
|
||
> operator decision rather than an implementer. That is the whole of the
|
||
> outstanding work.
|
||
>
|
||
> The lesson is the one this repo already has a memory for: check the code
|
||
> path before calling something a gap. A plan that says "FINAL -- ready to
|
||
> implement" describes the state when it was written, not now.
|
||
|
||
**Date:** 2026-09-01
|
||
**Follows:** `plans/conceptual-review-premise-and-execution.md` findings #1 and #2
|
||
**Status:** **FINAL — ready to implement.** Validated by replay against the
|
||
9,007 real chat decisions in `router.db`. Nothing here has been implemented;
|
||
`router.db` was not mutated. Both open decisions were resolved 2026-09-01 and
|
||
are now stated as instructions, not options:
|
||
|
||
1. `route_decisions.request_id` ships as part of this work — it is a
|
||
prerequisite for the validation in §5, not a follow-up.
|
||
2. Structural and `local_llm` verdicts **stop feeding proficiency** and remain
|
||
diagnostics only. See §3.
|
||
|
||
Section 6 is the file-by-file task list.
|
||
|
||
---
|
||
|
||
## 0. They are one problem seen from two sides
|
||
|
||
Both findings reduce to the same sentence: **absence of evidence is currently
|
||
treated as evidence of quality.**
|
||
|
||
- On the **ranking** side (#1), a model with no traffic evidence keeps its
|
||
saturated benchmark score and wins the band — `glm-5.2-fast` takes 2,243
|
||
decisions after the naive fold-in *because* it has never been measured on
|
||
real work.
|
||
- On the **selection** side (#2), a model excluded by a 3-sample benchmark score
|
||
never accumulates the evidence that would overturn it — `deepseek-v4-flash`
|
||
won 0 of 2,512 `tool_use_agentic` decisions and therefore has 0 outcome
|
||
samples there, forever.
|
||
|
||
So the fixes are complementary, not alternatives. Fixing the score without
|
||
adding exploration leaves cells that can never be filled. Adding exploration
|
||
without fixing the score feeds good data into a path that misreads it.
|
||
|
||
Ordering matters: **fix the ingestion path first**, then turn on exploration,
|
||
then spend the 961-row backlog.
|
||
|
||
---
|
||
|
||
## 1. What is actually wrong with the score
|
||
|
||
`feedback.py` and `eval_proficiency.py` both write through
|
||
`proficiency_store.add_self_eval` into a single `self_eval_score` running mean,
|
||
equal weight per sample. But they are measuring different quantities:
|
||
|
||
| | benchmark (`evals/tasks.yaml`) | client outcomes (`POST /outcome`) |
|
||
|---|---|---|
|
||
| scale | ~1.00 (saturated; 51% of rows sit at exactly 1.00) | ~0.80 |
|
||
| comparability | controlled — every model gets the same task | confounded — models get the requests routing sent them |
|
||
| discrimination | weak | strong |
|
||
| cost to add | a benchmark run | free |
|
||
|
||
Averaging a saturated absolute score with an uncontrolled success rate produces
|
||
a number whose value depends on the **mixing ratio**, and the mixing ratio is
|
||
set by how much the model has been used. That is the bug, stated exactly.
|
||
|
||
The fix is not to weight the two sources — it is to **stop treating the
|
||
benchmark as a level and start treating it as a prior**, with everything
|
||
expressed on the scale that matters: expected probability that a request
|
||
succeeds.
|
||
|
||
---
|
||
|
||
## 2. The scoring rule
|
||
|
||
Per (model, category), with `k` a prior strength in samples (`k = 20` below):
|
||
|
||
```
|
||
peer_rate = Σ successes / Σ samples # over models with traffic in this category
|
||
peer_bench = mean benchmark score # over that same set, so the ratio is calibrated
|
||
prior_m = min(1.0, peer_rate × bench_m / peer_bench)
|
||
score_m = (n_m × rate_m + k × prior_m) / (n_m + k)
|
||
```
|
||
|
||
Read it as: *the benchmark says where this model sits relative to its peers;
|
||
the peer traffic rate says what that position is worth in practice; a model's
|
||
own traffic pulls the estimate toward its observed rate in proportion to how
|
||
much of it there is.*
|
||
|
||
Properties that matter:
|
||
|
||
- **An unproven model lands at the peer average, not at the ceiling.** That is
|
||
the direct fix for #1. It is still admitted (consistent with the project's
|
||
"absent evidence does not disqualify" rule) — it just does not get to sit
|
||
above every measured model for free.
|
||
- **Scores become interpretable.** `blended_score` now means "expected pass
|
||
rate on your traffic." `quality_tolerance: 0.1` becomes "10 percentage points
|
||
of real success rate," which is a knob you can reason about. Today it bands
|
||
an abstract 0-1 quality index whose units are the benchmark's.
|
||
- **Cold start is unchanged.** A category with no traffic falls back to the
|
||
benchmark exactly as today; a model with neither returns `None` → neutral 0.5
|
||
downstream.
|
||
- **Shrinkage replaces the sample-count threshold.** `self_eval_min_samples`
|
||
gates a step change; `k` is continuous, which is the behaviour that section
|
||
was reaching for.
|
||
|
||
### Measured effect
|
||
|
||
Replaying all 9,007 real chat decisions through
|
||
`select_candidates` + `rank_candidates`, changing only the proficiency values:
|
||
|
||
| policy | est. cost | vs today |
|
||
|---|---|---|
|
||
| **A. today** (benchmark only) | $137.76 | 1.00x |
|
||
| **B. naive fold-in** (`feedback.py` as written) | $179.28 | **1.30x** |
|
||
| **D. proposed** (empirical Bayes, above) | **$109.35** | **0.79x** |
|
||
|
||
The proposed rule is **39% cheaper than the naive fold-in** and 21% cheaper
|
||
than today, and the traffic moves toward the models with the best *measured
|
||
real-world* pass rates rather than away from them:
|
||
|
||
```
|
||
today proposed
|
||
kimi-k2.7-code 2,722 deepseek-v4-flash 3,260
|
||
qwen3.6-35b 2,646 qwen3.6-35b 2,646
|
||
qwen3.6-35b-fast 1,236 kimi-k2.7-code 2,210
|
||
glm-5.3 745 glm-5.2-fast 558
|
||
deepseek-v4-flash 500 qwen3.6-35b-fast 133
|
||
```
|
||
|
||
`coding_refactor` under the rule — note the unproven rows now sit *below* or
|
||
level with the proven ones instead of above them:
|
||
|
||
| model | bench | n | rate | score |
|
||
|---|---|---|---|---|
|
||
| `glm-5.3` | 1.000 | 1 | 100% | 0.795 |
|
||
| `glm-5.2-flex` | 1.000 | 0 | — | 0.785 |
|
||
| `kimi-k3-fast` | 1.000 | 0 | — | 0.785 |
|
||
| `gemma-4-31b` | 0.997 | 0 | — | 0.783 |
|
||
| `deepseek-v4-flash` | 0.826 | 95 | 81% | 0.782 |
|
||
|
||
All inside one `quality_tolerance` band, so cost decides — and deepseek, the
|
||
one with 95 real samples at 81%, is by far the cheapest. That is the outcome
|
||
the review said was missing.
|
||
|
||
### One honest tradeoff
|
||
|
||
With this little traffic evidence, scores compress (0.78-0.93 in most
|
||
categories), so a 0.1 band covers much of the range and **cost decides more
|
||
often than it does today.** That is correct given the evidence — nothing is yet
|
||
*proven* better — but it is a real behavioural change, not a free win. Two
|
||
levers: raise `k` to lean harder on the benchmark while traffic is thin, or
|
||
narrow `quality_tolerance` now that its units are meaningful. Revisit both once
|
||
exploration has been running a few weeks.
|
||
|
||
### Rejected: a "proven ceiling" cap
|
||
|
||
The first formulation tried was: cap any model with `< 30` outcome samples at
|
||
the best *traffic-proven* score in its category. It was simulated and rejected —
|
||
it only bites when the best proven model happens to score low, so `coding_general`
|
||
collapsed to a single flat value while `coding_refactor` was untouched.
|
||
Inconsistent, and it needed a second threshold. The empirical-Bayes form gets
|
||
the same effect from one formula with no special case.
|
||
|
||
---
|
||
|
||
## 3. Implementation — scoring
|
||
|
||
### Schema (both via the existing `ensure_columns` migration pattern)
|
||
|
||
```sql
|
||
ALTER TABLE proficiency ADD COLUMN outcome_score REAL;
|
||
ALTER TABLE proficiency ADD COLUMN outcome_samples INTEGER DEFAULT 0;
|
||
|
||
ALTER TABLE route_decisions ADD COLUMN request_id TEXT; -- see note below
|
||
ALTER TABLE route_decisions ADD COLUMN exploration INTEGER DEFAULT 0;
|
||
```
|
||
|
||
`config/schema.sql` gains the same columns for fresh installs. Mirror
|
||
`proficiency_store.ensure_columns` / `_ensure_route_decisions_table` so a live
|
||
`router.db` migrates on load and on write, idempotently.
|
||
|
||
**`route_decisions.request_id` is a prerequisite, not a nice-to-have.** There is
|
||
currently no exact join from a client outcome back to the decision that produced
|
||
it — the review had to approximate with `session_key` + a 5-second window, which
|
||
is why the tier/outcome table in §5 there is marked indicative. `report_outcome`
|
||
already resolves `request_id` against `energy_observations`; recording it on the
|
||
decision closes the loop and is what makes the validation in §5 below possible.
|
||
|
||
### `proficiency.py` — one new pure function
|
||
|
||
Stays I/O-free; the per-category aggregates are passed in by the caller.
|
||
|
||
```python
|
||
def expected_success_rate(
|
||
benchmark_score: float | None,
|
||
outcome_score: float | None,
|
||
outcome_samples: int,
|
||
*,
|
||
peer_rate: float | None, # None when the category has no traffic yet
|
||
peer_benchmark: float | None,
|
||
prior_strength: int,
|
||
) -> tuple[float | None, Source | None]:
|
||
...
|
||
```
|
||
|
||
Returns `(None, None)` when there is neither benchmark nor outcome data, so the
|
||
neutral-0.5 path downstream is unchanged. Add `"outcome_blended"` to `Source`
|
||
so provenance stays inspectable the way `self_eval_thin` already is.
|
||
|
||
### `proficiency_store.py` — a second writer, and a category recompute
|
||
|
||
- New `add_outcome(conn, cfg, model_id, provider, category, scores)` writing
|
||
`outcome_score` / `outcome_samples` through the existing `accumulate`. Keep
|
||
the 0-1 clamp — the comment about a harness bug pushing a score to 1.50 and
|
||
raising `best` for every candidate applies verbatim here.
|
||
- **`blended_score` becomes a category-level computation**, because `peer_rate`
|
||
and `peer_benchmark` are aggregates over the category. `_write` cannot produce
|
||
the final value from one row any more. Split it:
|
||
- `_write` keeps writing `leaderboard_score` / `self_eval_score` /
|
||
`self_eval_samples` (and now `outcome_*`), and leaves `blended_score` at the
|
||
benchmark value from `blend()` so a single write is never internally
|
||
inconsistent.
|
||
- New `recompute_category(conn, cfg, category)` reads every row in the
|
||
category, re-derives each row's benchmark score by calling the **existing**
|
||
`blend()` on its stored components, computes `peer_rate` / `peer_benchmark`
|
||
from `outcome_score` / `outcome_samples`, then writes the final
|
||
`blended_score` + `source` per row.
|
||
|
||
Deriving the benchmark half from the stored components rather than caching it
|
||
means **no fourth score column and no drift** — the invariant this module
|
||
exists to hold (it is the only writer, so `blended_score` and `source` can
|
||
never disagree with their inputs) survives intact, and
|
||
`dispatcher.load_candidates` stays untouched on the hot path.
|
||
- Every writer calls `recompute_category` at the end of its transaction:
|
||
`add_self_eval`, `add_outcome`, `propagate_to_variants`, and the
|
||
`leaderboard.py` importer.
|
||
- `propagate_to_variants` must copy `outcome_score` / `outcome_samples` too, and
|
||
the `inherited_from` guard applies unchanged.
|
||
- `ensure_columns` gains the two new columns. No backfill is needed or wanted:
|
||
`outcome_samples` defaults to 0, which is exactly true of every existing row.
|
||
(Contrast `inherited_from`, where a NULL default was actively wrong and needed
|
||
`_backfill_inherited`.)
|
||
|
||
### `config.py` / `config/config.yaml` — the prior strength is a knob
|
||
|
||
`ProficiencyConfig` gains `outcome_prior_strength: int` (the `k` in §2), and it
|
||
**must be stated in `config/config.yaml`**, not left to a Pydantic default. This
|
||
project has already been bitten once by a setting that existed only as a default
|
||
and was therefore invisible to anyone tuning it (`classifier.max_input_chars`,
|
||
shipped into the wrong section, accepted and discarded because it happened to
|
||
match the code default). Every knob belongs in the file.
|
||
|
||
```yaml
|
||
proficiency:
|
||
# Prior strength, in samples, for folding real client outcomes into a score.
|
||
# A model's own traffic outweighs the benchmark-derived prior once it has
|
||
# more than this many outcome samples. 20 was chosen because the categories
|
||
# with real evidence carry 25-163 samples, so it lets a well-measured model
|
||
# move while still holding a 5-sample cell near its peers.
|
||
outcome_prior_strength: 20
|
||
```
|
||
|
||
`self_eval_min_samples`, `leaderboard_weight` and `self_eval_weight` are
|
||
unchanged — they still govern `blend()`, which now produces the *benchmark half*
|
||
that feeds the prior rather than the final score.
|
||
|
||
### `feedback.py` — client outcomes only
|
||
|
||
`SUCCESS_VERDICTS` stays. The changes:
|
||
|
||
- Client outcomes route to `add_outcome` instead of `add_self_eval`.
|
||
`applied_at` idempotency is unchanged.
|
||
- **`FAILURE_VERDICTS` drops `truncated` and `malformed`** — structural and
|
||
`local_llm` verdicts stop feeding proficiency entirely and remain diagnostics.
|
||
They are *failure-only* contributors (a passing structural check is
|
||
deliberately not recorded), and you cannot form a rate from failures alone:
|
||
folding them in would bias `outcome_score` downward by exactly however often
|
||
the checker happened to fire, which is a property of the checker, not the
|
||
model. They are 134 rows against 987 client outcomes, and the structural
|
||
checker's one documented encounter with real agent traffic produced ~29 false
|
||
`malformed`s before the `has_tool_calls` fix. `failed` (the client-outcome
|
||
failure verdict) stays.
|
||
- `coverage()` keeps reporting all verdicts, so the diagnostics stay visible
|
||
where they belong. Its existing 80%-unverifiable warning is unaffected.
|
||
|
||
---
|
||
|
||
## 4. Implementation — exploration
|
||
|
||
### The rule
|
||
|
||
On an ε-share of requests, instead of the rank winner, dispatch to the
|
||
**hard-filter-eligible candidate with the fewest outcome samples in this
|
||
category**, tie-broken by lowest cost, and skipped entirely if its estimated
|
||
cost exceeds `max_cost_ratio ×` the winner's.
|
||
|
||
Every hard filter still applies — context window, tier floor, access level,
|
||
latency class, vision, JSON mode. Exploration only ever reorders *within the
|
||
eligible set*, so it can never produce a request the model cannot serve.
|
||
|
||
### Cost, measured
|
||
|
||
Replayed at ε = 0.03, `max_cost_ratio` 4.0, over the same 9,007 decisions:
|
||
|
||
```
|
||
explored 263 of 9,007 (2.9%)
|
||
exploit-only $137.76
|
||
with explore $139.12 (+$1.36, +1.0%)
|
||
```
|
||
|
||
**$1.36 over nine days.** What it buys:
|
||
|
||
| cell | new samples | had |
|
||
|---|---|---|
|
||
| `deepseek-v4-flash` / `tool_use_agentic` | **+66** | **0** |
|
||
| `kimi-k2.7-code-fast` / `coding_refactor` | +29 | 6 |
|
||
| `kimi-k3` / `coding_general` | +23 | 0 |
|
||
| `glm-5.2-fast` / `coding_refactor` | +22 | 0 |
|
||
| `kimi-k3` / `coding_refactor` | +19 | 0 |
|
||
| `qwen3.6-35b-fast` / `coding_general` | +18 | 0 |
|
||
| `qwen3.6-35b` / `docs_writing` | +16 | 0 |
|
||
| …plus 5 more cells currently at zero | | |
|
||
|
||
The "fewest samples first" rule targets empty cells without being told to, and
|
||
66 samples is enough to confirm or kill the 3-task 0.333 that currently bans
|
||
the cheapest capable model from the largest category of traffic.
|
||
|
||
### Module shape
|
||
|
||
New `src/exploration.py`, following `circuit_breaker.py` / `session_cache.py`:
|
||
pure, module-level state only if needed, **injected RNG** the way
|
||
`circuit_breaker` injects time, and it never imports `dispatcher` or `config`.
|
||
|
||
```python
|
||
def choose(
|
||
ranked: Sequence[dict],
|
||
sample_counts: dict[str, int],
|
||
*,
|
||
epsilon: float,
|
||
max_cost_ratio: float,
|
||
rng: random.Random,
|
||
) -> tuple[dict, bool]: # (row, was_exploration)
|
||
```
|
||
|
||
`dispatcher.load_candidates` already LEFT JOINs `proficiency`; add
|
||
`p.outcome_samples` to the SELECT and pass the counts in. The call site is
|
||
immediately after `rank_candidates`, before `apply_flex_preference` — a flex
|
||
swap is a serving-class decision and should apply to whatever was chosen.
|
||
|
||
### Config
|
||
|
||
```yaml
|
||
exploration:
|
||
# Routes a small share of requests to the least-evidenced eligible candidate
|
||
# so proficiency scores can be corrected by evidence rather than frozen by
|
||
# the first 3-sample benchmark that touched them. Measured on 9,007 real
|
||
# decisions: 2.9% of traffic, +$1.36 (+1.0%), and it fills 11 (model,
|
||
# category) cells that currently hold ZERO outcome samples -- including
|
||
# deepseek-v4-flash / tool_use_agentic, which the router cannot otherwise
|
||
# ever measure because its own ranking excludes it.
|
||
enabled: true
|
||
epsilon: 0.03
|
||
# Never explore into something more than this multiple of the winner's cost.
|
||
max_cost_ratio: 4.0
|
||
# Exploration is for gathering evidence, not for gambling on high-stakes
|
||
# work. Tier 3 is excluded; 2,034 of 2,512 tool_use_agentic decisions are
|
||
# tier 2, so this costs almost no coverage.
|
||
max_tier: 2
|
||
```
|
||
|
||
`ExplorationConfig(StrictModel)` in `config.py`, registered on `RouterConfig`,
|
||
mirroring `CircuitBreakerConfig`'s shape (defaults on the model *and* stated in
|
||
the file). Validate `0.0 <= epsilon <= 1.0` and `max_cost_ratio >= 1.0` — an
|
||
epsilon above 1 or a ratio below 1 are both silently self-defeating rather than
|
||
loud.
|
||
|
||
There is deliberately no `min_samples_target` knob. "Fewest outcome samples
|
||
first" needs no threshold: it targets empty cells on its own, and once a cell
|
||
fills, the next-emptiest becomes the target automatically.
|
||
|
||
**Shipping this `enabled: true` breaks the project's usual "new knob ships off"
|
||
convention, and does so deliberately.** Off, it changes nothing and
|
||
`config.yaml`'s own stated experiment ("run with it off, let `POST /outcome`
|
||
report real pass/fail, and compare `tool_use_agentic` proficiency for deepseek
|
||
before and after") stays unrunnable — which is precisely the condition the
|
||
review flagged. The convention exists to stop unproven knobs changing behaviour
|
||
silently; this one has a measured cost (+1.0%), a measured benefit (11 empty
|
||
cells), and a hard cost cap (`max_cost_ratio`). It is the exception that earns
|
||
itself.
|
||
|
||
---
|
||
|
||
## 5. Sequencing
|
||
|
||
Four commits, in this order. Steps 1-2 are inert until step 3 runs, so the
|
||
service can be restarted between any of them.
|
||
|
||
1. **Schema + `request_id`.** Migrations only; no behaviour change. Verify a
|
||
live `router.db` migrates and existing rows are intact.
|
||
2. **Scoring path.** `expected_success_rate`, `add_outcome`,
|
||
`recompute_category`, `feedback.py` retargeted. Still inert — no outcome rows
|
||
have been applied yet, so every `outcome_samples` is 0 and every score
|
||
reproduces today's value. That is the acceptance test for this step: **after
|
||
step 2 and before step 3, replaying the decision stream must still produce
|
||
$137.76.**
|
||
3. **Spend the backlog.** `python -m feedback --dry-run`, then
|
||
`python -m feedback`, on the 961 unapplied rows through the new path.
|
||
Expected: the §2 table — traffic shifts toward `deepseek-v4-flash` and total
|
||
estimated cost falls to ~$109.
|
||
4. **Exploration on.** Then re-run `baseline_report.py` weekly.
|
||
|
||
### Validation, at ~2 weeks
|
||
|
||
The check that matters: `deepseek-v4-flash` / `tool_use_agentic` should hold
|
||
≳60 real outcome samples. Compare its measured rate against the 0.333 benchmark
|
||
score that currently bans it from 2,512 decisions. Either the benchmark was
|
||
right and the exclusion is now *earned*, or it was a 3-sample artifact costing
|
||
roughly 3x on the largest category of traffic. Both answers are worth $1.36.
|
||
|
||
Then revisit `quality_tolerance` and `outcome_prior_strength`, with scores that
|
||
finally mean something and enough evidence to set them from.
|
||
|
||
### The report this unlocks
|
||
|
||
Once `route_decisions.exploration` and `request_id` exist, the confound named
|
||
in review §4 becomes addressable: pass rates computed **on explored requests
|
||
only** are unconfounded by the routing policy, because assignment was random
|
||
within the eligible set. That is the first genuinely causal comparison this
|
||
project will have been able to make, and it is the thing that settles whether
|
||
the expensive models are worth their price.
|
||
|
||
---
|
||
|
||
## 6. File-by-file task list
|
||
|
||
**Commit 1 — schema**
|
||
|
||
- `config/schema.sql` — add `outcome_score REAL`, `outcome_samples INTEGER
|
||
DEFAULT 0` to `proficiency`; add `request_id TEXT`, `exploration INTEGER
|
||
DEFAULT 0` to the `route_decisions` DDL.
|
||
- `src/proficiency_store.py::ensure_columns` — the two `proficiency` columns.
|
||
No backfill (0 is correct for every existing row).
|
||
- `src/dispatcher.py::ensure_route_decisions` + `_ensure_route_decisions_table`
|
||
— the two `route_decisions` columns, same idempotent pattern.
|
||
- `src/dispatcher.py::persist_route_decision` — write `request_id` on every
|
||
decision, and `exploration` (0 for now). `request_id` is already in scope on
|
||
the chat path; on `/route` (which spends nothing upstream) it stays NULL.
|
||
|
||
**Commit 2 — scoring**
|
||
|
||
- `src/proficiency.py` — add `expected_success_rate(...)` per §2 and
|
||
`"outcome_blended"` to the `Source` literal. Keep it pure: peer aggregates
|
||
are arguments.
|
||
- `src/proficiency_store.py` — split `_write` (benchmark half only, plus the
|
||
new columns); add `add_outcome`; add `recompute_category`; call it from
|
||
`add_self_eval`, `add_outcome`, `propagate_to_variants`; extend
|
||
`propagate_to_variants` to copy `outcome_score` / `outcome_samples`.
|
||
- `src/leaderboard.py` — importer calls `recompute_category` after its writes.
|
||
- `src/config.py` — `ProficiencyConfig.outcome_prior_strength: int`.
|
||
- `config/config.yaml` — the `outcome_prior_strength` block from §3.
|
||
- `src/feedback.py` — client outcomes → `add_outcome`; drop `truncated` and
|
||
`malformed` from `FAILURE_VERDICTS`; docstring updated to say why.
|
||
|
||
**Commit 3 — exploration**
|
||
|
||
- `src/exploration.py` — new, pure, injected RNG, no `config`/`dispatcher`
|
||
import.
|
||
- `src/config.py` — `ExplorationConfig(StrictModel)` + field on `RouterConfig`.
|
||
- `config/config.yaml` — the `exploration:` block from §4.
|
||
- `src/dispatcher.py` — add `p.outcome_samples` to `load_candidates`'s SELECT;
|
||
call `exploration.choose` after `rank_candidates` and before
|
||
`apply_flex_preference`; gate on `cfg.exploration.enabled` and
|
||
`task_tier <= max_tier`; set `exploration=1` on the decision row and add
|
||
`explore=` to the existing `logs.info` routing line.
|
||
|
||
**Tests** (`tests/test_proficiency.py`, `tests/test_feedback.py`, new
|
||
`tests/test_exploration.py`) — pin the properties, not the arithmetic:
|
||
|
||
- an unproven model does not outrank a proven one in the same category;
|
||
- a category with no outcome data reproduces today's benchmark score exactly
|
||
(this is the step-2 acceptance test in miniature);
|
||
- a model with neither source still yields `None` → neutral 0.5 downstream;
|
||
- `recompute_category` is idempotent and leaves untouched categories alone;
|
||
- structural/`local_llm` verdicts no longer move any score;
|
||
- exploration never returns a row that fails a hard filter, never exceeds
|
||
`max_cost_ratio`, and returns the winner unchanged when `epsilon` is 0;
|
||
- with a seeded RNG, the explore share lands within tolerance of `epsilon`.
|
||
|
||
**Docs** — `docs/data-model.md` (four new columns), `docs/routing.md`
|
||
(scoring section: the score is now an expected pass rate, and what
|
||
`quality_tolerance` means in those units), `docs/evaluation.md` (benchmark is
|
||
now a prior, not the score). `CLAUDE.md`'s "Proficiency: category now changes
|
||
routing" and "The only ground truth" sections both need the new story once
|
||
step 3 has run and the numbers are real.
|
||
|
||
**Check while you are in here** — the admin portal's "apply feedback"
|
||
operational trigger runs `feedback.py`; confirm it still works after the
|
||
retarget, and that the models page shows the new `source` value rather than
|
||
blanking on an unrecognized string.
|
||
|
||
---
|
||
|
||
## 7. Summary
|
||
|
||
| | fix | cost | effect |
|
||
|---|---|---|---|
|
||
| **#1** | benchmark becomes a prior; outcomes are a separate, shrunk, peer-relative signal on the traffic scale | one formula + 2 columns | $179.28 → **$109.35** on replay; unproven models stop winning |
|
||
| **#2** | ε=3% exploration to the least-evidenced eligible candidate | **+$1.36 / 9 days (+1.0%)** | 11 empty cells filled, incl. +66 `deepseek`/`tool_use_agentic`; makes the table correctable |
|
||
|
||
Neither is a large change. Together they convert the router from something
|
||
carefully hand-tuned against your traffic into something that learns from it —
|
||
which is what the design said it was for.
|
||
|
||
---
|
||
|
||
## Appendix — reproduction
|
||
|
||
Scripts used to produce every number above are read-only and replay against
|
||
`router.db` without mutating it. `fix_sim2.py` (empirical Bayes) and
|
||
`explore_sim.py` (ε-greedy pricing) are the two that matter; both follow the
|
||
same shape as the appendix scripts in
|
||
`plans/conceptual-review-premise-and-execution.md` — load `models` +
|
||
`proficiency` + `verifications`, rebuild each decision's eligible set with
|
||
`routing.select_candidates`, rank with `routing.rank_candidates`, and sum
|
||
`estimated_cost`. Constants used: `k = 20`, `ε = 0.03`,
|
||
`max_cost_ratio = 4.0`, RNG seed 7, and the shipped
|
||
`quality_tolerance = 0.1` / `assumed_cache_rate = 0.917` /
|
||
`assumed_completion_tokens = 500`.
|
||
|
||
### Script D — empirical-Bayes replay (§2)
|
||
|
||
```python
|
||
"""Empirical-Bayes version: shrink toward a benchmark-informed peer prior,
|
||
all expressed on the TRAFFIC scale (expected pass rate).
|
||
|
||
prior_m = peer_rate x (bench_m / peer_bench) # benchmark sets relative position
|
||
score_m = (n_m * rate_m + k * prior_m) / (n_m + k)
|
||
|
||
An unproven model lands at the category's average traffic performance, adjusted
|
||
by where the benchmark puts it -- not at the benchmark ceiling.
|
||
"""
|
||
import sqlite3, sys
|
||
sys.path.insert(0, "src")
|
||
from config import load_config
|
||
from routing import select_candidates, rank_candidates
|
||
|
||
cfg = load_config("config/config.yaml"); K = 20
|
||
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
|
||
models = [dict(r) for r in conn.execute("select * from models")]
|
||
bench = {(r["model_id"], r["category"]): r["blended_score"]
|
||
for r in conn.execute("select model_id,category,blended_score from proficiency")}
|
||
out = {(r["m"], r["c"]): (r["s"], r["n"]) for r in conn.execute(
|
||
"""select model_id m, task_category c, sum(verdict='succeeded') s, count(*) n
|
||
from verifications where kind='client_outcome' and task_category is not null
|
||
group by 1,2""")}
|
||
decisions = conn.execute("""select task_category c, task_tier t,
|
||
required_context_tokens rc, latency_tolerance lt from route_decisions
|
||
where kind='chat' and task_category is not null
|
||
and required_context_tokens is not null and task_tier is not null""").fetchall()
|
||
conn.close()
|
||
|
||
def eb_scores(cat):
|
||
obs = {m: (s, n) for (m, c), (s, n) in out.items() if c == cat and n > 0}
|
||
ids = {m["model_id"] for m in models}
|
||
if not obs:
|
||
return {m: bench.get((m, cat)) for m in ids}
|
||
peer_rate = sum(s for s, n in obs.values()) / sum(n for s, n in obs.values())
|
||
bl = [bench[(m, cat)] for m in obs if (m, cat) in bench and bench[(m, cat)] is not None]
|
||
peer_bench = sum(bl) / len(bl) if bl else 1.0
|
||
sc = {}
|
||
for m in ids:
|
||
b = bench.get((m, cat))
|
||
if b is None:
|
||
sc[m] = None; continue
|
||
prior = min(1.0, peer_rate * (b / peer_bench)) if peer_bench else peer_rate
|
||
s, n = obs.get(m, (0, 0))
|
||
sc[m] = (n * (s / n) + K * prior) / (n + K) if n else prior
|
||
return sc
|
||
|
||
def replay(fn, label, show=8):
|
||
cache, tot, mix = {}, 0.0, {}
|
||
for d in decisions:
|
||
sc = cache.setdefault(d["c"], fn(d["c"]))
|
||
rows = [{**m, "proficiency": sc.get(m["model_id"])} for m in models]
|
||
cand = select_candidates(rows, required_context_tokens=d["rc"], required_tier=d["t"],
|
||
latency_tolerance=d["lt"] or "interactive",
|
||
allowed_access_levels=cfg.routing.allowed_access_levels,
|
||
exclude_stale=cfg.freshness.exclude_stale,
|
||
exclude_deprecated=cfg.freshness.exclude_deprecated)
|
||
rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
|
||
prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
|
||
cache_rate=cfg.objective.assumed_cache_rate)
|
||
if rk:
|
||
tot += rk[0]["cost"] or 0
|
||
mix[rk[0]["model_id"]] = mix.get(rk[0]["model_id"], 0) + 1
|
||
print(f"\n=== {label} === ${tot:,.2f}")
|
||
for m, n in sorted(mix.items(), key=lambda kv: -kv[1])[:show]:
|
||
print(f" {m:<26}{n:>6}")
|
||
return tot
|
||
|
||
a = replay(lambda c: {m["model_id"]: bench.get((m["model_id"], c)) for m in models}, "A. today")
|
||
d = replay(eb_scores, "D. empirical-Bayes on traffic scale")
|
||
print(f"\nA ${a:,.2f} B $179.28 (naive fold-in) D ${d:,.2f} D/A {d/a:.2f}x D/B {d/179.28:.2f}x")
|
||
|
||
for cat in ("coding_refactor", "coding_general", "tool_use_agentic"):
|
||
sc = eb_scores(cat)
|
||
print(f"\n{cat} (peer traffic rate anchors the prior)")
|
||
print(f" {'model':<24}{'bench':>7}{'n':>5}{'rate':>7}{'score':>8}")
|
||
for m in sorted(sc, key=lambda m: -(sc[m] if sc[m] is not None else -1))[:8]:
|
||
b = bench.get((m, cat)); s, n = out.get((m, cat), (0, 0))
|
||
if b is None: continue
|
||
print(f" {m:<24}{b:>7.3f}{n:>5}{(f'{s/n*100:.0f}%' if n else '-'):>7}{sc[m]:>8.3f}")
|
||
```
|
||
|
||
### Script E — epsilon-greedy exploration pricing (§4)
|
||
|
||
```python
|
||
"""Price an epsilon-greedy exploration budget on the real decision stream."""
|
||
import sqlite3, sys, random
|
||
sys.path.insert(0, "src")
|
||
from config import load_config
|
||
from routing import select_candidates, rank_candidates, estimated_cost
|
||
|
||
cfg = load_config("config/config.yaml")
|
||
conn = sqlite3.connect("router.db"); conn.row_factory = sqlite3.Row
|
||
models = [dict(r) for r in conn.execute("select * from models")]
|
||
bench = {(r["model_id"], r["category"]): r["blended_score"]
|
||
for r in conn.execute("select model_id,category,blended_score from proficiency")}
|
||
out = {(r["m"], r["c"]): r["n"] for r in conn.execute(
|
||
"""select model_id m, task_category c, count(*) n from verifications
|
||
where kind='client_outcome' and task_category is not null group by 1,2""")}
|
||
ds = conn.execute("""select task_category c, task_tier t, required_context_tokens rc,
|
||
latency_tolerance lt from route_decisions where kind='chat'
|
||
and task_category is not null and required_context_tokens is not null
|
||
and task_tier is not null""").fetchall()
|
||
conn.close()
|
||
|
||
EPS, MAX_RATIO = 0.03, 4.0
|
||
rng = random.Random(7)
|
||
exploit_cost = explore_cost = 0.0
|
||
n_explore = 0; gained = {}
|
||
for d in ds:
|
||
rows = [{**m, "proficiency": bench.get((m["model_id"], d["c"]))} for m in models]
|
||
cand = select_candidates(rows, required_context_tokens=d["rc"], required_tier=d["t"],
|
||
latency_tolerance=d["lt"] or "interactive",
|
||
allowed_access_levels=cfg.routing.allowed_access_levels,
|
||
exclude_stale=cfg.freshness.exclude_stale,
|
||
exclude_deprecated=cfg.freshness.exclude_deprecated)
|
||
rk = rank_candidates(cand, quality_tolerance=cfg.objective.quality_tolerance,
|
||
prompt_tokens=d["rc"], completion_tokens=cfg.objective.assumed_completion_tokens,
|
||
cache_rate=cfg.objective.assumed_cache_rate)
|
||
if not rk: continue
|
||
win = rk[0]; wc = win["cost"] or 0
|
||
exploit_cost += wc
|
||
# explore: fewest outcome samples in this category, cost-capped, cheapest tiebreak
|
||
if rng.random() < EPS and len(rk) > 1:
|
||
pool = [r for r in rk if (r["cost"] or 0) <= MAX_RATIO * max(wc, 1e-9)]
|
||
pool = [r for r in pool if r["model_id"] != win["model_id"]]
|
||
if pool:
|
||
pick = min(pool, key=lambda r: (out.get((r["model_id"], d["c"]), 0), r["cost"] or 0))
|
||
explore_cost += pick["cost"] or 0
|
||
n_explore += 1
|
||
gained[(pick["model_id"], d["c"])] = gained.get((pick["model_id"], d["c"]), 0) + 1
|
||
continue
|
||
explore_cost += wc
|
||
|
||
print(f"decisions {len(ds):,} explored {n_explore:,} ({n_explore/len(ds)*100:.1f}%)")
|
||
print(f"exploit-only cost ${exploit_cost:,.2f}")
|
||
print(f"with exploration ${explore_cost:,.2f} (+${explore_cost-exploit_cost:,.2f}, "
|
||
f"{(explore_cost/exploit_cost-1)*100:+.1f}%)")
|
||
print(f"\nnew outcome samples this window would have bought (top 12):")
|
||
for (m, c), n in sorted(gained.items(), key=lambda kv: -kv[1])[:12]:
|
||
have = out.get((m, c), 0)
|
||
print(f" {m:<24}{c:<18}+{n:>4} (had {have})")
|
||
```
|