Files
6krrt/tests/test_proficiency.py
adlee-was-taken 9c7119ad4f fix(proficiency): peer_rate is the sample-weighted pooled rate per plan §2
recompute_category took the unweighted mean of per-model outcome rates,
letting a 1-sample 100% model dominate the category prior as much as a
95-sample 81% one — the exact small-sample noise the empirical-Bayes fix
exists to suppress. Now Σ(n·rate)/Σn over the trafficked set, as
plans/proficiency-exposure-bias-and-exploration.md §2 specifies. Measured
on the live replay: −2.2% est. cost vs the unweighted form, and the
debugging prior un-distorts by 0.18.

Lands before the backlog fold so the whole table re-derives contract-correct
on the fold's recompute. Retroactive to the 38 rows already folded 2026-09-07
(recompute_category re-derives from stored components; no migration needed).
2026-09-08 21:24:58 -04:00

34 KiB