fix(proficiency): exposure-bias fix + backlog fold — empirical-Bayes scoring, pooled peer_rate, exploration live #62

Merged
alee merged 1 commits from exploration-bias-fix into main 2026-09-09 02:00:44 +00:00
Owner

Landing the FINAL plan plans/proficiency-exposure-bias-and-exploration.md (exposure bias #1 + self-sealing exclusions #2). Reconciliation confirmed most of §6 had already shipped independently; this branch adds the one un-landed code fix and executes the irreversible backlog spend.

This branch

  • fix(proficiency): peer_rate is the sample-weighted pooled rate per plan §2
    • recompute_category was taking the unweighted mean of per-model outcome rates, letting a 1-sample 100% model dominate the category prior as much as a 95-sample 81% model. Now Σ(n·rate)/Σn over the trafficked set (the §2 contract formula).
    • Measured live: −2.2% est. routing cost vs the unweighted form; debugging prior un-distorts by 0.18.
    • New discriminator test; full suite 1806 pass.

Already landed on main (reconciled, not re-implemented)

  • Schema: proficiency.outcome_score/outcome_samples, route_decisions.request_id/exploration (both live-DB-migrated).
  • Scoring path: proficiency.expected_success_rate + "outcome_prior"/"outcome_blended" sources, proficiency_store.proficiency_outcome.add_outcome/recompute_category, leaderboard recompute hook.
  • feedback.py: client outcomes → add_outcome; FAILURE_VERDICTS = ("failed",) (structural/local_llm are diagnostics-only).
  • Exploration: exploration.choose wired after rank, before flex; exploration: block enabled: true in config.yaml.
  • Docs: data-model.md, routing.md, admin-portal.md, new admin/frontend/proficiency.html (already renders outcome_blended/outcome_prior).
  • Explorer verified LIVE on the DB: 142 explored decisions, concentrated on the zero-evidence cells the plan targeted.

Executed (DB mutations, backed up first)

  • router.db backed up to /tmp/opencode/router-pre-fold-20260908.db (33MB, integrity ok) before the fold.
  • python -m feedback folded 1,139 client-outcome samples across 31 (model, category) cells through the new add_outcome path. 0 rows unapplied post-fold.
  • Scores now show the empirical-Bayes shape: heavy-traffic cells pulled toward measured rates (e.g. kimi-k2.7-code/coding_general 0.746 over 460 samples; deepseek-v4-flash/debugging 0.701 over 42).
  • Note: the plan's absolute §2 targets ($137.76 / $109.35) can't be reproduced verbatim — the DB has grown from 9,007 to 23,927 eligible chat decisions and the OpenRouter allowlist catalog (z-ai/glm-5.3-flash, nemotron...:free) is now ingested, so only the behavior (not the absolutes) transfers.

Not done (hard constraint respected)

  • feedback.py was NOT run against the backlog until the scoring fix was implemented AND tested — the exact failure the plan exists to prevent.
  • config/config.local.yaml untouched (gitignored, unrecoverable).

Follow-up for the operator

  • Restart the service (systemctl --user restart llm-router.service, non-/route traffic) to pick up the pooled fix; the fold data is already in place and routing picks it up on the next recompute.
  • Re-run baseline_report.py weekly; watch deepseek-v4-flash/tool_use_agentic (explorer now feeds it).
Landing the FINAL plan `plans/proficiency-exposure-bias-and-exploration.md` (exposure bias #1 + self-sealing exclusions #2). Reconciliation confirmed most of §6 had already shipped independently; this branch adds the one un-landed code fix and executes the irreversible backlog spend. ## This branch - **`fix(proficiency): peer_rate is the sample-weighted pooled rate per plan §2`** - `recompute_category` was taking the unweighted mean of per-model outcome rates, letting a 1-sample 100% model dominate the category prior as much as a 95-sample 81% model. Now `Σ(n·rate)/Σn` over the trafficked set (the §2 contract formula). - Measured live: −2.2% est. routing cost vs the unweighted form; `debugging` prior un-distorts by 0.18. - New discriminator test; full suite 1806 pass. ## Already landed on main (reconciled, not re-implemented) - Schema: `proficiency.outcome_score/outcome_samples`, `route_decisions.request_id/exploration` (both live-DB-migrated). - Scoring path: `proficiency.expected_success_rate` + `"outcome_prior"/"outcome_blended"` sources, `proficiency_store.proficiency_outcome.add_outcome/recompute_category`, `leaderboard` recompute hook. - `feedback.py`: client outcomes → `add_outcome`; `FAILURE_VERDICTS = ("failed",)` (structural/local_llm are diagnostics-only). - Exploration: `exploration.choose` wired after rank, before flex; `exploration:` block **enabled: true** in config.yaml. - Docs: `data-model.md`, `routing.md`, `admin-portal.md`, new `admin/frontend/proficiency.html` (already renders `outcome_blended`/`outcome_prior`). - Explorer verified LIVE on the DB: 142 explored decisions, concentrated on the zero-evidence cells the plan targeted. ## Executed (DB mutations, backed up first) - `router.db` backed up to `/tmp/opencode/router-pre-fold-20260908.db` (33MB, integrity ok) before the fold. - `python -m feedback` folded 1,139 client-outcome samples across 31 (model, category) cells through the new `add_outcome` path. 0 rows unapplied post-fold. - Scores now show the empirical-Bayes shape: heavy-traffic cells pulled toward measured rates (e.g. `kimi-k2.7-code`/`coding_general` 0.746 over 460 samples; `deepseek-v4-flash`/`debugging` 0.701 over 42). - Note: the plan's absolute §2 targets ($137.76 / $109.35) can't be reproduced verbatim — the DB has grown from 9,007 to 23,927 eligible chat decisions and the OpenRouter allowlist catalog (`z-ai/glm-5.3-flash`, `nemotron...:free`) is now ingested, so only the behavior (not the absolutes) transfers. ## Not done (hard constraint respected) - `feedback.py` was NOT run against the backlog until the scoring fix was implemented AND tested — the exact failure the plan exists to prevent. - `config/config.local.yaml` untouched (gitignored, unrecoverable). ## Follow-up for the operator - Restart the service (`systemctl --user restart llm-router.service`, non-`/route` traffic) to pick up the pooled fix; the fold data is already in place and routing picks it up on the next recompute. - Re-run `baseline_report.py` weekly; watch `deepseek-v4-flash`/`tool_use_agentic` (explorer now feeds it).
alee added 1 commit 2026-09-09 01:25:51 +00:00
recompute_category took the unweighted mean of per-model outcome rates,
letting a 1-sample 100% model dominate the category prior as much as a
95-sample 81% one — the exact small-sample noise the empirical-Bayes fix
exists to suppress. Now Σ(n·rate)/Σn over the trafficked set, as
plans/proficiency-exposure-bias-and-exploration.md §2 specifies. Measured
on the live replay: −2.2% est. cost vs the unweighted form, and the
debugging prior un-distorts by 0.18.

Lands before the backlog fold so the whole table re-derives contract-correct
on the fold's recompute. Retroactive to the 38 rows already folded 2026-09-07
(recompute_category re-derives from stored components; no migration needed).
alee merged commit 362a06dd27 into main 2026-09-09 02:00:44 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#62