# Admin portal uplift: proficiency, profiles, and gaming mode Status: done -- PRs #26-#36 **Status: FINAL — decision-complete.** Written 2026-09-05 against `main` at `12615d9`, from a live audit of the portal against the config surface and a browser pass over `profiles.html`. Three independent pieces of work, deliberately in one plan because they share one premise — the portal has fallen behind what the router does. Ship them in any order; only §3 has a bug attached, so it goes first if you sequence. ## The audit that produced this The portal is **5 pages** (`index`, `models`, `profiles`, `decisions`, `controls`) and **9 editable config keys**, against **21 config sections**. Seven sections have no portal presence at all: `exploration`, `escalation`, `iteration`, `local_dispatch_models`, `local_vision`, `freshness`, and effectively `classifier`. The user's own bar, stated 2026-09-04: *the portal does not need 100% coverage of every lever, but the main features should each have basic functionality and configurability there.* This plan closes the three gaps that fail that bar hardest. It does **not** try to close all seven. --- ## 1. Proficiency is invisible, and it is the table that decides routing `CLAUDE.md`: *"`proficiency_score` is the ONLY category-dependent term in the ranking."* Measured on the live DB: **148 rows, 17 models × 11 categories**, calibrated against 1,059 client outcomes. The portal's entire surface for it is `metrics.top_proficiency` — a top-N list by `blended_score` for one category, on the dashboard. An operator can see *which* model won a decision and not *why*. ### What is missing that matters `proficiency` carries twelve columns and the portal shows one. The ones an operator needs and cannot get: | column | why it decides something | |---|---| | `source` | 38 `outcome_blended`, 71 `outcome_prior`, 39 `self_eval_thin` on the live DB. A **prior is inherited, not measured** — treating those three as the same number is the mistake this project already made once with `-flex` rows | | `outcome_samples` / `self_eval_samples` | a 1.00 at n=2 and a 0.97 at n=14 are not comparable, and `docs_writing` already proved that at n=2 the ceiling was a sampling artifact | | `inherited_from` | which family a variant borrowed its score from | | `last_updated` | whether a score reflects recent traffic or a months-old sweep | ### Build A `proficiency.html` page: a **model × category matrix** of `blended_score`, each cell carrying source and sample count (tooltip or a compact marker), with the ability to sort/filter by category and by source. - **Distinguish measured from inherited visually.** `outcome_prior` and `self_eval_thin` must not look like `outcome_blended`. This is the single most important requirement on the page — the whole point is telling apart a score that traffic earned from one that was copied in. - Flag cells below `self_eval_min_samples` as thin. - Read-only. **Do NOT add editing.** Proficiency is derived from evaluation and client outcomes; a hand-edited score is a fabricated measurement, the same failure as the empty `leaderboards.yaml` and the provider's `static_fallback` carbon constant this project already excludes. - New endpoint `GET /admin/api/proficiency` returning the rows. `metrics.py` gains the query; it must not import `dispatcher`. ### Also missing, and cheap alongside it Nothing in the portal shows **outcomes**. `POST /outcome` is described in `CLAUDE.md` as *"the only ground truth the router gets"*, and there is no view of what has been reported. Worse, PR #34 added `model_attributable = 0` for degraded-classification outcomes — so some reports are deliberately excluded from folding and **an operator has no way to see that is happening**. Add to the proficiency page (or `decisions.html`, implementer's choice — state which and why): recent `verifications` rows of `kind='client_outcome'`, with verdict and **whether it was attributable**. A silently-excluded outcome is exactly the "recorded but never surfaced" pattern this project has now hit four times. --- ## 2. The profiles page is a dead end, and its zero-admit alarm cries wolf ### 2a. Nothing on the page is editable, and it never says why All five profiles are `source=builtin`. The render logic is: ```js const actions = isBuiltin ? '' : `