Files
6krrt/tests/test_admin_proficiency.py
adlee-was-taken e82f758b92 feat(admin): a proficiency page, because the table that decides routing was invisible
proficiency_score is the only category-dependent term in the ranking. The
portal's entire surface for it was metrics.top_proficiency -- a top-N list
for ONE category on the dashboard. An operator could see which model won a
decision and not why.

GET /admin/api/proficiency serves all 148 rows as a model x category matrix,
plus the four columns without which blended_score is not comparable across
rows: source, self_eval_samples / outcome_samples, inherited_from, and
last_updated. `thin` is computed server-side against
proficiency.self_eval_min_samples so the threshold cannot drift into a
hardcoded number in JavaScript. Queries live in metrics.py, which still does
not import dispatcher.

The page's one load-bearing requirement is that a score traffic EARNED must
not look like one that was COPIED IN, and getting that right needed a
correction found by looking at the real data: a -flex row carries its
family's number verbatim, source column included, so rows exist reading
source='outcome_blended' with inherited_from set. Colouring on source alone
painted those green, as measurements. Inheritance now outranks source in
cellClass, and the corner wedge stays on top of it -- two marks for the claim
that matters most. Hue carries provenance and nothing else; value is left to
the digits, because a heatmap of blended_score would put the eye on the
number and hide exactly the distinction the page is for.

Read-only, and there is deliberately no write endpoint: proficiency is
derived from evaluation and client outcomes, so a hand-edited score is a
fabricated measurement -- the same failure as the empty leaderboards.yaml and
the provider's static_fallback carbon constant. A test asserts POST/PUT/
PATCH/DELETE all 405.

Client outcomes ship on the same page rather than on decisions.html, because
the exclusion they need to expose is a proficiency fact: PR #34 records
degraded-classification outcomes with model_attributable = 0 and excludes
them from folding, and nothing showed that was happening. Each row carries
its verdict, whether it was counted or excluded, and whether feedback.py has
folded it yet. "Recorded but never surfaced" is a pattern this project has
now hit four times.

Verified against a COPY of the live DB on a throwaway 8081 instance (8080 is
the live router): 148 rows, 17 models x 11 categories, 30 inherited rows all
correctly drawn as copies, 50 outcomes, no console errors. Two rendering
bugs found and fixed that way -- Tabler already owns `.legend` as a 0.75rem
chart marker, which collapsed the legend to a single swatch under the table;
and long agent shell strings in `detail` need overflow-wrap plus truncation,
since max-width alone does nothing to an unbreakable string.

Also from that pass: a gated profile card read "1 admitted / 0 interactive"
because only the first stat came from the per-category probe. Both do now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-05 21:02:23 -04:00

11 KiB