proficiency_score is the only category-dependent term in the ranking. The portal's entire surface for it was metrics.top_proficiency -- a top-N list for ONE category on the dashboard. An operator could see which model won a decision and not why. GET /admin/api/proficiency serves all 148 rows as a model x category matrix, plus the four columns without which blended_score is not comparable across rows: source, self_eval_samples / outcome_samples, inherited_from, and last_updated. `thin` is computed server-side against proficiency.self_eval_min_samples so the threshold cannot drift into a hardcoded number in JavaScript. Queries live in metrics.py, which still does not import dispatcher. The page's one load-bearing requirement is that a score traffic EARNED must not look like one that was COPIED IN, and getting that right needed a correction found by looking at the real data: a -flex row carries its family's number verbatim, source column included, so rows exist reading source='outcome_blended' with inherited_from set. Colouring on source alone painted those green, as measurements. Inheritance now outranks source in cellClass, and the corner wedge stays on top of it -- two marks for the claim that matters most. Hue carries provenance and nothing else; value is left to the digits, because a heatmap of blended_score would put the eye on the number and hide exactly the distinction the page is for. Read-only, and there is deliberately no write endpoint: proficiency is derived from evaluation and client outcomes, so a hand-edited score is a fabricated measurement -- the same failure as the empty leaderboards.yaml and the provider's static_fallback carbon constant. A test asserts POST/PUT/ PATCH/DELETE all 405. Client outcomes ship on the same page rather than on decisions.html, because the exclusion they need to expose is a proficiency fact: PR #34 records degraded-classification outcomes with model_attributable = 0 and excludes them from folding, and nothing showed that was happening. Each row carries its verdict, whether it was counted or excluded, and whether feedback.py has folded it yet. "Recorded but never surfaced" is a pattern this project has now hit four times. Verified against a COPY of the live DB on a throwaway 8081 instance (8080 is the live router): 148 rows, 17 models x 11 categories, 30 inherited rows all correctly drawn as copies, 50 outcomes, no console errors. Two rendering bugs found and fixed that way -- Tabler already owns `.legend` as a 0.75rem chart marker, which collapsed the legend to a single swatch under the table; and long agent shell strings in `detail` need overflow-wrap plus truncation, since max-width alone does nothing to an unbreakable string. Also from that pass: a gated profile card read "1 admitted / 0 interactive" because only the first stat came from the per-category probe. Both do now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
11 KiB
11 KiB