Files
6krrt/plans/admin-portal-uplift.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

14 KiB
Raw Permalink Blame History

Admin portal uplift: proficiency, profiles, and gaming mode

Status: done -- PRs #26-#36

Status: FINAL — decision-complete. Written 2026-09-05 against main at 12615d9, from a live audit of the portal against the config surface and a browser pass over profiles.html.

Three independent pieces of work, deliberately in one plan because they share one premise — the portal has fallen behind what the router does. Ship them in any order; only §3 has a bug attached, so it goes first if you sequence.

The audit that produced this

The portal is 5 pages (index, models, profiles, decisions, controls) and 9 editable config keys, against 21 config sections. Seven sections have no portal presence at all: exploration, escalation, iteration, local_dispatch_models, local_vision, freshness, and effectively classifier.

The user's own bar, stated 2026-09-04: the portal does not need 100% coverage of every lever, but the main features should each have basic functionality and configurability there. This plan closes the three gaps that fail that bar hardest. It does not try to close all seven.


1. Proficiency is invisible, and it is the table that decides routing

CLAUDE.md: "proficiency_score is the ONLY category-dependent term in the ranking." Measured on the live DB: 148 rows, 17 models × 11 categories, calibrated against 1,059 client outcomes.

The portal's entire surface for it is metrics.top_proficiency — a top-N list by blended_score for one category, on the dashboard. An operator can see which model won a decision and not why.

What is missing that matters

proficiency carries twelve columns and the portal shows one. The ones an operator needs and cannot get:

column why it decides something
source 38 outcome_blended, 71 outcome_prior, 39 self_eval_thin on the live DB. A prior is inherited, not measured — treating those three as the same number is the mistake this project already made once with -flex rows
outcome_samples / self_eval_samples a 1.00 at n=2 and a 0.97 at n=14 are not comparable, and docs_writing already proved that at n=2 the ceiling was a sampling artifact
inherited_from which family a variant borrowed its score from
last_updated whether a score reflects recent traffic or a months-old sweep

Build

A proficiency.html page: a model × category matrix of blended_score, each cell carrying source and sample count (tooltip or a compact marker), with the ability to sort/filter by category and by source.

  • Distinguish measured from inherited visually. outcome_prior and self_eval_thin must not look like outcome_blended. This is the single most important requirement on the page — the whole point is telling apart a score that traffic earned from one that was copied in.
  • Flag cells below self_eval_min_samples as thin.
  • Read-only. Do NOT add editing. Proficiency is derived from evaluation and client outcomes; a hand-edited score is a fabricated measurement, the same failure as the empty leaderboards.yaml and the provider's static_fallback carbon constant this project already excludes.
  • New endpoint GET /admin/api/proficiency returning the rows. metrics.py gains the query; it must not import dispatcher.

Also missing, and cheap alongside it

Nothing in the portal shows outcomes. POST /outcome is described in CLAUDE.md as "the only ground truth the router gets", and there is no view of what has been reported. Worse, PR #34 added model_attributable = 0 for degraded-classification outcomes — so some reports are deliberately excluded from folding and an operator has no way to see that is happening.

Add to the proficiency page (or decisions.html, implementer's choice — state which and why): recent verifications rows of kind='client_outcome', with verdict and whether it was attributable. A silently-excluded outcome is exactly the "recorded but never surfaced" pattern this project has now hit four times.


2. The profiles page is a dead end, and its zero-admit alarm cries wolf

2a. Nothing on the page is editable, and it never says why

All five profiles are source=builtin. The render logic is:

const actions = isBuiltin ? '' : `<button data-action="edit">…<button data-action="delete">…`

Built-ins get no action buttons at all — by design, settled in plans/admin-profile-management.md. So the page is five cards wearing lock icons. The edit and delete machinery works fine (POST /api/profiles/{name}, DELETE, both already called by the page); there is simply nothing config-defined to point it at.

The obvious operator action — "edit onlycheaps to change the cost bar" — is impossible, and the page offers no alternative.

Add a Duplicate button to built-in cards. It opens the existing create modal pre-filled with that built-in's definition under a new name (onlycheaps-copy or similar), producing an editable config profile.

  • No backend change. POST /api/profiles/ already does everything needed.
  • The name must differ from the built-in: a config profile colliding with a built-in name is rejected at config load, deliberately.
  • Built-ins stay non-editable and non-deletable. This adds a path around that rule, it does not weaken it.

2b. locality reports "admits 0 models" and that is misleading

Measured live:

locality   definition: {"provider": "ollama-local"}   admitted: []
ollama-local rows:  qwen2.5-coder-router:14b  active  tier 1
                    eligible_categories = file_summarization,diff_checking

There is an active local model. _profile_probe (src/admin.py:849) calls select_candidates with task_category=None, and the model is category-gated, so it is excluded. locality genuinely admits 1 model for the two categories it exists to serve, and 0 for a category-less probe.

So the zero-admit badge — added because an empty candidate set is the shape of incident #3 — is firing on an artifact of how the panel asks the question. That is worse than cosmetic: it tells an operator a working profile is broken, which trains people to ignore the badge on the day it is real.

Fix by making the question honest. Probe per category and report coverage ("admits 1 model for 2 of 9 categories"), and reserve the warning badge for a profile that admits zero across all categories. If per-category probing is too expensive to do on every page load, state that and fall back to labelling the existing number "for an unspecified category" — but do not leave a bare admits 0 models warning on a profile that works.

A test must pin that locality does not raise the zero-admit alarm while an active ollama-local row exists.


3. Gaming mode — and the backoff that never fires

3a. The bug, first

_last_classifier_failure (src/dispatcher.py:521) is written and never read. _record_failure() and _record_success_cooldown() set it; _classify_cascade gates only the cloud step, and on the separate _last_account_refusal. The circuit breaker does not break the circuit.

Consequence: nothing stops the router retrying a known-dead local classifier on every request. A stopped Ollama refuses the connection immediately so the cost is small; a hung Ollama, or a VPN-bound one that black-holes, costs up to classifier.timeout_seconds — 120s on this deployment — per request, indefinitely.

Fix independently of the rest of this plan: _classify_cascade (or classify()) must skip the local attempt while time.time() - _last_classifier_failure < cfg.classifier.cooldown_seconds. A test must prove the local client is not constructed during the window, and that a success clears it.

3b. The feature

The user stops Ollama to free the GPU for games. The router should be told, rather than discovering it 120s at a time.

Add local_compute.enabled (default true) and a toggle on controls.html. Every local-model call site reads it:

call site file
classifier dispatcher.py:756 _classifier_client()
health probe dispatcher.py:1820
local verification the verification.local_llm_enabled path
local vision fallback dispatcher.py:~2648
local dispatch models _local_dispatch_config_for
local energy metering cfg.local_energy.enabled gates

One flag the code reads, NOT a macro that writes five keys. A macro is hard to undo cleanly, drifts the moment a sixth site appears, and leaves the operator unable to answer "why isn't the classifier running?" from one place. verification.local_llm_enabled and local_energy.enabled keep their own meanings; local_compute.enabled is an outer gate over all of them.

Skip, do not fail. With the flag off, the classifier must not attempt a call and handle the error — it goes straight to the cascade. That is the difference between "gaming mode" and "Ollama happens to be down", and it is why this is not just a config macro.

3c. Gaming mode REQUIRES a cloud classifier — it does not merely warn

Decided by the user 2026-09-05: gaming mode forces a cloud classifier. Without that, turning it on silently drops every request to stale-session → session-history → static guess, which is a worse router, not a differently-hosted one.

State plainly what is and is not already built, because it is easy to assume this is done:

state
CloudFallbackConfig, _cloud_classifier_client, the call and parsing built (PR #34)
account-refusal gate, per-session cost bounding via the put() gate built
a cloud_fallback block in config/config.yaml commented out — cfg.classifier.cloud_fallback is None, so step 3 returns early today
cloud as a primary classifier not built. It is step 3 of a cascade, reached only after local fails AND stale-session misses AND session-history misses

So this is a promotion, not a wiring job.

The rule:

  • local_compute.enabled = false requires classifier.cloud_fallback to be configured. The toggle refuses to enable without it, naming the config key. A gaming mode that quietly degrades classification is the thing this clause exists to prevent.
  • With the mode on, the local attempt is skipped entirely and the cloud classifier is the classifier — its results carry source="classifier_cloud", which is already in the attributable set, so outcomes still train proficiency normally.

Keep cascade steps 1 and 2 ahead of the cloud call, and say why in the UI. A stale session classification is free and was a real classification of that same session; paying a cloud call to re-derive an answer already held is spending money for nothing. "Force cloud" means cloud replaces the local model, not that it replaces free correct answers. If the operator genuinely wants every request classified fresh in the cloud, that is a separate knob and a separate justification — do not fold it in here.

Do not auto-write a cloud_fallback block from the toggle. Refusing with a clear message is honest; silently configuring a paid classifier on an account measured at $0.014 on 2026-09-05 is not.

metrics.classifier_degradation_warning already reports the degraded share, so the ongoing state stays visible once the mode is on.

Uncommenting the shipped example is the intended setup path. It points at deepseek-v4-flash, which is the cheapest routable model and the right default for a short prompt with a short JSON answer.


Non-goals

  • No editing of proficiency, ever. A hand-edited score is a fabricated measurement.
  • Do not make built-in profiles editable or deletable. §2a adds a path around that rule, not through it.
  • Do not close the other four uncovered config sections (exploration, escalation, iteration, freshness) here. They are a separate decision about how much of the lever surface belongs in a loopback portal.
  • Do not auto-configure cloud_fallback from the gaming-mode toggle.
  • Do not change routing, scoring, or proficiency computation. This plan is observability and operator control only — with the single exception of the §3a backoff fix, which is a defect.

Success criteria

  • A proficiency page shows model × category blended_score with source and sample count, and outcome_prior / self_eval_thin are visually distinguishable from outcome_blended.
  • Client outcomes are visible somewhere, including whether each was attributable — a model_attributable = 0 exclusion must not be silent.
  • GET /admin/api/proficiency exists; metrics.py still does not import dispatcher.
  • A built-in profile card offers Duplicate, which opens the create modal pre-filled; built-ins remain non-editable and non-deletable.
  • locality does not raise the zero-admit alarm while an active ollama-local row exists — pinned by a test.
  • _last_classifier_failure is read: no local classifier call is made inside the cooldown window after a failure, proven by a test asserting the client is not constructed; a success clears the window.
  • local_compute.enabled gates every local call site; with it off, no local client is constructed on any path.
  • Turning gaming mode on is refused when classifier.cloud_fallback is absent, with a message naming the key — not merely warned about.
  • With gaming mode on and a cloud classifier configured, classifications carry source="classifier_cloud" and therefore remain attributable, so proficiency keeps training normally.
  • Cascade steps 1 and 2 still run ahead of the cloud call: a free, real, same-session classification is not worth paying to re-derive.
  • Full suite green with local_energy.enabled both true and false.
  • The user's config/config.local.yaml is byte-identical after the run (local_energy.enabled: true, tariff_usd_per_kwh: 0.159). It is gitignored and NOT in git history — the unrecoverable file. config/config.yaml is clean and tracked, so editing it is an ordinary commit.