plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
14 KiB
Admin portal uplift: proficiency, profiles, and gaming mode
Status: done -- PRs #26-#36
Status: FINAL — decision-complete. Written 2026-09-05 against main at
12615d9, from a live audit of the portal against the config surface and a
browser pass over profiles.html.
Three independent pieces of work, deliberately in one plan because they share one premise — the portal has fallen behind what the router does. Ship them in any order; only §3 has a bug attached, so it goes first if you sequence.
The audit that produced this
The portal is 5 pages (index, models, profiles, decisions,
controls) and 9 editable config keys, against 21 config sections.
Seven sections have no portal presence at all: exploration, escalation,
iteration, local_dispatch_models, local_vision, freshness, and
effectively classifier.
The user's own bar, stated 2026-09-04: the portal does not need 100% coverage of every lever, but the main features should each have basic functionality and configurability there. This plan closes the three gaps that fail that bar hardest. It does not try to close all seven.
1. Proficiency is invisible, and it is the table that decides routing
CLAUDE.md: "proficiency_score is the ONLY category-dependent term in the
ranking." Measured on the live DB: 148 rows, 17 models × 11 categories,
calibrated against 1,059 client outcomes.
The portal's entire surface for it is metrics.top_proficiency — a top-N list
by blended_score for one category, on the dashboard. An operator can see
which model won a decision and not why.
What is missing that matters
proficiency carries twelve columns and the portal shows one. The ones an
operator needs and cannot get:
| column | why it decides something |
|---|---|
source |
38 outcome_blended, 71 outcome_prior, 39 self_eval_thin on the live DB. A prior is inherited, not measured — treating those three as the same number is the mistake this project already made once with -flex rows |
outcome_samples / self_eval_samples |
a 1.00 at n=2 and a 0.97 at n=14 are not comparable, and docs_writing already proved that at n=2 the ceiling was a sampling artifact |
inherited_from |
which family a variant borrowed its score from |
last_updated |
whether a score reflects recent traffic or a months-old sweep |
Build
A proficiency.html page: a model × category matrix of blended_score,
each cell carrying source and sample count (tooltip or a compact marker), with
the ability to sort/filter by category and by source.
- Distinguish measured from inherited visually.
outcome_priorandself_eval_thinmust not look likeoutcome_blended. This is the single most important requirement on the page — the whole point is telling apart a score that traffic earned from one that was copied in. - Flag cells below
self_eval_min_samplesas thin. - Read-only. Do NOT add editing. Proficiency is derived from evaluation
and client outcomes; a hand-edited score is a fabricated measurement, the
same failure as the empty
leaderboards.yamland the provider'sstatic_fallbackcarbon constant this project already excludes. - New endpoint
GET /admin/api/proficiencyreturning the rows.metrics.pygains the query; it must not importdispatcher.
Also missing, and cheap alongside it
Nothing in the portal shows outcomes. POST /outcome is described in
CLAUDE.md as "the only ground truth the router gets", and there is no view
of what has been reported. Worse, PR #34 added model_attributable = 0 for
degraded-classification outcomes — so some reports are deliberately excluded
from folding and an operator has no way to see that is happening.
Add to the proficiency page (or decisions.html, implementer's choice — state
which and why): recent verifications rows of kind='client_outcome', with
verdict and whether it was attributable. A silently-excluded outcome is
exactly the "recorded but never surfaced" pattern this project has now hit
four times.
2. The profiles page is a dead end, and its zero-admit alarm cries wolf
2a. Nothing on the page is editable, and it never says why
All five profiles are source=builtin. The render logic is:
const actions = isBuiltin ? '' : `<button data-action="edit">…<button data-action="delete">…`
Built-ins get no action buttons at all — by design, settled in
plans/admin-profile-management.md. So the page is five cards wearing lock
icons. The edit and delete machinery works fine (POST /api/profiles/{name},
DELETE, both already called by the page); there is simply nothing
config-defined to point it at.
The obvious operator action — "edit onlycheaps to change the cost bar" —
is impossible, and the page offers no alternative.
Add a Duplicate button to built-in cards. It opens the existing create
modal pre-filled with that built-in's definition under a new name
(onlycheaps-copy or similar), producing an editable config profile.
- No backend change.
POST /api/profiles/already does everything needed. - The name must differ from the built-in: a config profile colliding with a built-in name is rejected at config load, deliberately.
- Built-ins stay non-editable and non-deletable. This adds a path around that rule, it does not weaken it.
2b. locality reports "admits 0 models" and that is misleading
Measured live:
locality definition: {"provider": "ollama-local"} admitted: []
ollama-local rows: qwen2.5-coder-router:14b active tier 1
eligible_categories = file_summarization,diff_checking
There is an active local model. _profile_probe (src/admin.py:849)
calls select_candidates with task_category=None, and the model is
category-gated, so it is excluded. locality genuinely admits 1 model for the
two categories it exists to serve, and 0 for a category-less probe.
So the zero-admit badge — added because an empty candidate set is the shape of incident #3 — is firing on an artifact of how the panel asks the question. That is worse than cosmetic: it tells an operator a working profile is broken, which trains people to ignore the badge on the day it is real.
Fix by making the question honest. Probe per category and report coverage
("admits 1 model for 2 of 9 categories"), and reserve the warning badge for a
profile that admits zero across all categories. If per-category probing is
too expensive to do on every page load, state that and fall back to labelling
the existing number "for an unspecified category" — but do not leave a bare
admits 0 models warning on a profile that works.
A test must pin that locality does not raise the zero-admit alarm while an
active ollama-local row exists.
3. Gaming mode — and the backoff that never fires
3a. The bug, first
_last_classifier_failure (src/dispatcher.py:521) is written and never
read. _record_failure() and _record_success_cooldown() set it;
_classify_cascade gates only the cloud step, and on the separate
_last_account_refusal. The circuit breaker does not break the circuit.
Consequence: nothing stops the router retrying a known-dead local classifier
on every request. A stopped Ollama refuses the connection immediately so the
cost is small; a hung Ollama, or a VPN-bound one that black-holes, costs up
to classifier.timeout_seconds — 120s on this deployment — per request,
indefinitely.
Fix independently of the rest of this plan: _classify_cascade (or
classify()) must skip the local attempt while
time.time() - _last_classifier_failure < cfg.classifier.cooldown_seconds.
A test must prove the local client is not constructed during the window, and
that a success clears it.
3b. The feature
The user stops Ollama to free the GPU for games. The router should be told, rather than discovering it 120s at a time.
Add local_compute.enabled (default true) and a toggle on
controls.html. Every local-model call site reads it:
| call site | file |
|---|---|
| classifier | dispatcher.py:756 _classifier_client() |
| health probe | dispatcher.py:1820 |
| local verification | the verification.local_llm_enabled path |
| local vision fallback | dispatcher.py:~2648 |
| local dispatch models | _local_dispatch_config_for |
| local energy metering | cfg.local_energy.enabled gates |
One flag the code reads, NOT a macro that writes five keys. A macro is hard
to undo cleanly, drifts the moment a sixth site appears, and leaves the
operator unable to answer "why isn't the classifier running?" from one place.
verification.local_llm_enabled and local_energy.enabled keep their own
meanings; local_compute.enabled is an outer gate over all of them.
Skip, do not fail. With the flag off, the classifier must not attempt a call and handle the error — it goes straight to the cascade. That is the difference between "gaming mode" and "Ollama happens to be down", and it is why this is not just a config macro.
3c. Gaming mode REQUIRES a cloud classifier — it does not merely warn
Decided by the user 2026-09-05: gaming mode forces a cloud classifier. Without that, turning it on silently drops every request to stale-session → session-history → static guess, which is a worse router, not a differently-hosted one.
State plainly what is and is not already built, because it is easy to assume this is done:
| state | |
|---|---|
CloudFallbackConfig, _cloud_classifier_client, the call and parsing |
built (PR #34) |
account-refusal gate, per-session cost bounding via the put() gate |
built |
a cloud_fallback block in config/config.yaml |
commented out — cfg.classifier.cloud_fallback is None, so step 3 returns early today |
| cloud as a primary classifier | not built. It is step 3 of a cascade, reached only after local fails AND stale-session misses AND session-history misses |
So this is a promotion, not a wiring job.
The rule:
local_compute.enabled = falserequiresclassifier.cloud_fallbackto be configured. The toggle refuses to enable without it, naming the config key. A gaming mode that quietly degrades classification is the thing this clause exists to prevent.- With the mode on, the local attempt is skipped entirely and the cloud
classifier is the classifier — its results carry
source="classifier_cloud", which is already in the attributable set, so outcomes still train proficiency normally.
Keep cascade steps 1 and 2 ahead of the cloud call, and say why in the UI. A stale session classification is free and was a real classification of that same session; paying a cloud call to re-derive an answer already held is spending money for nothing. "Force cloud" means cloud replaces the local model, not that it replaces free correct answers. If the operator genuinely wants every request classified fresh in the cloud, that is a separate knob and a separate justification — do not fold it in here.
Do not auto-write a cloud_fallback block from the toggle. Refusing with
a clear message is honest; silently configuring a paid classifier on an
account measured at $0.014 on 2026-09-05 is not.
metrics.classifier_degradation_warning already reports the degraded share,
so the ongoing state stays visible once the mode is on.
Uncommenting the shipped example is the intended setup path. It points at
deepseek-v4-flash, which is the cheapest routable model and the right
default for a short prompt with a short JSON answer.
Non-goals
- No editing of proficiency, ever. A hand-edited score is a fabricated measurement.
- Do not make built-in profiles editable or deletable. §2a adds a path around that rule, not through it.
- Do not close the other four uncovered config sections (
exploration,escalation,iteration,freshness) here. They are a separate decision about how much of the lever surface belongs in a loopback portal. - Do not auto-configure
cloud_fallbackfrom the gaming-mode toggle. - Do not change routing, scoring, or proficiency computation. This plan is observability and operator control only — with the single exception of the §3a backoff fix, which is a defect.
Success criteria
- A proficiency page shows model × category
blended_scorewith source and sample count, andoutcome_prior/self_eval_thinare visually distinguishable fromoutcome_blended. - Client outcomes are visible somewhere, including whether each was
attributable — a
model_attributable = 0exclusion must not be silent. GET /admin/api/proficiencyexists;metrics.pystill does not importdispatcher.- A built-in profile card offers Duplicate, which opens the create modal pre-filled; built-ins remain non-editable and non-deletable.
localitydoes not raise the zero-admit alarm while an activeollama-localrow exists — pinned by a test._last_classifier_failureis read: no local classifier call is made inside the cooldown window after a failure, proven by a test asserting the client is not constructed; a success clears the window.local_compute.enabledgates every local call site; with it off, no local client is constructed on any path.- Turning gaming mode on is refused when
classifier.cloud_fallbackis absent, with a message naming the key — not merely warned about. - With gaming mode on and a cloud classifier configured, classifications carry
source="classifier_cloud"and therefore remain attributable, so proficiency keeps training normally. - Cascade steps 1 and 2 still run ahead of the cloud call: a free, real, same-session classification is not worth paying to re-derive.
- Full suite green with
local_energy.enabledboth true and false. - The user's
config/config.local.yamlis byte-identical after the run (local_energy.enabled: true,tariff_usd_per_kwh: 0.159). It is gitignored and NOT in git history — the unrecoverable file.config/config.yamlis clean and tracked, so editing it is an ordinary commit.