Files
6krrt/plans/admin-portal-uplift.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

289 lines
14 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Admin portal uplift: proficiency, profiles, and gaming mode
Status: done -- PRs #26-#36
**Status: FINAL — decision-complete.** Written 2026-09-05 against `main` at
`12615d9`, from a live audit of the portal against the config surface and a
browser pass over `profiles.html`.
Three independent pieces of work, deliberately in one plan because they share
one premise — the portal has fallen behind what the router does. Ship them in
any order; only §3 has a bug attached, so it goes first if you sequence.
## The audit that produced this
The portal is **5 pages** (`index`, `models`, `profiles`, `decisions`,
`controls`) and **9 editable config keys**, against **21 config sections**.
Seven sections have no portal presence at all: `exploration`, `escalation`,
`iteration`, `local_dispatch_models`, `local_vision`, `freshness`, and
effectively `classifier`.
The user's own bar, stated 2026-09-04: *the portal does not need 100% coverage
of every lever, but the main features should each have basic functionality and
configurability there.* This plan closes the three gaps that fail that bar
hardest. It does **not** try to close all seven.
---
## 1. Proficiency is invisible, and it is the table that decides routing
`CLAUDE.md`: *"`proficiency_score` is the ONLY category-dependent term in the
ranking."* Measured on the live DB: **148 rows, 17 models × 11 categories**,
calibrated against 1,059 client outcomes.
The portal's entire surface for it is `metrics.top_proficiency` — a top-N list
by `blended_score` for one category, on the dashboard. An operator can see
*which* model won a decision and not *why*.
### What is missing that matters
`proficiency` carries twelve columns and the portal shows one. The ones an
operator needs and cannot get:
| column | why it decides something |
|---|---|
| `source` | 38 `outcome_blended`, 71 `outcome_prior`, 39 `self_eval_thin` on the live DB. A **prior is inherited, not measured** — treating those three as the same number is the mistake this project already made once with `-flex` rows |
| `outcome_samples` / `self_eval_samples` | a 1.00 at n=2 and a 0.97 at n=14 are not comparable, and `docs_writing` already proved that at n=2 the ceiling was a sampling artifact |
| `inherited_from` | which family a variant borrowed its score from |
| `last_updated` | whether a score reflects recent traffic or a months-old sweep |
### Build
A `proficiency.html` page: a **model × category matrix** of `blended_score`,
each cell carrying source and sample count (tooltip or a compact marker), with
the ability to sort/filter by category and by source.
- **Distinguish measured from inherited visually.** `outcome_prior` and
`self_eval_thin` must not look like `outcome_blended`. This is the single
most important requirement on the page — the whole point is telling apart a
score that traffic earned from one that was copied in.
- Flag cells below `self_eval_min_samples` as thin.
- Read-only. **Do NOT add editing.** Proficiency is derived from evaluation
and client outcomes; a hand-edited score is a fabricated measurement, the
same failure as the empty `leaderboards.yaml` and the provider's
`static_fallback` carbon constant this project already excludes.
- New endpoint `GET /admin/api/proficiency` returning the rows. `metrics.py`
gains the query; it must not import `dispatcher`.
### Also missing, and cheap alongside it
Nothing in the portal shows **outcomes**. `POST /outcome` is described in
`CLAUDE.md` as *"the only ground truth the router gets"*, and there is no view
of what has been reported. Worse, PR #34 added `model_attributable = 0` for
degraded-classification outcomes — so some reports are deliberately excluded
from folding and **an operator has no way to see that is happening**.
Add to the proficiency page (or `decisions.html`, implementer's choice — state
which and why): recent `verifications` rows of `kind='client_outcome'`, with
verdict and **whether it was attributable**. A silently-excluded outcome is
exactly the "recorded but never surfaced" pattern this project has now hit
four times.
---
## 2. The profiles page is a dead end, and its zero-admit alarm cries wolf
### 2a. Nothing on the page is editable, and it never says why
All five profiles are `source=builtin`. The render logic is:
```js
const actions = isBuiltin ? '' : `<button data-action="edit">…<button data-action="delete">…`
```
Built-ins get **no action buttons at all** — by design, settled in
`plans/admin-profile-management.md`. So the page is five cards wearing lock
icons. The edit and delete machinery works fine (`POST /api/profiles/{name}`,
`DELETE`, both already called by the page); there is simply nothing
config-defined to point it at.
The obvious operator action — *"edit `onlycheaps` to change the cost bar"* —
is impossible, and the page offers no alternative.
**Add a Duplicate button to built-in cards.** It opens the existing create
modal pre-filled with that built-in's definition under a new name
(`onlycheaps-copy` or similar), producing an editable config profile.
- **No backend change.** `POST /api/profiles/` already does everything needed.
- The name must differ from the built-in: a config profile colliding with a
built-in name is rejected at config load, deliberately.
- Built-ins stay non-editable and non-deletable. This adds a path *around*
that rule, it does not weaken it.
### 2b. `locality` reports "admits 0 models" and that is misleading
Measured live:
```
locality definition: {"provider": "ollama-local"} admitted: []
ollama-local rows: qwen2.5-coder-router:14b active tier 1
eligible_categories = file_summarization,diff_checking
```
There **is** an active local model. `_profile_probe` (`src/admin.py:849`)
calls `select_candidates` with `task_category=None`, and the model is
category-gated, so it is excluded. `locality` genuinely admits 1 model for the
two categories it exists to serve, and 0 for a category-less probe.
So the zero-admit badge — added because an empty candidate set is the shape of
incident #3 — is firing on an artifact of how the panel asks the question. That
is worse than cosmetic: it tells an operator a working profile is broken, which
trains people to ignore the badge on the day it is real.
**Fix by making the question honest.** Probe per category and report coverage
("admits 1 model for 2 of 9 categories"), and reserve the warning badge for a
profile that admits zero across **all** categories. If per-category probing is
too expensive to do on every page load, state that and fall back to labelling
the existing number *"for an unspecified category"* — but do not leave a bare
`admits 0 models` warning on a profile that works.
A test must pin that `locality` does not raise the zero-admit alarm while an
active `ollama-local` row exists.
---
## 3. Gaming mode — and the backoff that never fires
### 3a. The bug, first
`_last_classifier_failure` (`src/dispatcher.py:521`) is **written and never
read**. `_record_failure()` and `_record_success_cooldown()` set it;
`_classify_cascade` gates only the *cloud* step, and on the separate
`_last_account_refusal`. The circuit breaker does not break the circuit.
Consequence: nothing stops the router retrying a known-dead local classifier
on every request. A *stopped* Ollama refuses the connection immediately so the
cost is small; a **hung** Ollama, or a VPN-bound one that black-holes, costs up
to `classifier.timeout_seconds` — **120s on this deployment** — per request,
indefinitely.
Fix independently of the rest of this plan: `_classify_cascade` (or
`classify()`) must skip the local attempt while
`time.time() - _last_classifier_failure < cfg.classifier.cooldown_seconds`.
A test must prove the local client is not constructed during the window, and
that a success clears it.
### 3b. The feature
The user stops Ollama to free the GPU for games. The router should be told,
rather than discovering it 120s at a time.
**Add `local_compute.enabled` (default true) and a toggle on
`controls.html`.** Every local-model call site reads it:
| call site | file |
|---|---|
| classifier | `dispatcher.py:756` `_classifier_client()` |
| health probe | `dispatcher.py:1820` |
| local verification | the `verification.local_llm_enabled` path |
| local vision fallback | `dispatcher.py:~2648` |
| local dispatch models | `_local_dispatch_config_for` |
| local energy metering | `cfg.local_energy.enabled` gates |
**One flag the code reads, NOT a macro that writes five keys.** A macro is hard
to undo cleanly, drifts the moment a sixth site appears, and leaves the
operator unable to answer "why isn't the classifier running?" from one place.
`verification.local_llm_enabled` and `local_energy.enabled` keep their own
meanings; `local_compute.enabled` is an outer gate over all of them.
**Skip, do not fail.** With the flag off, the classifier must not attempt a
call and handle the error — it goes straight to the cascade. That is the
difference between "gaming mode" and "Ollama happens to be down", and it is
why this is not just a config macro.
### 3c. Gaming mode REQUIRES a cloud classifier — it does not merely warn
**Decided by the user 2026-09-05: gaming mode forces a cloud classifier.**
Without that, turning it on silently drops every request to
stale-session → session-history → static guess, which is a worse router, not
a differently-hosted one.
State plainly what is and is not already built, because it is easy to assume
this is done:
| | state |
|---|---|
| `CloudFallbackConfig`, `_cloud_classifier_client`, the call and parsing | **built** (PR #34) |
| account-refusal gate, per-session cost bounding via the `put()` gate | **built** |
| a `cloud_fallback` block in `config/config.yaml` | **commented out** — `cfg.classifier.cloud_fallback` is `None`, so step 3 returns early today |
| cloud as a *primary* classifier | **not built.** It is step 3 of a cascade, reached only after local fails AND stale-session misses AND session-history misses |
So this is a promotion, not a wiring job.
**The rule:**
- `local_compute.enabled = false` **requires** `classifier.cloud_fallback` to
be configured. The toggle **refuses to enable** without it, naming the
config key. A gaming mode that quietly degrades classification is the thing
this clause exists to prevent.
- With the mode on, the local attempt is skipped entirely and the cloud
classifier is the classifier — its results carry `source="classifier_cloud"`,
which is already in the attributable set, so outcomes still train
proficiency normally.
**Keep cascade steps 1 and 2 ahead of the cloud call, and say why in the UI.**
A stale session classification is free and was a *real* classification of that
same session; paying a cloud call to re-derive an answer already held is
spending money for nothing. "Force cloud" means cloud replaces the **local
model**, not that it replaces free correct answers. If the operator genuinely
wants every request classified fresh in the cloud, that is a separate knob and
a separate justification — do not fold it in here.
Do **not** auto-write a `cloud_fallback` block from the toggle. Refusing with
a clear message is honest; silently configuring a paid classifier on an
account measured at **$0.014** on 2026-09-05 is not.
`metrics.classifier_degradation_warning` already reports the degraded share,
so the ongoing state stays visible once the mode is on.
**Uncommenting the shipped example is the intended setup path.** It points at
`deepseek-v4-flash`, which is the cheapest routable model and the right
default for a short prompt with a short JSON answer.
---
## Non-goals
- No editing of proficiency, ever. A hand-edited score is a fabricated
measurement.
- Do not make built-in profiles editable or deletable. §2a adds a path around
that rule, not through it.
- Do not close the other four uncovered config sections (`exploration`,
`escalation`, `iteration`, `freshness`) here. They are a separate decision
about how much of the lever surface belongs in a loopback portal.
- Do not auto-configure `cloud_fallback` from the gaming-mode toggle.
- Do not change routing, scoring, or proficiency computation. This plan is
observability and operator control only — with the single exception of the
§3a backoff fix, which is a defect.
## Success criteria
- A proficiency page shows model × category `blended_score` with **source and
sample count**, and `outcome_prior` / `self_eval_thin` are visually
distinguishable from `outcome_blended`.
- Client outcomes are visible somewhere, including whether each was
**attributable** — a `model_attributable = 0` exclusion must not be silent.
- `GET /admin/api/proficiency` exists; `metrics.py` still does not import
`dispatcher`.
- A built-in profile card offers Duplicate, which opens the create modal
pre-filled; built-ins remain non-editable and non-deletable.
- `locality` does not raise the zero-admit alarm while an active
`ollama-local` row exists — pinned by a test.
- `_last_classifier_failure` is **read**: no local classifier call is made
inside the cooldown window after a failure, proven by a test asserting the
client is not constructed; a success clears the window.
- `local_compute.enabled` gates every local call site; with it off, no local
client is constructed on any path.
- Turning gaming mode on is **refused** when `classifier.cloud_fallback` is
absent, with a message naming the key — not merely warned about.
- With gaming mode on and a cloud classifier configured, classifications carry
`source="classifier_cloud"` and therefore remain attributable, so proficiency
keeps training normally.
- Cascade steps 1 and 2 still run ahead of the cloud call: a free, real,
same-session classification is not worth paying to re-derive.
- Full suite green with `local_energy.enabled` both true and false.
- The user's `config/config.local.yaml` is byte-identical after the run
(`local_energy.enabled: true`, `tariff_usd_per_kwh: 0.159`). It is gitignored
and NOT in git history — the unrecoverable file. `config/config.yaml` is
clean and tracked, so editing it is an ordinary commit.