plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
289 lines
14 KiB
Markdown
289 lines
14 KiB
Markdown
# Admin portal uplift: proficiency, profiles, and gaming mode
|
||
|
||
Status: done -- PRs #26-#36
|
||
|
||
**Status: FINAL — decision-complete.** Written 2026-09-05 against `main` at
|
||
`12615d9`, from a live audit of the portal against the config surface and a
|
||
browser pass over `profiles.html`.
|
||
|
||
Three independent pieces of work, deliberately in one plan because they share
|
||
one premise — the portal has fallen behind what the router does. Ship them in
|
||
any order; only §3 has a bug attached, so it goes first if you sequence.
|
||
|
||
## The audit that produced this
|
||
|
||
The portal is **5 pages** (`index`, `models`, `profiles`, `decisions`,
|
||
`controls`) and **9 editable config keys**, against **21 config sections**.
|
||
Seven sections have no portal presence at all: `exploration`, `escalation`,
|
||
`iteration`, `local_dispatch_models`, `local_vision`, `freshness`, and
|
||
effectively `classifier`.
|
||
|
||
The user's own bar, stated 2026-09-04: *the portal does not need 100% coverage
|
||
of every lever, but the main features should each have basic functionality and
|
||
configurability there.* This plan closes the three gaps that fail that bar
|
||
hardest. It does **not** try to close all seven.
|
||
|
||
---
|
||
|
||
## 1. Proficiency is invisible, and it is the table that decides routing
|
||
|
||
`CLAUDE.md`: *"`proficiency_score` is the ONLY category-dependent term in the
|
||
ranking."* Measured on the live DB: **148 rows, 17 models × 11 categories**,
|
||
calibrated against 1,059 client outcomes.
|
||
|
||
The portal's entire surface for it is `metrics.top_proficiency` — a top-N list
|
||
by `blended_score` for one category, on the dashboard. An operator can see
|
||
*which* model won a decision and not *why*.
|
||
|
||
### What is missing that matters
|
||
|
||
`proficiency` carries twelve columns and the portal shows one. The ones an
|
||
operator needs and cannot get:
|
||
|
||
| column | why it decides something |
|
||
|---|---|
|
||
| `source` | 38 `outcome_blended`, 71 `outcome_prior`, 39 `self_eval_thin` on the live DB. A **prior is inherited, not measured** — treating those three as the same number is the mistake this project already made once with `-flex` rows |
|
||
| `outcome_samples` / `self_eval_samples` | a 1.00 at n=2 and a 0.97 at n=14 are not comparable, and `docs_writing` already proved that at n=2 the ceiling was a sampling artifact |
|
||
| `inherited_from` | which family a variant borrowed its score from |
|
||
| `last_updated` | whether a score reflects recent traffic or a months-old sweep |
|
||
|
||
### Build
|
||
|
||
A `proficiency.html` page: a **model × category matrix** of `blended_score`,
|
||
each cell carrying source and sample count (tooltip or a compact marker), with
|
||
the ability to sort/filter by category and by source.
|
||
|
||
- **Distinguish measured from inherited visually.** `outcome_prior` and
|
||
`self_eval_thin` must not look like `outcome_blended`. This is the single
|
||
most important requirement on the page — the whole point is telling apart a
|
||
score that traffic earned from one that was copied in.
|
||
- Flag cells below `self_eval_min_samples` as thin.
|
||
- Read-only. **Do NOT add editing.** Proficiency is derived from evaluation
|
||
and client outcomes; a hand-edited score is a fabricated measurement, the
|
||
same failure as the empty `leaderboards.yaml` and the provider's
|
||
`static_fallback` carbon constant this project already excludes.
|
||
- New endpoint `GET /admin/api/proficiency` returning the rows. `metrics.py`
|
||
gains the query; it must not import `dispatcher`.
|
||
|
||
### Also missing, and cheap alongside it
|
||
|
||
Nothing in the portal shows **outcomes**. `POST /outcome` is described in
|
||
`CLAUDE.md` as *"the only ground truth the router gets"*, and there is no view
|
||
of what has been reported. Worse, PR #34 added `model_attributable = 0` for
|
||
degraded-classification outcomes — so some reports are deliberately excluded
|
||
from folding and **an operator has no way to see that is happening**.
|
||
|
||
Add to the proficiency page (or `decisions.html`, implementer's choice — state
|
||
which and why): recent `verifications` rows of `kind='client_outcome'`, with
|
||
verdict and **whether it was attributable**. A silently-excluded outcome is
|
||
exactly the "recorded but never surfaced" pattern this project has now hit
|
||
four times.
|
||
|
||
---
|
||
|
||
## 2. The profiles page is a dead end, and its zero-admit alarm cries wolf
|
||
|
||
### 2a. Nothing on the page is editable, and it never says why
|
||
|
||
All five profiles are `source=builtin`. The render logic is:
|
||
|
||
```js
|
||
const actions = isBuiltin ? '' : `<button data-action="edit">…<button data-action="delete">…`
|
||
```
|
||
|
||
Built-ins get **no action buttons at all** — by design, settled in
|
||
`plans/admin-profile-management.md`. So the page is five cards wearing lock
|
||
icons. The edit and delete machinery works fine (`POST /api/profiles/{name}`,
|
||
`DELETE`, both already called by the page); there is simply nothing
|
||
config-defined to point it at.
|
||
|
||
The obvious operator action — *"edit `onlycheaps` to change the cost bar"* —
|
||
is impossible, and the page offers no alternative.
|
||
|
||
**Add a Duplicate button to built-in cards.** It opens the existing create
|
||
modal pre-filled with that built-in's definition under a new name
|
||
(`onlycheaps-copy` or similar), producing an editable config profile.
|
||
|
||
- **No backend change.** `POST /api/profiles/` already does everything needed.
|
||
- The name must differ from the built-in: a config profile colliding with a
|
||
built-in name is rejected at config load, deliberately.
|
||
- Built-ins stay non-editable and non-deletable. This adds a path *around*
|
||
that rule, it does not weaken it.
|
||
|
||
### 2b. `locality` reports "admits 0 models" and that is misleading
|
||
|
||
Measured live:
|
||
|
||
```
|
||
locality definition: {"provider": "ollama-local"} admitted: []
|
||
ollama-local rows: qwen2.5-coder-router:14b active tier 1
|
||
eligible_categories = file_summarization,diff_checking
|
||
```
|
||
|
||
There **is** an active local model. `_profile_probe` (`src/admin.py:849`)
|
||
calls `select_candidates` with `task_category=None`, and the model is
|
||
category-gated, so it is excluded. `locality` genuinely admits 1 model for the
|
||
two categories it exists to serve, and 0 for a category-less probe.
|
||
|
||
So the zero-admit badge — added because an empty candidate set is the shape of
|
||
incident #3 — is firing on an artifact of how the panel asks the question. That
|
||
is worse than cosmetic: it tells an operator a working profile is broken, which
|
||
trains people to ignore the badge on the day it is real.
|
||
|
||
**Fix by making the question honest.** Probe per category and report coverage
|
||
("admits 1 model for 2 of 9 categories"), and reserve the warning badge for a
|
||
profile that admits zero across **all** categories. If per-category probing is
|
||
too expensive to do on every page load, state that and fall back to labelling
|
||
the existing number *"for an unspecified category"* — but do not leave a bare
|
||
`admits 0 models` warning on a profile that works.
|
||
|
||
A test must pin that `locality` does not raise the zero-admit alarm while an
|
||
active `ollama-local` row exists.
|
||
|
||
---
|
||
|
||
## 3. Gaming mode — and the backoff that never fires
|
||
|
||
### 3a. The bug, first
|
||
|
||
`_last_classifier_failure` (`src/dispatcher.py:521`) is **written and never
|
||
read**. `_record_failure()` and `_record_success_cooldown()` set it;
|
||
`_classify_cascade` gates only the *cloud* step, and on the separate
|
||
`_last_account_refusal`. The circuit breaker does not break the circuit.
|
||
|
||
Consequence: nothing stops the router retrying a known-dead local classifier
|
||
on every request. A *stopped* Ollama refuses the connection immediately so the
|
||
cost is small; a **hung** Ollama, or a VPN-bound one that black-holes, costs up
|
||
to `classifier.timeout_seconds` — **120s on this deployment** — per request,
|
||
indefinitely.
|
||
|
||
Fix independently of the rest of this plan: `_classify_cascade` (or
|
||
`classify()`) must skip the local attempt while
|
||
`time.time() - _last_classifier_failure < cfg.classifier.cooldown_seconds`.
|
||
A test must prove the local client is not constructed during the window, and
|
||
that a success clears it.
|
||
|
||
### 3b. The feature
|
||
|
||
The user stops Ollama to free the GPU for games. The router should be told,
|
||
rather than discovering it 120s at a time.
|
||
|
||
**Add `local_compute.enabled` (default true) and a toggle on
|
||
`controls.html`.** Every local-model call site reads it:
|
||
|
||
| call site | file |
|
||
|---|---|
|
||
| classifier | `dispatcher.py:756` `_classifier_client()` |
|
||
| health probe | `dispatcher.py:1820` |
|
||
| local verification | the `verification.local_llm_enabled` path |
|
||
| local vision fallback | `dispatcher.py:~2648` |
|
||
| local dispatch models | `_local_dispatch_config_for` |
|
||
| local energy metering | `cfg.local_energy.enabled` gates |
|
||
|
||
**One flag the code reads, NOT a macro that writes five keys.** A macro is hard
|
||
to undo cleanly, drifts the moment a sixth site appears, and leaves the
|
||
operator unable to answer "why isn't the classifier running?" from one place.
|
||
`verification.local_llm_enabled` and `local_energy.enabled` keep their own
|
||
meanings; `local_compute.enabled` is an outer gate over all of them.
|
||
|
||
**Skip, do not fail.** With the flag off, the classifier must not attempt a
|
||
call and handle the error — it goes straight to the cascade. That is the
|
||
difference between "gaming mode" and "Ollama happens to be down", and it is
|
||
why this is not just a config macro.
|
||
|
||
### 3c. Gaming mode REQUIRES a cloud classifier — it does not merely warn
|
||
|
||
**Decided by the user 2026-09-05: gaming mode forces a cloud classifier.**
|
||
Without that, turning it on silently drops every request to
|
||
stale-session → session-history → static guess, which is a worse router, not
|
||
a differently-hosted one.
|
||
|
||
State plainly what is and is not already built, because it is easy to assume
|
||
this is done:
|
||
|
||
| | state |
|
||
|---|---|
|
||
| `CloudFallbackConfig`, `_cloud_classifier_client`, the call and parsing | **built** (PR #34) |
|
||
| account-refusal gate, per-session cost bounding via the `put()` gate | **built** |
|
||
| a `cloud_fallback` block in `config/config.yaml` | **commented out** — `cfg.classifier.cloud_fallback` is `None`, so step 3 returns early today |
|
||
| cloud as a *primary* classifier | **not built.** It is step 3 of a cascade, reached only after local fails AND stale-session misses AND session-history misses |
|
||
|
||
So this is a promotion, not a wiring job.
|
||
|
||
**The rule:**
|
||
|
||
- `local_compute.enabled = false` **requires** `classifier.cloud_fallback` to
|
||
be configured. The toggle **refuses to enable** without it, naming the
|
||
config key. A gaming mode that quietly degrades classification is the thing
|
||
this clause exists to prevent.
|
||
- With the mode on, the local attempt is skipped entirely and the cloud
|
||
classifier is the classifier — its results carry `source="classifier_cloud"`,
|
||
which is already in the attributable set, so outcomes still train
|
||
proficiency normally.
|
||
|
||
**Keep cascade steps 1 and 2 ahead of the cloud call, and say why in the UI.**
|
||
A stale session classification is free and was a *real* classification of that
|
||
same session; paying a cloud call to re-derive an answer already held is
|
||
spending money for nothing. "Force cloud" means cloud replaces the **local
|
||
model**, not that it replaces free correct answers. If the operator genuinely
|
||
wants every request classified fresh in the cloud, that is a separate knob and
|
||
a separate justification — do not fold it in here.
|
||
|
||
Do **not** auto-write a `cloud_fallback` block from the toggle. Refusing with
|
||
a clear message is honest; silently configuring a paid classifier on an
|
||
account measured at **$0.014** on 2026-09-05 is not.
|
||
|
||
`metrics.classifier_degradation_warning` already reports the degraded share,
|
||
so the ongoing state stays visible once the mode is on.
|
||
|
||
**Uncommenting the shipped example is the intended setup path.** It points at
|
||
`deepseek-v4-flash`, which is the cheapest routable model and the right
|
||
default for a short prompt with a short JSON answer.
|
||
|
||
---
|
||
|
||
## Non-goals
|
||
|
||
- No editing of proficiency, ever. A hand-edited score is a fabricated
|
||
measurement.
|
||
- Do not make built-in profiles editable or deletable. §2a adds a path around
|
||
that rule, not through it.
|
||
- Do not close the other four uncovered config sections (`exploration`,
|
||
`escalation`, `iteration`, `freshness`) here. They are a separate decision
|
||
about how much of the lever surface belongs in a loopback portal.
|
||
- Do not auto-configure `cloud_fallback` from the gaming-mode toggle.
|
||
- Do not change routing, scoring, or proficiency computation. This plan is
|
||
observability and operator control only — with the single exception of the
|
||
§3a backoff fix, which is a defect.
|
||
|
||
## Success criteria
|
||
|
||
- A proficiency page shows model × category `blended_score` with **source and
|
||
sample count**, and `outcome_prior` / `self_eval_thin` are visually
|
||
distinguishable from `outcome_blended`.
|
||
- Client outcomes are visible somewhere, including whether each was
|
||
**attributable** — a `model_attributable = 0` exclusion must not be silent.
|
||
- `GET /admin/api/proficiency` exists; `metrics.py` still does not import
|
||
`dispatcher`.
|
||
- A built-in profile card offers Duplicate, which opens the create modal
|
||
pre-filled; built-ins remain non-editable and non-deletable.
|
||
- `locality` does not raise the zero-admit alarm while an active
|
||
`ollama-local` row exists — pinned by a test.
|
||
- `_last_classifier_failure` is **read**: no local classifier call is made
|
||
inside the cooldown window after a failure, proven by a test asserting the
|
||
client is not constructed; a success clears the window.
|
||
- `local_compute.enabled` gates every local call site; with it off, no local
|
||
client is constructed on any path.
|
||
- Turning gaming mode on is **refused** when `classifier.cloud_fallback` is
|
||
absent, with a message naming the key — not merely warned about.
|
||
- With gaming mode on and a cloud classifier configured, classifications carry
|
||
`source="classifier_cloud"` and therefore remain attributable, so proficiency
|
||
keeps training normally.
|
||
- Cascade steps 1 and 2 still run ahead of the cloud call: a free, real,
|
||
same-session classification is not worth paying to re-derive.
|
||
- Full suite green with `local_energy.enabled` both true and false.
|
||
- The user's `config/config.local.yaml` is byte-identical after the run
|
||
(`local_energy.enabled: true`, `tariff_usd_per_kwh: 0.159`). It is gitignored
|
||
and NOT in git history — the unrecoverable file. `config/config.yaml` is
|
||
clean and tracked, so editing it is an ordinary commit.
|