docs: refresh CLAUDE.md for PRs #26-#34, and spec the admin portal uplift #35

Merged
alee merged 3 commits from docs/refresh-claude-md into main 2026-09-06 02:32:00 +00:00
2 changed files with 392 additions and 4 deletions

110
CLAUDE.md
View File

@@ -221,8 +221,10 @@ rather than from months of history.
- `seed_local_dispatch_energy.py` — standalone reference-shape sweep for `ollama-local` rows; derives per-token USD rates through the user's tariff and OLS on measured GPU draw. [architecture](docs/architecture.md).
- `poller.py` — also seeds/updates `provider='ollama-local'` rows from `config.yaml` each poll so local rows stay current even when NeuralWatt is unreachable.
- `logs.py` — per-request trace id (ContextVar), logfmt, journald priority prefixes; `logs.bind()` survives StreamingResponse generators. [operations](docs/operations.md).
- `metrics.py` / `GET /metrics` — read-only observability; takes `(conn, cfg)`, never imports `dispatcher`. [api](docs/api.md).
- `tui.py` — Textual dashboard over `/metrics` + `/events/decisions`; live feed, category→model panel, detail popup; data layer split into `tui_model.py`. [architecture](docs/architecture.md).
- `metrics.py` / `GET /metrics` — read-only observability; takes `(conn, cfg)`, never imports `dispatcher`. Also carries the three detectors added after the incidents below: capability sub-ceilings, the reactive rejection detector, and the classifier-degradation share. [api](docs/api.md).
- `tui.py` — Textual dashboard over `/metrics` + `/events/decisions`; live feed, category→model panel, detail popup; data layer split into `tui_model.py`. The decision table leads with a `time` column and carries `profile` plus an `E` flag for exploratory picks; the quota panel leads with balance and runway. [architecture](docs/architecture.md).
- `tests/test_tui_schema_drift.py` — the tripwire that keeps the two honest. A new `route_decisions` column must be registered as surfaced or deliberately-not, or the test fails **naming the column**. Five columns had already reached the schema without reaching the dashboard; `ROUTE_DECISIONS_COLUMNS` in `tests/test_route_decisions.py` had itself drifted.
- `tests/test_tui_warnings.py` — the same idea for warnings. Every class `/metrics` can emit must render in `#warnings-panel`, and every emitted warning must be registered — the second failing with the RAW text, because the point is that nobody knew the class existed. **Its fixture is a coupled system**: adding a seed can silence an existing class (a small-context seed once killed the escalation hazard by dragging the p95 down), which is why both directions are asserted.
- `router_cli.py` — one-shot `/route` probe (no spend), raw JSON with `--json`. [api](docs/api.md).
- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — loopback `/admin` portal: dashboard, models overrides, decisions log, controls. [admin-portal](docs/admin-portal.md).
- `tests/` — 1155 tests across 40+ files, offline, verified on Python 3.10 and 3.14. [README](README.md).
@@ -635,6 +637,60 @@ upstream token is requested — the classifier is the latency floor. Four settin
`max_retries=0` (SDK retry → 3x wall-clock), `max_output_tokens: 1024` (bounds reasoning), `max_input_chars: 8000` (doesn't need the document; head+tail clamp), `fallback_tier`/`fallback_category` (mid tier, not 502/503), plus `temperature: 0` (deterministic tier). "Local" means your hardware, not this machine: bind to the VPN address, not `0.0.0.0` (Ollama has no auth).
Full measurements: [docs/local-models.md](docs/local-models.md#the-classifier-is-the-latency-floor).
### A classifier failure no longer collapses to one fixed guess
`fallback_tier`/`fallback_category` above is now the LAST step, not the only
one. When the local classifier fails, `classify()` walks a cascade:
| # | step | cost |
|---|---|---|
| 1 | this session's cached classification, **staleness ignored** | free |
| 2 | this session's history in `route_decisions` | free |
| 3 | `classifier.cloud_fallback`, if configured | money |
| 4 | `fallback_tier` / `fallback_category` | free |
Steps 1 and 2 reuse a *real* classification rather than guessing. Step 2 is
what survives a restart, when the in-memory cache is gone but the session's
decisions are still on disk. Both filter degraded sources, so one outage's
guess cannot propagate through every later turn of a session and end up
looking like a measurement.
**Step 3 is absent by default and the whole feature costs nothing until it is
configured.** Three things bound it once it is:
- `classifier.cooldown_seconds` (30) is a **global** backoff, so a sustained
outage buys one cloud attempt per window across ALL requests, not one per
request;
- an account-level provider refusal suppresses step 3 for the same window —
an out-of-credit account makes that call a guaranteed wasted request;
- the `session_cache.put()` gate admits `classifier_cloud`, bounding a session
to one cloud call per staleness window instead of one per turn.
`session_stale` and `session_history` are deliberately NOT cached: borrowed
answers must not renew a staleness clock they never earned.
**The trap this had to avoid, and it is the reason to read this section.** A
degraded classification is recorded with `task_category = general_chat` —
which is a *fully scored* category (13 proficiency rows) as well as the
configured `fallback_category`. Without a guard, `POST /outcome` reports on
mislabelled outage traffic fold into `proficiency(model, general_chat)` and
drag real scores toward whatever happened to be flowing while the classifier
was down. So `report_outcome` marks outcomes of degraded-source decisions
`model_attributable = 0` — kept as a record, excluded from folding, the same
treatment the tool-call false-failures got. `feedback.py` is untouched; its
existing `AND model_attributable = 1` already does the work.
Attributable: `classifier`, `cached`, `override`, `classifier_cloud`. Not:
`fallback`, `session_stale`, `session_history`. The asymmetry is deliberate —
a degraded source is good enough to route one visibly-flagged request, but a
proficiency score is consulted by every future request, so attribution must
not trust more than routing does. Unknown decision rows **fail open**; an
over-applied exclusion already starved feedback once (`client_capped`).
`/metrics` warns when the degraded share of the last 24h crosses
`classifier.degraded_warn_threshold` over at least `degraded_warn_min`
decisions. A survivable failure is exactly the kind that goes unnoticed for
weeks.
## Local dispatch model
A second local model can now be dispatched directly for specific categories.
@@ -670,8 +726,26 @@ task_category)` provider-agnostic record as a cloud one.
## What's NOT built yet — pick up here
Built: session-directory attribution, the local energy ledger, and local model
dispatch (listed under "What's built and working" above). The three items below
remain open.
dispatch (listed under "What's built and working" above).
**As of 2026-09-05 the plan queue is empty.** Nine PRs (#26-#34) landed the
gitignored config overlay, admin profile CRUD writing to it, the provider
literal cleanup, quota balance/burn/runway, capability-aware ceiling and
rejection warnings, the TUI schema catch-up, and the classifier fallback
cascade. `plans/multi-provider-support.md` is PARKED on provider selection —
Z.ai was the recommendation and is no longer settled; the coupling surface in
it is measured and still valid.
Known follow-ups recorded but not specced, both small:
- The TUI decision table renders the literal `"None"` in the `ctx` cell when
`required_context_tokens` is absent — the same defect the `profile` cell was
written to avoid. See `.omo/notepads/tui-overhaul/issues.md`.
- A NULL `required_context_tokens` raises `TypeError` inside the
demand-ceiling comparison, and the column is nullable. Latent only: zero
such rows exist today, checked on the live DB.
The three items below remain open.
1. **Leaderboard priors are unfilled.** `leaderboards.yaml` ships empty on
purpose — inventing plausible-looking benchmark numbers would put
@@ -812,6 +886,34 @@ T. That is silent on all three tiers today and fires on the outage state
(94,196 vs an observed max of 268,168). See
`plans/catalog-staleness-and-poller-failure-modes.md` §4.4.
**This recurred on 2026-09-04 through a dimension the detector did not model,
and both halves of the fix are now in `metrics.py`.** Admin deprecations took
out `kimi-k3*` — the only vision-capable rows with enough context — so a
242,486-token image request 422'd while every existing check stayed silent,
because the *all-models* tier-1 ceiling was still 782,324. The vision-capable
ceiling had collapsed to 192,500.
- **Predictive:** `capability_ceilings` / `capability_demand_warnings` compute
`vision` and `json_mode` sub-ceilings and compare each against demand
actually observed for requests carrying images / requesting JSON. Two extra
series, not a bucket per capability combination.
- **Reactive:** `rejection_warnings` watches `route_decisions` for rows with
`selected_model IS NULL`. This is the more valuable half and the simpler
one — it catches the *next* dimension nobody predicted, at the cost of
firing after the first failure rather than before.
Two details in the reactive detector are load-bearing and easy to undo by
accident. It groups by `(task_tier, digit-normalized reason)` using the
**structured column**, because normalizing digits alone merges `tier >= 1`
and `tier >= 3` rejections into one group and hides whether the broadest or
the frontier candidate set went empty. And the signal is **novelty OR rate**,
never mere presence: measured on the live DB, routine rejections run ~3/hr
while the 2026-09-04 incident was only n=2 — *below* the noise floor — so no
single count threshold can both catch it and stay quiet. A group absent from
the 24h baseline warns at n≥2; a familiar group warns at the configured
count. Zero rejections warn about nothing: a genuinely impossible request
SHOULD 422.
The service holds a billable API key and has **no auth of its own**. Loopback
bind is the only thing standing between the open internet and your allowance;
add auth before widening `--host`.

View File

@@ -0,0 +1,286 @@
# Admin portal uplift: proficiency, profiles, and gaming mode
**Status: FINAL — decision-complete.** Written 2026-09-05 against `main` at
`12615d9`, from a live audit of the portal against the config surface and a
browser pass over `profiles.html`.
Three independent pieces of work, deliberately in one plan because they share
one premise — the portal has fallen behind what the router does. Ship them in
any order; only §3 has a bug attached, so it goes first if you sequence.
## The audit that produced this
The portal is **5 pages** (`index`, `models`, `profiles`, `decisions`,
`controls`) and **9 editable config keys**, against **21 config sections**.
Seven sections have no portal presence at all: `exploration`, `escalation`,
`iteration`, `local_dispatch_models`, `local_vision`, `freshness`, and
effectively `classifier`.
The user's own bar, stated 2026-09-04: *the portal does not need 100% coverage
of every lever, but the main features should each have basic functionality and
configurability there.* This plan closes the three gaps that fail that bar
hardest. It does **not** try to close all seven.
---
## 1. Proficiency is invisible, and it is the table that decides routing
`CLAUDE.md`: *"`proficiency_score` is the ONLY category-dependent term in the
ranking."* Measured on the live DB: **148 rows, 17 models × 11 categories**,
calibrated against 1,059 client outcomes.
The portal's entire surface for it is `metrics.top_proficiency` — a top-N list
by `blended_score` for one category, on the dashboard. An operator can see
*which* model won a decision and not *why*.
### What is missing that matters
`proficiency` carries twelve columns and the portal shows one. The ones an
operator needs and cannot get:
| column | why it decides something |
|---|---|
| `source` | 38 `outcome_blended`, 71 `outcome_prior`, 39 `self_eval_thin` on the live DB. A **prior is inherited, not measured** — treating those three as the same number is the mistake this project already made once with `-flex` rows |
| `outcome_samples` / `self_eval_samples` | a 1.00 at n=2 and a 0.97 at n=14 are not comparable, and `docs_writing` already proved that at n=2 the ceiling was a sampling artifact |
| `inherited_from` | which family a variant borrowed its score from |
| `last_updated` | whether a score reflects recent traffic or a months-old sweep |
### Build
A `proficiency.html` page: a **model × category matrix** of `blended_score`,
each cell carrying source and sample count (tooltip or a compact marker), with
the ability to sort/filter by category and by source.
- **Distinguish measured from inherited visually.** `outcome_prior` and
`self_eval_thin` must not look like `outcome_blended`. This is the single
most important requirement on the page — the whole point is telling apart a
score that traffic earned from one that was copied in.
- Flag cells below `self_eval_min_samples` as thin.
- Read-only. **Do NOT add editing.** Proficiency is derived from evaluation
and client outcomes; a hand-edited score is a fabricated measurement, the
same failure as the empty `leaderboards.yaml` and the provider's
`static_fallback` carbon constant this project already excludes.
- New endpoint `GET /admin/api/proficiency` returning the rows. `metrics.py`
gains the query; it must not import `dispatcher`.
### Also missing, and cheap alongside it
Nothing in the portal shows **outcomes**. `POST /outcome` is described in
`CLAUDE.md` as *"the only ground truth the router gets"*, and there is no view
of what has been reported. Worse, PR #34 added `model_attributable = 0` for
degraded-classification outcomes — so some reports are deliberately excluded
from folding and **an operator has no way to see that is happening**.
Add to the proficiency page (or `decisions.html`, implementer's choice — state
which and why): recent `verifications` rows of `kind='client_outcome'`, with
verdict and **whether it was attributable**. A silently-excluded outcome is
exactly the "recorded but never surfaced" pattern this project has now hit
four times.
---
## 2. The profiles page is a dead end, and its zero-admit alarm cries wolf
### 2a. Nothing on the page is editable, and it never says why
All five profiles are `source=builtin`. The render logic is:
```js
const actions = isBuiltin ? '' : `<button data-action="edit">…<button data-action="delete">…`
```
Built-ins get **no action buttons at all** — by design, settled in
`plans/admin-profile-management.md`. So the page is five cards wearing lock
icons. The edit and delete machinery works fine (`POST /api/profiles/{name}`,
`DELETE`, both already called by the page); there is simply nothing
config-defined to point it at.
The obvious operator action — *"edit `onlycheaps` to change the cost bar"* —
is impossible, and the page offers no alternative.
**Add a Duplicate button to built-in cards.** It opens the existing create
modal pre-filled with that built-in's definition under a new name
(`onlycheaps-copy` or similar), producing an editable config profile.
- **No backend change.** `POST /api/profiles/` already does everything needed.
- The name must differ from the built-in: a config profile colliding with a
built-in name is rejected at config load, deliberately.
- Built-ins stay non-editable and non-deletable. This adds a path *around*
that rule, it does not weaken it.
### 2b. `locality` reports "admits 0 models" and that is misleading
Measured live:
```
locality definition: {"provider": "ollama-local"} admitted: []
ollama-local rows: qwen2.5-coder-router:14b active tier 1
eligible_categories = file_summarization,diff_checking
```
There **is** an active local model. `_profile_probe` (`src/admin.py:849`)
calls `select_candidates` with `task_category=None`, and the model is
category-gated, so it is excluded. `locality` genuinely admits 1 model for the
two categories it exists to serve, and 0 for a category-less probe.
So the zero-admit badge — added because an empty candidate set is the shape of
incident #3 — is firing on an artifact of how the panel asks the question. That
is worse than cosmetic: it tells an operator a working profile is broken, which
trains people to ignore the badge on the day it is real.
**Fix by making the question honest.** Probe per category and report coverage
("admits 1 model for 2 of 9 categories"), and reserve the warning badge for a
profile that admits zero across **all** categories. If per-category probing is
too expensive to do on every page load, state that and fall back to labelling
the existing number *"for an unspecified category"* — but do not leave a bare
`admits 0 models` warning on a profile that works.
A test must pin that `locality` does not raise the zero-admit alarm while an
active `ollama-local` row exists.
---
## 3. Gaming mode — and the backoff that never fires
### 3a. The bug, first
`_last_classifier_failure` (`src/dispatcher.py:521`) is **written and never
read**. `_record_failure()` and `_record_success_cooldown()` set it;
`_classify_cascade` gates only the *cloud* step, and on the separate
`_last_account_refusal`. The circuit breaker does not break the circuit.
Consequence: nothing stops the router retrying a known-dead local classifier
on every request. A *stopped* Ollama refuses the connection immediately so the
cost is small; a **hung** Ollama, or a VPN-bound one that black-holes, costs up
to `classifier.timeout_seconds` — **120s on this deployment** — per request,
indefinitely.
Fix independently of the rest of this plan: `_classify_cascade` (or
`classify()`) must skip the local attempt while
`time.time() - _last_classifier_failure < cfg.classifier.cooldown_seconds`.
A test must prove the local client is not constructed during the window, and
that a success clears it.
### 3b. The feature
The user stops Ollama to free the GPU for games. The router should be told,
rather than discovering it 120s at a time.
**Add `local_compute.enabled` (default true) and a toggle on
`controls.html`.** Every local-model call site reads it:
| call site | file |
|---|---|
| classifier | `dispatcher.py:756` `_classifier_client()` |
| health probe | `dispatcher.py:1820` |
| local verification | the `verification.local_llm_enabled` path |
| local vision fallback | `dispatcher.py:~2648` |
| local dispatch models | `_local_dispatch_config_for` |
| local energy metering | `cfg.local_energy.enabled` gates |
**One flag the code reads, NOT a macro that writes five keys.** A macro is hard
to undo cleanly, drifts the moment a sixth site appears, and leaves the
operator unable to answer "why isn't the classifier running?" from one place.
`verification.local_llm_enabled` and `local_energy.enabled` keep their own
meanings; `local_compute.enabled` is an outer gate over all of them.
**Skip, do not fail.** With the flag off, the classifier must not attempt a
call and handle the error — it goes straight to the cascade. That is the
difference between "gaming mode" and "Ollama happens to be down", and it is
why this is not just a config macro.
### 3c. Gaming mode REQUIRES a cloud classifier — it does not merely warn
**Decided by the user 2026-09-05: gaming mode forces a cloud classifier.**
Without that, turning it on silently drops every request to
stale-session → session-history → static guess, which is a worse router, not
a differently-hosted one.
State plainly what is and is not already built, because it is easy to assume
this is done:
| | state |
|---|---|
| `CloudFallbackConfig`, `_cloud_classifier_client`, the call and parsing | **built** (PR #34) |
| account-refusal gate, per-session cost bounding via the `put()` gate | **built** |
| a `cloud_fallback` block in `config/config.yaml` | **commented out** — `cfg.classifier.cloud_fallback` is `None`, so step 3 returns early today |
| cloud as a *primary* classifier | **not built.** It is step 3 of a cascade, reached only after local fails AND stale-session misses AND session-history misses |
So this is a promotion, not a wiring job.
**The rule:**
- `local_compute.enabled = false` **requires** `classifier.cloud_fallback` to
be configured. The toggle **refuses to enable** without it, naming the
config key. A gaming mode that quietly degrades classification is the thing
this clause exists to prevent.
- With the mode on, the local attempt is skipped entirely and the cloud
classifier is the classifier — its results carry `source="classifier_cloud"`,
which is already in the attributable set, so outcomes still train
proficiency normally.
**Keep cascade steps 1 and 2 ahead of the cloud call, and say why in the UI.**
A stale session classification is free and was a *real* classification of that
same session; paying a cloud call to re-derive an answer already held is
spending money for nothing. "Force cloud" means cloud replaces the **local
model**, not that it replaces free correct answers. If the operator genuinely
wants every request classified fresh in the cloud, that is a separate knob and
a separate justification — do not fold it in here.
Do **not** auto-write a `cloud_fallback` block from the toggle. Refusing with
a clear message is honest; silently configuring a paid classifier on an
account measured at **$0.014** on 2026-09-05 is not.
`metrics.classifier_degradation_warning` already reports the degraded share,
so the ongoing state stays visible once the mode is on.
**Uncommenting the shipped example is the intended setup path.** It points at
`deepseek-v4-flash`, which is the cheapest routable model and the right
default for a short prompt with a short JSON answer.
---
## Non-goals
- No editing of proficiency, ever. A hand-edited score is a fabricated
measurement.
- Do not make built-in profiles editable or deletable. §2a adds a path around
that rule, not through it.
- Do not close the other four uncovered config sections (`exploration`,
`escalation`, `iteration`, `freshness`) here. They are a separate decision
about how much of the lever surface belongs in a loopback portal.
- Do not auto-configure `cloud_fallback` from the gaming-mode toggle.
- Do not change routing, scoring, or proficiency computation. This plan is
observability and operator control only — with the single exception of the
§3a backoff fix, which is a defect.
## Success criteria
- A proficiency page shows model × category `blended_score` with **source and
sample count**, and `outcome_prior` / `self_eval_thin` are visually
distinguishable from `outcome_blended`.
- Client outcomes are visible somewhere, including whether each was
**attributable** — a `model_attributable = 0` exclusion must not be silent.
- `GET /admin/api/proficiency` exists; `metrics.py` still does not import
`dispatcher`.
- A built-in profile card offers Duplicate, which opens the create modal
pre-filled; built-ins remain non-editable and non-deletable.
- `locality` does not raise the zero-admit alarm while an active
`ollama-local` row exists — pinned by a test.
- `_last_classifier_failure` is **read**: no local classifier call is made
inside the cooldown window after a failure, proven by a test asserting the
client is not constructed; a success clears the window.
- `local_compute.enabled` gates every local call site; with it off, no local
client is constructed on any path.
- Turning gaming mode on is **refused** when `classifier.cloud_fallback` is
absent, with a message naming the key — not merely warned about.
- With gaming mode on and a cloud classifier configured, classifications carry
`source="classifier_cloud"` and therefore remain attributable, so proficiency
keeps training normally.
- Cascade steps 1 and 2 still run ahead of the cloud call: a free, real,
same-session classification is not worth paying to re-derive.
- Full suite green with `local_energy.enabled` both true and false.
- The user's `config/config.local.yaml` is byte-identical after the run
(`local_energy.enabled: true`, `tariff_usd_per_kwh: 0.159`). It is gitignored
and NOT in git history — the unrecoverable file. `config/config.yaml` is
clean and tracked, so editing it is an ordinary commit.