From 3737bd4898d16f59791231f0c8bf69f5c43678a8 Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Sat, 5 Sep 2026 17:32:36 -0400 Subject: [PATCH 1/3] docs: bring CLAUDE.md up to date with PRs #26-#34 CLAUDE.md declares itself the file to trust on what is currently true, and no commit touched it during the nine merges. Four shipped features were absent entirely. Adds: the classifier fallback cascade and, more importantly, the trap it had to avoid -- a degraded classification lands on general_chat, which is a fully scored category AND the fallback_category, so unguarded outcome folding would drag real proficiency toward outage traffic. Records the attributable/not-attributable split and the fail-open rule. Adds the capability sub-ceiling and reactive rejection detectors to the incident-3 section, since that incident recurred on 2026-09-04 through a dimension the original detector did not model. Keeps the two details that are easy to undo by accident: grouping on the structured task_tier column rather than a normalized string, and the novelty-OR-rate signal, which exists because routine rejections run ~3/hr while the incident was n=2 -- below the noise floor, so no single threshold works. Notes the two drift tripwires and that the warning fixture is a coupled system where a new seed can silence an existing class. Marks the plan queue empty and records the two known follow-ups, including that the NULL demand crash is latent only -- zero such rows on the live DB, checked rather than assumed. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U --- CLAUDE.md | 110 ++++++++++++++++++++++++++++++++++++++++++++++++++++-- 1 file changed, 106 insertions(+), 4 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index b428b2b..4aca958 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -221,8 +221,10 @@ rather than from months of history. - `seed_local_dispatch_energy.py` — standalone reference-shape sweep for `ollama-local` rows; derives per-token USD rates through the user's tariff and OLS on measured GPU draw. [architecture](docs/architecture.md). - `poller.py` — also seeds/updates `provider='ollama-local'` rows from `config.yaml` each poll so local rows stay current even when NeuralWatt is unreachable. - `logs.py` — per-request trace id (ContextVar), logfmt, journald priority prefixes; `logs.bind()` survives StreamingResponse generators. [operations](docs/operations.md). -- `metrics.py` / `GET /metrics` — read-only observability; takes `(conn, cfg)`, never imports `dispatcher`. [api](docs/api.md). -- `tui.py` — Textual dashboard over `/metrics` + `/events/decisions`; live feed, category→model panel, detail popup; data layer split into `tui_model.py`. [architecture](docs/architecture.md). +- `metrics.py` / `GET /metrics` — read-only observability; takes `(conn, cfg)`, never imports `dispatcher`. Also carries the three detectors added after the incidents below: capability sub-ceilings, the reactive rejection detector, and the classifier-degradation share. [api](docs/api.md). +- `tui.py` — Textual dashboard over `/metrics` + `/events/decisions`; live feed, category→model panel, detail popup; data layer split into `tui_model.py`. The decision table leads with a `time` column and carries `profile` plus an `E` flag for exploratory picks; the quota panel leads with balance and runway. [architecture](docs/architecture.md). +- `tests/test_tui_schema_drift.py` — the tripwire that keeps the two honest. A new `route_decisions` column must be registered as surfaced or deliberately-not, or the test fails **naming the column**. Five columns had already reached the schema without reaching the dashboard; `ROUTE_DECISIONS_COLUMNS` in `tests/test_route_decisions.py` had itself drifted. +- `tests/test_tui_warnings.py` — the same idea for warnings. Every class `/metrics` can emit must render in `#warnings-panel`, and every emitted warning must be registered — the second failing with the RAW text, because the point is that nobody knew the class existed. **Its fixture is a coupled system**: adding a seed can silence an existing class (a small-context seed once killed the escalation hazard by dragging the p95 down), which is why both directions are asserted. - `router_cli.py` — one-shot `/route` probe (no spend), raw JSON with `--json`. [api](docs/api.md). - `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — loopback `/admin` portal: dashboard, models overrides, decisions log, controls. [admin-portal](docs/admin-portal.md). - `tests/` — 1155 tests across 40+ files, offline, verified on Python 3.10 and 3.14. [README](README.md). @@ -635,6 +637,60 @@ upstream token is requested — the classifier is the latency floor. Four settin `max_retries=0` (SDK retry → 3x wall-clock), `max_output_tokens: 1024` (bounds reasoning), `max_input_chars: 8000` (doesn't need the document; head+tail clamp), `fallback_tier`/`fallback_category` (mid tier, not 502/503), plus `temperature: 0` (deterministic tier). "Local" means your hardware, not this machine: bind to the VPN address, not `0.0.0.0` (Ollama has no auth). Full measurements: [docs/local-models.md](docs/local-models.md#the-classifier-is-the-latency-floor). +### A classifier failure no longer collapses to one fixed guess + +`fallback_tier`/`fallback_category` above is now the LAST step, not the only +one. When the local classifier fails, `classify()` walks a cascade: + +| # | step | cost | +|---|---|---| +| 1 | this session's cached classification, **staleness ignored** | free | +| 2 | this session's history in `route_decisions` | free | +| 3 | `classifier.cloud_fallback`, if configured | money | +| 4 | `fallback_tier` / `fallback_category` | free | + +Steps 1 and 2 reuse a *real* classification rather than guessing. Step 2 is +what survives a restart, when the in-memory cache is gone but the session's +decisions are still on disk. Both filter degraded sources, so one outage's +guess cannot propagate through every later turn of a session and end up +looking like a measurement. + +**Step 3 is absent by default and the whole feature costs nothing until it is +configured.** Three things bound it once it is: + +- `classifier.cooldown_seconds` (30) is a **global** backoff, so a sustained + outage buys one cloud attempt per window across ALL requests, not one per + request; +- an account-level provider refusal suppresses step 3 for the same window — + an out-of-credit account makes that call a guaranteed wasted request; +- the `session_cache.put()` gate admits `classifier_cloud`, bounding a session + to one cloud call per staleness window instead of one per turn. + `session_stale` and `session_history` are deliberately NOT cached: borrowed + answers must not renew a staleness clock they never earned. + +**The trap this had to avoid, and it is the reason to read this section.** A +degraded classification is recorded with `task_category = general_chat` — +which is a *fully scored* category (13 proficiency rows) as well as the +configured `fallback_category`. Without a guard, `POST /outcome` reports on +mislabelled outage traffic fold into `proficiency(model, general_chat)` and +drag real scores toward whatever happened to be flowing while the classifier +was down. So `report_outcome` marks outcomes of degraded-source decisions +`model_attributable = 0` — kept as a record, excluded from folding, the same +treatment the tool-call false-failures got. `feedback.py` is untouched; its +existing `AND model_attributable = 1` already does the work. + +Attributable: `classifier`, `cached`, `override`, `classifier_cloud`. Not: +`fallback`, `session_stale`, `session_history`. The asymmetry is deliberate — +a degraded source is good enough to route one visibly-flagged request, but a +proficiency score is consulted by every future request, so attribution must +not trust more than routing does. Unknown decision rows **fail open**; an +over-applied exclusion already starved feedback once (`client_capped`). + +`/metrics` warns when the degraded share of the last 24h crosses +`classifier.degraded_warn_threshold` over at least `degraded_warn_min` +decisions. A survivable failure is exactly the kind that goes unnoticed for +weeks. + ## Local dispatch model A second local model can now be dispatched directly for specific categories. @@ -670,8 +726,26 @@ task_category)` provider-agnostic record as a cloud one. ## What's NOT built yet — pick up here Built: session-directory attribution, the local energy ledger, and local model -dispatch (listed under "What's built and working" above). The three items below -remain open. +dispatch (listed under "What's built and working" above). + +**As of 2026-09-05 the plan queue is empty.** Nine PRs (#26-#34) landed the +gitignored config overlay, admin profile CRUD writing to it, the provider +literal cleanup, quota balance/burn/runway, capability-aware ceiling and +rejection warnings, the TUI schema catch-up, and the classifier fallback +cascade. `plans/multi-provider-support.md` is PARKED on provider selection — +Z.ai was the recommendation and is no longer settled; the coupling surface in +it is measured and still valid. + +Known follow-ups recorded but not specced, both small: + +- The TUI decision table renders the literal `"None"` in the `ctx` cell when + `required_context_tokens` is absent — the same defect the `profile` cell was + written to avoid. See `.omo/notepads/tui-overhaul/issues.md`. +- A NULL `required_context_tokens` raises `TypeError` inside the + demand-ceiling comparison, and the column is nullable. Latent only: zero + such rows exist today, checked on the live DB. + +The three items below remain open. 1. **Leaderboard priors are unfilled.** `leaderboards.yaml` ships empty on purpose — inventing plausible-looking benchmark numbers would put @@ -812,6 +886,34 @@ T. That is silent on all three tiers today and fires on the outage state (94,196 vs an observed max of 268,168). See `plans/catalog-staleness-and-poller-failure-modes.md` §4.4. +**This recurred on 2026-09-04 through a dimension the detector did not model, +and both halves of the fix are now in `metrics.py`.** Admin deprecations took +out `kimi-k3*` — the only vision-capable rows with enough context — so a +242,486-token image request 422'd while every existing check stayed silent, +because the *all-models* tier-1 ceiling was still 782,324. The vision-capable +ceiling had collapsed to 192,500. + +- **Predictive:** `capability_ceilings` / `capability_demand_warnings` compute + `vision` and `json_mode` sub-ceilings and compare each against demand + actually observed for requests carrying images / requesting JSON. Two extra + series, not a bucket per capability combination. +- **Reactive:** `rejection_warnings` watches `route_decisions` for rows with + `selected_model IS NULL`. This is the more valuable half and the simpler + one — it catches the *next* dimension nobody predicted, at the cost of + firing after the first failure rather than before. + +Two details in the reactive detector are load-bearing and easy to undo by +accident. It groups by `(task_tier, digit-normalized reason)` using the +**structured column**, because normalizing digits alone merges `tier >= 1` +and `tier >= 3` rejections into one group and hides whether the broadest or +the frontier candidate set went empty. And the signal is **novelty OR rate**, +never mere presence: measured on the live DB, routine rejections run ~3/hr +while the 2026-09-04 incident was only n=2 — *below* the noise floor — so no +single count threshold can both catch it and stay quiet. A group absent from +the 24h baseline warns at n≥2; a familiar group warns at the configured +count. Zero rejections warn about nothing: a genuinely impossible request +SHOULD 422. + The service holds a billable API key and has **no auth of its own**. Loopback bind is the only thing standing between the open internet and your allowance; add auth before widening `--host`. -- 2.49.1 From d6d3f949ea07f674b55a39bc2f1579bcb4f8b8f8 Mon Sep 17 00:00:00 2001 From: adlee-was-taken Date: Sat, 5 Sep 2026 19:42:16 -0400 Subject: [PATCH 2/3] docs(plans): admin portal uplift -- proficiency, profiles, gaming mode Three gaps found by auditing the portal against the config surface: 5 pages and 9 editable keys against 21 config sections, with seven sections having no presence at all. Proficiency is invisible despite being, per CLAUDE.md, the ONLY category-dependent term in the ranking -- 148 rows over 17 models and 11 categories, surfaced as a single top-N list. The page must distinguish measured scores from inherited ones: 71 of the 148 are outcome_prior and 39 self_eval_thin, and treating those as equal to outcome_blended is the mistake this project already made with -flex rows. Read-only by design; a hand-edited proficiency score is a fabricated measurement. The profiles page is a dead end -- all five profiles are built-ins, which render no action buttons, so nothing on the page is editable and it never says why. Duplicate-to-edit fixes it with no backend change. Its zero-admit alarm also cries wolf: locality reports "admits 0 models" while an active ollama-local row exists, because _profile_probe passes task_category=None and that model is gated to file_summarization and diff_checking. A false alarm on the badge that exists to catch incident #3 is worse than no badge. Gaming mode carries a real defect. _last_classifier_failure is written and never read, so the circuit breaker does not break the circuit and nothing stops the router retrying a dead local classifier -- up to 120s per request against a hung or black-holed Ollama. Records that the toggle should be one flag the code READS rather than a macro writing five keys, and that it must skip the local call rather than fail and recover. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U --- plans/admin-portal-uplift.md | 246 +++++++++++++++++++++++++++++++++++ 1 file changed, 246 insertions(+) create mode 100644 plans/admin-portal-uplift.md diff --git a/plans/admin-portal-uplift.md b/plans/admin-portal-uplift.md new file mode 100644 index 0000000..3fa9d87 --- /dev/null +++ b/plans/admin-portal-uplift.md @@ -0,0 +1,246 @@ +# Admin portal uplift: proficiency, profiles, and gaming mode + +**Status: FINAL — decision-complete.** Written 2026-09-05 against `main` at +`12615d9`, from a live audit of the portal against the config surface and a +browser pass over `profiles.html`. + +Three independent pieces of work, deliberately in one plan because they share +one premise — the portal has fallen behind what the router does. Ship them in +any order; only §3 has a bug attached, so it goes first if you sequence. + +## The audit that produced this + +The portal is **5 pages** (`index`, `models`, `profiles`, `decisions`, +`controls`) and **9 editable config keys**, against **21 config sections**. +Seven sections have no portal presence at all: `exploration`, `escalation`, +`iteration`, `local_dispatch_models`, `local_vision`, `freshness`, and +effectively `classifier`. + +The user's own bar, stated 2026-09-04: *the portal does not need 100% coverage +of every lever, but the main features should each have basic functionality and +configurability there.* This plan closes the three gaps that fail that bar +hardest. It does **not** try to close all seven. + +--- + +## 1. Proficiency is invisible, and it is the table that decides routing + +`CLAUDE.md`: *"`proficiency_score` is the ONLY category-dependent term in the +ranking."* Measured on the live DB: **148 rows, 17 models × 11 categories**, +calibrated against 1,059 client outcomes. + +The portal's entire surface for it is `metrics.top_proficiency` — a top-N list +by `blended_score` for one category, on the dashboard. An operator can see +*which* model won a decision and not *why*. + +### What is missing that matters + +`proficiency` carries twelve columns and the portal shows one. The ones an +operator needs and cannot get: + +| column | why it decides something | +|---|---| +| `source` | 38 `outcome_blended`, 71 `outcome_prior`, 39 `self_eval_thin` on the live DB. A **prior is inherited, not measured** — treating those three as the same number is the mistake this project already made once with `-flex` rows | +| `outcome_samples` / `self_eval_samples` | a 1.00 at n=2 and a 0.97 at n=14 are not comparable, and `docs_writing` already proved that at n=2 the ceiling was a sampling artifact | +| `inherited_from` | which family a variant borrowed its score from | +| `last_updated` | whether a score reflects recent traffic or a months-old sweep | + +### Build + +A `proficiency.html` page: a **model × category matrix** of `blended_score`, +each cell carrying source and sample count (tooltip or a compact marker), with +the ability to sort/filter by category and by source. + +- **Distinguish measured from inherited visually.** `outcome_prior` and + `self_eval_thin` must not look like `outcome_blended`. This is the single + most important requirement on the page — the whole point is telling apart a + score that traffic earned from one that was copied in. +- Flag cells below `self_eval_min_samples` as thin. +- Read-only. **Do NOT add editing.** Proficiency is derived from evaluation + and client outcomes; a hand-edited score is a fabricated measurement, the + same failure as the empty `leaderboards.yaml` and the provider's + `static_fallback` carbon constant this project already excludes. +- New endpoint `GET /admin/api/proficiency` returning the rows. `metrics.py` + gains the query; it must not import `dispatcher`. + +### Also missing, and cheap alongside it + +Nothing in the portal shows **outcomes**. `POST /outcome` is described in +`CLAUDE.md` as *"the only ground truth the router gets"*, and there is no view +of what has been reported. Worse, PR #34 added `model_attributable = 0` for +degraded-classification outcomes — so some reports are deliberately excluded +from folding and **an operator has no way to see that is happening**. + +Add to the proficiency page (or `decisions.html`, implementer's choice — state +which and why): recent `verifications` rows of `kind='client_outcome'`, with +verdict and **whether it was attributable**. A silently-excluded outcome is +exactly the "recorded but never surfaced" pattern this project has now hit +four times. + +--- + +## 2. The profiles page is a dead end, and its zero-admit alarm cries wolf + +### 2a. Nothing on the page is editable, and it never says why + +All five profiles are `source=builtin`. The render logic is: + +```js +const actions = isBuiltin ? '' : `