Files
6krrt/plans/cockpit-brainstorm.md
adlee-was-taken 69c969d104 plans: declare a valid Status on the six plans this branch adds
tests/test_plans_declare_status.py requires 'Status: <done|planned|in
progress|parked|reference> -- <reason>' in the first 8 lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-26 21:46:15 -04:00

398 lines
22 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Cockpit brainstorm: the router as an operator's instrument panel
Status: reference -- brainstorm, not a queue item; its six quick wins shipped in PR #101
**Date:** 2026-09-23
**Read against:** the repo tarball as of this date (code, config, plans,
screenshots). **Not** read against the live `router.db`, so every "gap" below
is a gap in what the code surfaces, not a claim about what the data shows.
Section 5 is the exception: it was read against the live portal.
## The framing
6krrt as a local LLM router with a cockpit. The operator:
1. adjusts dials and knobs in flight,
2. sees where tokens are wasted,
3. watches proficiency, and
4. sees poor results from a carrier or model surface, so they can decide
whether to preclude it (with the circuit breaker as the automatic version
of that decision for outages).
The router makes the per-request decision. The cockpit is where the operator
makes the slower decisions: which models and carriers are trusted, for what,
and at what settings. Under this framing, the question for every feature is
whether it helps the operator notice something and act on it.
---
## What the cockpit already has
| job | what exists | where |
|---|---|---|
| in-flight dials | ~11 runtime knobs (bool/float/int tables in `admin.py`), persisted config writes via `/api/config/{key}`, classifier card, profiles CRUD, gaming mode | Controls, Profiles |
| knob coverage | North Star rule #1 + `test_admin_knob_coverage.py` | tests |
| token waste | pinch savings card; `cache_rate_series` + warnings; incumbent cache pricing dial; `cost_estimate_calibration` in `/metrics` | Dashboard, Controls, `/metrics` |
| proficiency | model × category matrix with evidence status (measured / inherited / thin, n per cell); client outcome log with attributable + applied flags | Proficiency |
| poor results | `content_fault_warnings` (malformed output rate); `rejection_warnings`; availability override (active / deprecated / stale); allowlist editor | warnings bell, Models, Providers |
| breaker | passive per-(model, provider) breaker on 5xx; on/off knob | Controls (toggle only) |
The foundation is good. What's mostly missing is the link between a signal
and the action it should prompt, plus any record of which actions were taken.
---
## Cross-cutting ideas (these help all four jobs)
### X1. Flight recorder: a change log
**Gap.** `admin_schema.sql` has only `admin_model_overrides`. Nothing records
when a knob, override, allowlist entry, profile or feedback fold changed, what
it changed from, or why. Runtime knobs are "NEVER persisted", so a restart
silently reverts them, and nothing notes that the revert happened.
**Idea.**
- An `admin_changes` table: `(at, surface, key, old, new, reason, source)`,
where `source` is `runtime | persisted | override | allowlist | profile |
feedback_fold | restart`. Every admin write path appends a row.
- An `operator_epoch` id stamped on each `route_decisions` row, incremented
on each **operator action** (a row in `admin_changes`). Every metric can then
be split before and after a change without joining on timestamps. Scope it
to operator actions on purpose: "any change that affects routing" would also
cover catalog price polls and every feedback fold, so the epoch would tick
constantly and split nothing.
- Change markers drawn on the dashboard history charts.
- A drift badge when a runtime value differs from its persisted twin
("this will revert on restart").
- Restart reverts are loggable without persisting runtime knobs: the change
log already holds each knob's last runtime value, so at startup, compare it
with the loaded config and write one `restart` row per value discarded.
**Why first.** Adjusting knobs in flight produces no learning if you can't
later tell what a change did. The other ideas here lean on this one.
### X2. Structured warnings with actions attached
**Gap.** `/metrics` warnings are plain strings. The dashboard decides which
page fixes a warning, and how severe it is, by regex on the text
(`warningTarget` and `warningSeverity` in `index.html`). Rewording a warning
silently breaks its link and its severity.
**Idea.** Emit `{class, severity, subject: {model, provider, category?},
evidence_url, actions: [...]}`. The warning registry in
`test_tui_warnings.py` already lists every class, so it can serve as the
schema. The bell can then offer an action in place: "exclude from
`tool_use_agentic`", "quarantine 24h", "open decisions filtered to this
model".
### X3. Preview before apply (replay)
**Gap.** The reviews replayed thousands of real decisions to test a setting,
but only as one-off scripts pasted into plan documents.
**Idea.** For any change that affects ranking (`quality_tolerance`, the
incumbent dial, profile edits, availability overrides, graded exclusions
from P2 below), the portal replays the last N decisions through
`select_candidates` and `rank_candidates` with the proposed value, then shows:
- a winner-shift matrix: old winner → new winner, with counts,
- estimated cost delta,
- decisions that would become unroutable (the 2026-09-01 and 2026-09-04
incidents were both deprecations that emptied a candidate set).
This is cheap because the ranking modules are pure. Its output is an estimate
of routing change, not quality change, and should be labeled that way.
---
## 1. Dials in flight
- **D1. Timed changes.** "Apply for 2h, then revert." Turns a knob change into
a bounded experiment and lowers the risk of forgetting a test setting.
Needs X1 to record both the apply and the revert.
- **D2. Blast radius on hover.** How many of the last 24h of decisions this
knob would have touched. It's the X3 replay reduced to one number.
- **D3. Session-scoped A/B.** Assign new sessions to setting A or B and
compare cost and outcome rate per arm. It has to be per session, not per
turn, because a mid-session switch costs cache (token-waste Wave 2
measured 0.919 → 0.348). Larger build; only worth it once X1 exists and
outcome volume can support a comparison.
- **D4. Grouping by effect.** Group controls by what they move (cost,
quality, latency, safety, measurement-only) rather than by config section,
and mark the few that actually change routing. Knob coverage guarantees
every control exists; this makes them usable.
## 2. Token waste
**Gap.** Waste signals are spread across pinch, cache rate, calibration and
the verifications table, and none is in dollars in one place.
`cost_estimate_calibration` is in `/metrics` but not in the portal.
- **W1. Waste ledger.** One panel with dollars per waste class over a window,
each with a trend:
| class | how it's computed |
|---|---|
| switch cache loss | billed on switch turns − same prompt at the session's same-model cache rate |
| retries | re-billed prompt tokens from `iteration.py` attempts |
| paid for failure | billed cost of responses later reported `ok:false` via `/outcome` |
| empty-200 / malformed | billed cost of responses with a structural `malformed` verdict |
| truncation | responses stopped by `finish_reason: length` with no client cap |
| pinch (negative) | dollars saved, from `pinch_summary` |
"Paid for failure" counts only decisions that received an outcome, so it
shows n and outcome coverage (the share of billed decisions with a report)
next to the dollar figure.
- **W2. Why did it switch?** Record a switch reason on each decision: category
changed, breaker open, exploration, override removed the incumbent, profile
change. Then rank sessions by switch cost and drill into the turns. Without
a reason, a switch is only a cost; with one, it points to a knob.
- **W3. Calibration panel.** Surface `cost_estimate_calibration` per model,
headlining the **spread** (as its docstring argues), and flag pairs where
the estimator and the bill disagree on order. Where that happens, the cost
tiebreak picks the wrong model.
- **W4. Cost per successful outcome.** Per (model, provider, category):
billed dollars ÷ `ok:true` outcomes. A cheap model that fails often isn't
cheap. This is the one waste figure that includes quality.
Show n and outcome coverage next to it. Models that carry more traffic
collect more outcomes (the same exposure bias the feedback fold had to
correct), and if coverage stays hidden, a thinly reported cheap model will
look better than it is.
## 3. Proficiency monitoring
**Gap.** The matrix shows current state well. It doesn't show movement, and
it doesn't show which thin cells actually matter.
- **P1. Trend and drift per cell.** A sparkline of the outcome rate over time.
Compare a recent window against the long-run rate with an interval, and
flag drops that fall outside it. Carriers change serving setups (quant,
engine, hardware) without notice, and a drift flag is how that shows up.
- **P2. Evidence priority.** Rank cells by `traffic share × uncertainty`. A
thin cell that routing never consults doesn't matter; a thin cell deciding
30% of traffic does. The ranking tells you where to spend `eval_proficiency`
runs or exploration budget.
- **P3. The tolerance band per category.** Show which models sit within
`quality_tolerance` of the leader, so it's clear where cost is deciding and
where quality is.
- **P4. Label provenance per cell.** The share of each cell's outcomes whose
category came from a fresh classification versus a cached or borrowed one.
North Star rule #2 exists because borrowed labels trained the matrix;
this makes it visible cell by cell.
## 4. Poor results → operator preclusion (and the breaker)
**Gap.** The operator's only preclusion tools are binary: an availability
override (active / deprecated / stale) or removing a model from the
allowlist. The breaker is in-memory, trips only on 5xx, forgets everything
on restart, and shows nothing but an on/off toggle.
- **Q1. Carrier/model scorecard.** One sortable table per (model, provider)
over a window: outcome fail rate (with n), malformed / empty-200 rate,
5xx and timeout count, breaker trips, p95 latency, estimate-vs-bill error,
cache rate, cost per success (W4). This is the "who's misbehaving" view the
preclusion decision needs.
- **Q2. Same model, different carriers.** Group scorecard rows by base model,
e.g. `qwen3.6-35b` on NeuralWatt vs OpenRouter. This separates "the model is
bad at this" from "this carrier serves it badly", which calls for a
different fix: drop the carrier's row, or restrict the model.
- **Q3. Graded actions** in place of deprecate-or-not. Each action records a
reason and links its evidence in X1:
1. exclude from category X. This **cannot** reuse `eligible_categories`:
that field is restrict-only (`admin.py`, the probe docstring), and NULL
means "every category", so a deny needs its own column,
2. exclude when the request carries tools (the `deepseek-v4-flash` case).
This may be the way out of the frozen `tool_use_agentic` data:
`min_tool_proficiency` can only read scores that stopped moving on
2026-09-15, while a per-model manual rule lets the operator decide from
Q1 scorecard evidence instead,
3. restrict to batch / `-flex` only,
4. quarantine: exploration-only, so it keeps collecting evidence without
carrying traffic (see the open question below; deferred until Q1 shows
it is needed),
5. timed ban, e.g. 24h, auto-expiring and logged,
6. deprecate (what exists today).
Each action goes through at least the minimal X3 check (would this empty a
candidate set for any recent decision?) before committing. That check alone
would have caught both the 2026-09-01 and 2026-09-04 incidents; the full
winner-shift replay can come later.
- **Q4. Breaker visibility.** Show open circuits with `down_until`, current
cooldown and trip count; keep a trip history (persist it, since `_store`
is process-lifetime); add manual force-open (drain a model on purpose) and
force-close (reset after a known fix).
- **Q5. A quality breaker, separate from availability.** Trip on content
failures (empty-200, mangled output, a burst of `ok:false`) using the
novelty-or-rate rule already in `rejection_warnings`. This is roughly what
`plans/mangled-output-detection.md` specs (status: planned). Keep it a
separate breaker and panel so a quality trip doesn't read as an outage.
**One doc owns it:** fold Q5 into `mangled-output-detection.md` rather than
specifying it twice, and let this file point there.
- **Q6. Carrier-level breaker.** When several models on one provider trip
inside a short window, open the provider rather than walking its models one
cooldown at a time. Account exhaustion already has its own path; this is for
partial carrier outages.
---
## 5. Quality of life: ergonomics and readouts
Unlike the sections above, this one **was** read against the live portal
(8080, view-only, 2026-09-23), plus the frontend source. Each item names the
thing observed.
### Readouts that mislead today
- **R1. Home tiles use different windows.** Quota is "this period",
Decisions 7d, Models and Pinch 30d. Busiest models shows `kimi-k2.7-code`
at 10,028 while the Decisions tile says 2,887 for everything, and both are
correct. Label each tile's window on its face, or drive all of them from
the Activity card's 24h / 7d / 30d selector.
- **R2. The Decisions tile counts verifications, and mixes diagnostics with
ground truth.** It reads "2,887 verified in 7 days: 132 ok, 93 failed", but
`index.html` sums `verdict_mix`, which counts `verifications` rows, not
decisions. Checked read-only against the live DB on 2026-09-23:
| | tile | actually |
|---|---|---|
| headline | 2,887 | 2,678 route decisions; 2,887 is verification rows, 2,660 of them `unverifiable` |
| "ok" | 132 | 96 client `succeeded` + 29 `local_llm` ok + 7 structural ok |
| "failed" | 93 | 82 client `failed` + 11 `malformed` (local_llm + structural) |
The tile adds structural and `local_llm` verdicts, which CLAUDE.md calls
diagnostics only, to `/outcome` reports, the only ground truth. Show route
decisions as the headline, then client outcomes on their own: "178 client
reports (6.6%): 96 ok, 82 failed (46%)". Diagnostics, if shown at all, go on
a separate line.
- **R3. The Activity chart smooths across gaps.** The decisions series is a
spline through sparse points, so it draws continuous traffic across hours
that had none. Break the line at empty buckets or use bars. Separately,
billed requests exceed decisions at several peaks; the chart should say what
the excess is (cloud classifier calls? retries? unrouted pins?) rather than
leave two lines that disagree unexplained.
- **R4. Sparklines have no values.** The tile sparklines carry no axis and no
hover. Add a hover value, or min and max labels.
- **R5. The Classifier card flashes a wrong value.** On first paint the mode
select reads `local_llm` and only switches to `local_encoder` once config
loads. For a second, the card shows a value that isn't configured. Render
a loading state instead of the default.
- **R6. Config echo vs live state.** The Classifier card shows what the
overlay says (`local_encoder`, device `cpu`), not what the process actually
loaded. Show the resolved model, device and recent p50 latency from the
running classifier. (It also exposed doc drift: CLAUDE.md says the live
deployment runs `device: cuda`; the overlay says `cpu`.)
- **R7. Proficiency colour encodes evidence, not score.** Green / amber mean
measured / thin, so the best model in a column isn't visible without reading
every number. Mark each column's leader, outline the `quality_tolerance`
band (P3), make columns sortable, and add a "routable only" toggle so rows
that are deprecated or not allowlisted stop padding the matrix.
### Decisions page ergonomics
- **E1. Collapse runs.** The top 40 rows are one session: same category,
profile, tier, source and model, and only ctx and cost move. Fold
consecutive same-session, same-model rows into one expandable row: "38 turns,
ctx 75k to 98k, $0.27 total". Switches then stand out as row boundaries,
which is what W2 wants to surface.
- **E2. Session view.** A per-session rollup: turns, total cost, context
growth, switches, outcome count. Context climbing about 1k per turn is
visible in the raw table and invisible everywhere else.
- **E3. Filters in the URL.** `decisions.html` never reads `URLSearchParams`,
so no other page can link to "decisions for this model" or "this session".
X2 actions, the Q1 scorecard and the model modal all need that link.
- **E4. Row drill-down.** A row click does nothing. Open a panel with the
decision's rejected candidates, verification verdict, `/outcome` report,
and estimated vs billed cost from its energy observation.
- **E5. Server-side filtering.** The page loads 1,000 of 32,540 rows and
filters and searches only those, with a footnote saying so. Push filters to
the API so that "All" means all.
- **E6. Model, provider and source as filters.** Today they're reachable only
through free-text search.
### Controls page
- **C1. One knob table, not two.** Runtime Knobs and Persisted Config list
largely the same knobs in two columns, in different orders, with persisted
labels truncated (`objective.incumbent_cache_pri…`). Checking drift means
matching rows by eye. Use one table: knob, live value, persisted value,
layer (base / overlay), drift badge. That table is also X1's natural home.
- **C2. Separate Restart Service.** It sits in the same button row as Refresh
Catalog and Seed Energy. Move it apart, and have its confirmation list which
runtime values differ from persisted and will revert.
- **C3. Prose in cards.** Local Compute carries three paragraphs; the portal
style rule is no paragraphs. Keep one line plus a docs link or tooltip.
- **C4. Default and last change per knob.** Show the code default beside
each value, plus "changed 2h ago from 0.3" once X1 exists.
### Portal-wide
- **G1. Nav lives in eight files.** Each page hardcodes the same `<ul>` of
nav links. `navbar.js` already notes that adding Quota was an eight-file
edit; have it render the links too.
- **G2. Width.** Most pages sit in `container-xl` and use well under half of
a wide screen, while Decisions, the densest table, is the most cramped.
Proficiency already goes full-width. Go fluid on the data-heavy pages.
- **G3. Dismissals don't stick on rate warnings.** Dismissal is keyed on the
warning's exact text, on purpose (a changed condition resurfaces it). But
warnings that embed a live rate ("3.2x pace") change text every poll, so a
dismissal never holds. X2's `class` + `subject` is the right key; resurface
on a severity change, not a digit change.
- **G4. Keyboard.** `/` to focus search, `Esc` to close modals, and `j`/`k`
on the Decisions table. Cheap, and this is a page someone lives in.
---
## One thing the cockpit framing makes more urgent
The portal is loopback-only with no auth, and it already includes
`/api/restart-service`, `/api/apply-feedback` and provider deletes. The more
it becomes the place decisions are made, the stronger the pull to reach it
from other machines, and moving the router onto a Proxmox box on the LAN is
exactly that move. Add auth before binding anything but `127.0.0.1`. Same
warning CLAUDE.md already gives for the API, with more at stake.
---
## A possible order
| # | item | why here |
|---|---|---|
| 0 | portal auth | **only if** the router is leaving `127.0.0.1` (the Proxmox move); blocks that move, not this list |
| 1 | X1 change log + operator epoch | everything else reads it |
| 2 | Q1 scorecard + Q4 breaker visibility | the preclusion decision needs one view first |
| 3 | Q3 graded actions (via X1) + minimal X3 (empties-a-candidate-set check) | turns the scorecard into decisions without repeating 09-01 / 09-04 |
| 4 | X2 structured warnings | lets warnings trigger Q3 actions directly |
| 5 | W1 waste ledger + W2 switch reasons | token waste in dollars, with causes |
| 6 | X3 full replay preview (winner shift, cost delta) | makes knob changes safe to try |
| 7 | P1 drift, P2 evidence priority | proficiency monitoring over time |
| 8 | Q5 quality breaker (owned by `mangled-output-detection.md`), D1 timed changes | automate what the operator has been doing by hand |
| 9 | D3 session A/B, Q6 carrier breaker | only after the above prove out |
Section 5 sits outside this order: most of it is small and independent.
Quick wins that also unblock the numbered items: **E3** (URL filters, needed
by X2 and Q1), **R2** (outcome coverage, the same honesty W1 and W4 need),
**C1** (one knob table, X1's home), **G3** (fixed by X2's class keys). R5 and
G1 are fixes of a few lines each.
**North Star #1 applies to every item here.** Q5 thresholds, D1 durations,
timed-ban lengths and any quarantine settings are new config knobs, so each
ships with its admin control or `test_admin_knob_coverage.py` fails. Scope the
control into the item, not a follow-up.
### Open questions, with proposed answers
- **Is the change log append-only history, or also an undo stack?**
Proposed: append-only, plus a per-row "revert to old value" that writes one
key and appends its own row. That is not the shape CLAUDE.md rejected for
gaming mode: that concern was one flag writing five keys and drifting apart.
A one-key revert can't drift, and the history stays intact.
- **Should graded exclusions live in the DB or in `config.local.yaml`?**
Proposed: the DB, next to `admin_model_overrides` (extend it or add a
sibling table). Timed bans need an expiry, and an expiry belongs in a DB
column. The overlay is per-machine and not in git, so it's the wrong home
for judgments about a carrier. And availability overrides already live in
the DB. The category exclusion still needs a new deny column (see Q3.1).
- **Does quarantine conflict with session-scoped exploration?** Yes, and the
weight goes up: a quarantined model would carry whole sessions, not single
turns. There's a second problem too. `exploration.py` only chooses among
models that already passed the hard filters and were ranked, so quarantine
needs a new state, "eligible to explore, barred from winning", rather than
reusing anything existing. Defer it until Q1 shows it's needed.