tests/test_plans_declare_status.py requires 'Status: <done|planned|in progress|parked|reference> -- <reason>' in the first 8 lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
398 lines
22 KiB
Markdown
398 lines
22 KiB
Markdown
# Cockpit brainstorm: the router as an operator's instrument panel
|
||
|
||
Status: reference -- brainstorm, not a queue item; its six quick wins shipped in PR #101
|
||
|
||
**Date:** 2026-09-23
|
||
**Read against:** the repo tarball as of this date (code, config, plans,
|
||
screenshots). **Not** read against the live `router.db`, so every "gap" below
|
||
is a gap in what the code surfaces, not a claim about what the data shows.
|
||
Section 5 is the exception: it was read against the live portal.
|
||
|
||
## The framing
|
||
|
||
6krrt as a local LLM router with a cockpit. The operator:
|
||
|
||
1. adjusts dials and knobs in flight,
|
||
2. sees where tokens are wasted,
|
||
3. watches proficiency, and
|
||
4. sees poor results from a carrier or model surface, so they can decide
|
||
whether to preclude it (with the circuit breaker as the automatic version
|
||
of that decision for outages).
|
||
|
||
The router makes the per-request decision. The cockpit is where the operator
|
||
makes the slower decisions: which models and carriers are trusted, for what,
|
||
and at what settings. Under this framing, the question for every feature is
|
||
whether it helps the operator notice something and act on it.
|
||
|
||
---
|
||
|
||
## What the cockpit already has
|
||
|
||
| job | what exists | where |
|
||
|---|---|---|
|
||
| in-flight dials | ~11 runtime knobs (bool/float/int tables in `admin.py`), persisted config writes via `/api/config/{key}`, classifier card, profiles CRUD, gaming mode | Controls, Profiles |
|
||
| knob coverage | North Star rule #1 + `test_admin_knob_coverage.py` | tests |
|
||
| token waste | pinch savings card; `cache_rate_series` + warnings; incumbent cache pricing dial; `cost_estimate_calibration` in `/metrics` | Dashboard, Controls, `/metrics` |
|
||
| proficiency | model × category matrix with evidence status (measured / inherited / thin, n per cell); client outcome log with attributable + applied flags | Proficiency |
|
||
| poor results | `content_fault_warnings` (malformed output rate); `rejection_warnings`; availability override (active / deprecated / stale); allowlist editor | warnings bell, Models, Providers |
|
||
| breaker | passive per-(model, provider) breaker on 5xx; on/off knob | Controls (toggle only) |
|
||
|
||
The foundation is good. What's mostly missing is the link between a signal
|
||
and the action it should prompt, plus any record of which actions were taken.
|
||
|
||
---
|
||
|
||
## Cross-cutting ideas (these help all four jobs)
|
||
|
||
### X1. Flight recorder: a change log
|
||
|
||
**Gap.** `admin_schema.sql` has only `admin_model_overrides`. Nothing records
|
||
when a knob, override, allowlist entry, profile or feedback fold changed, what
|
||
it changed from, or why. Runtime knobs are "NEVER persisted", so a restart
|
||
silently reverts them, and nothing notes that the revert happened.
|
||
|
||
**Idea.**
|
||
- An `admin_changes` table: `(at, surface, key, old, new, reason, source)`,
|
||
where `source` is `runtime | persisted | override | allowlist | profile |
|
||
feedback_fold | restart`. Every admin write path appends a row.
|
||
- An `operator_epoch` id stamped on each `route_decisions` row, incremented
|
||
on each **operator action** (a row in `admin_changes`). Every metric can then
|
||
be split before and after a change without joining on timestamps. Scope it
|
||
to operator actions on purpose: "any change that affects routing" would also
|
||
cover catalog price polls and every feedback fold, so the epoch would tick
|
||
constantly and split nothing.
|
||
- Change markers drawn on the dashboard history charts.
|
||
- A drift badge when a runtime value differs from its persisted twin
|
||
("this will revert on restart").
|
||
- Restart reverts are loggable without persisting runtime knobs: the change
|
||
log already holds each knob's last runtime value, so at startup, compare it
|
||
with the loaded config and write one `restart` row per value discarded.
|
||
|
||
**Why first.** Adjusting knobs in flight produces no learning if you can't
|
||
later tell what a change did. The other ideas here lean on this one.
|
||
|
||
### X2. Structured warnings with actions attached
|
||
|
||
**Gap.** `/metrics` warnings are plain strings. The dashboard decides which
|
||
page fixes a warning, and how severe it is, by regex on the text
|
||
(`warningTarget` and `warningSeverity` in `index.html`). Rewording a warning
|
||
silently breaks its link and its severity.
|
||
|
||
**Idea.** Emit `{class, severity, subject: {model, provider, category?},
|
||
evidence_url, actions: [...]}`. The warning registry in
|
||
`test_tui_warnings.py` already lists every class, so it can serve as the
|
||
schema. The bell can then offer an action in place: "exclude from
|
||
`tool_use_agentic`", "quarantine 24h", "open decisions filtered to this
|
||
model".
|
||
|
||
### X3. Preview before apply (replay)
|
||
|
||
**Gap.** The reviews replayed thousands of real decisions to test a setting,
|
||
but only as one-off scripts pasted into plan documents.
|
||
|
||
**Idea.** For any change that affects ranking (`quality_tolerance`, the
|
||
incumbent dial, profile edits, availability overrides, graded exclusions
|
||
from P2 below), the portal replays the last N decisions through
|
||
`select_candidates` and `rank_candidates` with the proposed value, then shows:
|
||
- a winner-shift matrix: old winner → new winner, with counts,
|
||
- estimated cost delta,
|
||
- decisions that would become unroutable (the 2026-09-01 and 2026-09-04
|
||
incidents were both deprecations that emptied a candidate set).
|
||
|
||
This is cheap because the ranking modules are pure. Its output is an estimate
|
||
of routing change, not quality change, and should be labeled that way.
|
||
|
||
---
|
||
|
||
## 1. Dials in flight
|
||
|
||
- **D1. Timed changes.** "Apply for 2h, then revert." Turns a knob change into
|
||
a bounded experiment and lowers the risk of forgetting a test setting.
|
||
Needs X1 to record both the apply and the revert.
|
||
- **D2. Blast radius on hover.** How many of the last 24h of decisions this
|
||
knob would have touched. It's the X3 replay reduced to one number.
|
||
- **D3. Session-scoped A/B.** Assign new sessions to setting A or B and
|
||
compare cost and outcome rate per arm. It has to be per session, not per
|
||
turn, because a mid-session switch costs cache (token-waste Wave 2
|
||
measured 0.919 → 0.348). Larger build; only worth it once X1 exists and
|
||
outcome volume can support a comparison.
|
||
- **D4. Grouping by effect.** Group controls by what they move (cost,
|
||
quality, latency, safety, measurement-only) rather than by config section,
|
||
and mark the few that actually change routing. Knob coverage guarantees
|
||
every control exists; this makes them usable.
|
||
|
||
## 2. Token waste
|
||
|
||
**Gap.** Waste signals are spread across pinch, cache rate, calibration and
|
||
the verifications table, and none is in dollars in one place.
|
||
`cost_estimate_calibration` is in `/metrics` but not in the portal.
|
||
|
||
- **W1. Waste ledger.** One panel with dollars per waste class over a window,
|
||
each with a trend:
|
||
| class | how it's computed |
|
||
|---|---|
|
||
| switch cache loss | billed on switch turns − same prompt at the session's same-model cache rate |
|
||
| retries | re-billed prompt tokens from `iteration.py` attempts |
|
||
| paid for failure | billed cost of responses later reported `ok:false` via `/outcome` |
|
||
| empty-200 / malformed | billed cost of responses with a structural `malformed` verdict |
|
||
| truncation | responses stopped by `finish_reason: length` with no client cap |
|
||
| pinch (negative) | dollars saved, from `pinch_summary` |
|
||
|
||
"Paid for failure" counts only decisions that received an outcome, so it
|
||
shows n and outcome coverage (the share of billed decisions with a report)
|
||
next to the dollar figure.
|
||
- **W2. Why did it switch?** Record a switch reason on each decision: category
|
||
changed, breaker open, exploration, override removed the incumbent, profile
|
||
change. Then rank sessions by switch cost and drill into the turns. Without
|
||
a reason, a switch is only a cost; with one, it points to a knob.
|
||
- **W3. Calibration panel.** Surface `cost_estimate_calibration` per model,
|
||
headlining the **spread** (as its docstring argues), and flag pairs where
|
||
the estimator and the bill disagree on order. Where that happens, the cost
|
||
tiebreak picks the wrong model.
|
||
- **W4. Cost per successful outcome.** Per (model, provider, category):
|
||
billed dollars ÷ `ok:true` outcomes. A cheap model that fails often isn't
|
||
cheap. This is the one waste figure that includes quality.
|
||
|
||
Show n and outcome coverage next to it. Models that carry more traffic
|
||
collect more outcomes (the same exposure bias the feedback fold had to
|
||
correct), and if coverage stays hidden, a thinly reported cheap model will
|
||
look better than it is.
|
||
|
||
## 3. Proficiency monitoring
|
||
|
||
**Gap.** The matrix shows current state well. It doesn't show movement, and
|
||
it doesn't show which thin cells actually matter.
|
||
|
||
- **P1. Trend and drift per cell.** A sparkline of the outcome rate over time.
|
||
Compare a recent window against the long-run rate with an interval, and
|
||
flag drops that fall outside it. Carriers change serving setups (quant,
|
||
engine, hardware) without notice, and a drift flag is how that shows up.
|
||
- **P2. Evidence priority.** Rank cells by `traffic share × uncertainty`. A
|
||
thin cell that routing never consults doesn't matter; a thin cell deciding
|
||
30% of traffic does. The ranking tells you where to spend `eval_proficiency`
|
||
runs or exploration budget.
|
||
- **P3. The tolerance band per category.** Show which models sit within
|
||
`quality_tolerance` of the leader, so it's clear where cost is deciding and
|
||
where quality is.
|
||
- **P4. Label provenance per cell.** The share of each cell's outcomes whose
|
||
category came from a fresh classification versus a cached or borrowed one.
|
||
North Star rule #2 exists because borrowed labels trained the matrix;
|
||
this makes it visible cell by cell.
|
||
|
||
## 4. Poor results → operator preclusion (and the breaker)
|
||
|
||
**Gap.** The operator's only preclusion tools are binary: an availability
|
||
override (active / deprecated / stale) or removing a model from the
|
||
allowlist. The breaker is in-memory, trips only on 5xx, forgets everything
|
||
on restart, and shows nothing but an on/off toggle.
|
||
|
||
- **Q1. Carrier/model scorecard.** One sortable table per (model, provider)
|
||
over a window: outcome fail rate (with n), malformed / empty-200 rate,
|
||
5xx and timeout count, breaker trips, p95 latency, estimate-vs-bill error,
|
||
cache rate, cost per success (W4). This is the "who's misbehaving" view the
|
||
preclusion decision needs.
|
||
- **Q2. Same model, different carriers.** Group scorecard rows by base model,
|
||
e.g. `qwen3.6-35b` on NeuralWatt vs OpenRouter. This separates "the model is
|
||
bad at this" from "this carrier serves it badly", which calls for a
|
||
different fix: drop the carrier's row, or restrict the model.
|
||
- **Q3. Graded actions** in place of deprecate-or-not. Each action records a
|
||
reason and links its evidence in X1:
|
||
1. exclude from category X. This **cannot** reuse `eligible_categories`:
|
||
that field is restrict-only (`admin.py`, the probe docstring), and NULL
|
||
means "every category", so a deny needs its own column,
|
||
2. exclude when the request carries tools (the `deepseek-v4-flash` case).
|
||
This may be the way out of the frozen `tool_use_agentic` data:
|
||
`min_tool_proficiency` can only read scores that stopped moving on
|
||
2026-09-15, while a per-model manual rule lets the operator decide from
|
||
Q1 scorecard evidence instead,
|
||
3. restrict to batch / `-flex` only,
|
||
4. quarantine: exploration-only, so it keeps collecting evidence without
|
||
carrying traffic (see the open question below; deferred until Q1 shows
|
||
it is needed),
|
||
5. timed ban, e.g. 24h, auto-expiring and logged,
|
||
6. deprecate (what exists today).
|
||
|
||
Each action goes through at least the minimal X3 check (would this empty a
|
||
candidate set for any recent decision?) before committing. That check alone
|
||
would have caught both the 2026-09-01 and 2026-09-04 incidents; the full
|
||
winner-shift replay can come later.
|
||
- **Q4. Breaker visibility.** Show open circuits with `down_until`, current
|
||
cooldown and trip count; keep a trip history (persist it, since `_store`
|
||
is process-lifetime); add manual force-open (drain a model on purpose) and
|
||
force-close (reset after a known fix).
|
||
- **Q5. A quality breaker, separate from availability.** Trip on content
|
||
failures (empty-200, mangled output, a burst of `ok:false`) using the
|
||
novelty-or-rate rule already in `rejection_warnings`. This is roughly what
|
||
`plans/mangled-output-detection.md` specs (status: planned). Keep it a
|
||
separate breaker and panel so a quality trip doesn't read as an outage.
|
||
**One doc owns it:** fold Q5 into `mangled-output-detection.md` rather than
|
||
specifying it twice, and let this file point there.
|
||
- **Q6. Carrier-level breaker.** When several models on one provider trip
|
||
inside a short window, open the provider rather than walking its models one
|
||
cooldown at a time. Account exhaustion already has its own path; this is for
|
||
partial carrier outages.
|
||
|
||
---
|
||
|
||
## 5. Quality of life: ergonomics and readouts
|
||
|
||
Unlike the sections above, this one **was** read against the live portal
|
||
(8080, view-only, 2026-09-23), plus the frontend source. Each item names the
|
||
thing observed.
|
||
|
||
### Readouts that mislead today
|
||
|
||
- **R1. Home tiles use different windows.** Quota is "this period",
|
||
Decisions 7d, Models and Pinch 30d. Busiest models shows `kimi-k2.7-code`
|
||
at 10,028 while the Decisions tile says 2,887 for everything, and both are
|
||
correct. Label each tile's window on its face, or drive all of them from
|
||
the Activity card's 24h / 7d / 30d selector.
|
||
- **R2. The Decisions tile counts verifications, and mixes diagnostics with
|
||
ground truth.** It reads "2,887 verified in 7 days: 132 ok, 93 failed", but
|
||
`index.html` sums `verdict_mix`, which counts `verifications` rows, not
|
||
decisions. Checked read-only against the live DB on 2026-09-23:
|
||
| | tile | actually |
|
||
|---|---|---|
|
||
| headline | 2,887 | 2,678 route decisions; 2,887 is verification rows, 2,660 of them `unverifiable` |
|
||
| "ok" | 132 | 96 client `succeeded` + 29 `local_llm` ok + 7 structural ok |
|
||
| "failed" | 93 | 82 client `failed` + 11 `malformed` (local_llm + structural) |
|
||
The tile adds structural and `local_llm` verdicts, which CLAUDE.md calls
|
||
diagnostics only, to `/outcome` reports, the only ground truth. Show route
|
||
decisions as the headline, then client outcomes on their own: "178 client
|
||
reports (6.6%): 96 ok, 82 failed (46%)". Diagnostics, if shown at all, go on
|
||
a separate line.
|
||
- **R3. The Activity chart smooths across gaps.** The decisions series is a
|
||
spline through sparse points, so it draws continuous traffic across hours
|
||
that had none. Break the line at empty buckets or use bars. Separately,
|
||
billed requests exceed decisions at several peaks; the chart should say what
|
||
the excess is (cloud classifier calls? retries? unrouted pins?) rather than
|
||
leave two lines that disagree unexplained.
|
||
- **R4. Sparklines have no values.** The tile sparklines carry no axis and no
|
||
hover. Add a hover value, or min and max labels.
|
||
- **R5. The Classifier card flashes a wrong value.** On first paint the mode
|
||
select reads `local_llm` and only switches to `local_encoder` once config
|
||
loads. For a second, the card shows a value that isn't configured. Render
|
||
a loading state instead of the default.
|
||
- **R6. Config echo vs live state.** The Classifier card shows what the
|
||
overlay says (`local_encoder`, device `cpu`), not what the process actually
|
||
loaded. Show the resolved model, device and recent p50 latency from the
|
||
running classifier. (It also exposed doc drift: CLAUDE.md says the live
|
||
deployment runs `device: cuda`; the overlay says `cpu`.)
|
||
- **R7. Proficiency colour encodes evidence, not score.** Green / amber mean
|
||
measured / thin, so the best model in a column isn't visible without reading
|
||
every number. Mark each column's leader, outline the `quality_tolerance`
|
||
band (P3), make columns sortable, and add a "routable only" toggle so rows
|
||
that are deprecated or not allowlisted stop padding the matrix.
|
||
|
||
### Decisions page ergonomics
|
||
|
||
- **E1. Collapse runs.** The top 40 rows are one session: same category,
|
||
profile, tier, source and model, and only ctx and cost move. Fold
|
||
consecutive same-session, same-model rows into one expandable row: "38 turns,
|
||
ctx 75k to 98k, $0.27 total". Switches then stand out as row boundaries,
|
||
which is what W2 wants to surface.
|
||
- **E2. Session view.** A per-session rollup: turns, total cost, context
|
||
growth, switches, outcome count. Context climbing about 1k per turn is
|
||
visible in the raw table and invisible everywhere else.
|
||
- **E3. Filters in the URL.** `decisions.html` never reads `URLSearchParams`,
|
||
so no other page can link to "decisions for this model" or "this session".
|
||
X2 actions, the Q1 scorecard and the model modal all need that link.
|
||
- **E4. Row drill-down.** A row click does nothing. Open a panel with the
|
||
decision's rejected candidates, verification verdict, `/outcome` report,
|
||
and estimated vs billed cost from its energy observation.
|
||
- **E5. Server-side filtering.** The page loads 1,000 of 32,540 rows and
|
||
filters and searches only those, with a footnote saying so. Push filters to
|
||
the API so that "All" means all.
|
||
- **E6. Model, provider and source as filters.** Today they're reachable only
|
||
through free-text search.
|
||
|
||
### Controls page
|
||
|
||
- **C1. One knob table, not two.** Runtime Knobs and Persisted Config list
|
||
largely the same knobs in two columns, in different orders, with persisted
|
||
labels truncated (`objective.incumbent_cache_pri…`). Checking drift means
|
||
matching rows by eye. Use one table: knob, live value, persisted value,
|
||
layer (base / overlay), drift badge. That table is also X1's natural home.
|
||
- **C2. Separate Restart Service.** It sits in the same button row as Refresh
|
||
Catalog and Seed Energy. Move it apart, and have its confirmation list which
|
||
runtime values differ from persisted and will revert.
|
||
- **C3. Prose in cards.** Local Compute carries three paragraphs; the portal
|
||
style rule is no paragraphs. Keep one line plus a docs link or tooltip.
|
||
- **C4. Default and last change per knob.** Show the code default beside
|
||
each value, plus "changed 2h ago from 0.3" once X1 exists.
|
||
|
||
### Portal-wide
|
||
|
||
- **G1. Nav lives in eight files.** Each page hardcodes the same `<ul>` of
|
||
nav links. `navbar.js` already notes that adding Quota was an eight-file
|
||
edit; have it render the links too.
|
||
- **G2. Width.** Most pages sit in `container-xl` and use well under half of
|
||
a wide screen, while Decisions, the densest table, is the most cramped.
|
||
Proficiency already goes full-width. Go fluid on the data-heavy pages.
|
||
- **G3. Dismissals don't stick on rate warnings.** Dismissal is keyed on the
|
||
warning's exact text, on purpose (a changed condition resurfaces it). But
|
||
warnings that embed a live rate ("3.2x pace") change text every poll, so a
|
||
dismissal never holds. X2's `class` + `subject` is the right key; resurface
|
||
on a severity change, not a digit change.
|
||
- **G4. Keyboard.** `/` to focus search, `Esc` to close modals, and `j`/`k`
|
||
on the Decisions table. Cheap, and this is a page someone lives in.
|
||
|
||
---
|
||
|
||
## One thing the cockpit framing makes more urgent
|
||
|
||
The portal is loopback-only with no auth, and it already includes
|
||
`/api/restart-service`, `/api/apply-feedback` and provider deletes. The more
|
||
it becomes the place decisions are made, the stronger the pull to reach it
|
||
from other machines, and moving the router onto a Proxmox box on the LAN is
|
||
exactly that move. Add auth before binding anything but `127.0.0.1`. Same
|
||
warning CLAUDE.md already gives for the API, with more at stake.
|
||
|
||
---
|
||
|
||
## A possible order
|
||
|
||
| # | item | why here |
|
||
|---|---|---|
|
||
| 0 | portal auth | **only if** the router is leaving `127.0.0.1` (the Proxmox move); blocks that move, not this list |
|
||
| 1 | X1 change log + operator epoch | everything else reads it |
|
||
| 2 | Q1 scorecard + Q4 breaker visibility | the preclusion decision needs one view first |
|
||
| 3 | Q3 graded actions (via X1) + minimal X3 (empties-a-candidate-set check) | turns the scorecard into decisions without repeating 09-01 / 09-04 |
|
||
| 4 | X2 structured warnings | lets warnings trigger Q3 actions directly |
|
||
| 5 | W1 waste ledger + W2 switch reasons | token waste in dollars, with causes |
|
||
| 6 | X3 full replay preview (winner shift, cost delta) | makes knob changes safe to try |
|
||
| 7 | P1 drift, P2 evidence priority | proficiency monitoring over time |
|
||
| 8 | Q5 quality breaker (owned by `mangled-output-detection.md`), D1 timed changes | automate what the operator has been doing by hand |
|
||
| 9 | D3 session A/B, Q6 carrier breaker | only after the above prove out |
|
||
|
||
Section 5 sits outside this order: most of it is small and independent.
|
||
Quick wins that also unblock the numbered items: **E3** (URL filters, needed
|
||
by X2 and Q1), **R2** (outcome coverage, the same honesty W1 and W4 need),
|
||
**C1** (one knob table, X1's home), **G3** (fixed by X2's class keys). R5 and
|
||
G1 are fixes of a few lines each.
|
||
|
||
**North Star #1 applies to every item here.** Q5 thresholds, D1 durations,
|
||
timed-ban lengths and any quarantine settings are new config knobs, so each
|
||
ships with its admin control or `test_admin_knob_coverage.py` fails. Scope the
|
||
control into the item, not a follow-up.
|
||
|
||
### Open questions, with proposed answers
|
||
|
||
- **Is the change log append-only history, or also an undo stack?**
|
||
Proposed: append-only, plus a per-row "revert to old value" that writes one
|
||
key and appends its own row. That is not the shape CLAUDE.md rejected for
|
||
gaming mode: that concern was one flag writing five keys and drifting apart.
|
||
A one-key revert can't drift, and the history stays intact.
|
||
- **Should graded exclusions live in the DB or in `config.local.yaml`?**
|
||
Proposed: the DB, next to `admin_model_overrides` (extend it or add a
|
||
sibling table). Timed bans need an expiry, and an expiry belongs in a DB
|
||
column. The overlay is per-machine and not in git, so it's the wrong home
|
||
for judgments about a carrier. And availability overrides already live in
|
||
the DB. The category exclusion still needs a new deny column (see Q3.1).
|
||
- **Does quarantine conflict with session-scoped exploration?** Yes, and the
|
||
weight goes up: a quarantined model would carry whole sessions, not single
|
||
turns. There's a second problem too. `exploration.py` only chooses among
|
||
models that already passed the hard filters and were ranked, so quarantine
|
||
needs a new state, "eligible to explore, barred from winning", rather than
|
||
reusing anything existing. Defer it until Q1 shows it's needed.
|