tests/test_plans_declare_status.py requires 'Status: <done|planned|in progress|parked|reference> -- <reason>' in the first 8 lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
22 KiB
Cockpit brainstorm: the router as an operator's instrument panel
Status: reference -- brainstorm, not a queue item; its six quick wins shipped in PR #101
Date: 2026-09-23
Read against: the repo tarball as of this date (code, config, plans,
screenshots). Not read against the live router.db, so every "gap" below
is a gap in what the code surfaces, not a claim about what the data shows.
Section 5 is the exception: it was read against the live portal.
The framing
6krrt as a local LLM router with a cockpit. The operator:
- adjusts dials and knobs in flight,
- sees where tokens are wasted,
- watches proficiency, and
- sees poor results from a carrier or model surface, so they can decide whether to preclude it (with the circuit breaker as the automatic version of that decision for outages).
The router makes the per-request decision. The cockpit is where the operator makes the slower decisions: which models and carriers are trusted, for what, and at what settings. Under this framing, the question for every feature is whether it helps the operator notice something and act on it.
What the cockpit already has
| job | what exists | where |
|---|---|---|
| in-flight dials | ~11 runtime knobs (bool/float/int tables in admin.py), persisted config writes via /api/config/{key}, classifier card, profiles CRUD, gaming mode |
Controls, Profiles |
| knob coverage | North Star rule #1 + test_admin_knob_coverage.py |
tests |
| token waste | pinch savings card; cache_rate_series + warnings; incumbent cache pricing dial; cost_estimate_calibration in /metrics |
Dashboard, Controls, /metrics |
| proficiency | model × category matrix with evidence status (measured / inherited / thin, n per cell); client outcome log with attributable + applied flags | Proficiency |
| poor results | content_fault_warnings (malformed output rate); rejection_warnings; availability override (active / deprecated / stale); allowlist editor |
warnings bell, Models, Providers |
| breaker | passive per-(model, provider) breaker on 5xx; on/off knob | Controls (toggle only) |
The foundation is good. What's mostly missing is the link between a signal and the action it should prompt, plus any record of which actions were taken.
Cross-cutting ideas (these help all four jobs)
X1. Flight recorder: a change log
Gap. admin_schema.sql has only admin_model_overrides. Nothing records
when a knob, override, allowlist entry, profile or feedback fold changed, what
it changed from, or why. Runtime knobs are "NEVER persisted", so a restart
silently reverts them, and nothing notes that the revert happened.
Idea.
- An
admin_changestable:(at, surface, key, old, new, reason, source), wheresourceisruntime | persisted | override | allowlist | profile | feedback_fold | restart. Every admin write path appends a row. - An
operator_epochid stamped on eachroute_decisionsrow, incremented on each operator action (a row inadmin_changes). Every metric can then be split before and after a change without joining on timestamps. Scope it to operator actions on purpose: "any change that affects routing" would also cover catalog price polls and every feedback fold, so the epoch would tick constantly and split nothing. - Change markers drawn on the dashboard history charts.
- A drift badge when a runtime value differs from its persisted twin ("this will revert on restart").
- Restart reverts are loggable without persisting runtime knobs: the change
log already holds each knob's last runtime value, so at startup, compare it
with the loaded config and write one
restartrow per value discarded.
Why first. Adjusting knobs in flight produces no learning if you can't later tell what a change did. The other ideas here lean on this one.
X2. Structured warnings with actions attached
Gap. /metrics warnings are plain strings. The dashboard decides which
page fixes a warning, and how severe it is, by regex on the text
(warningTarget and warningSeverity in index.html). Rewording a warning
silently breaks its link and its severity.
Idea. Emit {class, severity, subject: {model, provider, category?}, evidence_url, actions: [...]}. The warning registry in
test_tui_warnings.py already lists every class, so it can serve as the
schema. The bell can then offer an action in place: "exclude from
tool_use_agentic", "quarantine 24h", "open decisions filtered to this
model".
X3. Preview before apply (replay)
Gap. The reviews replayed thousands of real decisions to test a setting, but only as one-off scripts pasted into plan documents.
Idea. For any change that affects ranking (quality_tolerance, the
incumbent dial, profile edits, availability overrides, graded exclusions
from P2 below), the portal replays the last N decisions through
select_candidates and rank_candidates with the proposed value, then shows:
- a winner-shift matrix: old winner → new winner, with counts,
- estimated cost delta,
- decisions that would become unroutable (the 2026-09-01 and 2026-09-04 incidents were both deprecations that emptied a candidate set).
This is cheap because the ranking modules are pure. Its output is an estimate of routing change, not quality change, and should be labeled that way.
1. Dials in flight
- D1. Timed changes. "Apply for 2h, then revert." Turns a knob change into a bounded experiment and lowers the risk of forgetting a test setting. Needs X1 to record both the apply and the revert.
- D2. Blast radius on hover. How many of the last 24h of decisions this knob would have touched. It's the X3 replay reduced to one number.
- D3. Session-scoped A/B. Assign new sessions to setting A or B and compare cost and outcome rate per arm. It has to be per session, not per turn, because a mid-session switch costs cache (token-waste Wave 2 measured 0.919 → 0.348). Larger build; only worth it once X1 exists and outcome volume can support a comparison.
- D4. Grouping by effect. Group controls by what they move (cost, quality, latency, safety, measurement-only) rather than by config section, and mark the few that actually change routing. Knob coverage guarantees every control exists; this makes them usable.
2. Token waste
Gap. Waste signals are spread across pinch, cache rate, calibration and
the verifications table, and none is in dollars in one place.
cost_estimate_calibration is in /metrics but not in the portal.
-
W1. Waste ledger. One panel with dollars per waste class over a window, each with a trend:
class how it's computed switch cache loss billed on switch turns − same prompt at the session's same-model cache rate retries re-billed prompt tokens from iteration.pyattemptspaid for failure billed cost of responses later reported ok:falsevia/outcomeempty-200 / malformed billed cost of responses with a structural malformedverdicttruncation responses stopped by finish_reason: lengthwith no client cappinch (negative) dollars saved, from pinch_summary"Paid for failure" counts only decisions that received an outcome, so it shows n and outcome coverage (the share of billed decisions with a report) next to the dollar figure.
-
W2. Why did it switch? Record a switch reason on each decision: category changed, breaker open, exploration, override removed the incumbent, profile change. Then rank sessions by switch cost and drill into the turns. Without a reason, a switch is only a cost; with one, it points to a knob.
-
W3. Calibration panel. Surface
cost_estimate_calibrationper model, headlining the spread (as its docstring argues), and flag pairs where the estimator and the bill disagree on order. Where that happens, the cost tiebreak picks the wrong model. -
W4. Cost per successful outcome. Per (model, provider, category): billed dollars ÷
ok:trueoutcomes. A cheap model that fails often isn't cheap. This is the one waste figure that includes quality.Show n and outcome coverage next to it. Models that carry more traffic collect more outcomes (the same exposure bias the feedback fold had to correct), and if coverage stays hidden, a thinly reported cheap model will look better than it is.
3. Proficiency monitoring
Gap. The matrix shows current state well. It doesn't show movement, and it doesn't show which thin cells actually matter.
- P1. Trend and drift per cell. A sparkline of the outcome rate over time. Compare a recent window against the long-run rate with an interval, and flag drops that fall outside it. Carriers change serving setups (quant, engine, hardware) without notice, and a drift flag is how that shows up.
- P2. Evidence priority. Rank cells by
traffic share × uncertainty. A thin cell that routing never consults doesn't matter; a thin cell deciding 30% of traffic does. The ranking tells you where to spendeval_proficiencyruns or exploration budget. - P3. The tolerance band per category. Show which models sit within
quality_toleranceof the leader, so it's clear where cost is deciding and where quality is. - P4. Label provenance per cell. The share of each cell's outcomes whose category came from a fresh classification versus a cached or borrowed one. North Star rule #2 exists because borrowed labels trained the matrix; this makes it visible cell by cell.
4. Poor results → operator preclusion (and the breaker)
Gap. The operator's only preclusion tools are binary: an availability override (active / deprecated / stale) or removing a model from the allowlist. The breaker is in-memory, trips only on 5xx, forgets everything on restart, and shows nothing but an on/off toggle.
-
Q1. Carrier/model scorecard. One sortable table per (model, provider) over a window: outcome fail rate (with n), malformed / empty-200 rate, 5xx and timeout count, breaker trips, p95 latency, estimate-vs-bill error, cache rate, cost per success (W4). This is the "who's misbehaving" view the preclusion decision needs.
-
Q2. Same model, different carriers. Group scorecard rows by base model, e.g.
qwen3.6-35bon NeuralWatt vs OpenRouter. This separates "the model is bad at this" from "this carrier serves it badly", which calls for a different fix: drop the carrier's row, or restrict the model. -
Q3. Graded actions in place of deprecate-or-not. Each action records a reason and links its evidence in X1:
- exclude from category X. This cannot reuse
eligible_categories: that field is restrict-only (admin.py, the probe docstring), and NULL means "every category", so a deny needs its own column, - exclude when the request carries tools (the
deepseek-v4-flashcase). This may be the way out of the frozentool_use_agenticdata:min_tool_proficiencycan only read scores that stopped moving on 2026-09-15, while a per-model manual rule lets the operator decide from Q1 scorecard evidence instead, - restrict to batch /
-flexonly, - quarantine: exploration-only, so it keeps collecting evidence without carrying traffic (see the open question below; deferred until Q1 shows it is needed),
- timed ban, e.g. 24h, auto-expiring and logged,
- deprecate (what exists today).
Each action goes through at least the minimal X3 check (would this empty a candidate set for any recent decision?) before committing. That check alone would have caught both the 2026-09-01 and 2026-09-04 incidents; the full winner-shift replay can come later.
- exclude from category X. This cannot reuse
-
Q4. Breaker visibility. Show open circuits with
down_until, current cooldown and trip count; keep a trip history (persist it, since_storeis process-lifetime); add manual force-open (drain a model on purpose) and force-close (reset after a known fix). -
Q5. A quality breaker, separate from availability. Trip on content failures (empty-200, mangled output, a burst of
ok:false) using the novelty-or-rate rule already inrejection_warnings. This is roughly whatplans/mangled-output-detection.mdspecs (status: planned). Keep it a separate breaker and panel so a quality trip doesn't read as an outage. One doc owns it: fold Q5 intomangled-output-detection.mdrather than specifying it twice, and let this file point there. -
Q6. Carrier-level breaker. When several models on one provider trip inside a short window, open the provider rather than walking its models one cooldown at a time. Account exhaustion already has its own path; this is for partial carrier outages.
5. Quality of life: ergonomics and readouts
Unlike the sections above, this one was read against the live portal (8080, view-only, 2026-09-23), plus the frontend source. Each item names the thing observed.
Readouts that mislead today
- R1. Home tiles use different windows. Quota is "this period",
Decisions 7d, Models and Pinch 30d. Busiest models shows
kimi-k2.7-codeat 10,028 while the Decisions tile says 2,887 for everything, and both are correct. Label each tile's window on its face, or drive all of them from the Activity card's 24h / 7d / 30d selector. - R2. The Decisions tile counts verifications, and mixes diagnostics with
ground truth. It reads "2,887 verified in 7 days: 132 ok, 93 failed", but
index.htmlsumsverdict_mix, which countsverificationsrows, not decisions. Checked read-only against the live DB on 2026-09-23:tile actually headline 2,887 2,678 route decisions; 2,887 is verification rows, 2,660 of them unverifiable"ok" 132 96 client succeeded+ 29local_llmok + 7 structural ok"failed" 93 82 client failed+ 11malformed(local_llm + structural)The tile adds structural and local_llmverdicts, which CLAUDE.md callsdiagnostics only, to /outcomereports, the only ground truth. Show routedecisions as the headline, then client outcomes on their own: "178 client reports (6.6%): 96 ok, 82 failed (46%)". Diagnostics, if shown at all, go on a separate line. - R3. The Activity chart smooths across gaps. The decisions series is a spline through sparse points, so it draws continuous traffic across hours that had none. Break the line at empty buckets or use bars. Separately, billed requests exceed decisions at several peaks; the chart should say what the excess is (cloud classifier calls? retries? unrouted pins?) rather than leave two lines that disagree unexplained.
- R4. Sparklines have no values. The tile sparklines carry no axis and no hover. Add a hover value, or min and max labels.
- R5. The Classifier card flashes a wrong value. On first paint the mode
select reads
local_llmand only switches tolocal_encoderonce config loads. For a second, the card shows a value that isn't configured. Render a loading state instead of the default. - R6. Config echo vs live state. The Classifier card shows what the
overlay says (
local_encoder, devicecpu), not what the process actually loaded. Show the resolved model, device and recent p50 latency from the running classifier. (It also exposed doc drift: CLAUDE.md says the live deployment runsdevice: cuda; the overlay sayscpu.) - R7. Proficiency colour encodes evidence, not score. Green / amber mean
measured / thin, so the best model in a column isn't visible without reading
every number. Mark each column's leader, outline the
quality_toleranceband (P3), make columns sortable, and add a "routable only" toggle so rows that are deprecated or not allowlisted stop padding the matrix.
Decisions page ergonomics
- E1. Collapse runs. The top 40 rows are one session: same category, profile, tier, source and model, and only ctx and cost move. Fold consecutive same-session, same-model rows into one expandable row: "38 turns, ctx 75k to 98k, $0.27 total". Switches then stand out as row boundaries, which is what W2 wants to surface.
- E2. Session view. A per-session rollup: turns, total cost, context growth, switches, outcome count. Context climbing about 1k per turn is visible in the raw table and invisible everywhere else.
- E3. Filters in the URL.
decisions.htmlnever readsURLSearchParams, so no other page can link to "decisions for this model" or "this session". X2 actions, the Q1 scorecard and the model modal all need that link. - E4. Row drill-down. A row click does nothing. Open a panel with the
decision's rejected candidates, verification verdict,
/outcomereport, and estimated vs billed cost from its energy observation. - E5. Server-side filtering. The page loads 1,000 of 32,540 rows and filters and searches only those, with a footnote saying so. Push filters to the API so that "All" means all.
- E6. Model, provider and source as filters. Today they're reachable only through free-text search.
Controls page
- C1. One knob table, not two. Runtime Knobs and Persisted Config list
largely the same knobs in two columns, in different orders, with persisted
labels truncated (
objective.incumbent_cache_pri…). Checking drift means matching rows by eye. Use one table: knob, live value, persisted value, layer (base / overlay), drift badge. That table is also X1's natural home. - C2. Separate Restart Service. It sits in the same button row as Refresh Catalog and Seed Energy. Move it apart, and have its confirmation list which runtime values differ from persisted and will revert.
- C3. Prose in cards. Local Compute carries three paragraphs; the portal style rule is no paragraphs. Keep one line plus a docs link or tooltip.
- C4. Default and last change per knob. Show the code default beside each value, plus "changed 2h ago from 0.3" once X1 exists.
Portal-wide
- G1. Nav lives in eight files. Each page hardcodes the same
<ul>of nav links.navbar.jsalready notes that adding Quota was an eight-file edit; have it render the links too. - G2. Width. Most pages sit in
container-xland use well under half of a wide screen, while Decisions, the densest table, is the most cramped. Proficiency already goes full-width. Go fluid on the data-heavy pages. - G3. Dismissals don't stick on rate warnings. Dismissal is keyed on the
warning's exact text, on purpose (a changed condition resurfaces it). But
warnings that embed a live rate ("3.2x pace") change text every poll, so a
dismissal never holds. X2's
class+subjectis the right key; resurface on a severity change, not a digit change. - G4. Keyboard.
/to focus search,Escto close modals, andj/kon the Decisions table. Cheap, and this is a page someone lives in.
One thing the cockpit framing makes more urgent
The portal is loopback-only with no auth, and it already includes
/api/restart-service, /api/apply-feedback and provider deletes. The more
it becomes the place decisions are made, the stronger the pull to reach it
from other machines, and moving the router onto a Proxmox box on the LAN is
exactly that move. Add auth before binding anything but 127.0.0.1. Same
warning CLAUDE.md already gives for the API, with more at stake.
A possible order
| # | item | why here |
|---|---|---|
| 0 | portal auth | only if the router is leaving 127.0.0.1 (the Proxmox move); blocks that move, not this list |
| 1 | X1 change log + operator epoch | everything else reads it |
| 2 | Q1 scorecard + Q4 breaker visibility | the preclusion decision needs one view first |
| 3 | Q3 graded actions (via X1) + minimal X3 (empties-a-candidate-set check) | turns the scorecard into decisions without repeating 09-01 / 09-04 |
| 4 | X2 structured warnings | lets warnings trigger Q3 actions directly |
| 5 | W1 waste ledger + W2 switch reasons | token waste in dollars, with causes |
| 6 | X3 full replay preview (winner shift, cost delta) | makes knob changes safe to try |
| 7 | P1 drift, P2 evidence priority | proficiency monitoring over time |
| 8 | Q5 quality breaker (owned by mangled-output-detection.md), D1 timed changes |
automate what the operator has been doing by hand |
| 9 | D3 session A/B, Q6 carrier breaker | only after the above prove out |
Section 5 sits outside this order: most of it is small and independent. Quick wins that also unblock the numbered items: E3 (URL filters, needed by X2 and Q1), R2 (outcome coverage, the same honesty W1 and W4 need), C1 (one knob table, X1's home), G3 (fixed by X2's class keys). R5 and G1 are fixes of a few lines each.
North Star #1 applies to every item here. Q5 thresholds, D1 durations,
timed-ban lengths and any quarantine settings are new config knobs, so each
ships with its admin control or test_admin_knob_coverage.py fails. Scope the
control into the item, not a follow-up.
Open questions, with proposed answers
- Is the change log append-only history, or also an undo stack? Proposed: append-only, plus a per-row "revert to old value" that writes one key and appends its own row. That is not the shape CLAUDE.md rejected for gaming mode: that concern was one flag writing five keys and drifting apart. A one-key revert can't drift, and the history stays intact.
- Should graded exclusions live in the DB or in
config.local.yaml? Proposed: the DB, next toadmin_model_overrides(extend it or add a sibling table). Timed bans need an expiry, and an expiry belongs in a DB column. The overlay is per-machine and not in git, so it's the wrong home for judgments about a carrier. And availability overrides already live in the DB. The category exclusion still needs a new deny column (see Q3.1). - Does quarantine conflict with session-scoped exploration? Yes, and the
weight goes up: a quarantined model would carry whole sessions, not single
turns. There's a second problem too.
exploration.pyonly chooses among models that already passed the hard filters and were ranked, so quarantine needs a new state, "eligible to explore, barred from winning", rather than reusing anything existing. Defer it until Q1 shows it's needed.