Files
6krrt/plans/cockpit-brainstorm.md
adlee-was-taken 69c969d104 plans: declare a valid Status on the six plans this branch adds
tests/test_plans_declare_status.py requires 'Status: <done|planned|in
progress|parked|reference> -- <reason>' in the first 8 lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-26 21:46:15 -04:00

22 KiB
Raw Permalink Blame History

Cockpit brainstorm: the router as an operator's instrument panel

Status: reference -- brainstorm, not a queue item; its six quick wins shipped in PR #101

Date: 2026-09-23 Read against: the repo tarball as of this date (code, config, plans, screenshots). Not read against the live router.db, so every "gap" below is a gap in what the code surfaces, not a claim about what the data shows. Section 5 is the exception: it was read against the live portal.

The framing

6krrt as a local LLM router with a cockpit. The operator:

  1. adjusts dials and knobs in flight,
  2. sees where tokens are wasted,
  3. watches proficiency, and
  4. sees poor results from a carrier or model surface, so they can decide whether to preclude it (with the circuit breaker as the automatic version of that decision for outages).

The router makes the per-request decision. The cockpit is where the operator makes the slower decisions: which models and carriers are trusted, for what, and at what settings. Under this framing, the question for every feature is whether it helps the operator notice something and act on it.


What the cockpit already has

job what exists where
in-flight dials ~11 runtime knobs (bool/float/int tables in admin.py), persisted config writes via /api/config/{key}, classifier card, profiles CRUD, gaming mode Controls, Profiles
knob coverage North Star rule #1 + test_admin_knob_coverage.py tests
token waste pinch savings card; cache_rate_series + warnings; incumbent cache pricing dial; cost_estimate_calibration in /metrics Dashboard, Controls, /metrics
proficiency model × category matrix with evidence status (measured / inherited / thin, n per cell); client outcome log with attributable + applied flags Proficiency
poor results content_fault_warnings (malformed output rate); rejection_warnings; availability override (active / deprecated / stale); allowlist editor warnings bell, Models, Providers
breaker passive per-(model, provider) breaker on 5xx; on/off knob Controls (toggle only)

The foundation is good. What's mostly missing is the link between a signal and the action it should prompt, plus any record of which actions were taken.


Cross-cutting ideas (these help all four jobs)

X1. Flight recorder: a change log

Gap. admin_schema.sql has only admin_model_overrides. Nothing records when a knob, override, allowlist entry, profile or feedback fold changed, what it changed from, or why. Runtime knobs are "NEVER persisted", so a restart silently reverts them, and nothing notes that the revert happened.

Idea.

  • An admin_changes table: (at, surface, key, old, new, reason, source), where source is runtime | persisted | override | allowlist | profile | feedback_fold | restart. Every admin write path appends a row.
  • An operator_epoch id stamped on each route_decisions row, incremented on each operator action (a row in admin_changes). Every metric can then be split before and after a change without joining on timestamps. Scope it to operator actions on purpose: "any change that affects routing" would also cover catalog price polls and every feedback fold, so the epoch would tick constantly and split nothing.
  • Change markers drawn on the dashboard history charts.
  • A drift badge when a runtime value differs from its persisted twin ("this will revert on restart").
  • Restart reverts are loggable without persisting runtime knobs: the change log already holds each knob's last runtime value, so at startup, compare it with the loaded config and write one restart row per value discarded.

Why first. Adjusting knobs in flight produces no learning if you can't later tell what a change did. The other ideas here lean on this one.

X2. Structured warnings with actions attached

Gap. /metrics warnings are plain strings. The dashboard decides which page fixes a warning, and how severe it is, by regex on the text (warningTarget and warningSeverity in index.html). Rewording a warning silently breaks its link and its severity.

Idea. Emit {class, severity, subject: {model, provider, category?}, evidence_url, actions: [...]}. The warning registry in test_tui_warnings.py already lists every class, so it can serve as the schema. The bell can then offer an action in place: "exclude from tool_use_agentic", "quarantine 24h", "open decisions filtered to this model".

X3. Preview before apply (replay)

Gap. The reviews replayed thousands of real decisions to test a setting, but only as one-off scripts pasted into plan documents.

Idea. For any change that affects ranking (quality_tolerance, the incumbent dial, profile edits, availability overrides, graded exclusions from P2 below), the portal replays the last N decisions through select_candidates and rank_candidates with the proposed value, then shows:

  • a winner-shift matrix: old winner → new winner, with counts,
  • estimated cost delta,
  • decisions that would become unroutable (the 2026-09-01 and 2026-09-04 incidents were both deprecations that emptied a candidate set).

This is cheap because the ranking modules are pure. Its output is an estimate of routing change, not quality change, and should be labeled that way.


1. Dials in flight

  • D1. Timed changes. "Apply for 2h, then revert." Turns a knob change into a bounded experiment and lowers the risk of forgetting a test setting. Needs X1 to record both the apply and the revert.
  • D2. Blast radius on hover. How many of the last 24h of decisions this knob would have touched. It's the X3 replay reduced to one number.
  • D3. Session-scoped A/B. Assign new sessions to setting A or B and compare cost and outcome rate per arm. It has to be per session, not per turn, because a mid-session switch costs cache (token-waste Wave 2 measured 0.919 → 0.348). Larger build; only worth it once X1 exists and outcome volume can support a comparison.
  • D4. Grouping by effect. Group controls by what they move (cost, quality, latency, safety, measurement-only) rather than by config section, and mark the few that actually change routing. Knob coverage guarantees every control exists; this makes them usable.

2. Token waste

Gap. Waste signals are spread across pinch, cache rate, calibration and the verifications table, and none is in dollars in one place. cost_estimate_calibration is in /metrics but not in the portal.

  • W1. Waste ledger. One panel with dollars per waste class over a window, each with a trend:

    class how it's computed
    switch cache loss billed on switch turns − same prompt at the session's same-model cache rate
    retries re-billed prompt tokens from iteration.py attempts
    paid for failure billed cost of responses later reported ok:false via /outcome
    empty-200 / malformed billed cost of responses with a structural malformed verdict
    truncation responses stopped by finish_reason: length with no client cap
    pinch (negative) dollars saved, from pinch_summary

    "Paid for failure" counts only decisions that received an outcome, so it shows n and outcome coverage (the share of billed decisions with a report) next to the dollar figure.

  • W2. Why did it switch? Record a switch reason on each decision: category changed, breaker open, exploration, override removed the incumbent, profile change. Then rank sessions by switch cost and drill into the turns. Without a reason, a switch is only a cost; with one, it points to a knob.

  • W3. Calibration panel. Surface cost_estimate_calibration per model, headlining the spread (as its docstring argues), and flag pairs where the estimator and the bill disagree on order. Where that happens, the cost tiebreak picks the wrong model.

  • W4. Cost per successful outcome. Per (model, provider, category): billed dollars ÷ ok:true outcomes. A cheap model that fails often isn't cheap. This is the one waste figure that includes quality.

    Show n and outcome coverage next to it. Models that carry more traffic collect more outcomes (the same exposure bias the feedback fold had to correct), and if coverage stays hidden, a thinly reported cheap model will look better than it is.

3. Proficiency monitoring

Gap. The matrix shows current state well. It doesn't show movement, and it doesn't show which thin cells actually matter.

  • P1. Trend and drift per cell. A sparkline of the outcome rate over time. Compare a recent window against the long-run rate with an interval, and flag drops that fall outside it. Carriers change serving setups (quant, engine, hardware) without notice, and a drift flag is how that shows up.
  • P2. Evidence priority. Rank cells by traffic share × uncertainty. A thin cell that routing never consults doesn't matter; a thin cell deciding 30% of traffic does. The ranking tells you where to spend eval_proficiency runs or exploration budget.
  • P3. The tolerance band per category. Show which models sit within quality_tolerance of the leader, so it's clear where cost is deciding and where quality is.
  • P4. Label provenance per cell. The share of each cell's outcomes whose category came from a fresh classification versus a cached or borrowed one. North Star rule #2 exists because borrowed labels trained the matrix; this makes it visible cell by cell.

4. Poor results → operator preclusion (and the breaker)

Gap. The operator's only preclusion tools are binary: an availability override (active / deprecated / stale) or removing a model from the allowlist. The breaker is in-memory, trips only on 5xx, forgets everything on restart, and shows nothing but an on/off toggle.

  • Q1. Carrier/model scorecard. One sortable table per (model, provider) over a window: outcome fail rate (with n), malformed / empty-200 rate, 5xx and timeout count, breaker trips, p95 latency, estimate-vs-bill error, cache rate, cost per success (W4). This is the "who's misbehaving" view the preclusion decision needs.

  • Q2. Same model, different carriers. Group scorecard rows by base model, e.g. qwen3.6-35b on NeuralWatt vs OpenRouter. This separates "the model is bad at this" from "this carrier serves it badly", which calls for a different fix: drop the carrier's row, or restrict the model.

  • Q3. Graded actions in place of deprecate-or-not. Each action records a reason and links its evidence in X1:

    1. exclude from category X. This cannot reuse eligible_categories: that field is restrict-only (admin.py, the probe docstring), and NULL means "every category", so a deny needs its own column,
    2. exclude when the request carries tools (the deepseek-v4-flash case). This may be the way out of the frozen tool_use_agentic data: min_tool_proficiency can only read scores that stopped moving on 2026-09-15, while a per-model manual rule lets the operator decide from Q1 scorecard evidence instead,
    3. restrict to batch / -flex only,
    4. quarantine: exploration-only, so it keeps collecting evidence without carrying traffic (see the open question below; deferred until Q1 shows it is needed),
    5. timed ban, e.g. 24h, auto-expiring and logged,
    6. deprecate (what exists today).

    Each action goes through at least the minimal X3 check (would this empty a candidate set for any recent decision?) before committing. That check alone would have caught both the 2026-09-01 and 2026-09-04 incidents; the full winner-shift replay can come later.

  • Q4. Breaker visibility. Show open circuits with down_until, current cooldown and trip count; keep a trip history (persist it, since _store is process-lifetime); add manual force-open (drain a model on purpose) and force-close (reset after a known fix).

  • Q5. A quality breaker, separate from availability. Trip on content failures (empty-200, mangled output, a burst of ok:false) using the novelty-or-rate rule already in rejection_warnings. This is roughly what plans/mangled-output-detection.md specs (status: planned). Keep it a separate breaker and panel so a quality trip doesn't read as an outage. One doc owns it: fold Q5 into mangled-output-detection.md rather than specifying it twice, and let this file point there.

  • Q6. Carrier-level breaker. When several models on one provider trip inside a short window, open the provider rather than walking its models one cooldown at a time. Account exhaustion already has its own path; this is for partial carrier outages.


5. Quality of life: ergonomics and readouts

Unlike the sections above, this one was read against the live portal (8080, view-only, 2026-09-23), plus the frontend source. Each item names the thing observed.

Readouts that mislead today

  • R1. Home tiles use different windows. Quota is "this period", Decisions 7d, Models and Pinch 30d. Busiest models shows kimi-k2.7-code at 10,028 while the Decisions tile says 2,887 for everything, and both are correct. Label each tile's window on its face, or drive all of them from the Activity card's 24h / 7d / 30d selector.
  • R2. The Decisions tile counts verifications, and mixes diagnostics with ground truth. It reads "2,887 verified in 7 days: 132 ok, 93 failed", but index.html sums verdict_mix, which counts verifications rows, not decisions. Checked read-only against the live DB on 2026-09-23:
    tile actually
    headline 2,887 2,678 route decisions; 2,887 is verification rows, 2,660 of them unverifiable
    "ok" 132 96 client succeeded + 29 local_llm ok + 7 structural ok
    "failed" 93 82 client failed + 11 malformed (local_llm + structural)
    The tile adds structural and local_llm verdicts, which CLAUDE.md calls
    diagnostics only, to /outcome reports, the only ground truth. Show route
    decisions as the headline, then client outcomes on their own: "178 client
    reports (6.6%): 96 ok, 82 failed (46%)". Diagnostics, if shown at all, go on
    a separate line.
  • R3. The Activity chart smooths across gaps. The decisions series is a spline through sparse points, so it draws continuous traffic across hours that had none. Break the line at empty buckets or use bars. Separately, billed requests exceed decisions at several peaks; the chart should say what the excess is (cloud classifier calls? retries? unrouted pins?) rather than leave two lines that disagree unexplained.
  • R4. Sparklines have no values. The tile sparklines carry no axis and no hover. Add a hover value, or min and max labels.
  • R5. The Classifier card flashes a wrong value. On first paint the mode select reads local_llm and only switches to local_encoder once config loads. For a second, the card shows a value that isn't configured. Render a loading state instead of the default.
  • R6. Config echo vs live state. The Classifier card shows what the overlay says (local_encoder, device cpu), not what the process actually loaded. Show the resolved model, device and recent p50 latency from the running classifier. (It also exposed doc drift: CLAUDE.md says the live deployment runs device: cuda; the overlay says cpu.)
  • R7. Proficiency colour encodes evidence, not score. Green / amber mean measured / thin, so the best model in a column isn't visible without reading every number. Mark each column's leader, outline the quality_tolerance band (P3), make columns sortable, and add a "routable only" toggle so rows that are deprecated or not allowlisted stop padding the matrix.

Decisions page ergonomics

  • E1. Collapse runs. The top 40 rows are one session: same category, profile, tier, source and model, and only ctx and cost move. Fold consecutive same-session, same-model rows into one expandable row: "38 turns, ctx 75k to 98k, $0.27 total". Switches then stand out as row boundaries, which is what W2 wants to surface.
  • E2. Session view. A per-session rollup: turns, total cost, context growth, switches, outcome count. Context climbing about 1k per turn is visible in the raw table and invisible everywhere else.
  • E3. Filters in the URL. decisions.html never reads URLSearchParams, so no other page can link to "decisions for this model" or "this session". X2 actions, the Q1 scorecard and the model modal all need that link.
  • E4. Row drill-down. A row click does nothing. Open a panel with the decision's rejected candidates, verification verdict, /outcome report, and estimated vs billed cost from its energy observation.
  • E5. Server-side filtering. The page loads 1,000 of 32,540 rows and filters and searches only those, with a footnote saying so. Push filters to the API so that "All" means all.
  • E6. Model, provider and source as filters. Today they're reachable only through free-text search.

Controls page

  • C1. One knob table, not two. Runtime Knobs and Persisted Config list largely the same knobs in two columns, in different orders, with persisted labels truncated (objective.incumbent_cache_pri…). Checking drift means matching rows by eye. Use one table: knob, live value, persisted value, layer (base / overlay), drift badge. That table is also X1's natural home.
  • C2. Separate Restart Service. It sits in the same button row as Refresh Catalog and Seed Energy. Move it apart, and have its confirmation list which runtime values differ from persisted and will revert.
  • C3. Prose in cards. Local Compute carries three paragraphs; the portal style rule is no paragraphs. Keep one line plus a docs link or tooltip.
  • C4. Default and last change per knob. Show the code default beside each value, plus "changed 2h ago from 0.3" once X1 exists.

Portal-wide

  • G1. Nav lives in eight files. Each page hardcodes the same <ul> of nav links. navbar.js already notes that adding Quota was an eight-file edit; have it render the links too.
  • G2. Width. Most pages sit in container-xl and use well under half of a wide screen, while Decisions, the densest table, is the most cramped. Proficiency already goes full-width. Go fluid on the data-heavy pages.
  • G3. Dismissals don't stick on rate warnings. Dismissal is keyed on the warning's exact text, on purpose (a changed condition resurfaces it). But warnings that embed a live rate ("3.2x pace") change text every poll, so a dismissal never holds. X2's class + subject is the right key; resurface on a severity change, not a digit change.
  • G4. Keyboard. / to focus search, Esc to close modals, and j/k on the Decisions table. Cheap, and this is a page someone lives in.

One thing the cockpit framing makes more urgent

The portal is loopback-only with no auth, and it already includes /api/restart-service, /api/apply-feedback and provider deletes. The more it becomes the place decisions are made, the stronger the pull to reach it from other machines, and moving the router onto a Proxmox box on the LAN is exactly that move. Add auth before binding anything but 127.0.0.1. Same warning CLAUDE.md already gives for the API, with more at stake.


A possible order

# item why here
0 portal auth only if the router is leaving 127.0.0.1 (the Proxmox move); blocks that move, not this list
1 X1 change log + operator epoch everything else reads it
2 Q1 scorecard + Q4 breaker visibility the preclusion decision needs one view first
3 Q3 graded actions (via X1) + minimal X3 (empties-a-candidate-set check) turns the scorecard into decisions without repeating 09-01 / 09-04
4 X2 structured warnings lets warnings trigger Q3 actions directly
5 W1 waste ledger + W2 switch reasons token waste in dollars, with causes
6 X3 full replay preview (winner shift, cost delta) makes knob changes safe to try
7 P1 drift, P2 evidence priority proficiency monitoring over time
8 Q5 quality breaker (owned by mangled-output-detection.md), D1 timed changes automate what the operator has been doing by hand
9 D3 session A/B, Q6 carrier breaker only after the above prove out

Section 5 sits outside this order: most of it is small and independent. Quick wins that also unblock the numbered items: E3 (URL filters, needed by X2 and Q1), R2 (outcome coverage, the same honesty W1 and W4 need), C1 (one knob table, X1's home), G3 (fixed by X2's class keys). R5 and G1 are fixes of a few lines each.

North Star #1 applies to every item here. Q5 thresholds, D1 durations, timed-ban lengths and any quarantine settings are new config knobs, so each ships with its admin control or test_admin_knob_coverage.py fails. Scope the control into the item, not a follow-up.

Open questions, with proposed answers

  • Is the change log append-only history, or also an undo stack? Proposed: append-only, plus a per-row "revert to old value" that writes one key and appends its own row. That is not the shape CLAUDE.md rejected for gaming mode: that concern was one flag writing five keys and drifting apart. A one-key revert can't drift, and the history stays intact.
  • Should graded exclusions live in the DB or in config.local.yaml? Proposed: the DB, next to admin_model_overrides (extend it or add a sibling table). Timed bans need an expiry, and an expiry belongs in a DB column. The overlay is per-machine and not in git, so it's the wrong home for judgments about a carrier. And availability overrides already live in the DB. The category exclusion still needs a new deny column (see Q3.1).
  • Does quarantine conflict with session-scoped exploration? Yes, and the weight goes up: a quarantined model would carry whole sessions, not single turns. There's a second problem too. exploration.py only chooses among models that already passed the hard filters and were ranked, so quarantine needs a new state, "eligible to explore, barred from winning", rather than reusing anything existing. Defer it until Q1 shows it's needed.