diff --git a/CLAUDE.md b/CLAUDE.md
index b2bec0a..2d7ce33 100644
--- a/CLAUDE.md
+++ b/CLAUDE.md
@@ -8,7 +8,7 @@ next steps, and is the one to trust on what is currently true.
## NORTH STAR GUIDELINES
-Three rules that outrank local cleverness. Each exists because it was broken
+Four rules that outrank local cleverness. Each exists because it was broken
first and the breakage was expensive to find. When a change conflicts with one
of these, the change is wrong — not the rule.
@@ -87,6 +87,36 @@ before measuring anything, and open it read-only:
`sqlite3.connect("file:...?mode=ro", uri=True)` — the `sqlite3` CLI here does
not accept `-uri`.
+### 4. Waste is surfaced in the admin portal, and stopping it is one click
+
+When the router can see money being wasted, it shows the operator where they
+already look, with the evidence and the lever to stop it side by side. A
+detector that only writes a log line, or a fix that needs a config edit, a
+restart or a long table scan, has not met this rule.
+
+"Waste" here means **spend with no concrete change landing**: an agent session
+looping, re-reading, or retrying the same failure, and a model that keeps
+producing such sessions. It does **not** mean steady spend. A healthy agent run
+can burn for hours, and a spend-rate alarm cannot tell the two apart.
+
+In practice:
+- **Stalled sessions are visible** in the portal with their evidence: turns and
+ $ since the last landed change, and the top repeated target. They also reach
+ the operator when nobody is watching (desktop alert).
+- **A model that keeps producing them is visible** as a per-model rollup, with
+ **Block model** beside the evidence. The block records its reason, shows in a
+ Blocked list, and is one click to undo.
+- **Automatic responses come after the visible one,** never instead of it, and
+ every automatic action shows up in the same place.
+
+**Why:** incident #8 (`docs/incidents.md`). Agent sessions looped for hours on
+2026-09-25/26: one worker read the same file 61 times, a planner re-read a spec
+68 times its length, and a model confabulated truncation that was not there.
+Every existing check stayed green. It was caught only by a human, or a Claude
+session, reading opencode's session store by hand. Pulling the model took a trip
+through a dropdown whose vocabulary is catalog `deprecated`. The detection
+existed nowhere, and the lever existed only for someone who already knew.
+
## What this is
A router that uses a local model (served via Ollama) to classify incoming
@@ -319,7 +349,12 @@ rather than from months of history.
- `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md).
- `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below.
- `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted.
-- `tests/` — 1155 tests across 40+ files, offline, verified on Python 3.10 and 3.14. [README](README.md).
+- `progress_detect.py` — loop-detection signals over a window of per-session probe calls: duplicate-bulk (`dup_min`), top-similarity (`top_min`/`top_min_ro`), slow-progress (`cum_min`) and coverage (`cover_min`) heuristics, gated on `min_calls`. See [watchdog](docs/watchdog.md).
+- `watchdog.py` — the per-session watchdog loop: every ~5 minutes it judges each session with a tool call since the last tick, on its full history, and writes a quiet `no_opencode` tick when opencode is not running. See [watchdog](docs/watchdog.md).
+- `notifier.py` — alert fan-out: desktop/`notify-send` plus per-channel `min_severity` and a rate limit. See [watchdog](docs/watchdog.md).
+- `watchdog_store.py` — `watchdog_ticks`, `watchdog_verdicts`, `watchdog_alerts`, `watchdog_channel_settings` (four tables + indexes). See [watchdog](docs/watchdog.md).
+- `router-link.js` — the opencode plugin; exported as a factory with `parentCache` as a property, because opencode 1.18.x rejects the whole plugin when any export is not a function. See [watchdog](docs/watchdog.md).
+- `tests/` — 2414 tests across 108 files, offline, verified on Python 3.10 and 3.14. [README](README.md).
## #45 — OpenRouter is an opt-in allowlist provider
@@ -1261,10 +1296,14 @@ be on.
### When the router goes unreachable, start at docs/incidents.md
-Four incidents so far, all sharing one shape: a change that looked local to the
-router silently degraded the agent depending on it, and none announced itself as
-a router problem. **`docs/incidents.md` carries the full write-ups plus a
-symptom -> one-line-check table**; read it rather than re-deriving a diagnosis.
+Eight incidents so far, nearly all sharing one shape: a change that looked local
+to the router silently degraded the agent depending on it, and none announced
+itself as a router problem. **`docs/incidents.md` carries the full write-ups plus
+a symptom -> one-line-check table**; read it rather than re-deriving a diagnosis.
+#8 is the exception worth knowing before an unattended agent run: the router
+worked perfectly while agent workers looped for hours with no progress, and no
+check noticed, because every check watched spend or availability rather than
+whether changes landed (`plans/no-progress-detection.md`).
Two conventions from those incidents that bind every session, and so stay here:
@@ -1283,6 +1322,14 @@ Recovery for an unreachable-but-`active` service is
`systemctl --user restart llm-router.service` -- a hung process was never in a
tracked stop job, so this issues a fresh cycle systemd does enforce a timeout on.
+The watchdog runs as its own systemd **timer** (`llm-router-watchdog.timer`),
+separate from the dispatcher, so a stuck router cannot silence the thing meant
+to notice it is stuck. Install and enable it with `systemctl --user enable
+--now llm-router-watchdog.timer` (the timer's unit file ships pointing at a
+placeholder home path and must be `sed`-repointed to the real one first), or
+run it once by hand with the `--once` flag. See [watchdog](docs/watchdog.md)
+for the signals it fires on, the alert lifecycle, and the known limits.
+
## Pointing a coding agent at it
The `/v1` endpoints are OpenAI-compatible, so any normal client works —
diff --git a/docs/admin-portal.md b/docs/admin-portal.md
index 2a4f632..b83d691 100644
--- a/docs/admin-portal.md
+++ b/docs/admin-portal.md
@@ -35,11 +35,30 @@ for the full model.
+**Loops** — the dashboard carries a Loops card (`admin/frontend/index.html#loops`)
+showing the watchdog's open alerts, one row per alert. It pulls from
+`GET /admin/api/watchdog/loops` (`watchdog_alerts` rows where `resolved_at` is
+null) via `loadLoops()`. Each row shows exactly three things: a severity badge
+(`critical` / `warning` / else `info`), the raw dedup key
+(`opencode-loop:`), and `opened_at`. That is all the panel
+shows today — it does **not** show the agent, model, top repeated target, calls,
+or $ since the last landed change; those live in `watchdog_verdicts` and have
+not shipped to this panel yet.
+
**Models** — `GET /admin/models` is the model-availability table. Each row shows
the serving class, tier, status, and an override dropdown (`active` /
-`deprecated` / `stale`) that writes through to the routing hard filters. By
-default the table hides deprecated and stale rows; enable the **Show deprecated
-/ stale** toggle to include them.
+`blocked` / `deprecated` / `stale`) that writes through to the routing hard
+filters. By default the table hides deprecated and stale rows; enable the
+**Show deprecated / stale** toggle to include them.
+
+`blocked` is a fourth value in that same per-model override dropdown, distinct
+from the catalog meaning of `deprecated`. It is an operator stop: a model an
+operator has switched off, not one the catalog flagged. Routing and pinned
+requests both exclude a blocked model (`_admin_excluded_models` in
+`dispatcher.py` and `metrics.py`). Undo is setting the dropdown back to
+`active` (clearing the override via `DELETE` on the availability path), not a
+separate unlock — there is no per-model Block button beside evidence and no
+Blocked list with reasons; those belong to the "not built yet" item below.
@@ -126,8 +145,8 @@ land in `config/config.local.yaml`; see
[config-local-overlay](config-local-overlay.md)). Changes are marked as dirty
and written only on save.
-It carries two dedicated cards, neither a row in the generic runtime-knob
-list, because both have cross-field structure a flat scalar/boolean input
+It carries three dedicated cards, none a row in the generic runtime-knob
+list, because each has cross-field structure a flat scalar/boolean input
can't safely represent:
- **Local Compute** — "gaming mode". See [Gaming mode](#gaming-mode) below.
@@ -145,6 +164,14 @@ can't safely represent:
anything reaches `config.local.yaml`, and mode plus its companion block are
written as one atomic change so an in-between invalid state is never even
written transiently.
+- **Watchdog** — the watchdog's live state (`controls.html#watchdog-card`,
+ `loadWatchdog()`): the last tick time (and a `no tick` badge before the
+ first one), how many sessions the last tick saw, how many verdicts it
+ flagged, and how many alerts are currently open. Below those, a set of
+ **notification channel** toggles (each channel with an enable/disable and a
+ minimum severity, persisted via `POST /admin/api/watchdog/channels`), a
+ **Refresh** button, and a **Send test alert** button (`POST
+ /admin/api/watchdog/test-alert`).
@@ -444,7 +471,15 @@ intermediate invalid state could land on disk. The GET response includes
**Model availability overrides**
`POST /admin/api/models/{model_id:path}/{provider}/availability` marks a model
-as `active`, `deprecated`, or `stale`. `DELETE` on the same path removes the
-override. The `{model_id:path}` converter accepts model ids that contain
-slashes, so providers like OpenRouter with slash-bearing ids are handled the
-same way as plain ids. Deprecation feeds into the routing hard filters.
+as `active`, `blocked`, `deprecated`, or `stale`. `DELETE` on the same path
+removes the override. The `{model_id:path}` converter accepts model ids that
+contain slashes, so providers like OpenRouter with slash-bearing ids are
+handled the same way as plain ids. Deprecation feeds into the routing hard
+filters.
+
+**Not built yet.** The Loops panel does not yet show the evidence columns an
+operator would act on — agent, model, top repeated target, calls, and $ since
+the last landed change — and there is no per-model stall rollup, no **Block
+model** button beside the evidence, and no Blocked list with reasons. Those are
+the remaining surface of [North Star #4](../CLAUDE.md) in CLAUDE.md, whose first
+half (conversation identity and the watchdog) is shipped.
diff --git a/docs/api.md b/docs/api.md
index 6309f58..207f6b8 100644
--- a/docs/api.md
+++ b/docs/api.md
@@ -173,6 +173,12 @@ These ids are what make `/outcome` attribution exact (below): an
`X-Router-Conversation` on the request sets the `session_key` row that a later
report can address by `conversation_id` without guessing `source` or ambiguity.
+**Plugin loader contract:** the opencode plugin that stamps these headers
+must export only functions — the opencode 1.18.x plugin loader rejects a
+plugin outright if any export is not a function (so the plugin's
+`parentCache` lives as a property on the factory rather than a named export).
+See [clients.md](clients.md) for the full contract.
+
**Capability 422s**: When no model survives the hard filters, the 422 names the
active constraints. That now includes "vision-capable model" or
"json-mode-capable model" when the request carried images or a JSON-mode
@@ -225,6 +231,65 @@ but a `true` report counts too. It's also the only quality signal that survives
streaming, since a retry can't reach a response whose bytes are already gone,
while a report arrives afterward and works either way.
+## Admin watchdog and availability API
+
+The admin portal's watchdog and model-availability operations sit under
+`/admin` (loopback-only, like the rest of the admin API). The watchdog
+endpoints back the dashboard's **Loops** card and the **Watchdog** control
+card; full behaviour is described in [watchdog.md](watchdog.md) and
+[admin-portal.md](admin-portal.md).
+
+| Method | Path | Description |
+|---|---|---|
+| `GET` | `/admin/api/watchdog/status` | Latest watchdog tick plus open-alert count |
+| `GET` | `/admin/api/watchdog/loops` | Open alert rows (one per looping session) |
+| `GET` | `/admin/api/watchdog/channels` | Notification-channel settings |
+| `POST` | `/admin/api/watchdog/channels` | Create or update a channel's enabled/min-severity |
+| `POST` | `/admin/api/watchdog/test-alert` | Deliver a test alert through the enabled channels |
+
+`GET /admin/api/watchdog/status` returns a single object:
+
+- `last_tick` — the most recent `watchdog_ticks` row, or `null` before the
+ first tick
+- `verdicts` — `{total, flagged}` for that tick's `watchdog_verdicts`, or
+ `null` when there is no tick yet
+- `open_alerts` — count of `watchdog_alerts` rows with `resolved_at` null
+
+`GET /admin/api/watchdog/loops` returns an array of `watchdog_alerts` rows
+where `resolved_at` is null, ordered by `opened_at` descending. Each element
+is the full alert row.
+
+`GET /admin/api/watchdog/channels` returns an array of
+`watchdog_channel_settings` rows ordered by `channel_name`.
+
+`POST /admin/api/watchdog/channels` upserts one channel. Body:
+
+- `channel_name` (required) — the channel to create or update; absence is a
+ 422
+- `enabled` (optional boolean) — when omitted the existing value is kept
+- `min_severity` (optional: `info`, `warning`, or `critical`) — any other
+ value is a 422
+
+Returns `{ok: true}`.
+
+`POST /admin/api/watchdog/test-alert` delivers a test alert through the
+currently enabled channels. Body:
+
+- `severity` (optional, default `warning`; `info`, `warning`, or `critical`)
+- `channel_name` (optional) — restrict to a single channel
+
+Returns `{sent: , severity: }`, or `{sent: 0, message: "no
+enabled channels"}` when nothing is enabled. An invalid `severity` is a 422.
+
+**Model availability** — `POST
+/admin/api/models/{model_id:path}/{provider}/availability` upserts an admin
+override for a model's availability. `availability` must be one of `active`,
+`blocked`, `deprecated`, or `stale` (anything else is a 422); `blocked` is an
+operator stop that routing and pinned requests exclude
+(`_admin_excluded_models`), distinct from the catalog meaning of `deprecated`.
+`DELETE` on the same path removes the override. The `{model_id:path}`
+converter accepts slash-bearing model ids (e.g. OpenRouter).
+
## Pinning and auto behavior for local models
`model: "auto"` will route eligible tasks to `qwen2.5-coder-router:14b` when the
diff --git a/docs/clients.md b/docs/clients.md
index dff01d9..7942d18 100644
--- a/docs/clients.md
+++ b/docs/clients.md
@@ -37,3 +37,11 @@ This resolved a real incident where reading a dependency's source 24 times
outweighed writing to the project dir 22 times: the write-like 3× weighting
puts the project directory that actually received the edits ahead of a
dependency's source that was merely read.
+
+**Plugin loader contract:** opencode 1.18.x (the observed version) rejects the
+entire plugin if *any* export is not a function. The `parentCache` map on
+`router-link.js` is therefore set as a property on the `RouterLink` factory
+function rather than as a named export (L221-224 of
+`deploy/opencode-plugin/router-link.js`). Confirm the plugin loaded by
+searching the log for `failed to load plugin`: `grep 'failed to load plugin'
+~/.local/share/opencode/log/opencode.log`.
diff --git a/docs/data-model.md b/docs/data-model.md
index 5634090..438f149 100644
--- a/docs/data-model.md
+++ b/docs/data-model.md
@@ -31,7 +31,7 @@ Three data tables plus one observability table, `PRAGMA foreign_keys = ON`:
| `access_level` | TEXT | `public` \| `preview` \| `canary` |
| `pricing_tbd` | INTEGER | |
| `deprecated` | INTEGER | |
-| `availability` | TEXT | `active` \| `deprecated` \| `stale` |
+| `availability` | TEXT | `active` \| `deprecated` \| `stale` \| `blocked` — the `blocked` value is an operator stop set via the admin override (see [admin-portal.md](admin-portal.md)); it is excluded from routing and pinned requests |
| `last_updated` | TEXT | ISO8601 |
**Serving class:** Neuralwatt ships ~6 base models as 19 catalog rows. The id
@@ -326,3 +326,71 @@ prompts, and answers are deliberately excluded. A test enforces that the write
path does not store prompt or answer text. The write is gated by
`logging.log_route_decisions` and is **best-effort**: a failed write is logged
at warning and swallowed so monitoring cannot slow or fail a request.
+
+## Watchdog Schema (SQLite)
+
+Four tables written by `watchdog.py` via `src/watchdog_store.py` (the code-side
+inline-create mirrors this schema for live databases). See
+[watchdog.md](watchdog.md) for how the loop reads them.
+
+### `watchdog_ticks` — one row per detection run
+
+| Column | Type | Notes |
+|---|---|---|
+| `id` | INTEGER | Autoincrement |
+| `ticked_at` | TEXT | ISO8601 run time |
+| `sessions_seen` | INTEGER | Count of sessions examined this run, default 0 |
+| `outcome` | TEXT | |
+
+### `watchdog_verdicts` — one row per session per tick
+
+| Column | Type | Notes |
+|---|---|---|
+| `id` | INTEGER | Autoincrement |
+| `tick_id` | INTEGER | FK → `watchdog_ticks.id`, NOT NULL |
+| `session_id` | TEXT | NOT NULL |
+| `session_root` | TEXT | |
+| `agent` | TEXT | |
+| `model_id` | TEXT | |
+| `provider` | TEXT | |
+| `flagged` | INTEGER | 0/1, default 0 |
+| `dup` | REAL | Duplicate-fraction signal |
+| `top` | INTEGER | |
+| `top_what` | TEXT | |
+| `landed` | INTEGER | |
+| `slow` | INTEGER | |
+| `coverage` | REAL | |
+| `calls_since_landed` | INTEGER | |
+| `cost_since_landed_usd` | REAL | |
+| `llm_second_opinion` | TEXT | |
+| `created_at` | TEXT | NOT NULL |
+
+Four indexes (`config/schema.sql` L482-485):
+
+| Index | On |
+|---|---|
+| `idx_watchdog_verdicts_model_id` | `watchdog_verdicts(model_id)` |
+| `idx_watchdog_verdicts_created_at` | `watchdog_verdicts(created_at)` |
+| `idx_watchdog_verdicts_flagged` | `watchdog_verdicts(flagged)` |
+| `idx_watchdog_verdicts_session_root` | `watchdog_verdicts(session_root)` |
+
+### `watchdog_alerts` — one row per open alert
+
+| Column | Type | Notes |
+|---|---|---|
+| `dedup_key` | TEXT | PRIMARY KEY; e.g. `opencode-loop:` |
+| `state` | TEXT | NOT NULL |
+| `severity` | TEXT | NOT NULL |
+| `flagged_ticks` | INTEGER | Default 0 |
+| `opened_at` | TEXT | NOT NULL |
+| `last_fired_at` | TEXT | NOT NULL |
+| `resolved_at` | TEXT | |
+
+### `watchdog_channel_settings` — per-channel notification config
+
+| Column | Type | Notes |
+|---|---|---|
+| `id` | INTEGER | Autoincrement |
+| `channel_name` | TEXT | NOT NULL, UNIQUE |
+| `enabled` | INTEGER | Default 1 |
+| `min_severity` | TEXT | Default `warning` |
diff --git a/docs/incidents.md b/docs/incidents.md
index 4ceea34..a4a0463 100644
--- a/docs/incidents.md
+++ b/docs/incidents.md
@@ -1,14 +1,18 @@
# Incidents: how this router has broken, and how to tell which one it is
-Seven times now, a change that looked local to the router has silently
-degraded either the agent depending on it or the operator trying to see it
-clearly. They share a shape worth naming: **none of them announce themselves
-as router problems.** Five of the seven presented as an opaque client-side
-error — a connection refused, an "Unprocessable Content", an "internal server
-error" — and diagnosing each meant knowing which log or table to look in. The
-other two didn't error toward the client at all: #5 destroyed data outright,
-and #6 just went quiet in one corner of the admin UI while `/health` stayed
-green the whole time.
+Eight times now, something around the router has silently degraded either the
+agent depending on it or the operator trying to see it clearly. They share a
+shape worth naming: **none of them announce themselves as router problems.**
+Five of the eight presented as an opaque client-side error — a connection
+refused, an "Unprocessable Content", an "internal server error" — and
+diagnosing each meant knowing which log or table to look in. The other three
+didn't error toward the client at all: #5 destroyed data outright, #6 just went
+quiet in one corner of the admin UI while `/health` stayed green the whole time,
+and #8 spent money for hours while every check stayed green.
+
+**#8 is the only one so far that cost money rather than capability**, and the
+router behaved as configured throughout. Read it before leaving
+an agent run unattended.
**#5 is the only one so far that destroyed data**, and the only one where
recovery depended on luck rather than design. Read it before running any
@@ -21,6 +25,15 @@ or removing any config key.
This page exists so the next one takes minutes rather than hours. Start with the
symptom table, then read only the relevant section.
+**How to write an entry (from #8 on).**
+- Separate **Evidence** (measured, with where it was measured) from **Theory**
+ (inferred). Anything not established by a measurement or a reproduction goes
+ under Theory. It states what supports it and what would confirm or refute
+ it, and stays there until that test is done.
+- Include a **Human response** timeline: what the operator noticed, decided and
+ did. Mark actions an assistant took, and at whose direction.
+- Entries #1 to #7 predate this convention.
+
`opencode.json` points opencode's own model traffic at `http://127.0.0.1:8080/v1`,
so on this machine the router *is* the coding agent's inference supply. Breaking
it breaks the thing you would use to fix it. That is why these keep happening,
@@ -38,6 +51,8 @@ and why the diagnostics below are worth having to hand.
| 500s + `no such table`, venv/.env gone | #5 `git clean -fdx` | `ls -la router.db .env .venv` — a 0-byte db and a missing `.venv` is conclusive |
| A newly-shipped admin knob just isn't there, `/health` is fine | #6 backend running stale code | `systemctl --user status llm-router` uptime vs. `git log -1 --format=%cd` on the commit that added the knob |
| `activating (auto-restart)`, `pydantic_core...ValidationError: ...extra_forbidden` in the journal | #7 renamed key vs. un-migrated overlay | `journalctl --user -u llm-router --since '5 min ago' \| grep extra_forbidden` then `grep -n config/config.local.yaml` |
+| Spend is high after an agent run, nothing errored, no warning fired | #8 agent looping with no progress | `sqlite3 router.db "select session_key, strftime('%H',observed_at,'localtime') h, round(sum(prompt_tokens)/1e6) mtok, round(sum(cost_usd),2) usd from energy_observations where observed_at > datetime('now','-12 hours') group by 1,2 order by usd desc limit 8;"` then check whether that run's commits actually landed |
+| opencode plugin does nothing | #8 plugin loader rejects non-function export | `grep 'failed to load plugin' ~/.local/share/opencode/log/opencode.log` |
That last row is not an incident yet — it is the silent-staleness failure
described in the "Run as a service" section of `CLAUDE.md`. An unpolled catalog
@@ -347,6 +362,253 @@ file is correct.
---
+## #8 — Agent sessions looped for hours with no progress, and nothing noticed (2026-09-25)
+
+**Symptom.** Nothing errored. The router stayed healthy, `/metrics` showed zero
+warnings, and the quota alarm read `kind: none`. The operator noticed early on
+2026-09-26 that an opencode run had "burned a bunch of tokens over the hours".
+
+### Evidence
+
+Measured read-only against the live `router.db` and opencode's own session
+store. The window is 2026-09-25 19:00 to 2026-09-26 02:00 local unless noted.
+
+**Spend.**
+
+| | |
+|---|---|
+| prompt tokens | 323M, 95% cached |
+| billed | about $7.81, roughly 3x the busiest of the previous ten days ($2.66 on 09-20) |
+| largest share | OpenRouter `z-ai/glm-5.3-flash`: 1,368 calls averaging 203k prompt tokens, $6.16 |
+| largest contexts | NeuralWatt `glm-5.3`: 47 calls averaging 673k, max 682,679, the only model whose window still fit |
+| 120k+ prompt calls | 921 of them, $6.57 of the total |
+| routing | every decision `classification_source = classifier`, default profile; nothing pinned |
+
+**What the agents did.** An Atlas run (`.omo/plans/cockpit-quick-wins.md`)
+delegated each item to a sub-agent worker:
+
+| worker session | tool calls | exact-duplicate calls | worst repeat | landed |
+|---|---|---|---|---|
+| item 2, 2nd attempt | 423 | 30% | `read navbar.js` x61 | nothing |
+| item 3, 1st attempt | 132 | 23% | `read test_admin_js_units.py` x16 | nothing |
+| item 4 | 134 | 19% | `read` of the plan x8 | code only |
+| item 2, 1st attempt | 307 | 17% | `read navbar.js` x21 | nothing |
+| healthy workers | 30 to 96 | 0 to 6% | 1 to 3 | yes |
+
+- Item 4's commit message claimed four test files it never wrote, and it
+ committed with `git add -A`.
+- Atlas later said its own turns were "just re-reading the plan and state over
+ and over instead of doing work".
+- A planning session described its reads as "heavily elided". The stored tool
+ outputs in opencode were intact.
+
+**Pinch was rewriting agent context.**
+- `pinch.budget_tokens: 50000` pruned 1,915 of 1,992 decisions over nine hours,
+ keeping 48% of the prompt on average.
+- The planning session's 527k-token prompts went out at about 125k.
+- `src/context_prune.py` replaces older tool results with
+ `[read: result omitted]`, or keeps a 1,500-character head and tail around a
+ `[N chars trimmed...]` marker. Results inside the protected window get the
+ same cut above `protected_max_chars` (20,000).
+- Pinch's config had not changed since 2026-09-06. It had pruned 90 to 100% of
+ requests daily for two weeks without incident. The input changed:
+
+| | 09-10 to 09-23 | 09-25 | 09-26 |
+|---|---|---|---|
+| avg prompt into pinch | 100k to 150k | 236k | 545k |
+| p90 prompt | 170k to 260k | 528k | 1.31M |
+
+- `opencode.json` advertised `auto` with `limit.context: 782324`, so opencode did
+ not compact until near that size.
+
+**Routing drifted with size.**
+- `deepseek/deepseek-v4-flash` has 384k effective context;
+ `z-ai/glm-5.3-flash` has 655k.
+- Their proficiency was tied (`diff_checking` 0.933 vs 0.935) and unchanged
+ since 09-10.
+- Deepseek went from 290 picks on 09-23 to 0 on 09-25, while appearing 851
+ times as runner-up.
+
+**A clean reproduction, 2026-09-26 04:04 to 04:14.** A fresh planning
+session, on the Lift A brief (188 lines):
+- **Conditions:** context about 88k tokens, pinch off, 0 of 29 tool parts
+ marked `compacted` by opencode, every stored read intact.
+- **What it did:** it still reported its reads "elided with `[...]`", wrote
+ fake "Earlier tool responses received ... (counts, not proof of task
+ completion)" blocks into its own replies (assistant text parts, not tool
+ output), invented session IDs, and called the brief "393 lines".
+- **Where the phrase is not:** in opencode, oh-my-openagent (JS bundle and
+ native binary), or this repo.
+- **Model:** `z-ai/glm-5.3-flash` served 27 of 34 chat decisions in the 15
+ minutes before it was stopped.
+
+**Found along the way.**
+- `deploy/opencode-plugin/router-outcome.js` has never read a real exit code.
+ On opencode 1.18, bash's exit is `output.metadata.exit`; the plugin reads
+ `output.exitCode`.
+- PR #99 (conversation identity) was merged to `origin/main` but not running.
+ Production ran from a local checkout 29 commits behind.
+- **Plugin loader rejected `router-link.js`.** The plugin exported `parentCache`
+ as a module-level const. Opencode 1.18's plugin loader discards any plugin
+ with a non-function export — the entire file is silently dropped. The fix was
+ to make `parentCache` a property on the factory function instead (L224,
+ `router-link.js` in PR #102, a24de1b). Confirmed by grepping
+ `~/.local/share/opencode/log/opencode.log` for `'failed to load plugin'`.
+- How much of the $7.81 was waste cannot be computed. The run did land six
+ reviewed commits, and the deployed router cannot tell conversations apart.
+
+### Theory
+
+These are inferences, not measurements. Each lists what would confirm or
+refute it.
+
+1. **Pinch cutting oversized sessions drove, or worsened, the re-read loops.**
+ Weakened by the clean reproduction, which looped with pinch off. As
+ sessions grew past about 250k, the fixed 50k budget removed most of each
+ agent's working memory, recent reads included. The agent saw its reads cut,
+ re-read them, grew the context, and got cut harder.
+ - Supported by: the agents' own descriptions ("elided", "drowning in
+ tool-output truncation"), intact stored outputs, the prune rates, and the
+ timing of the size jump.
+ - Not tested: no run has been observed with pinch off.
+ - Confirm: with pinch off and `auto` at 200k, the duplicate-read signature
+ should not recur on comparable runs. The interim watcher
+ (`plans/no-progress-detection-prototype.py`) and Phase 0 data can show it.
+2. **Sessions got that large because long unattended orchestration ran without
+ compaction.** Atlas made 668 tool calls, one worker 423, under a 782k limit,
+ and the re-read loop in theory 1 fed the growth. The limit is a fact; that it
+ explains the growth is theory.
+3. **Leading theory: `z-ai/glm-5.3-flash`, or its OpenRouter path,
+ confabulates truncation and then loops.** The clean reproduction above
+ removed pinch, context size and input corruption, and the behaviour stayed.
+ Theories 1 and 2 remain plausible aggravators, not the cause.
+ - Confirm: with glm excluded, comparable agent runs show no elision claims
+ and no duplicate-read signature. A direct A/B on one prompt would settle
+ model versus provider path.
+
+### Why nothing caught it
+
+Each existing guard answers a different question:
+- `circuit_breaker.py` trips on provider 5xx. There were none.
+- `runway_low_warning` fires when balance / burn drops under 6 hours. Burn is
+ averaged over a 24 h balance window: OpenRouter read $24.73 at $0.25/h, which
+ is 97 hours of runway. It asks "will I run out", not "is this being wasted".
+- Spend rate was not abnormal. The worst hour ($2.42) and worst 3-hour window
+ ($4.13) sit inside the prior 30 days' range (hourly p90 $1.46, max $4.83;
+ worst 3 h $9.65).
+- The router cannot see progress, and the deployed build could not tell
+ conversations apart.
+
+### Recurrences
+
+- **02:21 to about 02:56: Atlas looped after its plan was complete.** Every
+ checkbox was ticked, but `.omo/boulder.json` still carried `status: completed`
+ and `pr_url: .../pulls/77` from the previous plan (evidence: the file). The
+ continuation hook injected "continue" turns, 3 of 3 user turns in the window
+ (evidence: the session). Atlas re-derived its state each time: 54 turns,
+ about 19M cached-read tokens, `boulder.json` read 28 times, checkboxes grepped
+ 16 times. Cause: stale state, not pruning. The branch was never pushed, so PR
+ 77 was not touched.
+- **About 03:05 to 03:17: a read-only `explore` subagent** re-read
+ `router-outcome.js` 20 times (19% duplicate calls across 322). The production
+ sync had just deleted that file. Read-only agents never land changes, so the
+ detector needs a separate signal for them.
+
+### Human response
+
+Times are local, 2026-09-26.
+- **About 02:10.** Noticed the spend and asked for a mechanism in 6krrt to catch
+ it. Rejected a spend-rate alarm, since steady agent usage is legitimate, and
+ framed the target as "are concrete changes landing".
+- **About 02:20 to 02:40.** Decided the design for `plans/no-progress-detection.md`:
+ - response modes (`warn`, `auto_compact`, `auto_limit_context`, a two-stage
+ `auto_recover`), with every mode warning
+ - `warn` as the default
+ - at most 2 recoveries per session tree
+ - `notify-send` now, with pluggable SMS, RingCentral and PagerDuty channels
+ later
+ - a local-model watchdog timer, "so I don't have to rely on Claude"
+ - `auto`'s context limit lowered to 200,000 in the repo and global
+ `opencode.json`
+- **About 02:50.** Caught the post-completion Atlas loop by watching. At the
+ operator's direction, Claude aborted the session at 02:56 and cleared the
+ stale `pr_url` in `.omo/boulder.json`.
+- **About 03:10.** Merged PR #101. A first `git pull` aborted on local changes,
+ and the restart on that line reloaded the old code. Then ran the prepared sync
+ (back up, drop the changes already upstream, `reset --keep origin/main`,
+ re-apply the rest) and restarted onto `827c408` at 03:15.
+- **About 03:30.** Asked Claude to watch opencode for loops until the detector
+ ships (the interim watcher).
+- **About 03:45.** Reported the planning session's "truncation hell". At the
+ operator's direction ("hot fix that"), Claude set `pinch.enabled` false at
+ runtime and persisted it to `config/config.local.yaml` (backup
+ `config.local.yaml.bak-20260926-pinch`). The operator restarted the service
+ at 03:52 and chose not to open a PR for the switch flip.
+- **About 04:02.** Installed `router-link.js` in place of `router-outcome.js`
+ and restarted opencode (200k `auto` limit now active). Kicked off a fresh,
+ small Lift A planning session (`plans/no-progress-lift-a.md`).
+- **About 04:14.** After the fresh session reproduced the confabulation,
+ excluded `z-ai/glm-5.3-flash` from routing. It had no picks after 04:13:57,
+ and traffic moved to deepseek and mimo. At the operator's direction, Claude
+ aborted the confabulating session.
+- **~15:18.** Watchdog timer enabled (`systemctl --user enable --now
+ llm-router-watchdog.timer`).
+- **15:19.** Plugin-loader root cause found in `opencode.log`; plugin fix
+ installed and opencode restarted.
+- **15:19:57.** Router restarted.
+- **After 15:19.** Conversation identity confirmed arriving: 17 of 17 requests
+ after the restart carry `c:` keys and an agent.
+
+### Status
+
+**Partly mitigated; detection not built; leading theory (3) unconfirmed.**
+
+**Lift A shipped.** PR #102 (a24de1b) landed the surfacing half of North Star #4
+(conversation identity and watchdog), live ~15:18 EDT. **Surfacing half of North
+Star #4 is only partly shipped** — the identity half is confirmed (17 of 17
+requests after the restart carry `c:` keys), but the detection half still waits
+on Lift B (`plans/no-progress-detection.md`).
+
+**Confounded by design, and say so.** Three mitigations landed within about
+25 minutes:
+- pinch off at 03:50
+- `auto` capped at 200k at 04:02
+- `z-ai/glm-5.3-flash` excluded at about 04:14
+
+So "sessions look healthier since" cannot say which one mattered, or whether
+it is several interacting: model degradation near a full window, oversized
+sessions, pruning, and stale orchestration state. The one clean data point is
+the 04:04 reproduction (glm, pinch off, small context, still confabulated),
+which is why theory 3 leads.
+
+To attribute properly, change one variable at a time and watch the watchdog's
+stall rate. For example, re-admit glm with everything else held, or re-enable
+pinch with a budget above 200k. Until then, treat every theory here as
+unproven.
+Also missing: a **behavior breaker**. The availability breaker trips on 5xx and
+PR #76 trips on malformed output; nothing trips on well-formed output that
+confabulates or loops. Planned for Lift B (`plans/no-progress-detection.md`,
+section 8).
+- In place:
+ - pinch off, runtime and persisted
+ - `auto` limit 200,000, active for opencode sessions started after a restart
+ - #99 and #100 deployed at 03:15
+- Pending, operator:
+ - install `router-link.js` in place of `router-outcome.js`
+ - restart opencode
+- Pending, lift: `plans/no-progress-detection.md`, the watchdog first.
+- If pinch is re-enabled, its budget must sit above the 200k cap, and it must
+ never cut recent results.
+
+**Rule:** steady spend is not the failure; **spend with no concrete change
+landing is.** Until the detector ships:
+- Check an unattended agent run by whether commits are landing, not by the
+ quota page.
+- Verify each worker commit's `git show --stat` against its message; a worker's
+ "done" is a claim.
+
+---
+
## The pattern
The first five are defensible local decisions — free a port, stop a process,
@@ -373,6 +635,14 @@ file by mistake; #7 was a *correct*, deliberate, reviewed code change that
was still incomplete, because "every consumer" silently excluded a consumer
that git cannot see.
+#8 breaks the pattern from the other side: the router did what it was
+configured to do, routing each request to a model that fit. It was the one
+component that saw all the spend while being blind to whether any of it produced
+anything, so it measured cost when the thing that mattered was progress. Whether
+its own `pinch` pruning drove the loops is still a theory (see #8, Theory 1). If
+confirmed, #8 joins the family of defensible settings, a 50k budget chosen for
+~130k sessions, that turned harmful when the input changed shape.
+
The generalisable fix is the same each time: **compute the thing that is
actually true, and surface it where the operator is already looking.** That is
what the `/metrics` warnings and the inline admin warning are for, and it is the
diff --git a/docs/watchdog.md b/docs/watchdog.md
new file mode 100644
index 0000000..1f2665b
--- /dev/null
+++ b/docs/watchdog.md
@@ -0,0 +1,408 @@
+> Deep dive into the watchdog subsystem — opencode session loop detection. Back to
+> [README](../README.md).
+
+Watchdog is a periodic scanner that probes running opencode sessions for looping
+patterns — repeated tool calls without progress. It runs as a systemd oneshot
+every 5 minutes (`deploy/llm-router-watchdog.timer`) and, when a session flags,
+writes into a local SQLite table and fires desktop notifications through the
+configured channel pipeline.
+
+The detector itself is a pure module (`src/progress_detect.py`) that returns a
+boolean verdict plus a reason dict. The orchestrator (`src/watchdog.py`) reads
+opencode session data over HTTP, runs the detector, and manages the alert
+state machine. The notifier (`src/notifier.py`) routes events through
+configured channels to desktop alerts (`notify-send`).
+
+## What it catches and doesn't
+
+Watchdog looks for **looping** — sessions that make repeated tool calls against
+the same targets without landing meaningful changes. It detects four signal
+types:
+
+- **Dup signal**: a call appears more than 25% of the time in the sliding
+ window. The call matcher first normalises `bash` commands (strips env var
+ prefixes and `cd` prefixes, joins lines, collapses comments), then checks
+ tool name + canonicalised JSON args (sorted keys) for equality.
+
+- **Top signal**: a single target is called 12+ times in the window (or 8+
+ for read-only agents like explore/librarian/oracle). Targets merge calls on
+ the same `(tool, basename)` for file reads, `(bash, first-two-words)` for
+ commands, and `(tool, pattern[:60])` for grep/glob.
+
+- **Slow signal**: a single tool call accounts for 15+ of the session's total
+ calls, AND no call in the entire session history has landed (tree write or
+ git commit). This catches sessions stuck on a single read/analysis.
+
+- **Coverage signal**: the session rereads one file at 4x its line count
+ within the window. The coverage score divides total bytes read by file length,
+ capped at the file's actual lines. A file read 4x its length means the model
+ is rereading without making progress.
+
+**NOT steady spend.** A session that steadily calls different files on a
+difficult refactor is not flagged — each call lands a unique `(tool, args)`
+and the top target never crosses the threshold. The detector only fires when
+repetition outpaces the call budget, not when a session is quietly burning
+tokens across many distinct tools.
+
+### Landed calls
+
+A call counts as **landed** when it:
+- is an `edit`, `write`, or `patch` with a non-empty `metadata.diff` field, or
+- is a `bash` command containing `git commit` with `metadata.exit == 0`
+
+Landed calls break both the slow and coverage signals. A session that reads a
+file 4x, then edits it once, clears all signals immediately — the landed check
+exits early and the verdict returns `False`.
+
+### Ancestral trees
+
+Sessions can be parented: the opencode API returns `parentID` on child sessions.
+Watchdog resolves root sessions and evaluates EACH session in the tree on its
+own calls, NOT merged into the root. A child's flagged verdict bubbles up to the
+root level, which is what the alert dedup key uses (`opencode-loop:{root_session_id}`).
+
+For the landed-time check, a session sees NOT only its own landed calls but also
+its transitive descendants' landed calls. This prevents a parent session from
+being flagged just because its child landed — even though the parent may still
+be looping independently.
+
+## Signals and thresholds
+
+Thresholds live in `DetectConfig` (`src/progress_detect.py:38-53`) and are
+configurable under `watchdog.detector.*` in `config/config.yaml`.
+
+| Key | Default | Meaning |
+|---|---|---|
+| `watchdog.detector.dup_min` | 0.25 | Fraction: calls must repeat at this rate to flag |
+| `watchdog.detector.top_min` | 12 | Minimum calls to a single target (normal agents) |
+| `watchdog.detector.top_min_ro` | 8 | Minimum calls to a single target (read-only agents) |
+| `watchdog.detector.cum_min` | 15 | Same call must account for this many total calls |
+| `watchdog.detector.cover_min` | 4.0 | File must be reread this many times its length |
+| `watchdog.detector.window` | 60 | Sliding window in number of tool calls (not seconds) |
+| `watchdog.detector.min_calls` | 40 | Minimum total calls before evaluation triggers |
+| `watchdog.detector.read_only_agent_keywords` | `["explore", "librarian", "oracle"]` | Title keywords that make an agent read-only (lowered `top_min` from 12 to 8) |
+
+Read-only agents get a lowered top threshold because agents whose job is
+exploration or library work are expected to call the same targets repeatedly —
+the floor is 8 instead of 12.
+
+## Alert lifecycle
+
+The alert state machine is per-root-session and tracked in the
+`watchdog_alerts` table.
+
+### Trigger (initial detection)
+
+When the detector flags a root session's tree for the first time:
+
+1. The watchdog writes a `watchdog_alerts` row with `state='open'`.
+2. The `local_llm_enabled` path asks a local Ollama model (defaulting to
+ `verification.model` or `qwen2.5-coder-router:14b`) for a second opinion.
+ It receives the agent name and the reason dict as context, and the model
+ must respond with only "yes" or "no" — no explanation.
+3. If the local LLM answers "yes", severity is set to `critical`. If it
+ answers "no" or times out (returns `None`), severity is set to `warning`.
+4. The notifier dispatches an `AlertEvent` through every matching channel.
+5. The `flagged_ticks` counter starts at 1.
+
+A cap of 3 simultaneous LLM second-opinion calls (`_MAX_LLM = 3`) prevents a
+burst of flagged sessions from flooding the local model.
+
+### Escalate (persistent looping)
+
+On every subsequent tick where the session is still flagged:
+
+1. `flagged_ticks` is incremented.
+2. If `flagged_ticks >= 3` (15 minutes at 5-minute ticks) OR the LLM answers
+ "yes" again, the alert escalates to `severity='critical'` and the escalation
+ event is dispatched.
+3. If the alert is already critical, ticking continues but no duplicate
+ escalation fires.
+
+### Resolve (session is gone or fixed)
+
+A root session alert resolves in two conditions:
+
+- **Session disappeared**: the root session ID is no longer returned by the
+ opencode `/session` API. The watchdog writes `state='resolved'`,
+ `resolved_at=`, and `severity='info'`.
+- **Session fixed**: the root session's tree no longer has any flagged sessions
+ in the detector's verdict. The alert is resolved the same way and a resolve
+ event is dispatched.
+
+Resolve events always bypass the rate limit — they must clear the previous
+notification so the operator knows the alert is over.
+
+### No-opencode
+
+When no `rc-servers.json` or no opencode server answers, the tick writes an
+`outcome='no_opencode'` record and returns immediately with zero alerts. This
+uses the filesystem lock (`router.db/.watchdog.lock`) so two watchdog instances
+cannot run simultaneously.
+
+## Admin surfaces
+
+Watchdog data surfaces through three admin pages, each serving a different
+operational question.
+
+### Loops panel — index.html#loops
+
+On the main dashboard (`admin/frontend/index.html`), the **Loops** card
+(id="loops") shows all currently open alerts:
+
+- **Severity badge**: the alert's current severity (`critical` in red,
+ `warning` in orange-yellow, `info` in grey)
+- **Dedup key**: the raw `opencode-loop:{root_session_id}` string, so the
+ operator can cross-reference with `journalctl` or `router.db`
+- **Opened at**: `opened_at` ISO timestamp from the first trigger
+- State: `open` or `resolved`; the panel filters to `resolved_at IS NULL`
+
+The panel is client-side only — it polls `/api/watchdog/loops` which queries
+`watchdog_alerts` for all unresolved rows.
+
+### Watchdog card — controls.html
+
+The Controls page (`admin/frontend/controls.html`) includes a dedicated
+**Watchdog** card with operational status:
+
+- **Last tick**: timestamp of the most recent `watchdog_ticks` row
+- **Sessions seen**: how many opencode sessions were enumerated on that tick
+- **Flagged verdicts**: count of flagged + total from the latest tick's
+ `watchdog_verdicts` rows
+- **Open alerts**: count from `watchdog_alerts WHERE resolved_at IS NULL`
+- **Notification channels**: list of configured channels with toggle controls
+ (enabled/min_severity), editable via `POST /api/watchdog/channels`
+- **Refresh** button: reloads the watchdog status from the API
+- **Send test alert** button: fires a test `AlertEvent` through enabled channels
+
+The status endpoints are:
+- `GET /api/watchdog/status` — last tick, verdict counts, open alert count
+- `GET /api/watchdog/channels` — per-channel settings from `watchdog_channel_settings`
+- `POST /api/watchdog/channels` — upsert channel settings (enabled, min_severity)
+- `POST /api/watchdog/test-alert` — deliver a test event (severity + optional
+ channel_name filter) to a live Notifier instance
+
+### Per-model stall rollup — models.html
+
+The `watchdog_verdicts` table has `model_id`/`provider` columns and indexes on
+them, but no per-model stall rollup, no Block button beside evidence, and no
+Blocked list are built yet. `blocked` is only a value in the Models override
+dropdown (see docs/admin-portal.md).
+
+## Database schema
+
+Four tables live in `router.db`, created idempotently by both
+`config/schema.sql` and `watchdog_store.py`:
+
+```sql
+watchdog_ticks -- one row per watchdog tick
+watchdog_verdicts -- per-session verdict at each tick (joined to ticks)
+watchdog_alerts -- alert lifecycle state machine (dedup_key PK)
+watchdog_channel_settings -- per-channel toggle + severity gate
+```
+
+Indexes on `watchdog_verdicts` cover `model_id`, `created_at`, `flagged`, and
+`session_root` — the columns the admin endpoint queries most frequently.
+
+The ticks table carries `ticked_at`, `sessions_seen`, and `outcome`
+(`'ok'`, `'flagged'`, or `'no_opencode'`). The verdicts table joins to ticks
+via `tick_id` and carries the full reason dict fields: `dup`, `top`,
+`top_what`, `landed`, `slow`, `coverage`, plus `calls_since_landed`,
+`cost_since_landed_usd`, and the LLM second-opinion answer.
+
+Alerts are keyed by `dedup_key` (deduplicated per root session) and carry
+`opened_at`, `last_fired_at`, and `resolved_at`. The transition logic lives in
+the `_fire_alert()` helper: trigger inserts only if no open row exists,
+escalate increments `flagged_ticks` and bumps severity, and resolve writes
+`resolved_at` and downgrades to `severity='info'`.
+
+## Knobs
+
+All watchdog configuration lives under `watchdog:` in `config/config.yaml`:
+
+| Key | Default | Required |
+|---|---|---|
+| `watchdog.enabled` | `true` | Gates the entire subsystem |
+| `watchdog.local_llm_enabled` | `true` | Whether to call Ollama for second opinions |
+| `watchdog.model` | `null` (uses `verification.model`) | Local model name |
+| `watchdog.detector.window` | 60 | Sliding window size in calls |
+| `watchdog.detector.dup_min` | 0.25 | Dup threshold (fraction) |
+| `watchdog.detector.top_min` | 12 | Top target threshold (normal agents) |
+| `watchdog.detector.top_min_ro` | 8 | Top target threshold (read-only agents) |
+| `watchdog.detector.cum_min` | 15 | Slow-signal threshold |
+| `watchdog.detector.cover_min` | 4.0 | Coverage-signal multiplier |
+| `watchdog.detector.min_calls` | 40 | Minimum eval calls |
+| `watchdog.read_only_agents` | `["explore", "librarian", "oracle"]` | Agent title keywords for reduced threshold |
+| `watchdog.dashboard_base_url` | `"http://127.0.0.1:8080/admin"` | Base URL for alert links |
+
+All seven `watchdog.detector.*` keys plus `watchdog.enabled` and
+`watchdog.local_llm_enabled` (nine allowlisted keys total) are allowlisted for
+live editing through the
+admin portal (`POST /admin/api/config/{key}`), which writes to
+`config.local.yaml` (the gitignored machine-local overlay). The changes are
+validated by `RouterConfig` before reaching disk, and take effect at the next
+tick — no service restart required since `detect_config_from_pydantic()` reads
+cfg fresh each invocation.
+
+## Install
+
+The watchdog ships as two systemd user units:
+
+- `llm-router-watchdog.service` — a `Type=oneshot` that runs
+ `PYTHONPATH=%h/llm-router/src .venv/bin/python -m watchdog --once`
+- `llm-router-watchdog.timer` — fires 5 min after boot, then every 5 min
+
+Installation follows the same pattern as the other deploy units. The shipped
+unit files use `%h/llm-router` placeholders that need rewriting to your actual
+repo path:
+
+```bash
+REPO=$(pwd)
+for u in deploy/llm-router-watchdog.{service,timer}; do
+ sed "s|%h/llm-router|${REPO}|g" "$u" \
+ > ~/.config/systemd/user/"$(basename "$u")"
+done
+systemctl --user daemon-reload
+systemctl --user enable --now llm-router-watchdog.timer
+```
+
+The service unit has **no `EnvironmentFile`** — watchdog needs no provider API
+keys because it only queries the local opencode session APIs and the local
+Ollama instance.
+
+Check installation:
+
+```bash
+systemctl --user status llm-router-watchdog.timer
+journalctl --user -u llm-router-watchdog.service --no-pager
+```
+
+## Run --once
+
+For ad-hoc debugging or pre-deploy smoke test:
+
+```bash
+PYTHONPATH=src .venv/bin/python -m watchdog --once
+```
+
+The `--once` flag runs a single tick and exits 0 once the tick has run,
+however many alerts it fired (the count is logged as `watchdog tick done: N
+alert(s) fired`); systemd would otherwise mark the oneshot unit failed on every
+real alert. It connects to the same `router.db` that the timed
+service uses, acquires the filesystem lock, reads `rc-servers.json`, probes
+opencode servers, and follows the full evaluate-alert-resolve pipeline.
+
+Add `--config /path/to/config.yaml` to point at a non-default config file.
+
+### Debug output
+
+The `--once` run emits structured `watchdog=` log lines for every state
+transition:
+
+```
+2026-01-15 14:30:01 INFO watchdog: watchdog=2026-01-15T14:30:01+00:00 dedup='opencode-loop:ses_xxxxx' state=trigger severity=warning title='opencode loop: explore'
+```
+
+The `title` field carries the agent name derived from the session title. The
+`dedup` field is the root session's dedup key, matching `journalctl` traces
+and admin panel rows.
+
+## Backtest
+
+`scripts/progress_backtest.py` replays labelled sessions through the detector
+to validate threshold calibration. It reads from a fixture JSON file and runs
+a sliding-window evaluation with step 5.
+
+```bash
+PYTHONPATH=src .venv/bin/python -m scripts.progress_backtest --fixture \
+ tests/fixtures/progress/fixture.json
+```
+
+The fixture ships in `tests/fixtures/progress/fixture.json` and contains
+15 labelled sessions: 8 marked `must_flag` and 7 marked `must_not_flag`.
+Expected output after a successful run:
+
+```
+# backtest complete: 8/15 flagged
+```
+
+At least 8 sessions should trigger a flag. The fixture is generated from live
+opencode sessions (see [export fixture](#re-export-fixture) below) and scrubbed
+to protect file contents — content is SHA-1 hashed except for the few keys the
+detector actually inspects (`filePath`, `command`, `pattern`, etc.).
+
+## Re-export fixture
+
+The fixture generator pulls sessions from the opencode SQLite database
+(`~/.local/share/opencode/opencode.db`) and writes the scrubbed test fixture:
+
+```bash
+python scripts/export_progress_fixture.py
+```
+
+Output goes to `tests/fixtures/progress/fixture.json`. The script takes a
+session ID to label map (`LABELS` dict), extracts call history, resolves file
+line counts (via `git show` for deleted worktree paths), and scrubs all strings
+to SHA-1 prefixes, preserving only the detector's target keys.
+
+The fixture generation script hardcodes paths to the operator's home directory
+and a specific worktree commit (`WORKTREE_COMMIT = "827c408"`). When
+regenerating, update those paths to match the current environment.
+
+## Notifier pipeline
+
+Alert events flow through `src/notifier.py` before reaching the operator. Each
+alert is an `AlertEvent` dataclass mirroring PagerDuty Events v2 shape:
+`dedup_key`, `severity` (info/warning/critical), `state` (trigger/escalate/
+resolve), `title`, `summary`, `details`, and `source` (hardcoded to
+`"6krrt-watchdog"`).
+
+Each configured channel has its own `min_severity` gate — a critical event
+passes through a channel with `min_severity=warning`, but a warning event does
+not pass a channel with `min_severity=critical`. Each channel also has a rate
+limit window of 300 seconds (5 minutes). Resolve events always bypass the rate
+limit to ensure clear-the-alert notifications are delivered.
+
+Currently only `type=desktop` channels are implemented; they invoke
+`notify-send -u {urgency} {title} {summary}`. Unknown channel types are
+logged as warnings and skipped. Missing `notify-send` (no `DISPLAY` or
+`libnotify` installed) is logged at warning level — the notifier never raises.
+
+Channel settings (`watchdog_channel_settings`) are editable at runtime through
+the admin portal and stored per-channel in the database, separate from the
+base config.
+
+## Known limits
+
+- **Under 40 calls: invisible.** `min_calls=40` means the detector does not
+ evaluate sessions with fewer tool calls. Short troubleshooting sessions that
+ loop on 10 calls pass through undetected.
+
+- **Atlas at cover_min 4.0 exactly.** A session whose last 60 calls read a
+ single file exactly 4x its line count hits the coverage threshold at the
+ boundary. There is no margin: `>=` comparison means exactly 4.0x flags.
+ This matters for sessions that genuinely need to re-read a reference file
+ repeatedly — the coverage signal cannot be tuned per-file.
+
+- **Stale localhost:4096 probes.** Watchdog reads `rc-servers.json` from
+ `~/.local/share/opencode/rc-servers.json`, which may contain stale entries
+ for opencode sessions that have since ended. These produce empty call lists
+ and are silently skipped — but they add network round-trip latency to each
+ tick. Only the first answering server's sessions are evaluated (the watchdog
+ picks `answering[0]` from the list of servers that successfully return a
+ session list).
+
+- **No streaming path.** `POST /outcome` is the only ground truth signal for
+ streaming traffic; the watchdog has no equivalent endpoint to report streaming
+ session outcomes. It can only see tool-call repetition, not whether the answer
+ was useful.
+
+- **Local LLM gate.** Second-opinion calls are gated by both
+ `local_compute.enabled` (global) and `watchdog.local_llm_enabled`
+ (per-subsystem). If either is false, no LLM calls are made and severity stays
+ as `warning` on initial trigger regardless of model quality. The cap of 3
+ simultaneous LLM calls means that if 5 sessions flag, 2 wait without opinion.
+
+- **No provider calls.** Watchdog does not touch the dispatch path. It only
+ reads opencode session data and runs local heuristics. No provider calls, no
+ quota consume, no cost incurred.
diff --git a/opencode.json b/opencode.json
index 91db3f1..eb67d2d 100644
--- a/opencode.json
+++ b/opencode.json
@@ -13,7 +13,7 @@
"auto": {
"name": "auto (router picks, interactive)",
"limit": {
- "context": 782324,
+ "context": 200000,
"output": 16384
},
"modalities": {
diff --git a/plans/cockpit-brainstorm.md b/plans/cockpit-brainstorm.md
new file mode 100644
index 0000000..93d321f
--- /dev/null
+++ b/plans/cockpit-brainstorm.md
@@ -0,0 +1,397 @@
+# Cockpit brainstorm: the router as an operator's instrument panel
+
+Status: reference -- brainstorm, not a queue item; its six quick wins shipped in PR #101
+
+**Date:** 2026-09-23
+**Read against:** the repo tarball as of this date (code, config, plans,
+screenshots). **Not** read against the live `router.db`, so every "gap" below
+is a gap in what the code surfaces, not a claim about what the data shows.
+Section 5 is the exception: it was read against the live portal.
+
+## The framing
+
+6krrt as a local LLM router with a cockpit. The operator:
+
+1. adjusts dials and knobs in flight,
+2. sees where tokens are wasted,
+3. watches proficiency, and
+4. sees poor results from a carrier or model surface, so they can decide
+ whether to preclude it (with the circuit breaker as the automatic version
+ of that decision for outages).
+
+The router makes the per-request decision. The cockpit is where the operator
+makes the slower decisions: which models and carriers are trusted, for what,
+and at what settings. Under this framing, the question for every feature is
+whether it helps the operator notice something and act on it.
+
+---
+
+## What the cockpit already has
+
+| job | what exists | where |
+|---|---|---|
+| in-flight dials | ~11 runtime knobs (bool/float/int tables in `admin.py`), persisted config writes via `/api/config/{key}`, classifier card, profiles CRUD, gaming mode | Controls, Profiles |
+| knob coverage | North Star rule #1 + `test_admin_knob_coverage.py` | tests |
+| token waste | pinch savings card; `cache_rate_series` + warnings; incumbent cache pricing dial; `cost_estimate_calibration` in `/metrics` | Dashboard, Controls, `/metrics` |
+| proficiency | model × category matrix with evidence status (measured / inherited / thin, n per cell); client outcome log with attributable + applied flags | Proficiency |
+| poor results | `content_fault_warnings` (malformed output rate); `rejection_warnings`; availability override (active / deprecated / stale); allowlist editor | warnings bell, Models, Providers |
+| breaker | passive per-(model, provider) breaker on 5xx; on/off knob | Controls (toggle only) |
+
+The foundation is good. What's mostly missing is the link between a signal
+and the action it should prompt, plus any record of which actions were taken.
+
+---
+
+## Cross-cutting ideas (these help all four jobs)
+
+### X1. Flight recorder: a change log
+
+**Gap.** `admin_schema.sql` has only `admin_model_overrides`. Nothing records
+when a knob, override, allowlist entry, profile or feedback fold changed, what
+it changed from, or why. Runtime knobs are "NEVER persisted", so a restart
+silently reverts them, and nothing notes that the revert happened.
+
+**Idea.**
+- An `admin_changes` table: `(at, surface, key, old, new, reason, source)`,
+ where `source` is `runtime | persisted | override | allowlist | profile |
+ feedback_fold | restart`. Every admin write path appends a row.
+- An `operator_epoch` id stamped on each `route_decisions` row, incremented
+ on each **operator action** (a row in `admin_changes`). Every metric can then
+ be split before and after a change without joining on timestamps. Scope it
+ to operator actions on purpose: "any change that affects routing" would also
+ cover catalog price polls and every feedback fold, so the epoch would tick
+ constantly and split nothing.
+- Change markers drawn on the dashboard history charts.
+- A drift badge when a runtime value differs from its persisted twin
+ ("this will revert on restart").
+- Restart reverts are loggable without persisting runtime knobs: the change
+ log already holds each knob's last runtime value, so at startup, compare it
+ with the loaded config and write one `restart` row per value discarded.
+
+**Why first.** Adjusting knobs in flight produces no learning if you can't
+later tell what a change did. The other ideas here lean on this one.
+
+### X2. Structured warnings with actions attached
+
+**Gap.** `/metrics` warnings are plain strings. The dashboard decides which
+page fixes a warning, and how severe it is, by regex on the text
+(`warningTarget` and `warningSeverity` in `index.html`). Rewording a warning
+silently breaks its link and its severity.
+
+**Idea.** Emit `{class, severity, subject: {model, provider, category?},
+evidence_url, actions: [...]}`. The warning registry in
+`test_tui_warnings.py` already lists every class, so it can serve as the
+schema. The bell can then offer an action in place: "exclude from
+`tool_use_agentic`", "quarantine 24h", "open decisions filtered to this
+model".
+
+### X3. Preview before apply (replay)
+
+**Gap.** The reviews replayed thousands of real decisions to test a setting,
+but only as one-off scripts pasted into plan documents.
+
+**Idea.** For any change that affects ranking (`quality_tolerance`, the
+incumbent dial, profile edits, availability overrides, graded exclusions
+from P2 below), the portal replays the last N decisions through
+`select_candidates` and `rank_candidates` with the proposed value, then shows:
+- a winner-shift matrix: old winner → new winner, with counts,
+- estimated cost delta,
+- decisions that would become unroutable (the 2026-09-01 and 2026-09-04
+ incidents were both deprecations that emptied a candidate set).
+
+This is cheap because the ranking modules are pure. Its output is an estimate
+of routing change, not quality change, and should be labeled that way.
+
+---
+
+## 1. Dials in flight
+
+- **D1. Timed changes.** "Apply for 2h, then revert." Turns a knob change into
+ a bounded experiment and lowers the risk of forgetting a test setting.
+ Needs X1 to record both the apply and the revert.
+- **D2. Blast radius on hover.** How many of the last 24h of decisions this
+ knob would have touched. It's the X3 replay reduced to one number.
+- **D3. Session-scoped A/B.** Assign new sessions to setting A or B and
+ compare cost and outcome rate per arm. It has to be per session, not per
+ turn, because a mid-session switch costs cache (token-waste Wave 2
+ measured 0.919 → 0.348). Larger build; only worth it once X1 exists and
+ outcome volume can support a comparison.
+- **D4. Grouping by effect.** Group controls by what they move (cost,
+ quality, latency, safety, measurement-only) rather than by config section,
+ and mark the few that actually change routing. Knob coverage guarantees
+ every control exists; this makes them usable.
+
+## 2. Token waste
+
+**Gap.** Waste signals are spread across pinch, cache rate, calibration and
+the verifications table, and none is in dollars in one place.
+`cost_estimate_calibration` is in `/metrics` but not in the portal.
+
+- **W1. Waste ledger.** One panel with dollars per waste class over a window,
+ each with a trend:
+ | class | how it's computed |
+ |---|---|
+ | switch cache loss | billed on switch turns − same prompt at the session's same-model cache rate |
+ | retries | re-billed prompt tokens from `iteration.py` attempts |
+ | paid for failure | billed cost of responses later reported `ok:false` via `/outcome` |
+ | empty-200 / malformed | billed cost of responses with a structural `malformed` verdict |
+ | truncation | responses stopped by `finish_reason: length` with no client cap |
+ | pinch (negative) | dollars saved, from `pinch_summary` |
+
+ "Paid for failure" counts only decisions that received an outcome, so it
+ shows n and outcome coverage (the share of billed decisions with a report)
+ next to the dollar figure.
+- **W2. Why did it switch?** Record a switch reason on each decision: category
+ changed, breaker open, exploration, override removed the incumbent, profile
+ change. Then rank sessions by switch cost and drill into the turns. Without
+ a reason, a switch is only a cost; with one, it points to a knob.
+- **W3. Calibration panel.** Surface `cost_estimate_calibration` per model,
+ headlining the **spread** (as its docstring argues), and flag pairs where
+ the estimator and the bill disagree on order. Where that happens, the cost
+ tiebreak picks the wrong model.
+- **W4. Cost per successful outcome.** Per (model, provider, category):
+ billed dollars ÷ `ok:true` outcomes. A cheap model that fails often isn't
+ cheap. This is the one waste figure that includes quality.
+
+ Show n and outcome coverage next to it. Models that carry more traffic
+ collect more outcomes (the same exposure bias the feedback fold had to
+ correct), and if coverage stays hidden, a thinly reported cheap model will
+ look better than it is.
+
+## 3. Proficiency monitoring
+
+**Gap.** The matrix shows current state well. It doesn't show movement, and
+it doesn't show which thin cells actually matter.
+
+- **P1. Trend and drift per cell.** A sparkline of the outcome rate over time.
+ Compare a recent window against the long-run rate with an interval, and
+ flag drops that fall outside it. Carriers change serving setups (quant,
+ engine, hardware) without notice, and a drift flag is how that shows up.
+- **P2. Evidence priority.** Rank cells by `traffic share × uncertainty`. A
+ thin cell that routing never consults doesn't matter; a thin cell deciding
+ 30% of traffic does. The ranking tells you where to spend `eval_proficiency`
+ runs or exploration budget.
+- **P3. The tolerance band per category.** Show which models sit within
+ `quality_tolerance` of the leader, so it's clear where cost is deciding and
+ where quality is.
+- **P4. Label provenance per cell.** The share of each cell's outcomes whose
+ category came from a fresh classification versus a cached or borrowed one.
+ North Star rule #2 exists because borrowed labels trained the matrix;
+ this makes it visible cell by cell.
+
+## 4. Poor results → operator preclusion (and the breaker)
+
+**Gap.** The operator's only preclusion tools are binary: an availability
+override (active / deprecated / stale) or removing a model from the
+allowlist. The breaker is in-memory, trips only on 5xx, forgets everything
+on restart, and shows nothing but an on/off toggle.
+
+- **Q1. Carrier/model scorecard.** One sortable table per (model, provider)
+ over a window: outcome fail rate (with n), malformed / empty-200 rate,
+ 5xx and timeout count, breaker trips, p95 latency, estimate-vs-bill error,
+ cache rate, cost per success (W4). This is the "who's misbehaving" view the
+ preclusion decision needs.
+- **Q2. Same model, different carriers.** Group scorecard rows by base model,
+ e.g. `qwen3.6-35b` on NeuralWatt vs OpenRouter. This separates "the model is
+ bad at this" from "this carrier serves it badly", which calls for a
+ different fix: drop the carrier's row, or restrict the model.
+- **Q3. Graded actions** in place of deprecate-or-not. Each action records a
+ reason and links its evidence in X1:
+ 1. exclude from category X. This **cannot** reuse `eligible_categories`:
+ that field is restrict-only (`admin.py`, the probe docstring), and NULL
+ means "every category", so a deny needs its own column,
+ 2. exclude when the request carries tools (the `deepseek-v4-flash` case).
+ This may be the way out of the frozen `tool_use_agentic` data:
+ `min_tool_proficiency` can only read scores that stopped moving on
+ 2026-09-15, while a per-model manual rule lets the operator decide from
+ Q1 scorecard evidence instead,
+ 3. restrict to batch / `-flex` only,
+ 4. quarantine: exploration-only, so it keeps collecting evidence without
+ carrying traffic (see the open question below; deferred until Q1 shows
+ it is needed),
+ 5. timed ban, e.g. 24h, auto-expiring and logged,
+ 6. deprecate (what exists today).
+
+ Each action goes through at least the minimal X3 check (would this empty a
+ candidate set for any recent decision?) before committing. That check alone
+ would have caught both the 2026-09-01 and 2026-09-04 incidents; the full
+ winner-shift replay can come later.
+- **Q4. Breaker visibility.** Show open circuits with `down_until`, current
+ cooldown and trip count; keep a trip history (persist it, since `_store`
+ is process-lifetime); add manual force-open (drain a model on purpose) and
+ force-close (reset after a known fix).
+- **Q5. A quality breaker, separate from availability.** Trip on content
+ failures (empty-200, mangled output, a burst of `ok:false`) using the
+ novelty-or-rate rule already in `rejection_warnings`. This is roughly what
+ `plans/mangled-output-detection.md` specs (status: planned). Keep it a
+ separate breaker and panel so a quality trip doesn't read as an outage.
+ **One doc owns it:** fold Q5 into `mangled-output-detection.md` rather than
+ specifying it twice, and let this file point there.
+- **Q6. Carrier-level breaker.** When several models on one provider trip
+ inside a short window, open the provider rather than walking its models one
+ cooldown at a time. Account exhaustion already has its own path; this is for
+ partial carrier outages.
+
+---
+
+## 5. Quality of life: ergonomics and readouts
+
+Unlike the sections above, this one **was** read against the live portal
+(8080, view-only, 2026-09-23), plus the frontend source. Each item names the
+thing observed.
+
+### Readouts that mislead today
+
+- **R1. Home tiles use different windows.** Quota is "this period",
+ Decisions 7d, Models and Pinch 30d. Busiest models shows `kimi-k2.7-code`
+ at 10,028 while the Decisions tile says 2,887 for everything, and both are
+ correct. Label each tile's window on its face, or drive all of them from
+ the Activity card's 24h / 7d / 30d selector.
+- **R2. The Decisions tile counts verifications, and mixes diagnostics with
+ ground truth.** It reads "2,887 verified in 7 days: 132 ok, 93 failed", but
+ `index.html` sums `verdict_mix`, which counts `verifications` rows, not
+ decisions. Checked read-only against the live DB on 2026-09-23:
+ | | tile | actually |
+ |---|---|---|
+ | headline | 2,887 | 2,678 route decisions; 2,887 is verification rows, 2,660 of them `unverifiable` |
+ | "ok" | 132 | 96 client `succeeded` + 29 `local_llm` ok + 7 structural ok |
+ | "failed" | 93 | 82 client `failed` + 11 `malformed` (local_llm + structural) |
+ The tile adds structural and `local_llm` verdicts, which CLAUDE.md calls
+ diagnostics only, to `/outcome` reports, the only ground truth. Show route
+ decisions as the headline, then client outcomes on their own: "178 client
+ reports (6.6%): 96 ok, 82 failed (46%)". Diagnostics, if shown at all, go on
+ a separate line.
+- **R3. The Activity chart smooths across gaps.** The decisions series is a
+ spline through sparse points, so it draws continuous traffic across hours
+ that had none. Break the line at empty buckets or use bars. Separately,
+ billed requests exceed decisions at several peaks; the chart should say what
+ the excess is (cloud classifier calls? retries? unrouted pins?) rather than
+ leave two lines that disagree unexplained.
+- **R4. Sparklines have no values.** The tile sparklines carry no axis and no
+ hover. Add a hover value, or min and max labels.
+- **R5. The Classifier card flashes a wrong value.** On first paint the mode
+ select reads `local_llm` and only switches to `local_encoder` once config
+ loads. For a second, the card shows a value that isn't configured. Render
+ a loading state instead of the default.
+- **R6. Config echo vs live state.** The Classifier card shows what the
+ overlay says (`local_encoder`, device `cpu`), not what the process actually
+ loaded. Show the resolved model, device and recent p50 latency from the
+ running classifier. (It also exposed doc drift: CLAUDE.md says the live
+ deployment runs `device: cuda`; the overlay says `cpu`.)
+- **R7. Proficiency colour encodes evidence, not score.** Green / amber mean
+ measured / thin, so the best model in a column isn't visible without reading
+ every number. Mark each column's leader, outline the `quality_tolerance`
+ band (P3), make columns sortable, and add a "routable only" toggle so rows
+ that are deprecated or not allowlisted stop padding the matrix.
+
+### Decisions page ergonomics
+
+- **E1. Collapse runs.** The top 40 rows are one session: same category,
+ profile, tier, source and model, and only ctx and cost move. Fold
+ consecutive same-session, same-model rows into one expandable row: "38 turns,
+ ctx 75k to 98k, $0.27 total". Switches then stand out as row boundaries,
+ which is what W2 wants to surface.
+- **E2. Session view.** A per-session rollup: turns, total cost, context
+ growth, switches, outcome count. Context climbing about 1k per turn is
+ visible in the raw table and invisible everywhere else.
+- **E3. Filters in the URL.** `decisions.html` never reads `URLSearchParams`,
+ so no other page can link to "decisions for this model" or "this session".
+ X2 actions, the Q1 scorecard and the model modal all need that link.
+- **E4. Row drill-down.** A row click does nothing. Open a panel with the
+ decision's rejected candidates, verification verdict, `/outcome` report,
+ and estimated vs billed cost from its energy observation.
+- **E5. Server-side filtering.** The page loads 1,000 of 32,540 rows and
+ filters and searches only those, with a footnote saying so. Push filters to
+ the API so that "All" means all.
+- **E6. Model, provider and source as filters.** Today they're reachable only
+ through free-text search.
+
+### Controls page
+
+- **C1. One knob table, not two.** Runtime Knobs and Persisted Config list
+ largely the same knobs in two columns, in different orders, with persisted
+ labels truncated (`objective.incumbent_cache_pri…`). Checking drift means
+ matching rows by eye. Use one table: knob, live value, persisted value,
+ layer (base / overlay), drift badge. That table is also X1's natural home.
+- **C2. Separate Restart Service.** It sits in the same button row as Refresh
+ Catalog and Seed Energy. Move it apart, and have its confirmation list which
+ runtime values differ from persisted and will revert.
+- **C3. Prose in cards.** Local Compute carries three paragraphs; the portal
+ style rule is no paragraphs. Keep one line plus a docs link or tooltip.
+- **C4. Default and last change per knob.** Show the code default beside
+ each value, plus "changed 2h ago from 0.3" once X1 exists.
+
+### Portal-wide
+
+- **G1. Nav lives in eight files.** Each page hardcodes the same `
` of
+ nav links. `navbar.js` already notes that adding Quota was an eight-file
+ edit; have it render the links too.
+- **G2. Width.** Most pages sit in `container-xl` and use well under half of
+ a wide screen, while Decisions, the densest table, is the most cramped.
+ Proficiency already goes full-width. Go fluid on the data-heavy pages.
+- **G3. Dismissals don't stick on rate warnings.** Dismissal is keyed on the
+ warning's exact text, on purpose (a changed condition resurfaces it). But
+ warnings that embed a live rate ("3.2x pace") change text every poll, so a
+ dismissal never holds. X2's `class` + `subject` is the right key; resurface
+ on a severity change, not a digit change.
+- **G4. Keyboard.** `/` to focus search, `Esc` to close modals, and `j`/`k`
+ on the Decisions table. Cheap, and this is a page someone lives in.
+
+---
+
+## One thing the cockpit framing makes more urgent
+
+The portal is loopback-only with no auth, and it already includes
+`/api/restart-service`, `/api/apply-feedback` and provider deletes. The more
+it becomes the place decisions are made, the stronger the pull to reach it
+from other machines, and moving the router onto a Proxmox box on the LAN is
+exactly that move. Add auth before binding anything but `127.0.0.1`. Same
+warning CLAUDE.md already gives for the API, with more at stake.
+
+---
+
+## A possible order
+
+| # | item | why here |
+|---|---|---|
+| 0 | portal auth | **only if** the router is leaving `127.0.0.1` (the Proxmox move); blocks that move, not this list |
+| 1 | X1 change log + operator epoch | everything else reads it |
+| 2 | Q1 scorecard + Q4 breaker visibility | the preclusion decision needs one view first |
+| 3 | Q3 graded actions (via X1) + minimal X3 (empties-a-candidate-set check) | turns the scorecard into decisions without repeating 09-01 / 09-04 |
+| 4 | X2 structured warnings | lets warnings trigger Q3 actions directly |
+| 5 | W1 waste ledger + W2 switch reasons | token waste in dollars, with causes |
+| 6 | X3 full replay preview (winner shift, cost delta) | makes knob changes safe to try |
+| 7 | P1 drift, P2 evidence priority | proficiency monitoring over time |
+| 8 | Q5 quality breaker (owned by `mangled-output-detection.md`), D1 timed changes | automate what the operator has been doing by hand |
+| 9 | D3 session A/B, Q6 carrier breaker | only after the above prove out |
+
+Section 5 sits outside this order: most of it is small and independent.
+Quick wins that also unblock the numbered items: **E3** (URL filters, needed
+by X2 and Q1), **R2** (outcome coverage, the same honesty W1 and W4 need),
+**C1** (one knob table, X1's home), **G3** (fixed by X2's class keys). R5 and
+G1 are fixes of a few lines each.
+
+**North Star #1 applies to every item here.** Q5 thresholds, D1 durations,
+timed-ban lengths and any quarantine settings are new config knobs, so each
+ships with its admin control or `test_admin_knob_coverage.py` fails. Scope the
+control into the item, not a follow-up.
+
+### Open questions, with proposed answers
+
+- **Is the change log append-only history, or also an undo stack?**
+ Proposed: append-only, plus a per-row "revert to old value" that writes one
+ key and appends its own row. That is not the shape CLAUDE.md rejected for
+ gaming mode: that concern was one flag writing five keys and drifting apart.
+ A one-key revert can't drift, and the history stays intact.
+- **Should graded exclusions live in the DB or in `config.local.yaml`?**
+ Proposed: the DB, next to `admin_model_overrides` (extend it or add a
+ sibling table). Timed bans need an expiry, and an expiry belongs in a DB
+ column. The overlay is per-machine and not in git, so it's the wrong home
+ for judgments about a carrier. And availability overrides already live in
+ the DB. The category exclusion still needs a new deny column (see Q3.1).
+- **Does quarantine conflict with session-scoped exploration?** Yes, and the
+ weight goes up: a quarantined model would carry whole sessions, not single
+ turns. There's a second problem too. `exploration.py` only chooses among
+ models that already passed the hard filters and were ranked, so quarantine
+ needs a new state, "eligible to explore, barred from winning", rather than
+ reusing anything existing. Defer it until Q1 shows it's needed.
diff --git a/plans/cockpit-quick-wins.md b/plans/cockpit-quick-wins.md
new file mode 100644
index 0000000..f1bd843
--- /dev/null
+++ b/plans/cockpit-quick-wins.md
@@ -0,0 +1,217 @@
+# Cockpit quick wins: six small admin-portal fixes
+
+Status: done -- shipped in PR #101
+Date: 2026-09-23
+Source: `plans/cockpit-brainstorm.md` section 5 (items R2, R5, E3, C1, G1, G3).
+Each item below was observed on the live portal or checked read-only against
+the live `router.db` on 2026-09-23.
+
+## Ground rules (all items)
+
+- **Worktree, not the main checkout.** Create a worktree from `main` HEAD
+ (`c87e675` or later) on branch `feat/cockpit-quick-wins`. The main checkout
+ at `/home/alee/Sources/6krrt` has UNCOMMITTED edits from other work,
+ including `admin/frontend/controls.html` (a `confidence_threshold` ->
+ `confidence_min` rename around lines 1026 and 1150) and
+ `tests/test_admin_frontend.py`. Do not touch, stash, or commit them. C1 edits
+ a different region of `controls.html`; keep it that way so the later merge
+ is clean.
+- Commit by explicit path. Never `git add -A` or `git add .`.
+- **Never touch port 8080** (production). For visual checks, run a throwaway
+ instance on **8081** from the worktree, with a temp copy of the DB, never the
+ live `router.db` for writes.
+- Never edit `config/config.yaml` or `config/config.local.yaml`.
+- ASCII only in new code, comments, and UI strings. No middle-dot separators
+ (U+00B7); relate facts with layout, or with `:` or a rephrase.
+- No paragraphs of prose in the UI. One line per hint at most.
+- One commit per item, in the order listed. `pytest` (offline) and
+ `ruff check` stay green after each commit.
+- For every UI change, take a screenshot on 8081 and check the changed area
+ zoomed in, not just the item's checklist.
+
+---
+
+## 1. R5: Classifier card shows a wrong value on first paint
+
+**Problem.** `admin/frontend/controls.html:264-268`: the
+`#classifier-mode-select` `