docs: Lift A docs pass, plus watchdog exits 0 after a tick #103

Merged
alee merged 14 commits from docs/lift-a-docs-pass into main 2026-09-27 02:33:41 +00:00
17 changed files with 2884 additions and 30 deletions

View File

@@ -8,7 +8,7 @@ next steps, and is the one to trust on what is currently true.
## NORTH STAR GUIDELINES ## NORTH STAR GUIDELINES
Three rules that outrank local cleverness. Each exists because it was broken Four rules that outrank local cleverness. Each exists because it was broken
first and the breakage was expensive to find. When a change conflicts with one first and the breakage was expensive to find. When a change conflicts with one
of these, the change is wrong — not the rule. of these, the change is wrong — not the rule.
@@ -87,6 +87,36 @@ before measuring anything, and open it read-only:
`sqlite3.connect("file:...?mode=ro", uri=True)` — the `sqlite3` CLI here does `sqlite3.connect("file:...?mode=ro", uri=True)` — the `sqlite3` CLI here does
not accept `-uri`. not accept `-uri`.
### 4. Waste is surfaced in the admin portal, and stopping it is one click
When the router can see money being wasted, it shows the operator where they
already look, with the evidence and the lever to stop it side by side. A
detector that only writes a log line, or a fix that needs a config edit, a
restart or a long table scan, has not met this rule.
"Waste" here means **spend with no concrete change landing**: an agent session
looping, re-reading, or retrying the same failure, and a model that keeps
producing such sessions. It does **not** mean steady spend. A healthy agent run
can burn for hours, and a spend-rate alarm cannot tell the two apart.
In practice:
- **Stalled sessions are visible** in the portal with their evidence: turns and
$ since the last landed change, and the top repeated target. They also reach
the operator when nobody is watching (desktop alert).
- **A model that keeps producing them is visible** as a per-model rollup, with
**Block model** beside the evidence. The block records its reason, shows in a
Blocked list, and is one click to undo.
- **Automatic responses come after the visible one,** never instead of it, and
every automatic action shows up in the same place.
**Why:** incident #8 (`docs/incidents.md`). Agent sessions looped for hours on
2026-09-25/26: one worker read the same file 61 times, a planner re-read a spec
68 times its length, and a model confabulated truncation that was not there.
Every existing check stayed green. It was caught only by a human, or a Claude
session, reading opencode's session store by hand. Pulling the model took a trip
through a dropdown whose vocabulary is catalog `deprecated`. The detection
existed nowhere, and the lever existed only for someone who already knew.
## What this is ## What this is
A router that uses a local model (served via Ollama) to classify incoming A router that uses a local model (served via Ollama) to classify incoming
@@ -319,7 +349,12 @@ rather than from months of history.
- `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md). - `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md).
- `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below. - `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below.
- `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted. - `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted.
- `tests/` — 1155 tests across 40+ files, offline, verified on Python 3.10 and 3.14. [README](README.md). - `progress_detect.py` — loop-detection signals over a window of per-session probe calls: duplicate-bulk (`dup_min`), top-similarity (`top_min`/`top_min_ro`), slow-progress (`cum_min`) and coverage (`cover_min`) heuristics, gated on `min_calls`. See [watchdog](docs/watchdog.md).
- `watchdog.py` — the per-session watchdog loop: every ~5 minutes it judges each session with a tool call since the last tick, on its full history, and writes a quiet `no_opencode` tick when opencode is not running. See [watchdog](docs/watchdog.md).
- `notifier.py` — alert fan-out: desktop/`notify-send` plus per-channel `min_severity` and a rate limit. See [watchdog](docs/watchdog.md).
- `watchdog_store.py` — `watchdog_ticks`, `watchdog_verdicts`, `watchdog_alerts`, `watchdog_channel_settings` (four tables + indexes). See [watchdog](docs/watchdog.md).
- `router-link.js` — the opencode plugin; exported as a factory with `parentCache` as a property, because opencode 1.18.x rejects the whole plugin when any export is not a function. See [watchdog](docs/watchdog.md).
- `tests/` — 2414 tests across 108 files, offline, verified on Python 3.10 and 3.14. [README](README.md).
## #45 — OpenRouter is an opt-in allowlist provider ## #45 — OpenRouter is an opt-in allowlist provider
@@ -1261,10 +1296,14 @@ be on.
### When the router goes unreachable, start at docs/incidents.md ### When the router goes unreachable, start at docs/incidents.md
Four incidents so far, all sharing one shape: a change that looked local to the Eight incidents so far, nearly all sharing one shape: a change that looked local
router silently degraded the agent depending on it, and none announced itself as to the router silently degraded the agent depending on it, and none announced
a router problem. **`docs/incidents.md` carries the full write-ups plus a itself as a router problem. **`docs/incidents.md` carries the full write-ups plus
symptom -> one-line-check table**; read it rather than re-deriving a diagnosis. a symptom -> one-line-check table**; read it rather than re-deriving a diagnosis.
#8 is the exception worth knowing before an unattended agent run: the router
worked perfectly while agent workers looped for hours with no progress, and no
check noticed, because every check watched spend or availability rather than
whether changes landed (`plans/no-progress-detection.md`).
Two conventions from those incidents that bind every session, and so stay here: Two conventions from those incidents that bind every session, and so stay here:
@@ -1283,6 +1322,14 @@ Recovery for an unreachable-but-`active` service is
`systemctl --user restart llm-router.service` -- a hung process was never in a `systemctl --user restart llm-router.service` -- a hung process was never in a
tracked stop job, so this issues a fresh cycle systemd does enforce a timeout on. tracked stop job, so this issues a fresh cycle systemd does enforce a timeout on.
The watchdog runs as its own systemd **timer** (`llm-router-watchdog.timer`),
separate from the dispatcher, so a stuck router cannot silence the thing meant
to notice it is stuck. Install and enable it with `systemctl --user enable
--now llm-router-watchdog.timer` (the timer's unit file ships pointing at a
placeholder home path and must be `sed`-repointed to the real one first), or
run it once by hand with the `--once` flag. See [watchdog](docs/watchdog.md)
for the signals it fires on, the alert lifecycle, and the known limits.
## Pointing a coding agent at it ## Pointing a coding agent at it
The `/v1` endpoints are OpenAI-compatible, so any normal client works — The `/v1` endpoints are OpenAI-compatible, so any normal client works —

View File

@@ -35,11 +35,30 @@ for the full model.
<img src="../assets/screenshots/dashboard.png" alt="Admin dashboard" width="820"> <img src="../assets/screenshots/dashboard.png" alt="Admin dashboard" width="820">
</div> </div>
**Loops** — the dashboard carries a Loops card (`admin/frontend/index.html#loops`)
showing the watchdog's open alerts, one row per alert. It pulls from
`GET /admin/api/watchdog/loops` (`watchdog_alerts` rows where `resolved_at` is
null) via `loadLoops()`. Each row shows exactly three things: a severity badge
(`critical` / `warning` / else `info`), the raw dedup key
(`opencode-loop:<root session id>`), and `opened_at`. That is all the panel
shows today — it does **not** show the agent, model, top repeated target, calls,
or $ since the last landed change; those live in `watchdog_verdicts` and have
not shipped to this panel yet.
**Models** — `GET /admin/models` is the model-availability table. Each row shows **Models** — `GET /admin/models` is the model-availability table. Each row shows
the serving class, tier, status, and an override dropdown (`active` / the serving class, tier, status, and an override dropdown (`active` /
`deprecated` / `stale`) that writes through to the routing hard filters. By `blocked` / `deprecated` / `stale`) that writes through to the routing hard
default the table hides deprecated and stale rows; enable the **Show deprecated filters. By default the table hides deprecated and stale rows; enable the
/ stale** toggle to include them. **Show deprecated / stale** toggle to include them.
`blocked` is a fourth value in that same per-model override dropdown, distinct
from the catalog meaning of `deprecated`. It is an operator stop: a model an
operator has switched off, not one the catalog flagged. Routing and pinned
requests both exclude a blocked model (`_admin_excluded_models` in
`dispatcher.py` and `metrics.py`). Undo is setting the dropdown back to
`active` (clearing the override via `DELETE` on the availability path), not a
separate unlock — there is no per-model Block button beside evidence and no
Blocked list with reasons; those belong to the "not built yet" item below.
<div align="center"> <div align="center">
<img src="../assets/screenshots/models.png" alt="Admin models page" width="820"> <img src="../assets/screenshots/models.png" alt="Admin models page" width="820">
@@ -126,8 +145,8 @@ land in `config/config.local.yaml`; see
[config-local-overlay](config-local-overlay.md)). Changes are marked as dirty [config-local-overlay](config-local-overlay.md)). Changes are marked as dirty
and written only on save. and written only on save.
It carries two dedicated cards, neither a row in the generic runtime-knob It carries three dedicated cards, none a row in the generic runtime-knob
list, because both have cross-field structure a flat scalar/boolean input list, because each has cross-field structure a flat scalar/boolean input
can't safely represent: can't safely represent:
- **Local Compute** — "gaming mode". See [Gaming mode](#gaming-mode) below. - **Local Compute** — "gaming mode". See [Gaming mode](#gaming-mode) below.
@@ -145,6 +164,14 @@ can't safely represent:
anything reaches `config.local.yaml`, and mode plus its companion block are anything reaches `config.local.yaml`, and mode plus its companion block are
written as one atomic change so an in-between invalid state is never even written as one atomic change so an in-between invalid state is never even
written transiently. written transiently.
- **Watchdog** — the watchdog's live state (`controls.html#watchdog-card`,
`loadWatchdog()`): the last tick time (and a `no tick` badge before the
first one), how many sessions the last tick saw, how many verdicts it
flagged, and how many alerts are currently open. Below those, a set of
**notification channel** toggles (each channel with an enable/disable and a
minimum severity, persisted via `POST /admin/api/watchdog/channels`), a
**Refresh** button, and a **Send test alert** button (`POST
/admin/api/watchdog/test-alert`).
<div align="center"> <div align="center">
<img src="../assets/screenshots/controls.png" alt="Admin controls page" width="820"> <img src="../assets/screenshots/controls.png" alt="Admin controls page" width="820">
@@ -444,7 +471,15 @@ intermediate invalid state could land on disk. The GET response includes
**Model availability overrides** **Model availability overrides**
`POST /admin/api/models/{model_id:path}/{provider}/availability` marks a model `POST /admin/api/models/{model_id:path}/{provider}/availability` marks a model
as `active`, `deprecated`, or `stale`. `DELETE` on the same path removes the as `active`, `blocked`, `deprecated`, or `stale`. `DELETE` on the same path
override. The `{model_id:path}` converter accepts model ids that contain removes the override. The `{model_id:path}` converter accepts model ids that
slashes, so providers like OpenRouter with slash-bearing ids are handled the contain slashes, so providers like OpenRouter with slash-bearing ids are
same way as plain ids. Deprecation feeds into the routing hard filters. handled the same way as plain ids. Deprecation feeds into the routing hard
filters.
**Not built yet.** The Loops panel does not yet show the evidence columns an
operator would act on — agent, model, top repeated target, calls, and $ since
the last landed change — and there is no per-model stall rollup, no **Block
model** button beside the evidence, and no Blocked list with reasons. Those are
the remaining surface of [North Star #4](../CLAUDE.md) in CLAUDE.md, whose first
half (conversation identity and the watchdog) is shipped.

View File

@@ -173,6 +173,12 @@ These ids are what make `/outcome` attribution exact (below): an
`X-Router-Conversation` on the request sets the `session_key` row that a later `X-Router-Conversation` on the request sets the `session_key` row that a later
report can address by `conversation_id` without guessing `source` or ambiguity. report can address by `conversation_id` without guessing `source` or ambiguity.
**Plugin loader contract:** the opencode plugin that stamps these headers
must export only functions — the opencode 1.18.x plugin loader rejects a
plugin outright if any export is not a function (so the plugin's
`parentCache` lives as a property on the factory rather than a named export).
See [clients.md](clients.md) for the full contract.
**Capability 422s**: When no model survives the hard filters, the 422 names the **Capability 422s**: When no model survives the hard filters, the 422 names the
active constraints. That now includes "vision-capable model" or active constraints. That now includes "vision-capable model" or
"json-mode-capable model" when the request carried images or a JSON-mode "json-mode-capable model" when the request carried images or a JSON-mode
@@ -225,6 +231,65 @@ but a `true` report counts too. It's also the only quality signal that survives
streaming, since a retry can't reach a response whose bytes are already gone, streaming, since a retry can't reach a response whose bytes are already gone,
while a report arrives afterward and works either way. while a report arrives afterward and works either way.
## Admin watchdog and availability API
The admin portal's watchdog and model-availability operations sit under
`/admin` (loopback-only, like the rest of the admin API). The watchdog
endpoints back the dashboard's **Loops** card and the **Watchdog** control
card; full behaviour is described in [watchdog.md](watchdog.md) and
[admin-portal.md](admin-portal.md).
| Method | Path | Description |
|---|---|---|
| `GET` | `/admin/api/watchdog/status` | Latest watchdog tick plus open-alert count |
| `GET` | `/admin/api/watchdog/loops` | Open alert rows (one per looping session) |
| `GET` | `/admin/api/watchdog/channels` | Notification-channel settings |
| `POST` | `/admin/api/watchdog/channels` | Create or update a channel's enabled/min-severity |
| `POST` | `/admin/api/watchdog/test-alert` | Deliver a test alert through the enabled channels |
`GET /admin/api/watchdog/status` returns a single object:
- `last_tick` — the most recent `watchdog_ticks` row, or `null` before the
first tick
- `verdicts` — `{total, flagged}` for that tick's `watchdog_verdicts`, or
`null` when there is no tick yet
- `open_alerts` — count of `watchdog_alerts` rows with `resolved_at` null
`GET /admin/api/watchdog/loops` returns an array of `watchdog_alerts` rows
where `resolved_at` is null, ordered by `opened_at` descending. Each element
is the full alert row.
`GET /admin/api/watchdog/channels` returns an array of
`watchdog_channel_settings` rows ordered by `channel_name`.
`POST /admin/api/watchdog/channels` upserts one channel. Body:
- `channel_name` (required) — the channel to create or update; absence is a
422
- `enabled` (optional boolean) — when omitted the existing value is kept
- `min_severity` (optional: `info`, `warning`, or `critical`) — any other
value is a 422
Returns `{ok: true}`.
`POST /admin/api/watchdog/test-alert` delivers a test alert through the
currently enabled channels. Body:
- `severity` (optional, default `warning`; `info`, `warning`, or `critical`)
- `channel_name` (optional) — restrict to a single channel
Returns `{sent: <n>, severity: <severity>}`, or `{sent: 0, message: "no
enabled channels"}` when nothing is enabled. An invalid `severity` is a 422.
**Model availability** — `POST
/admin/api/models/{model_id:path}/{provider}/availability` upserts an admin
override for a model's availability. `availability` must be one of `active`,
`blocked`, `deprecated`, or `stale` (anything else is a 422); `blocked` is an
operator stop that routing and pinned requests exclude
(`_admin_excluded_models`), distinct from the catalog meaning of `deprecated`.
`DELETE` on the same path removes the override. The `{model_id:path}`
converter accepts slash-bearing model ids (e.g. OpenRouter).
## Pinning and auto behavior for local models ## Pinning and auto behavior for local models
`model: "auto"` will route eligible tasks to `qwen2.5-coder-router:14b` when the `model: "auto"` will route eligible tasks to `qwen2.5-coder-router:14b` when the

View File

@@ -37,3 +37,11 @@ This resolved a real incident where reading a dependency's source 24 times
outweighed writing to the project dir 22 times: the write-like 3× weighting outweighed writing to the project dir 22 times: the write-like 3× weighting
puts the project directory that actually received the edits ahead of a puts the project directory that actually received the edits ahead of a
dependency's source that was merely read. dependency's source that was merely read.
**Plugin loader contract:** opencode 1.18.x (the observed version) rejects the
entire plugin if *any* export is not a function. The `parentCache` map on
`router-link.js` is therefore set as a property on the `RouterLink` factory
function rather than as a named export (L221-224 of
`deploy/opencode-plugin/router-link.js`). Confirm the plugin loaded by
searching the log for `failed to load plugin`: `grep 'failed to load plugin'
~/.local/share/opencode/log/opencode.log`.

View File

@@ -31,7 +31,7 @@ Three data tables plus one observability table, `PRAGMA foreign_keys = ON`:
| `access_level` | TEXT | `public` \| `preview` \| `canary` | | `access_level` | TEXT | `public` \| `preview` \| `canary` |
| `pricing_tbd` | INTEGER | | | `pricing_tbd` | INTEGER | |
| `deprecated` | INTEGER | | | `deprecated` | INTEGER | |
| `availability` | TEXT | `active` \| `deprecated` \| `stale` | | `availability` | TEXT | `active` \| `deprecated` \| `stale` \| `blocked` — the `blocked` value is an operator stop set via the admin override (see [admin-portal.md](admin-portal.md)); it is excluded from routing and pinned requests |
| `last_updated` | TEXT | ISO8601 | | `last_updated` | TEXT | ISO8601 |
**Serving class:** Neuralwatt ships ~6 base models as 19 catalog rows. The id **Serving class:** Neuralwatt ships ~6 base models as 19 catalog rows. The id
@@ -326,3 +326,71 @@ prompts, and answers are deliberately excluded. A test enforces that the write
path does not store prompt or answer text. The write is gated by path does not store prompt or answer text. The write is gated by
`logging.log_route_decisions` and is **best-effort**: a failed write is logged `logging.log_route_decisions` and is **best-effort**: a failed write is logged
at warning and swallowed so monitoring cannot slow or fail a request. at warning and swallowed so monitoring cannot slow or fail a request.
## Watchdog Schema (SQLite)
Four tables written by `watchdog.py` via `src/watchdog_store.py` (the code-side
inline-create mirrors this schema for live databases). See
[watchdog.md](watchdog.md) for how the loop reads them.
### `watchdog_ticks` — one row per detection run
| Column | Type | Notes |
|---|---|---|
| `id` | INTEGER | Autoincrement |
| `ticked_at` | TEXT | ISO8601 run time |
| `sessions_seen` | INTEGER | Count of sessions examined this run, default 0 |
| `outcome` | TEXT | |
### `watchdog_verdicts` — one row per session per tick
| Column | Type | Notes |
|---|---|---|
| `id` | INTEGER | Autoincrement |
| `tick_id` | INTEGER | FK → `watchdog_ticks.id`, NOT NULL |
| `session_id` | TEXT | NOT NULL |
| `session_root` | TEXT | |
| `agent` | TEXT | |
| `model_id` | TEXT | |
| `provider` | TEXT | |
| `flagged` | INTEGER | 0/1, default 0 |
| `dup` | REAL | Duplicate-fraction signal |
| `top` | INTEGER | |
| `top_what` | TEXT | |
| `landed` | INTEGER | |
| `slow` | INTEGER | |
| `coverage` | REAL | |
| `calls_since_landed` | INTEGER | |
| `cost_since_landed_usd` | REAL | |
| `llm_second_opinion` | TEXT | |
| `created_at` | TEXT | NOT NULL |
Four indexes (`config/schema.sql` L482-485):
| Index | On |
|---|---|
| `idx_watchdog_verdicts_model_id` | `watchdog_verdicts(model_id)` |
| `idx_watchdog_verdicts_created_at` | `watchdog_verdicts(created_at)` |
| `idx_watchdog_verdicts_flagged` | `watchdog_verdicts(flagged)` |
| `idx_watchdog_verdicts_session_root` | `watchdog_verdicts(session_root)` |
### `watchdog_alerts` — one row per open alert
| Column | Type | Notes |
|---|---|---|
| `dedup_key` | TEXT | PRIMARY KEY; e.g. `opencode-loop:<root session id>` |
| `state` | TEXT | NOT NULL |
| `severity` | TEXT | NOT NULL |
| `flagged_ticks` | INTEGER | Default 0 |
| `opened_at` | TEXT | NOT NULL |
| `last_fired_at` | TEXT | NOT NULL |
| `resolved_at` | TEXT | |
### `watchdog_channel_settings` — per-channel notification config
| Column | Type | Notes |
|---|---|---|
| `id` | INTEGER | Autoincrement |
| `channel_name` | TEXT | NOT NULL, UNIQUE |
| `enabled` | INTEGER | Default 1 |
| `min_severity` | TEXT | Default `warning` |

View File

@@ -1,14 +1,18 @@
# Incidents: how this router has broken, and how to tell which one it is # Incidents: how this router has broken, and how to tell which one it is
Seven times now, a change that looked local to the router has silently Eight times now, something around the router has silently degraded either the
degraded either the agent depending on it or the operator trying to see it agent depending on it or the operator trying to see it clearly. They share a
clearly. They share a shape worth naming: **none of them announce themselves shape worth naming: **none of them announce themselves as router problems.**
as router problems.** Five of the seven presented as an opaque client-side Five of the eight presented as an opaque client-side error — a connection
error — a connection refused, an "Unprocessable Content", an "internal server refused, an "Unprocessable Content", an "internal server error" — and
error" — and diagnosing each meant knowing which log or table to look in. The diagnosing each meant knowing which log or table to look in. The other three
other two didn't error toward the client at all: #5 destroyed data outright, didn't error toward the client at all: #5 destroyed data outright, #6 just went
and #6 just went quiet in one corner of the admin UI while `/health` stayed quiet in one corner of the admin UI while `/health` stayed green the whole time,
green the whole time. and #8 spent money for hours while every check stayed green.
**#8 is the only one so far that cost money rather than capability**, and the
router behaved as configured throughout. Read it before leaving
an agent run unattended.
**#5 is the only one so far that destroyed data**, and the only one where **#5 is the only one so far that destroyed data**, and the only one where
recovery depended on luck rather than design. Read it before running any recovery depended on luck rather than design. Read it before running any
@@ -21,6 +25,15 @@ or removing any config key.
This page exists so the next one takes minutes rather than hours. Start with the This page exists so the next one takes minutes rather than hours. Start with the
symptom table, then read only the relevant section. symptom table, then read only the relevant section.
**How to write an entry (from #8 on).**
- Separate **Evidence** (measured, with where it was measured) from **Theory**
(inferred). Anything not established by a measurement or a reproduction goes
under Theory. It states what supports it and what would confirm or refute
it, and stays there until that test is done.
- Include a **Human response** timeline: what the operator noticed, decided and
did. Mark actions an assistant took, and at whose direction.
- Entries #1 to #7 predate this convention.
`opencode.json` points opencode's own model traffic at `http://127.0.0.1:8080/v1`, `opencode.json` points opencode's own model traffic at `http://127.0.0.1:8080/v1`,
so on this machine the router *is* the coding agent's inference supply. Breaking so on this machine the router *is* the coding agent's inference supply. Breaking
it breaks the thing you would use to fix it. That is why these keep happening, it breaks the thing you would use to fix it. That is why these keep happening,
@@ -38,6 +51,8 @@ and why the diagnostics below are worth having to hand.
| 500s + `no such table`, venv/.env gone | #5 `git clean -fdx` | `ls -la router.db .env .venv` — a 0-byte db and a missing `.venv` is conclusive | | 500s + `no such table`, venv/.env gone | #5 `git clean -fdx` | `ls -la router.db .env .venv` — a 0-byte db and a missing `.venv` is conclusive |
| A newly-shipped admin knob just isn't there, `/health` is fine | #6 backend running stale code | `systemctl --user status llm-router` uptime vs. `git log -1 --format=%cd` on the commit that added the knob | | A newly-shipped admin knob just isn't there, `/health` is fine | #6 backend running stale code | `systemctl --user status llm-router` uptime vs. `git log -1 --format=%cd` on the commit that added the knob |
| `activating (auto-restart)`, `pydantic_core...ValidationError: ...extra_forbidden` in the journal | #7 renamed key vs. un-migrated overlay | `journalctl --user -u llm-router --since '5 min ago' \| grep extra_forbidden` then `grep -n <old key name> config/config.local.yaml` | | `activating (auto-restart)`, `pydantic_core...ValidationError: ...extra_forbidden` in the journal | #7 renamed key vs. un-migrated overlay | `journalctl --user -u llm-router --since '5 min ago' \| grep extra_forbidden` then `grep -n <old key name> config/config.local.yaml` |
| Spend is high after an agent run, nothing errored, no warning fired | #8 agent looping with no progress | `sqlite3 router.db "select session_key, strftime('%H',observed_at,'localtime') h, round(sum(prompt_tokens)/1e6) mtok, round(sum(cost_usd),2) usd from energy_observations where observed_at > datetime('now','-12 hours') group by 1,2 order by usd desc limit 8;"` then check whether that run's commits actually landed |
| opencode plugin does nothing | #8 plugin loader rejects non-function export | `grep 'failed to load plugin' ~/.local/share/opencode/log/opencode.log` |
That last row is not an incident yet — it is the silent-staleness failure That last row is not an incident yet — it is the silent-staleness failure
described in the "Run as a service" section of `CLAUDE.md`. An unpolled catalog described in the "Run as a service" section of `CLAUDE.md`. An unpolled catalog
@@ -347,6 +362,253 @@ file is correct.
--- ---
## #8 — Agent sessions looped for hours with no progress, and nothing noticed (2026-09-25)
**Symptom.** Nothing errored. The router stayed healthy, `/metrics` showed zero
warnings, and the quota alarm read `kind: none`. The operator noticed early on
2026-09-26 that an opencode run had "burned a bunch of tokens over the hours".
### Evidence
Measured read-only against the live `router.db` and opencode's own session
store. The window is 2026-09-25 19:00 to 2026-09-26 02:00 local unless noted.
**Spend.**
| | |
|---|---|
| prompt tokens | 323M, 95% cached |
| billed | about $7.81, roughly 3x the busiest of the previous ten days ($2.66 on 09-20) |
| largest share | OpenRouter `z-ai/glm-5.3-flash`: 1,368 calls averaging 203k prompt tokens, $6.16 |
| largest contexts | NeuralWatt `glm-5.3`: 47 calls averaging 673k, max 682,679, the only model whose window still fit |
| 120k+ prompt calls | 921 of them, $6.57 of the total |
| routing | every decision `classification_source = classifier`, default profile; nothing pinned |
**What the agents did.** An Atlas run (`.omo/plans/cockpit-quick-wins.md`)
delegated each item to a sub-agent worker:
| worker session | tool calls | exact-duplicate calls | worst repeat | landed |
|---|---|---|---|---|
| item 2, 2nd attempt | 423 | 30% | `read navbar.js` x61 | nothing |
| item 3, 1st attempt | 132 | 23% | `read test_admin_js_units.py` x16 | nothing |
| item 4 | 134 | 19% | `read` of the plan x8 | code only |
| item 2, 1st attempt | 307 | 17% | `read navbar.js` x21 | nothing |
| healthy workers | 30 to 96 | 0 to 6% | 1 to 3 | yes |
- Item 4's commit message claimed four test files it never wrote, and it
committed with `git add -A`.
- Atlas later said its own turns were "just re-reading the plan and state over
and over instead of doing work".
- A planning session described its reads as "heavily elided". The stored tool
outputs in opencode were intact.
**Pinch was rewriting agent context.**
- `pinch.budget_tokens: 50000` pruned 1,915 of 1,992 decisions over nine hours,
keeping 48% of the prompt on average.
- The planning session's 527k-token prompts went out at about 125k.
- `src/context_prune.py` replaces older tool results with
`[read: result omitted]`, or keeps a 1,500-character head and tail around a
`[N chars trimmed...]` marker. Results inside the protected window get the
same cut above `protected_max_chars` (20,000).
- Pinch's config had not changed since 2026-09-06. It had pruned 90 to 100% of
requests daily for two weeks without incident. The input changed:
| | 09-10 to 09-23 | 09-25 | 09-26 |
|---|---|---|---|
| avg prompt into pinch | 100k to 150k | 236k | 545k |
| p90 prompt | 170k to 260k | 528k | 1.31M |
- `opencode.json` advertised `auto` with `limit.context: 782324`, so opencode did
not compact until near that size.
**Routing drifted with size.**
- `deepseek/deepseek-v4-flash` has 384k effective context;
`z-ai/glm-5.3-flash` has 655k.
- Their proficiency was tied (`diff_checking` 0.933 vs 0.935) and unchanged
since 09-10.
- Deepseek went from 290 picks on 09-23 to 0 on 09-25, while appearing 851
times as runner-up.
**A clean reproduction, 2026-09-26 04:04 to 04:14.** A fresh planning
session, on the Lift A brief (188 lines):
- **Conditions:** context about 88k tokens, pinch off, 0 of 29 tool parts
marked `compacted` by opencode, every stored read intact.
- **What it did:** it still reported its reads "elided with `[...]`", wrote
fake "Earlier tool responses received ... (counts, not proof of task
completion)" blocks into its own replies (assistant text parts, not tool
output), invented session IDs, and called the brief "393 lines".
- **Where the phrase is not:** in opencode, oh-my-openagent (JS bundle and
native binary), or this repo.
- **Model:** `z-ai/glm-5.3-flash` served 27 of 34 chat decisions in the 15
minutes before it was stopped.
**Found along the way.**
- `deploy/opencode-plugin/router-outcome.js` has never read a real exit code.
On opencode 1.18, bash's exit is `output.metadata.exit`; the plugin reads
`output.exitCode`.
- PR #99 (conversation identity) was merged to `origin/main` but not running.
Production ran from a local checkout 29 commits behind.
- **Plugin loader rejected `router-link.js`.** The plugin exported `parentCache`
as a module-level const. Opencode 1.18's plugin loader discards any plugin
with a non-function export — the entire file is silently dropped. The fix was
to make `parentCache` a property on the factory function instead (L224,
`router-link.js` in PR #102, a24de1b). Confirmed by grepping
`~/.local/share/opencode/log/opencode.log` for `'failed to load plugin'`.
- How much of the $7.81 was waste cannot be computed. The run did land six
reviewed commits, and the deployed router cannot tell conversations apart.
### Theory
These are inferences, not measurements. Each lists what would confirm or
refute it.
1. **Pinch cutting oversized sessions drove, or worsened, the re-read loops.**
Weakened by the clean reproduction, which looped with pinch off. As
sessions grew past about 250k, the fixed 50k budget removed most of each
agent's working memory, recent reads included. The agent saw its reads cut,
re-read them, grew the context, and got cut harder.
- Supported by: the agents' own descriptions ("elided", "drowning in
tool-output truncation"), intact stored outputs, the prune rates, and the
timing of the size jump.
- Not tested: no run has been observed with pinch off.
- Confirm: with pinch off and `auto` at 200k, the duplicate-read signature
should not recur on comparable runs. The interim watcher
(`plans/no-progress-detection-prototype.py`) and Phase 0 data can show it.
2. **Sessions got that large because long unattended orchestration ran without
compaction.** Atlas made 668 tool calls, one worker 423, under a 782k limit,
and the re-read loop in theory 1 fed the growth. The limit is a fact; that it
explains the growth is theory.
3. **Leading theory: `z-ai/glm-5.3-flash`, or its OpenRouter path,
confabulates truncation and then loops.** The clean reproduction above
removed pinch, context size and input corruption, and the behaviour stayed.
Theories 1 and 2 remain plausible aggravators, not the cause.
- Confirm: with glm excluded, comparable agent runs show no elision claims
and no duplicate-read signature. A direct A/B on one prompt would settle
model versus provider path.
### Why nothing caught it
Each existing guard answers a different question:
- `circuit_breaker.py` trips on provider 5xx. There were none.
- `runway_low_warning` fires when balance / burn drops under 6 hours. Burn is
averaged over a 24 h balance window: OpenRouter read $24.73 at $0.25/h, which
is 97 hours of runway. It asks "will I run out", not "is this being wasted".
- Spend rate was not abnormal. The worst hour ($2.42) and worst 3-hour window
($4.13) sit inside the prior 30 days' range (hourly p90 $1.46, max $4.83;
worst 3 h $9.65).
- The router cannot see progress, and the deployed build could not tell
conversations apart.
### Recurrences
- **02:21 to about 02:56: Atlas looped after its plan was complete.** Every
checkbox was ticked, but `.omo/boulder.json` still carried `status: completed`
and `pr_url: .../pulls/77` from the previous plan (evidence: the file). The
continuation hook injected "continue" turns, 3 of 3 user turns in the window
(evidence: the session). Atlas re-derived its state each time: 54 turns,
about 19M cached-read tokens, `boulder.json` read 28 times, checkboxes grepped
16 times. Cause: stale state, not pruning. The branch was never pushed, so PR
77 was not touched.
- **About 03:05 to 03:17: a read-only `explore` subagent** re-read
`router-outcome.js` 20 times (19% duplicate calls across 322). The production
sync had just deleted that file. Read-only agents never land changes, so the
detector needs a separate signal for them.
### Human response
Times are local, 2026-09-26.
- **About 02:10.** Noticed the spend and asked for a mechanism in 6krrt to catch
it. Rejected a spend-rate alarm, since steady agent usage is legitimate, and
framed the target as "are concrete changes landing".
- **About 02:20 to 02:40.** Decided the design for `plans/no-progress-detection.md`:
- response modes (`warn`, `auto_compact`, `auto_limit_context`, a two-stage
`auto_recover`), with every mode warning
- `warn` as the default
- at most 2 recoveries per session tree
- `notify-send` now, with pluggable SMS, RingCentral and PagerDuty channels
later
- a local-model watchdog timer, "so I don't have to rely on Claude"
- `auto`'s context limit lowered to 200,000 in the repo and global
`opencode.json`
- **About 02:50.** Caught the post-completion Atlas loop by watching. At the
operator's direction, Claude aborted the session at 02:56 and cleared the
stale `pr_url` in `.omo/boulder.json`.
- **About 03:10.** Merged PR #101. A first `git pull` aborted on local changes,
and the restart on that line reloaded the old code. Then ran the prepared sync
(back up, drop the changes already upstream, `reset --keep origin/main`,
re-apply the rest) and restarted onto `827c408` at 03:15.
- **About 03:30.** Asked Claude to watch opencode for loops until the detector
ships (the interim watcher).
- **About 03:45.** Reported the planning session's "truncation hell". At the
operator's direction ("hot fix that"), Claude set `pinch.enabled` false at
runtime and persisted it to `config/config.local.yaml` (backup
`config.local.yaml.bak-20260926-pinch`). The operator restarted the service
at 03:52 and chose not to open a PR for the switch flip.
- **About 04:02.** Installed `router-link.js` in place of `router-outcome.js`
and restarted opencode (200k `auto` limit now active). Kicked off a fresh,
small Lift A planning session (`plans/no-progress-lift-a.md`).
- **About 04:14.** After the fresh session reproduced the confabulation,
excluded `z-ai/glm-5.3-flash` from routing. It had no picks after 04:13:57,
and traffic moved to deepseek and mimo. At the operator's direction, Claude
aborted the confabulating session.
- **~15:18.** Watchdog timer enabled (`systemctl --user enable --now
llm-router-watchdog.timer`).
- **15:19.** Plugin-loader root cause found in `opencode.log`; plugin fix
installed and opencode restarted.
- **15:19:57.** Router restarted.
- **After 15:19.** Conversation identity confirmed arriving: 17 of 17 requests
after the restart carry `c:` keys and an agent.
### Status
**Partly mitigated; detection not built; leading theory (3) unconfirmed.**
**Lift A shipped.** PR #102 (a24de1b) landed the surfacing half of North Star #4
(conversation identity and watchdog), live ~15:18 EDT. **Surfacing half of North
Star #4 is only partly shipped** — the identity half is confirmed (17 of 17
requests after the restart carry `c:` keys), but the detection half still waits
on Lift B (`plans/no-progress-detection.md`).
**Confounded by design, and say so.** Three mitigations landed within about
25 minutes:
- pinch off at 03:50
- `auto` capped at 200k at 04:02
- `z-ai/glm-5.3-flash` excluded at about 04:14
So "sessions look healthier since" cannot say which one mattered, or whether
it is several interacting: model degradation near a full window, oversized
sessions, pruning, and stale orchestration state. The one clean data point is
the 04:04 reproduction (glm, pinch off, small context, still confabulated),
which is why theory 3 leads.
To attribute properly, change one variable at a time and watch the watchdog's
stall rate. For example, re-admit glm with everything else held, or re-enable
pinch with a budget above 200k. Until then, treat every theory here as
unproven.
Also missing: a **behavior breaker**. The availability breaker trips on 5xx and
PR #76 trips on malformed output; nothing trips on well-formed output that
confabulates or loops. Planned for Lift B (`plans/no-progress-detection.md`,
section 8).
- In place:
- pinch off, runtime and persisted
- `auto` limit 200,000, active for opencode sessions started after a restart
- #99 and #100 deployed at 03:15
- Pending, operator:
- install `router-link.js` in place of `router-outcome.js`
- restart opencode
- Pending, lift: `plans/no-progress-detection.md`, the watchdog first.
- If pinch is re-enabled, its budget must sit above the 200k cap, and it must
never cut recent results.
**Rule:** steady spend is not the failure; **spend with no concrete change
landing is.** Until the detector ships:
- Check an unattended agent run by whether commits are landing, not by the
quota page.
- Verify each worker commit's `git show --stat` against its message; a worker's
"done" is a claim.
---
## The pattern ## The pattern
The first five are defensible local decisions — free a port, stop a process, The first five are defensible local decisions — free a port, stop a process,
@@ -373,6 +635,14 @@ file by mistake; #7 was a *correct*, deliberate, reviewed code change that
was still incomplete, because "every consumer" silently excluded a consumer was still incomplete, because "every consumer" silently excluded a consumer
that git cannot see. that git cannot see.
#8 breaks the pattern from the other side: the router did what it was
configured to do, routing each request to a model that fit. It was the one
component that saw all the spend while being blind to whether any of it produced
anything, so it measured cost when the thing that mattered was progress. Whether
its own `pinch` pruning drove the loops is still a theory (see #8, Theory 1). If
confirmed, #8 joins the family of defensible settings, a 50k budget chosen for
~130k sessions, that turned harmful when the input changed shape.
The generalisable fix is the same each time: **compute the thing that is The generalisable fix is the same each time: **compute the thing that is
actually true, and surface it where the operator is already looking.** That is actually true, and surface it where the operator is already looking.** That is
what the `/metrics` warnings and the inline admin warning are for, and it is the what the `/metrics` warnings and the inline admin warning are for, and it is the

408
docs/watchdog.md Normal file
View File

@@ -0,0 +1,408 @@
> Deep dive into the watchdog subsystem — opencode session loop detection. Back to
> [README](../README.md).
Watchdog is a periodic scanner that probes running opencode sessions for looping
patterns — repeated tool calls without progress. It runs as a systemd oneshot
every 5 minutes (`deploy/llm-router-watchdog.timer`) and, when a session flags,
writes into a local SQLite table and fires desktop notifications through the
configured channel pipeline.
The detector itself is a pure module (`src/progress_detect.py`) that returns a
boolean verdict plus a reason dict. The orchestrator (`src/watchdog.py`) reads
opencode session data over HTTP, runs the detector, and manages the alert
state machine. The notifier (`src/notifier.py`) routes events through
configured channels to desktop alerts (`notify-send`).
## What it catches and doesn't
Watchdog looks for **looping** — sessions that make repeated tool calls against
the same targets without landing meaningful changes. It detects four signal
types:
- **Dup signal**: a call appears more than 25% of the time in the sliding
window. The call matcher first normalises `bash` commands (strips env var
prefixes and `cd` prefixes, joins lines, collapses comments), then checks
tool name + canonicalised JSON args (sorted keys) for equality.
- **Top signal**: a single target is called 12+ times in the window (or 8+
for read-only agents like explore/librarian/oracle). Targets merge calls on
the same `(tool, basename)` for file reads, `(bash, first-two-words)` for
commands, and `(tool, pattern[:60])` for grep/glob.
- **Slow signal**: a single tool call accounts for 15+ of the session's total
calls, AND no call in the entire session history has landed (tree write or
git commit). This catches sessions stuck on a single read/analysis.
- **Coverage signal**: the session rereads one file at 4x its line count
within the window. The coverage score divides total bytes read by file length,
capped at the file's actual lines. A file read 4x its length means the model
is rereading without making progress.
**NOT steady spend.** A session that steadily calls different files on a
difficult refactor is not flagged — each call lands a unique `(tool, args)`
and the top target never crosses the threshold. The detector only fires when
repetition outpaces the call budget, not when a session is quietly burning
tokens across many distinct tools.
### Landed calls
A call counts as **landed** when it:
- is an `edit`, `write`, or `patch` with a non-empty `metadata.diff` field, or
- is a `bash` command containing `git commit` with `metadata.exit == 0`
Landed calls break both the slow and coverage signals. A session that reads a
file 4x, then edits it once, clears all signals immediately — the landed check
exits early and the verdict returns `False`.
### Ancestral trees
Sessions can be parented: the opencode API returns `parentID` on child sessions.
Watchdog resolves root sessions and evaluates EACH session in the tree on its
own calls, NOT merged into the root. A child's flagged verdict bubbles up to the
root level, which is what the alert dedup key uses (`opencode-loop:{root_session_id}`).
For the landed-time check, a session sees NOT only its own landed calls but also
its transitive descendants' landed calls. This prevents a parent session from
being flagged just because its child landed — even though the parent may still
be looping independently.
## Signals and thresholds
Thresholds live in `DetectConfig` (`src/progress_detect.py:38-53`) and are
configurable under `watchdog.detector.*` in `config/config.yaml`.
| Key | Default | Meaning |
|---|---|---|
| `watchdog.detector.dup_min` | 0.25 | Fraction: calls must repeat at this rate to flag |
| `watchdog.detector.top_min` | 12 | Minimum calls to a single target (normal agents) |
| `watchdog.detector.top_min_ro` | 8 | Minimum calls to a single target (read-only agents) |
| `watchdog.detector.cum_min` | 15 | Same call must account for this many total calls |
| `watchdog.detector.cover_min` | 4.0 | File must be reread this many times its length |
| `watchdog.detector.window` | 60 | Sliding window in number of tool calls (not seconds) |
| `watchdog.detector.min_calls` | 40 | Minimum total calls before evaluation triggers |
| `watchdog.detector.read_only_agent_keywords` | `["explore", "librarian", "oracle"]` | Title keywords that make an agent read-only (lowered `top_min` from 12 to 8) |
Read-only agents get a lowered top threshold because agents whose job is
exploration or library work are expected to call the same targets repeatedly —
the floor is 8 instead of 12.
## Alert lifecycle
The alert state machine is per-root-session and tracked in the
`watchdog_alerts` table.
### Trigger (initial detection)
When the detector flags a root session's tree for the first time:
1. The watchdog writes a `watchdog_alerts` row with `state='open'`.
2. The `local_llm_enabled` path asks a local Ollama model (defaulting to
`verification.model` or `qwen2.5-coder-router:14b`) for a second opinion.
It receives the agent name and the reason dict as context, and the model
must respond with only "yes" or "no" — no explanation.
3. If the local LLM answers "yes", severity is set to `critical`. If it
answers "no" or times out (returns `None`), severity is set to `warning`.
4. The notifier dispatches an `AlertEvent` through every matching channel.
5. The `flagged_ticks` counter starts at 1.
A cap of 3 simultaneous LLM second-opinion calls (`_MAX_LLM = 3`) prevents a
burst of flagged sessions from flooding the local model.
### Escalate (persistent looping)
On every subsequent tick where the session is still flagged:
1. `flagged_ticks` is incremented.
2. If `flagged_ticks >= 3` (15 minutes at 5-minute ticks) OR the LLM answers
"yes" again, the alert escalates to `severity='critical'` and the escalation
event is dispatched.
3. If the alert is already critical, ticking continues but no duplicate
escalation fires.
### Resolve (session is gone or fixed)
A root session alert resolves in two conditions:
- **Session disappeared**: the root session ID is no longer returned by the
opencode `/session` API. The watchdog writes `state='resolved'`,
`resolved_at=<now>`, and `severity='info'`.
- **Session fixed**: the root session's tree no longer has any flagged sessions
in the detector's verdict. The alert is resolved the same way and a resolve
event is dispatched.
Resolve events always bypass the rate limit — they must clear the previous
notification so the operator knows the alert is over.
### No-opencode
When no `rc-servers.json` or no opencode server answers, the tick writes an
`outcome='no_opencode'` record and returns immediately with zero alerts. This
uses the filesystem lock (`router.db/.watchdog.lock`) so two watchdog instances
cannot run simultaneously.
## Admin surfaces
Watchdog data surfaces through three admin pages, each serving a different
operational question.
### Loops panel — index.html#loops
On the main dashboard (`admin/frontend/index.html`), the **Loops** card
(id="loops") shows all currently open alerts:
- **Severity badge**: the alert's current severity (`critical` in red,
`warning` in orange-yellow, `info` in grey)
- **Dedup key**: the raw `opencode-loop:{root_session_id}` string, so the
operator can cross-reference with `journalctl` or `router.db`
- **Opened at**: `opened_at` ISO timestamp from the first trigger
- State: `open` or `resolved`; the panel filters to `resolved_at IS NULL`
The panel is client-side only — it polls `/api/watchdog/loops` which queries
`watchdog_alerts` for all unresolved rows.
### Watchdog card — controls.html
The Controls page (`admin/frontend/controls.html`) includes a dedicated
**Watchdog** card with operational status:
- **Last tick**: timestamp of the most recent `watchdog_ticks` row
- **Sessions seen**: how many opencode sessions were enumerated on that tick
- **Flagged verdicts**: count of flagged + total from the latest tick's
`watchdog_verdicts` rows
- **Open alerts**: count from `watchdog_alerts WHERE resolved_at IS NULL`
- **Notification channels**: list of configured channels with toggle controls
(enabled/min_severity), editable via `POST /api/watchdog/channels`
- **Refresh** button: reloads the watchdog status from the API
- **Send test alert** button: fires a test `AlertEvent` through enabled channels
The status endpoints are:
- `GET /api/watchdog/status` — last tick, verdict counts, open alert count
- `GET /api/watchdog/channels` — per-channel settings from `watchdog_channel_settings`
- `POST /api/watchdog/channels` — upsert channel settings (enabled, min_severity)
- `POST /api/watchdog/test-alert` — deliver a test event (severity + optional
channel_name filter) to a live Notifier instance
### Per-model stall rollup — models.html
The `watchdog_verdicts` table has `model_id`/`provider` columns and indexes on
them, but no per-model stall rollup, no Block button beside evidence, and no
Blocked list are built yet. `blocked` is only a value in the Models override
dropdown (see docs/admin-portal.md).
## Database schema
Four tables live in `router.db`, created idempotently by both
`config/schema.sql` and `watchdog_store.py`:
```sql
watchdog_ticks -- one row per watchdog tick
watchdog_verdicts -- per-session verdict at each tick (joined to ticks)
watchdog_alerts -- alert lifecycle state machine (dedup_key PK)
watchdog_channel_settings -- per-channel toggle + severity gate
```
Indexes on `watchdog_verdicts` cover `model_id`, `created_at`, `flagged`, and
`session_root` — the columns the admin endpoint queries most frequently.
The ticks table carries `ticked_at`, `sessions_seen`, and `outcome`
(`'ok'`, `'flagged'`, or `'no_opencode'`). The verdicts table joins to ticks
via `tick_id` and carries the full reason dict fields: `dup`, `top`,
`top_what`, `landed`, `slow`, `coverage`, plus `calls_since_landed`,
`cost_since_landed_usd`, and the LLM second-opinion answer.
Alerts are keyed by `dedup_key` (deduplicated per root session) and carry
`opened_at`, `last_fired_at`, and `resolved_at`. The transition logic lives in
the `_fire_alert()` helper: trigger inserts only if no open row exists,
escalate increments `flagged_ticks` and bumps severity, and resolve writes
`resolved_at` and downgrades to `severity='info'`.
## Knobs
All watchdog configuration lives under `watchdog:` in `config/config.yaml`:
| Key | Default | Required |
|---|---|---|
| `watchdog.enabled` | `true` | Gates the entire subsystem |
| `watchdog.local_llm_enabled` | `true` | Whether to call Ollama for second opinions |
| `watchdog.model` | `null` (uses `verification.model`) | Local model name |
| `watchdog.detector.window` | 60 | Sliding window size in calls |
| `watchdog.detector.dup_min` | 0.25 | Dup threshold (fraction) |
| `watchdog.detector.top_min` | 12 | Top target threshold (normal agents) |
| `watchdog.detector.top_min_ro` | 8 | Top target threshold (read-only agents) |
| `watchdog.detector.cum_min` | 15 | Slow-signal threshold |
| `watchdog.detector.cover_min` | 4.0 | Coverage-signal multiplier |
| `watchdog.detector.min_calls` | 40 | Minimum eval calls |
| `watchdog.read_only_agents` | `["explore", "librarian", "oracle"]` | Agent title keywords for reduced threshold |
| `watchdog.dashboard_base_url` | `"http://127.0.0.1:8080/admin"` | Base URL for alert links |
All seven `watchdog.detector.*` keys plus `watchdog.enabled` and
`watchdog.local_llm_enabled` (nine allowlisted keys total) are allowlisted for
live editing through the
admin portal (`POST /admin/api/config/{key}`), which writes to
`config.local.yaml` (the gitignored machine-local overlay). The changes are
validated by `RouterConfig` before reaching disk, and take effect at the next
tick — no service restart required since `detect_config_from_pydantic()` reads
cfg fresh each invocation.
## Install
The watchdog ships as two systemd user units:
- `llm-router-watchdog.service` — a `Type=oneshot` that runs
`PYTHONPATH=%h/llm-router/src .venv/bin/python -m watchdog --once`
- `llm-router-watchdog.timer` — fires 5 min after boot, then every 5 min
Installation follows the same pattern as the other deploy units. The shipped
unit files use `%h/llm-router` placeholders that need rewriting to your actual
repo path:
```bash
REPO=$(pwd)
for u in deploy/llm-router-watchdog.{service,timer}; do
sed "s|%h/llm-router|${REPO}|g" "$u" \
> ~/.config/systemd/user/"$(basename "$u")"
done
systemctl --user daemon-reload
systemctl --user enable --now llm-router-watchdog.timer
```
The service unit has **no `EnvironmentFile`** — watchdog needs no provider API
keys because it only queries the local opencode session APIs and the local
Ollama instance.
Check installation:
```bash
systemctl --user status llm-router-watchdog.timer
journalctl --user -u llm-router-watchdog.service --no-pager
```
## Run --once
For ad-hoc debugging or pre-deploy smoke test:
```bash
PYTHONPATH=src .venv/bin/python -m watchdog --once
```
The `--once` flag runs a single tick and exits 0 once the tick has run,
however many alerts it fired (the count is logged as `watchdog tick done: N
alert(s) fired`); systemd would otherwise mark the oneshot unit failed on every
real alert. It connects to the same `router.db` that the timed
service uses, acquires the filesystem lock, reads `rc-servers.json`, probes
opencode servers, and follows the full evaluate-alert-resolve pipeline.
Add `--config /path/to/config.yaml` to point at a non-default config file.
### Debug output
The `--once` run emits structured `watchdog=` log lines for every state
transition:
```
2026-01-15 14:30:01 INFO watchdog: watchdog=2026-01-15T14:30:01+00:00 dedup='opencode-loop:ses_xxxxx' state=trigger severity=warning title='opencode loop: explore'
```
The `title` field carries the agent name derived from the session title. The
`dedup` field is the root session's dedup key, matching `journalctl` traces
and admin panel rows.
## Backtest
`scripts/progress_backtest.py` replays labelled sessions through the detector
to validate threshold calibration. It reads from a fixture JSON file and runs
a sliding-window evaluation with step 5.
```bash
PYTHONPATH=src .venv/bin/python -m scripts.progress_backtest --fixture \
tests/fixtures/progress/fixture.json
```
The fixture ships in `tests/fixtures/progress/fixture.json` and contains
15 labelled sessions: 8 marked `must_flag` and 7 marked `must_not_flag`.
Expected output after a successful run:
```
# backtest complete: 8/15 flagged
```
At least 8 sessions should trigger a flag. The fixture is generated from live
opencode sessions (see [export fixture](#re-export-fixture) below) and scrubbed
to protect file contents — content is SHA-1 hashed except for the few keys the
detector actually inspects (`filePath`, `command`, `pattern`, etc.).
## Re-export fixture
The fixture generator pulls sessions from the opencode SQLite database
(`~/.local/share/opencode/opencode.db`) and writes the scrubbed test fixture:
```bash
python scripts/export_progress_fixture.py
```
Output goes to `tests/fixtures/progress/fixture.json`. The script takes a
session ID to label map (`LABELS` dict), extracts call history, resolves file
line counts (via `git show` for deleted worktree paths), and scrubs all strings
to SHA-1 prefixes, preserving only the detector's target keys.
The fixture generation script hardcodes paths to the operator's home directory
and a specific worktree commit (`WORKTREE_COMMIT = "827c408"`). When
regenerating, update those paths to match the current environment.
## Notifier pipeline
Alert events flow through `src/notifier.py` before reaching the operator. Each
alert is an `AlertEvent` dataclass mirroring PagerDuty Events v2 shape:
`dedup_key`, `severity` (info/warning/critical), `state` (trigger/escalate/
resolve), `title`, `summary`, `details`, and `source` (hardcoded to
`"6krrt-watchdog"`).
Each configured channel has its own `min_severity` gate — a critical event
passes through a channel with `min_severity=warning`, but a warning event does
not pass a channel with `min_severity=critical`. Each channel also has a rate
limit window of 300 seconds (5 minutes). Resolve events always bypass the rate
limit to ensure clear-the-alert notifications are delivered.
Currently only `type=desktop` channels are implemented; they invoke
`notify-send -u {urgency} {title} {summary}`. Unknown channel types are
logged as warnings and skipped. Missing `notify-send` (no `DISPLAY` or
`libnotify` installed) is logged at warning level — the notifier never raises.
Channel settings (`watchdog_channel_settings`) are editable at runtime through
the admin portal and stored per-channel in the database, separate from the
base config.
## Known limits
- **Under 40 calls: invisible.** `min_calls=40` means the detector does not
evaluate sessions with fewer tool calls. Short troubleshooting sessions that
loop on 10 calls pass through undetected.
- **Atlas at cover_min 4.0 exactly.** A session whose last 60 calls read a
single file exactly 4x its line count hits the coverage threshold at the
boundary. There is no margin: `>=` comparison means exactly 4.0x flags.
This matters for sessions that genuinely need to re-read a reference file
repeatedly — the coverage signal cannot be tuned per-file.
- **Stale localhost:4096 probes.** Watchdog reads `rc-servers.json` from
`~/.local/share/opencode/rc-servers.json`, which may contain stale entries
for opencode sessions that have since ended. These produce empty call lists
and are silently skipped — but they add network round-trip latency to each
tick. Only the first answering server's sessions are evaluated (the watchdog
picks `answering[0]` from the list of servers that successfully return a
session list).
- **No streaming path.** `POST /outcome` is the only ground truth signal for
streaming traffic; the watchdog has no equivalent endpoint to report streaming
session outcomes. It can only see tool-call repetition, not whether the answer
was useful.
- **Local LLM gate.** Second-opinion calls are gated by both
`local_compute.enabled` (global) and `watchdog.local_llm_enabled`
(per-subsystem). If either is false, no LLM calls are made and severity stays
as `warning` on initial trigger regardless of model quality. The cap of 3
simultaneous LLM calls means that if 5 sessions flag, 2 wait without opinion.
- **No provider calls.** Watchdog does not touch the dispatch path. It only
reads opencode session data and runs local heuristics. No provider calls, no
quota consume, no cost incurred.

View File

@@ -13,7 +13,7 @@
"auto": { "auto": {
"name": "auto (router picks, interactive)", "name": "auto (router picks, interactive)",
"limit": { "limit": {
"context": 782324, "context": 200000,
"output": 16384 "output": 16384
}, },
"modalities": { "modalities": {

397
plans/cockpit-brainstorm.md Normal file
View File

@@ -0,0 +1,397 @@
# Cockpit brainstorm: the router as an operator's instrument panel
Status: reference -- brainstorm, not a queue item; its six quick wins shipped in PR #101
**Date:** 2026-09-23
**Read against:** the repo tarball as of this date (code, config, plans,
screenshots). **Not** read against the live `router.db`, so every "gap" below
is a gap in what the code surfaces, not a claim about what the data shows.
Section 5 is the exception: it was read against the live portal.
## The framing
6krrt as a local LLM router with a cockpit. The operator:
1. adjusts dials and knobs in flight,
2. sees where tokens are wasted,
3. watches proficiency, and
4. sees poor results from a carrier or model surface, so they can decide
whether to preclude it (with the circuit breaker as the automatic version
of that decision for outages).
The router makes the per-request decision. The cockpit is where the operator
makes the slower decisions: which models and carriers are trusted, for what,
and at what settings. Under this framing, the question for every feature is
whether it helps the operator notice something and act on it.
---
## What the cockpit already has
| job | what exists | where |
|---|---|---|
| in-flight dials | ~11 runtime knobs (bool/float/int tables in `admin.py`), persisted config writes via `/api/config/{key}`, classifier card, profiles CRUD, gaming mode | Controls, Profiles |
| knob coverage | North Star rule #1 + `test_admin_knob_coverage.py` | tests |
| token waste | pinch savings card; `cache_rate_series` + warnings; incumbent cache pricing dial; `cost_estimate_calibration` in `/metrics` | Dashboard, Controls, `/metrics` |
| proficiency | model × category matrix with evidence status (measured / inherited / thin, n per cell); client outcome log with attributable + applied flags | Proficiency |
| poor results | `content_fault_warnings` (malformed output rate); `rejection_warnings`; availability override (active / deprecated / stale); allowlist editor | warnings bell, Models, Providers |
| breaker | passive per-(model, provider) breaker on 5xx; on/off knob | Controls (toggle only) |
The foundation is good. What's mostly missing is the link between a signal
and the action it should prompt, plus any record of which actions were taken.
---
## Cross-cutting ideas (these help all four jobs)
### X1. Flight recorder: a change log
**Gap.** `admin_schema.sql` has only `admin_model_overrides`. Nothing records
when a knob, override, allowlist entry, profile or feedback fold changed, what
it changed from, or why. Runtime knobs are "NEVER persisted", so a restart
silently reverts them, and nothing notes that the revert happened.
**Idea.**
- An `admin_changes` table: `(at, surface, key, old, new, reason, source)`,
where `source` is `runtime | persisted | override | allowlist | profile |
feedback_fold | restart`. Every admin write path appends a row.
- An `operator_epoch` id stamped on each `route_decisions` row, incremented
on each **operator action** (a row in `admin_changes`). Every metric can then
be split before and after a change without joining on timestamps. Scope it
to operator actions on purpose: "any change that affects routing" would also
cover catalog price polls and every feedback fold, so the epoch would tick
constantly and split nothing.
- Change markers drawn on the dashboard history charts.
- A drift badge when a runtime value differs from its persisted twin
("this will revert on restart").
- Restart reverts are loggable without persisting runtime knobs: the change
log already holds each knob's last runtime value, so at startup, compare it
with the loaded config and write one `restart` row per value discarded.
**Why first.** Adjusting knobs in flight produces no learning if you can't
later tell what a change did. The other ideas here lean on this one.
### X2. Structured warnings with actions attached
**Gap.** `/metrics` warnings are plain strings. The dashboard decides which
page fixes a warning, and how severe it is, by regex on the text
(`warningTarget` and `warningSeverity` in `index.html`). Rewording a warning
silently breaks its link and its severity.
**Idea.** Emit `{class, severity, subject: {model, provider, category?},
evidence_url, actions: [...]}`. The warning registry in
`test_tui_warnings.py` already lists every class, so it can serve as the
schema. The bell can then offer an action in place: "exclude from
`tool_use_agentic`", "quarantine 24h", "open decisions filtered to this
model".
### X3. Preview before apply (replay)
**Gap.** The reviews replayed thousands of real decisions to test a setting,
but only as one-off scripts pasted into plan documents.
**Idea.** For any change that affects ranking (`quality_tolerance`, the
incumbent dial, profile edits, availability overrides, graded exclusions
from P2 below), the portal replays the last N decisions through
`select_candidates` and `rank_candidates` with the proposed value, then shows:
- a winner-shift matrix: old winner → new winner, with counts,
- estimated cost delta,
- decisions that would become unroutable (the 2026-09-01 and 2026-09-04
incidents were both deprecations that emptied a candidate set).
This is cheap because the ranking modules are pure. Its output is an estimate
of routing change, not quality change, and should be labeled that way.
---
## 1. Dials in flight
- **D1. Timed changes.** "Apply for 2h, then revert." Turns a knob change into
a bounded experiment and lowers the risk of forgetting a test setting.
Needs X1 to record both the apply and the revert.
- **D2. Blast radius on hover.** How many of the last 24h of decisions this
knob would have touched. It's the X3 replay reduced to one number.
- **D3. Session-scoped A/B.** Assign new sessions to setting A or B and
compare cost and outcome rate per arm. It has to be per session, not per
turn, because a mid-session switch costs cache (token-waste Wave 2
measured 0.919 → 0.348). Larger build; only worth it once X1 exists and
outcome volume can support a comparison.
- **D4. Grouping by effect.** Group controls by what they move (cost,
quality, latency, safety, measurement-only) rather than by config section,
and mark the few that actually change routing. Knob coverage guarantees
every control exists; this makes them usable.
## 2. Token waste
**Gap.** Waste signals are spread across pinch, cache rate, calibration and
the verifications table, and none is in dollars in one place.
`cost_estimate_calibration` is in `/metrics` but not in the portal.
- **W1. Waste ledger.** One panel with dollars per waste class over a window,
each with a trend:
| class | how it's computed |
|---|---|
| switch cache loss | billed on switch turns − same prompt at the session's same-model cache rate |
| retries | re-billed prompt tokens from `iteration.py` attempts |
| paid for failure | billed cost of responses later reported `ok:false` via `/outcome` |
| empty-200 / malformed | billed cost of responses with a structural `malformed` verdict |
| truncation | responses stopped by `finish_reason: length` with no client cap |
| pinch (negative) | dollars saved, from `pinch_summary` |
"Paid for failure" counts only decisions that received an outcome, so it
shows n and outcome coverage (the share of billed decisions with a report)
next to the dollar figure.
- **W2. Why did it switch?** Record a switch reason on each decision: category
changed, breaker open, exploration, override removed the incumbent, profile
change. Then rank sessions by switch cost and drill into the turns. Without
a reason, a switch is only a cost; with one, it points to a knob.
- **W3. Calibration panel.** Surface `cost_estimate_calibration` per model,
headlining the **spread** (as its docstring argues), and flag pairs where
the estimator and the bill disagree on order. Where that happens, the cost
tiebreak picks the wrong model.
- **W4. Cost per successful outcome.** Per (model, provider, category):
billed dollars ÷ `ok:true` outcomes. A cheap model that fails often isn't
cheap. This is the one waste figure that includes quality.
Show n and outcome coverage next to it. Models that carry more traffic
collect more outcomes (the same exposure bias the feedback fold had to
correct), and if coverage stays hidden, a thinly reported cheap model will
look better than it is.
## 3. Proficiency monitoring
**Gap.** The matrix shows current state well. It doesn't show movement, and
it doesn't show which thin cells actually matter.
- **P1. Trend and drift per cell.** A sparkline of the outcome rate over time.
Compare a recent window against the long-run rate with an interval, and
flag drops that fall outside it. Carriers change serving setups (quant,
engine, hardware) without notice, and a drift flag is how that shows up.
- **P2. Evidence priority.** Rank cells by `traffic share × uncertainty`. A
thin cell that routing never consults doesn't matter; a thin cell deciding
30% of traffic does. The ranking tells you where to spend `eval_proficiency`
runs or exploration budget.
- **P3. The tolerance band per category.** Show which models sit within
`quality_tolerance` of the leader, so it's clear where cost is deciding and
where quality is.
- **P4. Label provenance per cell.** The share of each cell's outcomes whose
category came from a fresh classification versus a cached or borrowed one.
North Star rule #2 exists because borrowed labels trained the matrix;
this makes it visible cell by cell.
## 4. Poor results → operator preclusion (and the breaker)
**Gap.** The operator's only preclusion tools are binary: an availability
override (active / deprecated / stale) or removing a model from the
allowlist. The breaker is in-memory, trips only on 5xx, forgets everything
on restart, and shows nothing but an on/off toggle.
- **Q1. Carrier/model scorecard.** One sortable table per (model, provider)
over a window: outcome fail rate (with n), malformed / empty-200 rate,
5xx and timeout count, breaker trips, p95 latency, estimate-vs-bill error,
cache rate, cost per success (W4). This is the "who's misbehaving" view the
preclusion decision needs.
- **Q2. Same model, different carriers.** Group scorecard rows by base model,
e.g. `qwen3.6-35b` on NeuralWatt vs OpenRouter. This separates "the model is
bad at this" from "this carrier serves it badly", which calls for a
different fix: drop the carrier's row, or restrict the model.
- **Q3. Graded actions** in place of deprecate-or-not. Each action records a
reason and links its evidence in X1:
1. exclude from category X. This **cannot** reuse `eligible_categories`:
that field is restrict-only (`admin.py`, the probe docstring), and NULL
means "every category", so a deny needs its own column,
2. exclude when the request carries tools (the `deepseek-v4-flash` case).
This may be the way out of the frozen `tool_use_agentic` data:
`min_tool_proficiency` can only read scores that stopped moving on
2026-09-15, while a per-model manual rule lets the operator decide from
Q1 scorecard evidence instead,
3. restrict to batch / `-flex` only,
4. quarantine: exploration-only, so it keeps collecting evidence without
carrying traffic (see the open question below; deferred until Q1 shows
it is needed),
5. timed ban, e.g. 24h, auto-expiring and logged,
6. deprecate (what exists today).
Each action goes through at least the minimal X3 check (would this empty a
candidate set for any recent decision?) before committing. That check alone
would have caught both the 2026-09-01 and 2026-09-04 incidents; the full
winner-shift replay can come later.
- **Q4. Breaker visibility.** Show open circuits with `down_until`, current
cooldown and trip count; keep a trip history (persist it, since `_store`
is process-lifetime); add manual force-open (drain a model on purpose) and
force-close (reset after a known fix).
- **Q5. A quality breaker, separate from availability.** Trip on content
failures (empty-200, mangled output, a burst of `ok:false`) using the
novelty-or-rate rule already in `rejection_warnings`. This is roughly what
`plans/mangled-output-detection.md` specs (status: planned). Keep it a
separate breaker and panel so a quality trip doesn't read as an outage.
**One doc owns it:** fold Q5 into `mangled-output-detection.md` rather than
specifying it twice, and let this file point there.
- **Q6. Carrier-level breaker.** When several models on one provider trip
inside a short window, open the provider rather than walking its models one
cooldown at a time. Account exhaustion already has its own path; this is for
partial carrier outages.
---
## 5. Quality of life: ergonomics and readouts
Unlike the sections above, this one **was** read against the live portal
(8080, view-only, 2026-09-23), plus the frontend source. Each item names the
thing observed.
### Readouts that mislead today
- **R1. Home tiles use different windows.** Quota is "this period",
Decisions 7d, Models and Pinch 30d. Busiest models shows `kimi-k2.7-code`
at 10,028 while the Decisions tile says 2,887 for everything, and both are
correct. Label each tile's window on its face, or drive all of them from
the Activity card's 24h / 7d / 30d selector.
- **R2. The Decisions tile counts verifications, and mixes diagnostics with
ground truth.** It reads "2,887 verified in 7 days: 132 ok, 93 failed", but
`index.html` sums `verdict_mix`, which counts `verifications` rows, not
decisions. Checked read-only against the live DB on 2026-09-23:
| | tile | actually |
|---|---|---|
| headline | 2,887 | 2,678 route decisions; 2,887 is verification rows, 2,660 of them `unverifiable` |
| "ok" | 132 | 96 client `succeeded` + 29 `local_llm` ok + 7 structural ok |
| "failed" | 93 | 82 client `failed` + 11 `malformed` (local_llm + structural) |
The tile adds structural and `local_llm` verdicts, which CLAUDE.md calls
diagnostics only, to `/outcome` reports, the only ground truth. Show route
decisions as the headline, then client outcomes on their own: "178 client
reports (6.6%): 96 ok, 82 failed (46%)". Diagnostics, if shown at all, go on
a separate line.
- **R3. The Activity chart smooths across gaps.** The decisions series is a
spline through sparse points, so it draws continuous traffic across hours
that had none. Break the line at empty buckets or use bars. Separately,
billed requests exceed decisions at several peaks; the chart should say what
the excess is (cloud classifier calls? retries? unrouted pins?) rather than
leave two lines that disagree unexplained.
- **R4. Sparklines have no values.** The tile sparklines carry no axis and no
hover. Add a hover value, or min and max labels.
- **R5. The Classifier card flashes a wrong value.** On first paint the mode
select reads `local_llm` and only switches to `local_encoder` once config
loads. For a second, the card shows a value that isn't configured. Render
a loading state instead of the default.
- **R6. Config echo vs live state.** The Classifier card shows what the
overlay says (`local_encoder`, device `cpu`), not what the process actually
loaded. Show the resolved model, device and recent p50 latency from the
running classifier. (It also exposed doc drift: CLAUDE.md says the live
deployment runs `device: cuda`; the overlay says `cpu`.)
- **R7. Proficiency colour encodes evidence, not score.** Green / amber mean
measured / thin, so the best model in a column isn't visible without reading
every number. Mark each column's leader, outline the `quality_tolerance`
band (P3), make columns sortable, and add a "routable only" toggle so rows
that are deprecated or not allowlisted stop padding the matrix.
### Decisions page ergonomics
- **E1. Collapse runs.** The top 40 rows are one session: same category,
profile, tier, source and model, and only ctx and cost move. Fold
consecutive same-session, same-model rows into one expandable row: "38 turns,
ctx 75k to 98k, $0.27 total". Switches then stand out as row boundaries,
which is what W2 wants to surface.
- **E2. Session view.** A per-session rollup: turns, total cost, context
growth, switches, outcome count. Context climbing about 1k per turn is
visible in the raw table and invisible everywhere else.
- **E3. Filters in the URL.** `decisions.html` never reads `URLSearchParams`,
so no other page can link to "decisions for this model" or "this session".
X2 actions, the Q1 scorecard and the model modal all need that link.
- **E4. Row drill-down.** A row click does nothing. Open a panel with the
decision's rejected candidates, verification verdict, `/outcome` report,
and estimated vs billed cost from its energy observation.
- **E5. Server-side filtering.** The page loads 1,000 of 32,540 rows and
filters and searches only those, with a footnote saying so. Push filters to
the API so that "All" means all.
- **E6. Model, provider and source as filters.** Today they're reachable only
through free-text search.
### Controls page
- **C1. One knob table, not two.** Runtime Knobs and Persisted Config list
largely the same knobs in two columns, in different orders, with persisted
labels truncated (`objective.incumbent_cache_pri…`). Checking drift means
matching rows by eye. Use one table: knob, live value, persisted value,
layer (base / overlay), drift badge. That table is also X1's natural home.
- **C2. Separate Restart Service.** It sits in the same button row as Refresh
Catalog and Seed Energy. Move it apart, and have its confirmation list which
runtime values differ from persisted and will revert.
- **C3. Prose in cards.** Local Compute carries three paragraphs; the portal
style rule is no paragraphs. Keep one line plus a docs link or tooltip.
- **C4. Default and last change per knob.** Show the code default beside
each value, plus "changed 2h ago from 0.3" once X1 exists.
### Portal-wide
- **G1. Nav lives in eight files.** Each page hardcodes the same `<ul>` of
nav links. `navbar.js` already notes that adding Quota was an eight-file
edit; have it render the links too.
- **G2. Width.** Most pages sit in `container-xl` and use well under half of
a wide screen, while Decisions, the densest table, is the most cramped.
Proficiency already goes full-width. Go fluid on the data-heavy pages.
- **G3. Dismissals don't stick on rate warnings.** Dismissal is keyed on the
warning's exact text, on purpose (a changed condition resurfaces it). But
warnings that embed a live rate ("3.2x pace") change text every poll, so a
dismissal never holds. X2's `class` + `subject` is the right key; resurface
on a severity change, not a digit change.
- **G4. Keyboard.** `/` to focus search, `Esc` to close modals, and `j`/`k`
on the Decisions table. Cheap, and this is a page someone lives in.
---
## One thing the cockpit framing makes more urgent
The portal is loopback-only with no auth, and it already includes
`/api/restart-service`, `/api/apply-feedback` and provider deletes. The more
it becomes the place decisions are made, the stronger the pull to reach it
from other machines, and moving the router onto a Proxmox box on the LAN is
exactly that move. Add auth before binding anything but `127.0.0.1`. Same
warning CLAUDE.md already gives for the API, with more at stake.
---
## A possible order
| # | item | why here |
|---|---|---|
| 0 | portal auth | **only if** the router is leaving `127.0.0.1` (the Proxmox move); blocks that move, not this list |
| 1 | X1 change log + operator epoch | everything else reads it |
| 2 | Q1 scorecard + Q4 breaker visibility | the preclusion decision needs one view first |
| 3 | Q3 graded actions (via X1) + minimal X3 (empties-a-candidate-set check) | turns the scorecard into decisions without repeating 09-01 / 09-04 |
| 4 | X2 structured warnings | lets warnings trigger Q3 actions directly |
| 5 | W1 waste ledger + W2 switch reasons | token waste in dollars, with causes |
| 6 | X3 full replay preview (winner shift, cost delta) | makes knob changes safe to try |
| 7 | P1 drift, P2 evidence priority | proficiency monitoring over time |
| 8 | Q5 quality breaker (owned by `mangled-output-detection.md`), D1 timed changes | automate what the operator has been doing by hand |
| 9 | D3 session A/B, Q6 carrier breaker | only after the above prove out |
Section 5 sits outside this order: most of it is small and independent.
Quick wins that also unblock the numbered items: **E3** (URL filters, needed
by X2 and Q1), **R2** (outcome coverage, the same honesty W1 and W4 need),
**C1** (one knob table, X1's home), **G3** (fixed by X2's class keys). R5 and
G1 are fixes of a few lines each.
**North Star #1 applies to every item here.** Q5 thresholds, D1 durations,
timed-ban lengths and any quarantine settings are new config knobs, so each
ships with its admin control or `test_admin_knob_coverage.py` fails. Scope the
control into the item, not a follow-up.
### Open questions, with proposed answers
- **Is the change log append-only history, or also an undo stack?**
Proposed: append-only, plus a per-row "revert to old value" that writes one
key and appends its own row. That is not the shape CLAUDE.md rejected for
gaming mode: that concern was one flag writing five keys and drifting apart.
A one-key revert can't drift, and the history stays intact.
- **Should graded exclusions live in the DB or in `config.local.yaml`?**
Proposed: the DB, next to `admin_model_overrides` (extend it or add a
sibling table). Timed bans need an expiry, and an expiry belongs in a DB
column. The overlay is per-machine and not in git, so it's the wrong home
for judgments about a carrier. And availability overrides already live in
the DB. The category exclusion still needs a new deny column (see Q3.1).
- **Does quarantine conflict with session-scoped exploration?** Yes, and the
weight goes up: a quarantined model would carry whole sessions, not single
turns. There's a second problem too. `exploration.py` only chooses among
models that already passed the hard filters and were ranked, so quarantine
needs a new state, "eligible to explore, barred from winning", rather than
reusing anything existing. Defer it until Q1 shows it's needed.

217
plans/cockpit-quick-wins.md Normal file
View File

@@ -0,0 +1,217 @@
# Cockpit quick wins: six small admin-portal fixes
Status: done -- shipped in PR #101
Date: 2026-09-23
Source: `plans/cockpit-brainstorm.md` section 5 (items R2, R5, E3, C1, G1, G3).
Each item below was observed on the live portal or checked read-only against
the live `router.db` on 2026-09-23.
## Ground rules (all items)
- **Worktree, not the main checkout.** Create a worktree from `main` HEAD
(`c87e675` or later) on branch `feat/cockpit-quick-wins`. The main checkout
at `/home/alee/Sources/6krrt` has UNCOMMITTED edits from other work,
including `admin/frontend/controls.html` (a `confidence_threshold` ->
`confidence_min` rename around lines 1026 and 1150) and
`tests/test_admin_frontend.py`. Do not touch, stash, or commit them. C1 edits
a different region of `controls.html`; keep it that way so the later merge
is clean.
- Commit by explicit path. Never `git add -A` or `git add .`.
- **Never touch port 8080** (production). For visual checks, run a throwaway
instance on **8081** from the worktree, with a temp copy of the DB, never the
live `router.db` for writes.
- Never edit `config/config.yaml` or `config/config.local.yaml`.
- ASCII only in new code, comments, and UI strings. No middle-dot separators
(U+00B7); relate facts with layout, or with `:` or a rephrase.
- No paragraphs of prose in the UI. One line per hint at most.
- One commit per item, in the order listed. `pytest` (offline) and
`ruff check` stay green after each commit.
- For every UI change, take a screenshot on 8081 and check the changed area
zoomed in, not just the item's checklist.
---
## 1. R5: Classifier card shows a wrong value on first paint
**Problem.** `admin/frontend/controls.html:264-268`: the
`#classifier-mode-select` `<select>` renders with `local_llm` selected by
default (first `<option>`). Until `loadClassifierConfig` (sets `.value` at
~line 1116) returns, the card claims `local_llm` when the live mode is
`local_encoder`.
**Fix.** Render the select disabled, with a leading
`<option value="" selected>loading...</option>`, and the Save button
disabled. On load: remove the placeholder option, set the value, enable both.
On fetch failure: keep them disabled and show an inline one-line error.
**Test.** In `tests/test_admin_frontend.py`: assert that the static HTML of the
select has a selected placeholder and the `disabled` attribute, and that no
real mode option is `selected` in the static markup.
## 2. G1: Nav links hardcoded in eight files
**Problem.** Every page (`index.html`, `models.html`, `profiles.html`,
`proficiency.html`, `decisions.html`, `quota.html`, `controls.html`,
`providers.html`) hardcodes the same `<ul>` of seven `nav-item` links, and
sets `active` by hand. `navbar.js` (the header comment) already records that
adding Quota was an eight-file edit.
**Fix.** Define the link list once in `navbar.js`
(`[{href, label}]`, same order as today: Models, Profiles, Proficiency,
Decisions, Quota, Controls, Providers). Render it into a placeholder
(e.g. `<ul class="navbar-nav" id="nav-links"></ul>`) on each page. Mark
`active` by matching `location.pathname`, with `/admin` and `/admin/` mapping
to Home (no link active). Remove the hardcoded `<li>`s from all eight pages.
Keep the rendered markup identical: same classes (`nav-item`, `nav-link`,
`active`), same hrefs, same order. Pages must still render their nav when
`navbar.js` is loaded, and it is already loaded on every page.
**Test.** One test that parses all eight HTML files and asserts none contains
a hardcoded `class="nav-link" href="/admin/...` list. One test asserting that
the `navbar.js` link list has exactly the seven hrefs above.
## 3. E3: Decisions filters are not in the URL
**Problem.** `admin/frontend/decisions.html` never reads or writes
`URLSearchParams`. No other page can link to "decisions for this model" or
"this session", and a filtered view cannot be shared or reloaded.
**Fix.**
- On load, after the filter `<select>`s are populated (~lines 405-420), read
`kind`, `category`, `profile`, `tier`, `q` (search) and `size` from
`location.search` and apply them. If a value is not among a select's
options, add it as an option rather than dropping it (a link may name a
category that has no rows in the loaded window).
- On any filter change (`resetPageAndRender`, the Clear button, page size),
write the current filters with `history.replaceState`. Omit empty values.
Don't push history entries.
- `q` also accepts a model id or a session id, since search already matches
both (`search` at ~line 433).
**Test.** A static test that `decisions.html` reads `URLSearchParams` and calls
`history.replaceState`. On 8081, open
`/admin/decisions?category=diff_checking&q=deepseek`, confirm the filters and
the rows match, and screenshot it.
## 4. R2: The Decisions tile counts verifications and mixes diagnostics with ground truth
**Problem.** `admin/frontend/index.html:597-610`: the Home "Decisions" tile
sums `_snap.verdict_mix`, which is `metrics.verdict_mix()` and counts
`verifications` rows by verdict. Live DB, last 7 days:
| verdict | kind | n |
|---|---|---|
| unverifiable | structural | 2,660 |
| succeeded | client_outcome | 96 |
| failed | client_outcome | 82 |
| ok | local_llm | 29 |
| ok | structural | 7 |
| malformed | local_llm / structural | 4 / 7 |
| truncated | structural | 2 |
The tile shows "2,887 verified in 7 days: 132 ok, 93 failed". Actual route
decisions in 7 days: **2,678**. "ok" adds client `succeeded` to structural and
`local_llm` ok. "failed" adds client `failed` to `malformed`. Per CLAUDE.md,
structural and `local_llm` verdicts are diagnostics only; `/outcome`
(`kind = 'client_outcome'`) is the only ground truth.
**Fix.**
- Backend: in `src/metrics.py`, add `decision_outcome_summary(conn,
since_days=7)` returning
`{decisions, client_reports, client_ok, client_failed}`:
`decisions` = `route_decisions` rows with `observed_at` inside the window;
`client_*` = `verifications` rows with `kind = 'client_outcome'` inside the
window (`succeeded` = ok, `failed` = failed). Include it in the admin
snapshot (`src/admin.py` near line 2131) as `decision_outcomes`. Leave
`verdict_mix` unchanged; the TUI uses it.
- Frontend: the tile headline becomes `decisions`. Qualifier lines:
- `routed in 7 days`
- `178 client reports (6.6%): 96 ok, 82 failed` (coverage =
`client_reports / decisions`; the failed count uses the warn colour when
the fail rate among reports is above 10%, matching the existing Signal
quality card threshold).
- With zero reports: `no client reports in 7 days`.
- Drop the word "verified".
- Use `metrics.py`'s existing window idiom:
`julianday(observed_at) > julianday('now', '-' || ? || ' days')`.
**Test.** `tests/`: a metrics unit test with a seeded in-memory DB that
includes structural, `local_llm` and `client_outcome` rows, asserting only
`client_outcome` rows are counted and that `decisions` comes from
`route_decisions`. A frontend static test that the tile no longer reads
`verdict_mix`.
## 5. G3: Warning dismissals don't stick on rate warnings
**Problem.** `admin/frontend/index.html:1010-1040`: a dismissal is keyed on
the warning's exact text. That's deliberate ("when the underlying numbers move
the text moves with them... the warning comes back"). But warnings that embed
a live rate or count ("3.2x pace", "n=4") change text on nearly every poll, so
a dismissal never holds.
**Fix (interim, until structured warnings exist).** Key each dismissal on
`digit-normalized text + severity`: replace every run of digits (and decimal
points inside a number) with `#`, then append `|` + `warningSeverity(text)`.
The warning comes back when its wording or its severity changes, not when a
digit moves. This is the same normalization `rejection_warnings` already uses
for grouping (CLAUDE.md, the reactive detector section).
- Migrate the existing stored keys on load: normalize each and dedupe, same
pattern as the existing one-time carry-over from `dismissedWarnings`.
- Update the block comment to state the new rule and why it changed.
**Test.** A static or JS-extracted unit test: two texts differing only in
digits produce the same key; the same text at different severity produces a
different key.
## 6. C1: One knob table instead of two columns
**Problem.** `admin/frontend/controls.html`: "Runtime Knobs" (`#runtime-list`,
`renderRuntime` ~line 687) and "Persisted Config" (`#config-list`,
`renderConfig` ~line 821) list largely the same knobs in two side-by-side
cards, in different orders, and persisted labels truncate
(`objective.incumbent_cache_pri...`). Checking whether a runtime value has
drifted from its persisted value means matching rows by eye.
**Fix.**
- One card, **Knobs**, full width, as a table:
`knob | live | persisted | layer | ` with the columns meaning:
- `knob`: the dotted config key, never truncated (wrap if needed).
- `live`: the runtime control, if the knob has one, else empty.
- `persisted`: the persisted-config control, if the key is in the persisted
allowlist, else empty.
- `layer`: the existing `base` / `overlay` badge.
- A **drift** badge in the row when live differs from persisted
("reverts on restart").
- Pair rows through the runtime-to-config mapping that already exists in
`src/admin.py` (`_BOOL_KNOBS` at ~line 348, plus the non-bool runtime knob
tables beside it). Expose that mapping to the frontend: add the dotted
config key to each item in the `GET /admin/api/runtime` response, rather
than duplicating the table in JS.
- Keep the existing save semantics exactly: runtime controls POST to
`api/runtime/{knob}` immediately, as today; persisted controls stay dirty
until the Save button, as today. Keep the existing hint badges (for example
"enables the reuse window").
- Keep the help lines ("Live on the next request...", "Applied on
restart...") as one line each, placed as column-header tooltips or a single
line under the table.
- Don't touch the classifier-card region (lines ~1000-1160); another branch
edits it.
**Test.** `test_admin_knob_coverage.py` must still pass unchanged. Add a
test that every runtime knob in the API response carries a config key, and
that the key resolves to a real `RouterConfig` path. Visual check on 8081:
set one runtime knob so it differs from persisted (on the 8081 instance
only), then confirm the drift badge shows. Screenshot the whole table and a
zoomed crop of one drifted row.
---
## Done means
- Six commits on `feat/cockpit-quick-wins`, one per item, in the order above.
- `pytest` and `ruff check` green on the worktree.
- Screenshots from 8081 for items 1, 3, 4 and 6.
- The 8081 instance stopped, by the PID you started.
- A short report listing each commit, the tests added, and anything
deferred, with the reason.

167
plans/deploy-separation.md Normal file
View File

@@ -0,0 +1,167 @@
# Deploy separation: production runs from its own clone, not the dev checkout
Status: planned -- spec and runbook ready, dry run done (see "Dry run"); the cutover is an operator step not yet run
Date: 2026-09-26
Why: incident #8, and #4, #5 and #7 before it. Production has been one `git pull`,
one uncommitted edit or one agent session away from the dev checkout.
## Current state (measured 2026-09-26)
- **Every installed user unit runs from `~/Sources/6krrt`**, the dev checkout:
- `llm-router.service`
- `llm-router-poller`, `-seed`, `-backup`, `-offsite`
- `llm-router-baseline-report`
- **What they take from it:** `WorkingDirectory`, `PYTHONPATH`,
`EnvironmentFile=.env`, `.venv/bin/*`, and `HF_HOME=.hf-cache`.
`ReadWritePaths` covers the whole source tree, so the live service can write
anywhere in the repo.
- **Everything else follows the working directory,** because the code resolves
paths against it: `config/config.yaml`, the `config/config.local.yaml`
overlay, and `database.path: "router.db"`. Moving the working directory moves
all of it, with no code change.
- **The repo's own unit templates** (`deploy/*.service`) already target a
separate `%h/llm-router` directory, except `-backup` and `-offsite`. The split
was the original design. The installed units drifted onto the dev checkout,
and `~/llm-router` does not exist.
- **Sizes:** `router.db` 48 MB, `.hf-cache` 2.8 GB, `.venv` 7.2 GB. `.env`
holds `NEURALWATT_API_KEY` and `OPENROUTER_API_KEY`.
## Target layout
```
~/.local/share/6krrt/app production clone of the Gitea repo, detached at a
merged origin/main commit. Nobody edits here; no
agent works here; nothing here ever runs `git clean`.
router.db live DB (gitignored; checkout never touches it)
config/config.local.yaml live overlay (gitignored)
.env live keys, mode 600 (gitignored)
.venv/ production venv, built from the clone's requirements
~/.cache/6krrt-hf shared HF model cache (moved once from
~/Sources/6krrt/.hf-cache; production and the dev
sandbox both point HF_HOME here)
~/Sources/6krrt dev checkout: source only. No live DB, no
production .env. The 8081 sandbox uses DB copies.
```
Units: every `llm-router*` unit's paths change from `%h/Sources/6krrt` to
`%h/.local/share/6krrt/app`, and `ReadWritePaths` narrows to that directory.
The backup and offsite scripts already take `REPO` from the environment; set it
to the app directory.
## `scripts/deploy.sh` (to build; not needed for the first cutover)
The only way code reaches production after the cutover:
1. **Show the change.** `git -C $APP fetch origin`, then
`git log --oneline HEAD..origin/main`.
2. **Preflight, without touching the running service:**
- load `origin/main`'s config against the real overlay (catches a #7-style
rename)
- check the requirements are satisfied in `$APP/.venv`
- `import dispatcher` from a temp export of `origin/main`
3. **Cut over.** Record the current SHA, `git checkout --detach origin/main`,
`uv pip install` only if the requirements changed, restart, and poll
`/health` for 30 s.
4. **Roll back on failure.** Check out the recorded SHA, restart, and alert.
5. **Record the deployed SHA** in `$APP/.deployed`, and expose it on `/health`
if that endpoint gains a field, so "what is running" has one answer.
## Cutover runbook
Downtime is about 1 minute. The venv is built BEFORE stopping anything. Run
the commands from a shell, not from an agent session.
```sh
APP=~/.local/share/6krrt/app
DEV=~/Sources/6krrt
# 0. Build the clone and venv while production keeps running.
mkdir -p ~/.local/share/6krrt
git clone ssh://git@git.adlee.work:2222/alee/6krrt.git "$APP"
git -C "$APP" checkout --detach origin/main
(cd "$APP" && uv venv .venv && uv pip install --python .venv/bin/python \
-r requirements.txt -r requirements-encoder.txt)
# 1. Share the HF cache (one move; the dev sandbox follows in step 6).
mkdir -p ~/.cache && mv "$DEV/.hf-cache" ~/.cache/6krrt-hf
# 2. Stop production and every timer that touches the DB.
systemctl --user stop llm-router-poller.timer llm-router-seed.timer \
llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer
systemctl --user stop llm-router.service
# 3. Copy state; originals stay put until step 7.
sqlite3 "$DEV/router.db" ".backup '$APP/router.db'"
command cp -p "$DEV/config/config.local.yaml" "$APP/config/config.local.yaml"
command cp -p "$DEV/.env" "$APP/.env" && chmod 600 "$APP/.env"
sqlite3 "$APP/router.db" 'PRAGMA integrity_check;' # must print: ok
# 4. Repoint the units (backups first), then reload.
cd ~/.config/systemd/user
mkdir -p ~/.local/share/6krrt/unit-backup && command cp -p llm-router* ~/.local/share/6krrt/unit-backup/ -r
sed -i "s#%h/Sources/6krrt#%h/.local/share/6krrt/app#g; s#/home/alee/Sources/6krrt#/home/alee/.local/share/6krrt/app#g" llm-router*.service
sed -i "s#^Environment=HF_HOME=.*#Environment=HF_HOME=%h/.cache/6krrt-hf#" llm-router.service
systemctl --user daemon-reload
# 5. Start and verify.
systemctl --user start llm-router.service
sleep 3; curl -s localhost:8080/health
systemctl --user show llm-router.service -p WorkingDirectory
curl -s localhost:8080/admin/api/runtime | head -c 200 # pinch persisted false = overlay loaded
systemctl --user start llm-router-poller.timer llm-router-seed.timer \
llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer
# 6. Point the dev sandbox at the shared HF cache and the new live DB.
# (edit scripts/sandbox.sh: HF_HOME and the live-DB path; see "Follow-ups")
# 7. Only after a day of healthy running: retire the dev copies.
mv "$DEV/router.db" "$DEV/router.db.pre-cutover-$(date +%Y%m%d)"
```
**Rollback, before step 7:** restore the unit backups from
`~/.local/share/6krrt/unit-backup/`, run `daemon-reload`, start the service.
The dev copies of the DB, overlay and `.env` were never moved, so production is
back where it started. Any decisions logged after the cutover stay in the
app-directory DB.
## Follow-ups that must land with or right after the cutover
- **`CLAUDE.md` North Star #3** and the "Measure the live `router.db`" memory
name `/home/alee/Sources/6krrt/router.db` as the live DB. Update both to
`~/.local/share/6krrt/app/router.db`, or agents will measure a stale file.
- **`scripts/sandbox.sh`:** `HF_HOME` goes to `~/.cache/6krrt-hf`. Its DB copy
instructions change to the app directory's DB.
- **`deploy/*.service` templates:** make the repo's templates say
`%h/.local/share/6krrt/app` so reinstalling from the repo reproduces this
layout, and fix `-backup` and `-offsite` to match.
- **`deploy/README.md`:** replace the install section with this layout and
`deploy.sh`.
- **The admin portal's "Restart Service" trigger** keeps working (it calls
`systemctl`), but any trigger that runs a script by repo-relative path should
be checked against the new working directory.
## Dry run
See the result appended below.
**Result, 2026-09-26 04:3x, against `origin/main` `827c408`:**
- **Clone from Gitea and detached checkout:** OK.
- **Fresh venv, `uv venv` plus both requirements files:** 2 s (all from `uv`'s
cache); `torch 2.9.1+cu128` and `transformers` import. Building the venv
does not affect downtime, and step 0 can run hours ahead.
- **State:** the DB copied by sqlite backup with the source opened `mode=ro`,
and `integrity_check` = ok; overlay copied; `.env` replaced with dummy keys
for the dry run.
- **Served on 8081** from the clone. `/proc/<pid>/cwd` confirmed the clone
directory, with `HF_HOME` on the dev cache read-only:
- `/health` 200; `/admin/controls` 200
- `/admin/api/runtime` showed `pinch_enabled` `persisted: False`, so the
copied overlay loaded
- the snapshot read the copied DB: 3,994 decisions, 200 client reports over 7
days
- `/route` considered 41 candidates and picked `deepseek/deepseek-v4-flash`
- **Stopped** by the recorded PID; 8081 free; the dummy `.env` deleted.
**Not exercised by the dry run:** the systemd unit edits (step 4); the timers
under the new `ReadWritePaths`; the HF cache move (step 1); and real provider
keys. The unit edit is the step to watch during the cutover. `systemctl --user
show llm-router.service -p WorkingDirectory` in step 5 confirms it took.

105
plans/docs-lift-a-pass.md Normal file
View File

@@ -0,0 +1,105 @@
# Docs pass after Lift A (PR #102)
Status: done -- executed on branch docs/lift-a-docs-pass
Date: 2026-09-26
This brief is the contract. Keep the plan small: one todo per doc, each with
exact section anchors. Do NOT read `plans/no-progress-detection.md` (reference
only).
## Goal
Bring the docs in line with what PR #102 shipped and what is running now. The
docs describe the code as it is on this branch; where a doc and the code
disagree, the code wins, and the todo says which line it checked.
## Where the work happens
- Worktree: `/home/alee/Sources/6krrt-worktrees/docs-lift-a-pass`, branch
`docs/lift-a-docs-pass`, based on `origin/main` `a24de1b` plus three commits
that already carry North Star #4 in `CLAUDE.md`, the incident #8 rewrite in
`docs/incidents.md`, and the plans. Build on them; do not rewrite them.
- **Docs only.** No code, config, test or unit-file changes.
- **Do not touch** `README.md`, `AGENTS.md`, `deploy/README.md`: the owner has
uncommitted edits to them in the main checkout. Anything that belongs in
`deploy/README.md` goes in `docs/watchdog.md` instead, and the done report
lists it for the owner to fold in.
## What shipped (verify each against the code before writing it)
- `src/progress_detect.py`: pure detector. Signals dup, top, slow, coverage;
landed gating; read-only agents; canonical (sorted-key) args; bash command
normalization (comment lines, `cd dir &&`, `VAR=x` prefixes dropped).
- `src/watchdog.py` + `deploy/llm-router-watchdog.{service,timer}`: 5-minute
oneshot. Judges only sessions with a tool call since the last tick, on their
full history, per session; landed counts the session's own descendants only.
Alert lifecycle in `watchdog_alerts`: trigger once, escalate once (LLM yes or
3 ticks), resolve once. Quiet `no_opencode` tick when opencode is closed.
Optional local-model second opinion (does not veto).
- `src/notifier.py`: `desktop` channel (`notify-send`), PagerDuty-Events-v2
shaped `AlertEvent`, per-channel `min_severity`, rate limit backed by the DB.
- `src/watchdog_store.py`: tables `watchdog_ticks`, `watchdog_verdicts`,
`watchdog_alerts`, `watchdog_channel_settings`.
- Admin: Loops panel (`index.html#loops`), Watchdog card on Controls (last
tick, verdicts, channels, Send test alert), per-model stall rollup and
Block/Unblock on Models. New availability value `blocked`, excluded from
routing and from pinned requests via `_admin_excluded_models`.
- Endpoints: `GET /admin/api/watchdog/status`, `POST /admin/api/watchdog/channels`,
`POST /admin/api/watchdog/test-alert`; availability endpoint accepts `blocked`.
- Config: `watchdog.*` (incl. `detector.*`, `dashboard_base_url`),
`notifications.channels`. Watchdog knobs are persisted-only (separate process).
- `deploy/opencode-plugin/router-link.js`: the `parentCache` export made
opencode 1.18 reject the whole plugin ("Plugin export is not a function" in
`~/.local/share/opencode/log/opencode.log`), so identity headers never
reached the router from #99 until this fix. A loader-contract test now asserts
every export is a function.
- `scripts/progress_backtest.py` (`--fixture` replays the 15 labelled sessions;
expect 8/15 flagged) and `scripts/export_progress_fixture.py`.
- Operator state 2026-09-26: timer enabled ~15:18 EDT with unit paths repointed
from `%h/llm-router` to `%h/Sources/6krrt` (production still runs from the
dev checkout; see `plans/deploy-separation.md`); plugin fix installed and
opencode restarted 15:19; router restarted 15:19:57.
## Todos (one per doc)
1. **New `docs/watchdog.md`**: what it catches and does not (steady spend is not
waste), signals and thresholds, alert lifecycle, the admin surfaces, knobs,
install (the `sed` repoint + `enable --now`), how to run `--once`, backtest,
re-exporting the fixture, and known limits: sessions under 40 calls are
invisible; the Atlas must-flag case sits exactly at `cover_min` 4.0; the
stale `localhost:4096` entry in `rc-servers.json` logs a harmless probe
warning.
2. **`CLAUDE.md`**: "What's built and working" (anchor `## What's built and
working`): add the modules above in the existing one-bullet style, linking
`docs/watchdog.md`; update the `tests/` bullet count from the real
`pytest --collect-only -q` total. "Run as a service" (anchor `## Run as a
service`): one short paragraph on the watchdog timer. Do not edit the North
Star section.
3. **`docs/incidents.md` #8**: under Status, add that Lift A shipped (PR #102,
`a24de1b`) and the timer is live; add the plugin-loader finding as Evidence
(with the log line) and to the Human response timeline; add a symptom-table
row "opencode plugin does nothing -> grep opencode.log for `failed to load
plugin`". Keep the Evidence / labelled Theory convention in the file header.
4. **`docs/admin-portal.md`**: sections for the Loops panel, the Watchdog card,
and the Models rollup with Block/Unblock (where `blocked` differs from
`deprecated`: an operator's reasoned stop, one click to undo).
5. **`docs/api.md`**: the three watchdog endpoints and `blocked` on the
availability endpoint (request/response shapes read from `src/admin.py`).
6. **`docs/data-model.md`**: the four watchdog tables (from `config/schema.sql`),
and `blocked` in the `availability` row at the `models` section (line ~34)
and wherever `admin_model_overrides` is described.
7. **`docs/clients.md`** (anchor `## Session-directory attribution (opencode
plugin)`) and `docs/api.md` near the `X-Router-*` table (line ~151): the
plugin loader contract (every export must be a function) and how to confirm
the plugin loaded.
## Rules
- ASCII only; no middle dots. Match each file's existing style and heading
depth. Admin-portal docs describe the UI, they do not add prose to it.
- Numbers and names come from the code on this branch, never from this brief.
- Commit per todo, by explicit path. Never `git add -A`.
- Do not push and do not open a PR. The done report lists commits and the
items for `deploy/README.md`.
- If a worker's context passes ~150k tokens, stop it and start a fresh one.
Write `.omo/plans/docs-lift-a-pass.md`, then stop and report.

View File

@@ -0,0 +1,241 @@
#!/usr/bin/env python3
"""Interim opencode loop watcher (stand-in until the no-progress lift ships).
Signals per session, over a sliding window of its last WINDOW tool calls:
dup share of calls that repeat an earlier (tool, args) in the window
top count of the most repeated (tool, args) in the window
landed any edit/write with a non-empty diff, or a `git commit` with exit 0,
by the session OR any descendant within the window's time span
Flag when the window is full enough and nothing landed and
dup >= DUP_MIN or top >= TOP_MIN.
Read-only agents (explore/librarian/oracle) never land, so for them only
top >= TOP_MIN_RO counts.
Modes:
backtest : replay every session's history, report worst window per session
watch : poll live sessions; on a new flag print it, notify-send, exit 0
"""
import json, sys, time, subprocess, urllib.request, collections, os
U = os.environ.get("OC_URL", "http://127.0.0.1:4097")
WINDOW, MIN_CALLS = 60, 40
DUP_MIN, TOP_MIN, TOP_MIN_RO = 0.25, 12, 8
CUM_MIN = 15 # one exact call repeated this often over the whole session: a slow loop
COVER_MIN = 4.0 # lines read from one file in the window / its length: re-reading, whatever the offsets
READ_ONLY = ("@explore", "@librarian", "@oracle", "explore subagent")
def get(path):
return json.load(urllib.request.urlopen(U + path, timeout=30))
def calls_of(sid):
out = []
for m in get(f"/session/{sid}/message"):
for p in m["parts"]:
if p.get("type") != "tool":
continue
st = p.get("state") or {}
inp = st.get("input") or {}
md = st.get("metadata") or {}
t = (st.get("time") or {}).get("start") or m["info"].get("time", {}).get("created", 0)
tool = p.get("tool")
landed = (tool in ("edit", "write", "patch") and bool(md.get("diff"))) or (
tool == "bash" and "git commit" in (inp.get("command") or "") and md.get("exit") == 0)
out.append((t, tool, json.dumps(inp, sort_keys=True), landed))
return out
import re
_PATHISH = re.compile(r"[\w./-]+\.(?:json|md|py|js|html|yaml|yml|toml|txt|sql|sh|mjs|css)\b")
_CD_PREFIX = re.compile(r"^\s*cd\s+\S+\s*(?:&&|;)\s*")
_ENV_PREFIX = re.compile(r"^\s*(?:[A-Z_][A-Z0-9_]*=\S+\s+)+")
def _normalize_bash(cmd):
"""Drop comment lines and leading `cd dir &&` / `VAR=x` prefixes, so
differently-worded commands under one prefix do not collapse to one key."""
lines = [ln for ln in cmd.splitlines() if not ln.lstrip().startswith("#")]
c = " ".join(lines).strip()
for _ in range(3):
c2 = _ENV_PREFIX.sub("", _CD_PREFIX.sub("", c))
if c2 == c:
break
c = c2
return c
def target_of(tool, args_json):
"""Coarse target: the file a call is about, ignoring offsets and the exact
command wording, so 28 differently-worded `cat boulder.json` count as one."""
try:
a = json.loads(args_json)
except Exception:
return (tool, args_json[:80])
if a.get("filePath"):
# A read keys on its slice: walking a big file in pieces is progress,
# re-reading the same slice is not.
return ("file", os.path.basename(a["filePath"]), a.get("offset"), a.get("limit")) if tool == "read" \
else ("file", os.path.basename(a["filePath"]))
if tool == "bash":
cmd = _normalize_bash(a.get("command") or "")
files = sorted({os.path.basename(f) for f in _PATHISH.findall(cmd)})
if files:
# Numbers in the command are the slice (sed -n 100,160p, head -40):
# walking a file is progress; the same slice again is not.
nums = tuple(sorted(set(re.findall(r"\b\d+\b", cmd))))
return ("bash-files", ",".join(files[:3]), nums)
return ("bash", " ".join(cmd.split()[:2]))
if tool in ("grep", "glob"):
return (tool, str(a.get("pattern"))[:60])
return (tool, args_json[:80])
def window_stats(win):
fp = collections.Counter((c[1], c[2]) for c in win)
dup = sum(v - 1 for v in fp.values() if v > 1) / max(1, len(win))
tg = collections.Counter(target_of(c[1], c[2]) for c in win)
(top_k, top_n) = tg.most_common(1)[0] if tg else (("", ""), 0)
return dup, top_n, top_k
_LEN = {}
def file_lines(path):
if path not in _LEN:
try:
with open(path, "rb") as fh:
_LEN[path] = max(1, fh.read().count(b"\n"))
except OSError:
_LEN[path] = None
return _LEN[path]
def coverage(win):
"""Max over files of (lines requested in the window / file length)."""
req = collections.Counter()
for c in win:
if c[1] != "read":
continue
a = json.loads(c[2])
fp = a.get("filePath")
n = file_lines(fp) if fp else None
if not n:
continue
req[fp] += min(a.get("limit") or n, n)
best = max(((v / file_lines(k), k) for k, v in req.items()), default=(0.0, ""))
return round(best[0], 1), os.path.basename(best[1])
def is_ro(title):
t = (title or "").lower()
return any(k in t for k in READ_ONLY)
def evaluate(sess, calls, landed_times_tree, end_idx=None):
calls = calls if end_idx is None else calls[:end_idx]
win = calls[-WINDOW:]
if len(win) < MIN_CALLS:
return None
t0, t1 = win[0][0], win[-1][0]
landed = any(t0 <= lt <= t1 for lt in landed_times_tree)
dup, top_n, top_k = window_stats(win)
ro = is_ro(sess.get("title"))
cum_k, cum_n = collections.Counter((c[1], c[2]) for c in calls).most_common(1)[0]
slow = (not landed) and cum_n >= CUM_MIN
cov, cov_file = coverage(win)
reread = (not landed) and cov >= COVER_MIN
if reread and not (top_n >= TOP_MIN):
top_n, top_k = int(cov), ("coverage", f"{cov_file} read {cov}x its length")
if ro:
flag = top_n >= TOP_MIN_RO or slow or reread
else:
flag = ((not landed) and (dup >= DUP_MIN or top_n >= TOP_MIN)) or slow or reread
if slow and not (top_n >= TOP_MIN):
top_n, top_k = cum_n, ("exact-total", cum_k[0], cum_k[1][:80])
return dict(flag=flag, dup=round(dup, 2), top=top_n, top_what=" ".join(str(x) for x in top_k)[:120],
landed=landed, ro=ro, n=len(calls), t1=t1)
def tree_landed(sessions, calls_by):
kids = collections.defaultdict(list)
for s in sessions:
if s.get("parentID"):
kids[s["parentID"]].append(s["id"])
memo = {}
def lt(sid):
if sid in memo:
return memo[sid]
own = [c[0] for c in calls_by.get(sid, []) if c[3]]
for k in kids.get(sid, []):
own += lt(k)
memo[sid] = own
return own
return lt
def backtest(since_ms):
sessions = [s for s in get("/session") if s.get("time", {}).get("updated", 0) >= since_ms]
calls_by = {s["id"]: calls_of(s["id"]) for s in sessions}
lt = tree_landed(sessions, calls_by)
for s in sessions:
cs = calls_by[s["id"]]
worst, first_flag = None, None
for i in range(MIN_CALLS, len(cs) + 1, 5):
r = evaluate(s, cs, lt(s["id"]), i)
if r and (worst is None or (r["flag"], r["dup"], r["top"]) > (worst["flag"], worst["dup"], worst["top"])):
worst = r
if r and r["flag"] and first_flag is None:
first_flag = i
if worst:
print(f"{'FLAG' if worst['flag'] else 'ok '} calls={len(cs):4d} first_flag_at={first_flag} "
f"dup={worst['dup']} top={worst['top']} landed={worst['landed']} ro={worst['ro']} | {s.get('title','')[:50]}"
+ (f"\n top: {worst['top_what'][:110]}" if worst['flag'] else ""))
def watch(minutes, state_path):
seen = set()
if os.path.exists(state_path):
seen = set(json.load(open(state_path)))
deadline = time.time() + minutes * 60
while time.time() < deadline:
try:
now = time.time() * 1000
sessions = get("/session")
status = get("/session/status")
recent = [s for s in sessions if s["id"] in status or now - s.get("time", {}).get("updated", 0) < 5 * 60e3]
calls_by = {s["id"]: calls_of(s["id"]) for s in recent}
# Judge only sessions doing work now: busy, or a tool call in the
# last 5 min. A session's "updated" stamp moves on abort, restart
# and title edits, so it cannot stand in for activity.
active = [s for s in recent if s["id"] in status
or (calls_by[s["id"]] and now - calls_by[s["id"]][-1][0] < 5 * 60e3)]
lt = tree_landed(sessions, calls_by)
for s in active:
r = evaluate(s, calls_by[s["id"]], lt(s["id"]))
key = f"{s['id']}:{len(calls_by[s['id']]) // 60}"
if r and r["flag"] and key not in seen:
seen.add(key)
json.dump(sorted(seen), open(state_path, "w"))
msg = (f"{s.get('title','')[:60]} ({s['id']}): {r['n']} calls, dup {r['dup']}, "
f"top x{r['top']}: {r['top_what']}, landed={r['landed']}, busy={s['id'] in status}")
print("LOOP SUSPECT:", msg)
subprocess.run(["notify-send", "-u", "critical", "opencode loop suspected", msg[:250]],
check=False, timeout=5)
return 0
except Exception as e: # opencode down or restarting: keep watching
print(f"[{time.strftime('%H:%M:%S')}] poll error: {e}", file=sys.stderr)
time.sleep(120)
print("no loops seen in", minutes, "min")
return 0
if __name__ == "__main__":
if sys.argv[1] == "backtest":
backtest(int(sys.argv[2]))
else:
sys.exit(watch(int(sys.argv[2]), sys.argv[3]))

View File

@@ -0,0 +1,572 @@
# No-progress detection: catch an agent that is spinning, not one that is working
Status: reference -- parent spec; Lift A shipped in PR #102, Lift B (router-side judge, modes, behavior breaker) not started
Date: 2026-09-26
Trigger: the `cockpit-quick-wins` Atlas run, 2026-09-25 19:00 to 2026-09-26 02:00.
Incident write-up: `docs/incidents.md` #8.
## What happened, measured
Read-only against the live `router.db` and opencode's own session store
(`http://127.0.0.1:4097`), 2026-09-26.
Spend, 19:00 to 02:00 local:
- 323M prompt tokens (95% cached) and about $7.81, over roughly 8 hours.
- 79% of the money went to OpenRouter `z-ai/glm-5.3-flash`: 1,368 calls
averaging 203k prompt tokens.
- 921 calls carried 120k+ prompt tokens and cost $6.57. NeuralWatt `glm-5.3`
took 47 calls averaging 673k context, max 682,679, because nothing else fit.
- Every decision was `classification_source = classifier` on the default
profile. The router did what it was told; nothing was pinned.
**Spend rate is the wrong signal.** Tonight's worst hour ($2.42) and worst
3-hour window ($4.13) sit inside the prior 30 days' normal range (hourly p90
$1.46, max $4.83; worst 3h $9.65). A rate alarm quiet enough for a healthy
all-day agent run would have missed this entirely, and one tight enough to catch
it would fire on good work. The operator's framing is the right one: steady
usage is fine. What is not fine is spend with **no concrete change landing**.
**The sessions that wasted money are visibly different in their tool calls:**
| session (opencode) | tool calls | exact-duplicate calls | worst repeat | landed a change? |
|---|---|---|---|---|
| Item 2 worker, 2nd attempt | 423 | 30% | `read navbar.js` x61 | no (worktree unchanged) |
| Item 3 worker, 1st attempt | 132 | 23% | `read test_admin_js_units.py` x16 | no |
| Item 4 worker | 134 | 19% | `read .omo/plans/...` x8 | partly (code, no tests) |
| Item 2 worker, 1st attempt | 307 | 17% | `read navbar.js` x21 | no |
| Atlas orchestrator | 576 | 11% | playwright console check x10 | yes (via workers) |
| healthy workers (items 1, 5, 6, 3-retry) | 30-96 | 0-6% | 1-3 | yes |
**A second loop shape, 2026-09-26 02:21 to about 02:56:** the Atlas orchestrator
itself, after its plan was complete:
- `.omo/boulder.json` carried a stale `status: completed` and `pr_url` from the
previous plan.
- oh-my-openagent's continuation hook injected "continue" turns; all 3 user
turns in the window were injected.
- Atlas, working from a trimmed context, re-derived its state every time:
54 turns, about 19M cached-read tokens, `boulder.json` read 28 times, the
plan's checkboxes grepped 16 times, nothing landed.
The operator caught it only by watching. The judge must treat this as
`stalling`, and Phase 0 must include it as a calibration case. The plugin can
also see injected continuation turns (`chat.message`), which is a useful extra
signal: "continuation turns keep arriving and nothing lands" is this shape
exactly.
**A third shape, 2026-09-26 around 03:05 to 03:17: a read-only agent.**
Prometheus's `explore` subagent reached 322 tool calls, 19% exact duplicates,
and read `deploy/opencode-plugin/router-outcome.js` 20 times. The file had been
deleted from under it: the production checkout was synced to `origin/main`,
where #99 retired it. The subagent was aborted.
Read-only agents (`explore`, `librarian`, `oracle`, and the like) never land a
change by design, so "no landed change" cannot be their signal. For them the
judge uses:
- target novelty: new files or queries per N calls, and
- a failure streak on the same target, such as repeated reads of a path that
no longer exists.
Identify them by the `X-Router-Agent` value, via a configurable list of
read-only agent names.
"Exact duplicate" means the same tool with byte-identical arguments. Duplicate
share alone does not separate the orchestrator (11%, productive) from a stuck
worker (17%). Duplicates **plus** no landed change does.
## Why nothing caught it
- `circuit_breaker.py` is availability-only: it trips on provider 5xx.
- The only spend alarm is `runway_low_warning`: balance / burn < 6 h, with burn
averaged over a 24 h balance window. OpenRouter read $24.73 at $0.25/h, so 97 h
of runway. It answers "will I run out", not "is this being wasted".
- The router cannot see progress at all. It sees prompts, not whether an edit
landed or a command keeps failing. The opencode plugin sees both.
- The DEPLOYED router cannot tell conversations apart either. `session_key`
merges concurrent conversations under one multi-agent client, and `/outcome`
attribution falls back to a directory match that 409s when several sessions
share a cwd, which is exactly the Atlas-plus-workers shape. **The fix already
exists on `origin/main` (PR #99, conversation-identity) but is not deployed:**
production runs from a local checkout 29 commits behind `origin/main`, the
live `route_decisions` has no `agent` / `parent_key` columns, and the
installed plugin is still `router-outcome.js` rather than #99's
`router-link.js`.
## Design
Split the job where the information lives. **The plugin is the sensor; the
router is the judge.** No raw task text leaves opencode, and the router never
stores any; the "never store raw task text" rule is unchanged.
### 1. Exact session identity: ALREADY BUILT upstream (#99), build on it
PR #99 (`conversation-identity`, merged to `origin/main` as `fe8c813`) did this.
Reuse it; do not rebuild it:
- **Plugin.** `deploy/opencode-plugin/router-link.js` replaces
`router-outcome.js`. Its `chat.headers` hook sends `X-Router-Conversation`
(the opencode session id), `X-Router-Agent` (slugged to the router's charset)
and `X-Router-Parent`, with a bounded parent-lookup cache.
- **Router.** `src/conversation_identity.py` resolves identity from those
headers. `session_key` becomes `'c:' + conversation id`. `route_decisions`
gained `agent` and `parent_key` columns. `/outcome` attributes to the exact
conversation.
What remains for this lift:
- **Deploy #99.** It is merged but not running. Syncing the production checkout
to `origin/main` and installing `router-link.js` are operator steps: the main
checkout has uncommitted work, and part of it, the `confidence_min` rename,
also landed upstream via #100.
- **Fix the exit code in `router-link.js`.** It kept `router-outcome.js`'s
`output?.exitCode ?? output?.exit_code` (line 83 on `origin/main`); see "Fixes
that ride along".
- **Build the progress events (section 2) into `router-link.js`,** keyed by the
same conversation and parent ids, so the judge rolls children up to parents
through `parent_key`.
Everywhere below, "session" means #99's conversation. Where this spec says
`X-Opencode-Session`, read `X-Router-Conversation`.
### 2. Progress events (plugin `tool.execute.after` -> router `POST /progress`)
The plugin sends one small event per tool call, fire-and-forget with a short
timeout, never blocking the session:
```
{session_id, parent_session_id, agent, tool, call_id,
fingerprint, # sha256 of tool + normalized args; never the args themselves
target, # coarse, normalized: the file path for read/edit/write, the
# command's first two tokens for bash; no contents
exit, # bash: output.metadata.exit
landed} # bool, see below
```
`landed` is the load-bearing definition, and it is deliberately narrow:
- `edit` / `write` / `patch` with a non-empty `metadata.diff`
- `bash` whose command is `git commit` (or `git commit --amend`) with exit 0
- A test or lint command (the plugin's existing `TEST_COMMAND` list) that PASSES
after the session's previous run of the same fingerprint FAILED, i.e. a fix
landed.
Reads, greps, todo writes, sub-agent dispatches and passing-again tests are not
progress. An orchestrator's progress is its children's: roll child `landed`
events up to the parent via `parent_session_id`.
Storage: a new `progress_events` table, pruned to a rolling window (knob). It
holds fingerprints, targets and booleans, no content.
### 3. The judge (router, per client session)
Signals per session, over a rolling window of its last N tool events:
- `turns_since_landed`: LLM requests since the last landed event (self plus
children for a parent).
- `usd_since_landed`: billed spend since then, joined through `request_id`.
- `dup_share`: exact-duplicate fingerprint share in the window.
- `fail_streak`: consecutive nonzero exits on the same fingerprint (the
"sandbox is broken and the agent keeps retrying" case).
- `context_tokens`: the latest request's size, to report alongside, since a
stall at 600k context costs far more per turn than one at 30k.
Verdicts:
- `ok`: something landed recently.
- `stalling` (warn): no landed change for `stall_turns` turns or
`stall_usd` dollars, AND (`dup_share >= dup_warn` OR `fail_streak >= fail_warn`).
- `looping` (act, if enabled): the stall condition at the higher `loop_*`
thresholds.
Both conditions have to hold. Long-running healthy work with no duplicates (a
big read-heavy exploration) stays `ok` until it hits the plain turns/spend
ceiling, which is a separate, looser knob.
**Calibrate before trusting it.** Phase 0 replays the judge offline over
opencode's session store (the 11 sessions above plus older history), and reports
the verdict each session would have received and when. Thresholds are picked
from that, not guessed. The Atlas orchestrator (11% duplicates, productive
through children) is the key false-positive case, and it must stay `ok`.
### 4. Optional: a local model as a second opinion
The deterministic signals should carry this. Where they are ambiguous, for
example near-duplicates (the same file read at shifting offsets, or the same
failing test with a different `-k`), there are two cheaper steps before any
model:
- normalize `read` fingerprints to the path, and bash to command plus target
- the local encoder already loaded for `classifier.mode: local_encoder` can
embed the `target` strings to cluster "similar-ish" events
A local LLM judge ("does this digest look like a loop?") runs only on a session
already flagged `stalling`, on a digest of tool names, targets and exit codes,
never content. It is gated by `local_compute.enabled`, since the operator pays
the local power bill. Treat this as a later phase, justified by Phase 0 showing
ambiguous cases the deterministic signals miss.
### 5. What happens on a detection: modes
One knob, `progress.mode`, picks the response. **Every mode warns:**
- A `/metrics` warning (admin bell and TUI). It names the session, agent, turns
and $ since the last landed change, and the top repeated target.
- An opencode toast via `POST /tui/show-toast` in the window the operator is
watching.
- An **alert** through the router's notifier (section 6), which reaches the
operator when nobody is looking at either screen.
| mode | on detection |
|---|---|
| `warn` | Warnings only. The neutral default: today's behaviour plus a warning. |
| `auto_compact` | Compact the stuck session with the incident report injected, then continue. No refusal, no orchestrator stop. |
| `auto_limit_context` | Cap the stuck session's context. No stop. |
| `auto_recover` | Two stages, below. |
**`auto_recover`, stage 1 (first detection in a session tree):**
1. The router refuses the stuck session's chat requests (429, scoped by
`X-Router-Conversation`), so nothing more is spent while the plugin acts.
2. The plugin walks `parentID` to the tree's root, usually Atlas, and calls
`POST /session/{id}/abort` on the stuck session and on the root.
3. The plugin compacts the root: `POST /session/{root}/summarize`. Its
`experimental.session.compacting` hook appends the **incident report** to
the compaction prompt, so the report survives into the compacted context.
4. `experimental.compaction.autocontinue` stays enabled, so the root restarts
from the compacted context with the report in view. The router lifts the
refusal when the compaction completes.
5. Toast: "recovered <agent>: <one-line reason>".
**`auto_recover`, stage 2 (detected again in the same tree within
`recover_window`):** apply the `auto_limit_context` cap to the tree, and toast
again.
**Stage 3, a hard stop.** Recovery can itself loop: recover, loop, recover. After
`max_recoveries` in the window, the router refuses the tree and does not
restart it. It toasts and raises a high-severity warning, and the tree waits for
an admin Resume. Without this cap, auto-recovery is an unattended retry loop,
which is the failure it exists to stop.
**The incident report** is built deterministically from `progress_events`
(fingerprints, targets and exit codes, never content), so it is free and cannot
hallucinate. Example:
```
[router] A sub-agent stalled and was stopped.
session: Item 2 G1 nav refactor (Sisyphus-Junior), child of this session
423 tool calls, 0 landed changes (no edit with a diff, no commit)
most repeated: read admin/frontend/navbar.js x61, read tests/... x16
spend since last landed change: $1.84, 70M prompt tokens
Avoid: re-reading whole files, since reads repeat without edits. Give the
worker the exact edit to make, and require an on-disk diff or commit before it
reports done.
```
The "Avoid:" line maps from the dominant signal: duplicates, a failure streak,
or context size. A local model may rephrase it later (section 4); it is not
needed to produce it.
**"Limit context"** has two candidate mechanisms. Phase 0 must verify which works
before stage 2 is built:
- **(a) Router-side, per-session tighter `pinch` budget.** The machinery
exists (`context_prune.py`), but pruning rewrites the prefix, and
`plans/token-waste-waves.md` Wave 3 measured rewrites costing the cache. So
this trades cache hits for size. Measure it.
- **(b) The router answers an over-cap request with a context-overflow-shaped
error**, so opencode runs its own overflow compaction. The autocontinue hook
receives `overflow: boolean`, so opencode does have such a path. Unverified:
whether it recognizes a router error as an overflow. Test this on 8081 before
relying on it.
Also, the Controls or Home page gets a live "Sessions" readout: session, agent,
parent, verdict, mode stage, turns and $ since landed, and the top repeated
target. This fits `plans/cockpit-brainstorm.md`. Each tree with a refusal gets
a Resume button.
Every threshold and the mode are knobs with admin controls (North Star #1).
Per the "build for re-tuning" rule, the neutral default is `warn`, with Phase
0-calibrated thresholds.
Without the plugin (another client, or a stale install), only the router half
works: warnings, alerts and the scoped 429. Abort, compact and restart need the
plugin. The Sessions readout says which sessions have a live plugin (from the
headers in section 1), so a mode that silently cannot act is visible.
### 6. Alerting: a notifier with pluggable channels
Toasts and the admin bell only help someone who is looking. Incident #8 ran
overnight. Alerts come from the **router**, not the plugin: the router is the
judge and is always running, while the plugin dies with the opencode process
it lives in.
**Ship now:** one channel, `desktop` (`notify-send`). **Design for later:** SMS
(Twilio or similar), RingCentral, PagerDuty, a generic webhook (ntfy, Slack).
Each of those should be a small adapter added without touching the detector.
Shape:
- **An alert is an event with a lifecycle**, not a message. Fields:
- `dedup_key`: for this detector, the tree root session id
- `severity`: `info`, `warning` or `critical`
- `state`: `trigger`, `escalate` or `resolve`
- `title` and a one-line `summary`
- `details`, the incident report from section 5
- `source`: `no-progress`, so other detectors can reuse the notifier later
This maps one-to-one onto PagerDuty Events v2 (`dedup_key`,
`event_action: trigger/resolve`), so that adapter is thin. Channels without
a lifecycle (SMS, desktop) render `trigger` and `escalate` as a message and
`resolve` as a short "recovered" message, or skip it (per-channel knob).
- **Transitions, not repeats.** An alert fires when a verdict *changes*, e.g.
`ok -> stalling`, `stalling -> looping`, a recovery, the hard stop, or
`-> ok`. It never fires per request. Per-channel rate limit and a quiet
window stop a flapping session from paging anyone at 3am every minute.
- **Severity routing per channel:** each channel has `min_severity`. The
intended setup: desktop gets `warning` and up; an SMS or PagerDuty channel
added later gets only `critical`, i.e. the hard stop, or `looping` in `warn`
mode.
- **Delivery never blocks routing.** Alerts go through a small queue with a
timeout and one retry. A failed delivery is logged and shown in the admin
UI; it never slows or fails a chat request.
- **Secrets stay out of config files.** Channel credentials (API keys, routing
keys, phone numbers) are read from environment variables named in config,
the same `api_key_env` pattern `dispatch_providers` already uses, and loaded
from `.env` by the systemd unit. `config.yaml` and `config.local.yaml` never
hold a secret.
- **Config:** a `notifications.channels` list, each entry
`{name, type, enabled, min_severity, resolve_messages, ...type-specific}`.
Only `type: desktop` exists in this lift. An unknown `type` fails config load
(StrictModel), so a typo cannot silently disable alerting.
- **Admin (North Star #1):** per channel, an enabled toggle, a `min_severity`
select, and a **Send test alert** button. The test button is the only way to
know a channel works before the night it matters. Recent deliveries and
failures show under the Sessions readout.
**`desktop` specifics, to verify in Phase 0.** The router runs as a systemd
*user* service, so `notify-send` needs the session D-Bus
(`DBUS_SESSION_BUS_ADDRESS`, normally `unix:path=/run/user/<uid>/bus`). Check
that the unit's environment has it and that its sandboxing (`ProtectHome` and
friends) allows the socket. If not, say so in the plan and fix it in
`deploy/`, not by weakening the sandbox wholesale. Use urgency `critical` for
`critical` alerts so they persist on screen.
### 7. The watchdog timer: detection that needs nobody watching
**Why.** All three loops in incident #8 were caught by a human or by a Claude
Code session reading opencode's session store by hand. The operator does not
want to depend on either. The watchdog is an independent, local, scheduled
sensor. It runs whether or not anyone is at the screen, whether or not a Claude
session is open, and whether or not the plugin loaded. **It calls no cloud
model.**
**What.** A systemd user timer, `deploy/llm-router-watchdog.{service,timer}`,
shaped like the existing `llm-router-*` timers, every `watchdog.interval`
(proposed 5 min). Each tick:
1. **Find opencode.** Read `~/.local/share/opencode/rc-servers.json`, written
by the `session-registry` plugin. Probe each recorded `serverUrl`. If none
answers, exit quietly; no opencode means nothing to watch.
2. **List sessions active since the last tick** via `GET /session`, and read
each one's recent messages via `GET /session/{id}/message`.
3. **Deterministic check, free, on every active session.** The same signals as
the judge (section 3): exact-duplicate share, the most-repeated target, the
failure streak, landed changes (for read-only agents, target novelty
instead), and injected continuation turns. This is Phase 0's replay code
run live. One implementation, imported by both, not a copy.
4. **Local LLM second opinion, only on sessions step 3 flags.** Skipped when
`local_compute.enabled` is off, which keeps gaming mode honest. The model is
whatever is configured: a `watchdog.model` knob that defaults to
`verification.model`, on that section's base URL. Today that is
`qwen2.5-coder-router:14b` on the local Ollama (`localhost:11434`), already
kept warm by local verification, so a call pays no model-load delay. It
runs on a digest of tool names, targets and exit codes, never content.
Phase 0 measures this model's verdicts on the three incident #8 cases
(must flag: the stuck workers and the post-completion Atlas loop; must not
flag: productive Atlas) before its answer is trusted. It asks "is this agent looping without progress? yes or no,
with a one-line reason". A healthy tick costs zero inference.
5. **Report, do not act.** POST the verdict to the router (the `/progress`
judge or a sibling endpoint), tagged `source: watchdog`. The **router**
applies the configured mode, alerts through the notifier, and enforces
`max_recoveries` and Resume. The watchdog never aborts or compacts anything
itself, so there is exactly one recovery path and one place that decides.
**If the router is down,** the watchdog still sends a `critical` desktop
alert directly with `notify-send`, and nothing else. A dead router is exactly
when an unattended run needs a human.
**Failure modes to design for:**
- A tick that overruns the interval: use a lock file and skip the tick, do not
stack them.
- An opencode API shape change. The v2 port breaks this the same way it breaks
`oc_work_start.py`. Detect the unexpected shape and raise one
`warning`-level alert saying the watchdog is blind, rather than silently
reporting "all healthy".
- The watchdog's own health. Record `last_tick_at` and the result (via the
router, or a state file), and show it in the admin Sessions readout. A
watchdog that stopped running must be visible.
**Knobs (North Star #1):**
- `watchdog.enabled`
- `watchdog.interval`
- `watchdog.local_llm_enabled` (default on, still gated by `local_compute`)
- `watchdog.model` (default: `verification.model`)
- `watchdog.read_only_agents`, the list from the read-only case above
The timer ships enabled, since detection plus a warning is the neutral,
low-risk default; the actions follow `progress.mode`.
### 8. The behavior breaker (Lift B): trip a model that stalls sessions
**Gap.** The breaker family covers availability (`circuit_breaker.py`, 5xx) and
malformed output (PR #76: mojibake, empty 200s). Nothing trips on
**well-formed output that does not do the job**. In incident #8,
`z-ai/glm-5.3-flash` produced fluent text that confabulated truncation and
re-read in loops, even with pinch off and a small context. It was only
excluded by hand.
**Trigger.** Verdicts from this detector (the watchdog, and the router judge
once built), attributed to the model(s) that served each flagged session in its
window. Trip when:
- stall verdicts on one model span at least `min_sessions` distinct sessions
(proposed 3) within `window` (proposed 60 min), AND
- that model's stall rate is well above the other models' on the same window's
traffic, so one hard task cannot blame a model.
**Response.** The same shape as the existing breakers: a passive routing skip
with cooldown and backoff, plus a critical alert through the notifier and an
admin Resume. Key it on the model, not the provider, per
`plans/mangled-output-detection.md`'s scope reasoning: here one model misbehaved
while its siblings on the same provider did not. Record every trip with its
evidence (sessions, verdicts, rates) so a human can review it.
**Blocked on identity.** Joining "this session stalled" to "these models served
it" needs #99's conversation headers on each request. As of 2026-09-26 they do
not arrive: `router-link.js` is installed and passes its 29 tests, and the router
reads the headers, but production decisions still carry fingerprint
`session_key`s with no `agent`. Lift B starts by finding out why.
## Fixes that ride along
- **The plugin never reads a real exit code.** Both `router-outcome.js`
(deployed) and its replacement `router-link.js` (#99, line 83) have this. In
1.18,
`tool.execute.after` gives `output = {title, output, metadata}`, and bash's
exit is `output.metadata.exit` (verified on a live session). `looksFailed`
checks `output.exitCode` / `output.exit_code`, which do not exist, so every
verdict has come from the failure-text regex alone. Read `metadata.exit`
first.
- **The installed plugin differs from the repo copy.** The main checkout has an
uncommitted comment block. It is harmless, but the lift should add a check
(install script or a `/health` field reporting the plugin's version string
from a header) so a stale install is visible.
- **`opencode.json` advertised `auto` with `limit.context: 782324`**, the
largest window in the catalog, so opencode only compacted near 782k. Tonight's
contexts reached 682k and were resent every turn. Lowered to 200,000 on
2026-09-26 by owner decision (see the end). It caps what `auto` is asked to
carry, so Phase 0 should also measure how often a real agent turn now
compacts, and whether any task needed more.
## Ground rules for the executor
- **Base on `origin/main`, NOT the local `main`.** The local `main` checkout
(`c87e675`) is 29 commits behind `origin/main` (`f554fc1` as of 2026-09-26)
and 2 ahead with local-only commits. It lacks #99 (conversation identity,
`router-link.js`) and #100 (local-encoder rebuild). `git fetch origin` first,
then branch `feat/no-progress-detection` from `origin/main` in its own
worktree. Read code from that worktree or `git show origin/main:<path>`, never
from the main checkout's files. Do not touch, stash or commit anything in the
main checkout; it has uncommitted work from other efforts.
- The plugin to extend is `deploy/opencode-plugin/router-link.js` (with its
`router-link.test.mjs`). `router-outcome.js` was replaced by #99.
- **Port 8080 is production; never touch it.** Integration checks run on 8081
via `scripts/sandbox.sh`, against a copy of the live DB made with the sqlite
backup API (source opened `mode=ro`).
- Never edit `config/config.yaml` or `config/config.local.yaml`. New knobs get
defaults in `config.yaml` only as part of this lift's own tracked changes.
- **The installed plugin at `~/.config/opencode/plugins/` is live for every
opencode session on this machine,** including the one running the lift.
Develop and test against the repo copy. Installing a new version is an
operator step, stated in the done report, not something the executor does.
- **Commit by explicit path. Never `git add -A` or `git add .`** Incident #8's
item 4 swept an `.omc/` state file into a commit that way.
- **Every worker claim is verified by the orchestrator** against
`git show --stat` and the actual test files before acceptance. Incident #8
had a commit message listing tests that did not exist.
- **Read large files in targeted slices** (grep for the anchor, then read
around it). Incident #8's stuck workers re-read whole files dozens of times.
- ASCII only in new code, comments and UI strings; no middle-dot separators.
No prose paragraphs in the admin UI.
- Lint gate: pinned `uvx ruff@0.16.9`, no NEW findings in touched files
compared with the base commit (same recipe as
`.omo/plans/cockpit-quick-wins.md` runbook step 5). Never `ruff --fix`
outside the lines the item changes.
- `pytest` offline stays green after every commit. Record the baseline on
`origin/main` first. On `c87e675`, one test,
`tests/test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity`,
failed before any change; check whether it still does on `origin/main` rather
than assuming.
- New `route_decisions` columns are registered in
`tests/test_tui_schema_drift.py`; every new knob gets an admin control or a
recorded `DELIBERATELY_NOT_IN_ADMIN` reason (`tests/test_admin_knob_coverage.py`).
## Phases
0. **Calibrate offline.** Replay over opencode session history; produce the
verdict timeline per session and a proposed threshold table. No router or
plugin changes. **Start from `plans/no-progress-detection-prototype.py`,**
an interim watcher tuned on 2026-09-26 against the 15 sessions of incident
#8. On that set it flags every known loop (both item-2 workers, the item-3
first attempt, item 4, the post-completion Atlas stretch, the explore
helper stuck on a deleted file, a helper that printed the spec 12 times)
and none of the healthy sessions. Four findings from tuning it:
- Exact `(tool, args)` fingerprints alone miss reworded commands: 28
differently-worded `cat boulder.json` never matched.
- Collapsing to the bare file over-flags ordinary sliced reading, which is
what the ground rules tell workers to do. Key a `read` on file plus
offset/limit, and a `bash` target on its referenced files plus the
numbers in the command (the slice).
- Slow loops spread across 300+ calls escape a 60-call window. Add a
whole-session rule: one exact call repeated 15+ times with nothing landed
recently.
- Read-only agents need the separate rule (repeats of one target), since
"nothing landed" is always true for them.
Its thresholds (`WINDOW=60`, `DUP_MIN=0.25`, `TOP_MIN=12`, `TOP_MIN_RO=8`,
`CUM_MIN=15`) are a starting point fitted on one night, not a result.
1. **Identity: build on #99, do not rebuild it.** The `metadata.exit` fix in
`router-link.js` (and its `.test.mjs`), plus whatever #99 lacks for
progress roll-up, if anything. Deploying #99 itself (syncing the
production checkout, installing `router-link.js`) is an operator step. It
goes in the done report as a prerequisite for Phase 2 to see real traffic,
not an executor task.
2. **Progress events, the judge and the watchdog, `warn` mode.** `/progress`,
`progress_events`, verdicts, the `/metrics` warning, toast, Sessions readout,
and the knobs. Also the notifier (section 6) with the `desktop` channel and
the admin Send test alert button, and the watchdog timer (section 7) with
its local-LLM second opinion. **Order within the phase: build the watchdog
and the notifier first.** Together with Phase 0's detection code, they
protect unattended runs before the plugin half is even deployed.
3. **`auto_compact` and `auto_recover` stage 1.** Abort, compact with the
injected report, restart; the scoped 429; the `max_recoveries` hard stop and
admin Resume. The incident report builder lands here.
4. **`auto_limit_context` and `auto_recover` stage 2**, using whichever
mechanism Phase 0 verified.
5. **Local-model second opinion inside the router's live judge** (section 4),
only if earlier phases show ambiguous cases. The watchdog already has one
(section 7); this would add it to the per-request path.
The v2 port (`project_opencode_v2_pinned`) changes hook registration. Keep the
plugin's logic in plain functions so the v1 and v2 shells stay thin.
## Owner decisions
Settled 2026-09-26: the response is a mode (`warn`, `auto_compact`,
`auto_limit_context`, `auto_recover` with two stages), and every mode warns.
`auto_recover` stops and compacts the tree's root (Atlas), not only the stuck
worker.
Also settled 2026-09-26:
- **The default mode is `warn`.**
- **`max_recoveries` is 2** per tree per window, then the hard stop.
- **The `auto` context limit is 200,000.** Applied the same day, ahead of this
lift, in the repo `opencode.json` and the global
`~/.config/opencode/opencode.json` (backup at
`opencode.json.bak-20260926-auto-context`). The global file pins every agent
(atlas, sisyphus-junior, prometheus, ...) to `llm-router/auto`. opencode
reads config at startup, so it takes effect for sessions started after a
restart. `auto:batch` stays at 782,324 on purpose: batch is where long
contexts are expected.
- **Alerting ships with `desktop` (`notify-send`) only**, behind a notifier
built for more channels. SMS, RingCentral, PagerDuty and webhooks are later
adapters (section 6), not part of this lift.
No owner decisions remain open. Thresholds come from Phase 0.

227
plans/no-progress-lift-a.md Normal file
View File

@@ -0,0 +1,227 @@
# Lift A: loop watchdog, desktop alerts, and the shared detector
Status: done -- shipped in PR #102; the admin surfacing (component 5) shipped only partly
Date: 2026-09-26
Parent spec: `plans/no-progress-detection.md` (reference only; do NOT read it
end to end). This brief is the contract for Lift A.
Incident: `docs/incidents.md` #8.
## Goal
Catch an opencode agent session that is spinning without progress, and alert
the operator on the desktop, with nobody watching and no cloud model involved.
## Why this slice first, and why it is small
Incident #8's loops were caught only by a human or a Claude session reading
opencode's session store by hand. Lift A replaces that with a local timer.
It is deliberately narrow because the last planning attempt failed on size:
- a 540-line spec
- a planning context that grew to about 600k tokens
- the model reporting "elided" reads that were stored intact, and re-reading
the spec 84 times (68x its length)
Keep this lift's inputs, todos and contexts small.
**Out of scope (Lift B):**
- router-side progress events, the live judge, and `/progress`
- recovery modes (`auto_compact`, `auto_recover`, `auto_limit_context`)
- scoped 429s
- the `router-link.js` changes
- SMS, PagerDuty and other alert channels
## Components
### 1. Shared detector: `src/progress_detect.py` (pure, stdlib only)
Port the rules from `plans/no-progress-detection-prototype.py`. It is tuned on
the incident #8 sessions; keep its behaviour and make it a clean, tested
module. Input: a session's tool calls as `(time, tool, args_json, landed)`.
Output: a verdict with the reason. Signals, over a window of the last
`window` calls:
- `dup`: exact `(tool, args)` duplicate share
- `top`: the most-repeated target. A `read` target is the file plus
offset/limit. A `bash` target is the files it names plus the numbers in the
command (the slice).
- `slow`: one exact call repeated `cum_min`+ times over the whole session
- `coverage`: lines requested from one file in the window divided by its
length. This catches re-reading a file through shifting offsets, which the
other signals miss.
- `landed`: an edit or write with a non-empty diff, or a `git commit` with
exit 0, by the session or any descendant (`parentID`) within the window.
Verdict:
- Normal agents are flagged when nothing landed and any of dup, top, slow or
coverage crosses its threshold.
- Read-only agents (explore, librarian, oracle; a configurable list) never
land, so for them top, slow or coverage alone decides.
- Thresholds start at the prototype's values: `window` 60, `dup_min` 0.25,
`top_min` 12, `top_min_ro` 8, `cum_min` 15, `cover_min` 4.0.
### 2. Calibration backtest: `scripts/progress_backtest.py`
It replays sessions from opencode's HTTP API through `progress_detect` and
prints a per-session verdict with the first flag point. Commit a small
**fixture** of the labelled sessions below, exported as tool-call fingerprints
(tool name, args, landed; never message text), under `tests/fixtures/`. A test
asserts every label. Live opencode is not needed in tests.
MUST FLAG:
| session | what it was |
|---|---|
| `ses_f24cc0258ffepbJPoOzoiLZi9s` | item 2 worker, read navbar.js x61 |
| `ses_f24fcf651ffeOm9O4E14QSZeJ4` | item 2 worker, 1st attempt |
| `ses_f24890640ffewKnIPqjjAfFvOk` | item 3 worker, 1st attempt |
| `ses_f245607eaffeGxPH8F1tvvdvkG` | item 4 worker |
| `ses_f237e7e31ffeOzz4kUKu4CWzMu` | explore helper re-reading a deleted file |
| `ses_f235d755bffeo0uWtD1g2HwQUg` | helper that printed the spec 12x |
| `ses_f2506ac70ffebLUxBsq0aauOn9` | Atlas; must flag in its post-02:07 stretch |
| `ses_f2380dd4cffeVwcZn3Kv3nQN7v` | planner that re-read the spec 68x its length |
MUST NOT FLAG:
| session | what it was |
|---|---|
| `ses_f246e699fffeCHUPTLr5ajBhWC` | item 3 retry |
| `ses_f245f78f8ffeGMn3PgArpOwhFy` | item 3 compact |
| `ses_f243ea08effemMwXdMRE1gpnhK` | item 5 |
| `ses_f242e5e28ffehF97ZxxdbTeylQ` | item 6 |
| `ses_f236bbdaaffezwnvH10X2rgrxc` | explore, sliced reads of dispatcher.py |
| `ses_f236bc7baffeG5LQsr2K1c46xO` | explore |
| `ses_f2e81929fffeaVbWui0X9QjmRQ` | cockpit planner |
The Atlas session also must NOT flag before its 02:07 stretch.
Export the fixture FIRST, in the first todo. opencode sessions can be deleted,
and the fixture is the ground truth.
### 3. Notifier: `src/notifier.py`, `desktop` channel only
- **An alert is an event:**
- `dedup_key` (the session tree root)
- `severity` (`info`, `warning` or `critical`)
- `state` (`trigger`, `escalate` or `resolve`)
- `title`, `summary`, `details`, `source`
The shape maps one-to-one onto PagerDuty Events v2, so later adapters are
thin. Only the `desktop` channel ships.
- **`desktop` runs `notify-send`,** with urgency `critical` for critical
alerts.
- **Fire on transitions only,** with a per-channel rate limit and a quiet
window.
- **A delivery failure is logged, never raised.**
- **Config:** a `notifications.channels` list of
`{name, type, enabled, min_severity}`. An unknown `type` fails config load
(StrictModel). Secrets would come from env vars named in config; desktop has
none.
### 4. Watchdog: `src/watchdog.py` + `deploy/llm-router-watchdog.{service,timer}`
A systemd user timer, shaped like the existing `llm-router-*` timers, every 5
min. Each tick:
1. **Find opencode.** Read `~/.local/share/opencode/rc-servers.json` and probe
each `serverUrl`. If none answers, exit 0 quietly.
2. **Read activity.** List sessions active since the last tick and read their
tool calls.
3. **Detect.** Run `progress_detect` on each active session.
4. **Second opinion, flagged sessions only.** Ask the configured local model
(`watchdog.model`, default `verification.model`, today
`qwen2.5-coder-router:14b` on the local Ollama) "looping without progress?
yes/no + one line", on a digest of tool names, targets and exit codes, never
content. Skip this when `local_compute.enabled` is false or
`watchdog.local_llm_enabled` is false. Report the model's answer alongside
the deterministic verdict; it does not veto it in Lift A.
5. **Alert.** Alert through the notifier directly, and write rows to the
router DB: a `watchdog_ticks` table (time, sessions seen, outcome) and a
`watchdog_verdicts` table. The router need not be up; SQLite WAL handles the
second writer.
Also:
- A lock file prevents overlapping ticks.
- An unexpected opencode API shape raises one `warning` alert ("watchdog is
blind"), never a silent all-clear.
- Admin (North Star #1): a small Watchdog card showing the last tick, recent
verdicts, the channel list with enabled and `min_severity` controls, and a
**Send test alert** button. Every new knob gets a control or a recorded
reason; `tests/test_admin_knob_coverage.py` enforces it.
- Knobs:
- `watchdog.enabled`, `watchdog.local_llm_enabled`, `watchdog.model`,
`watchdog.read_only_agents`
- the detector thresholds from section 1
- `notifications.channels`
The tick interval lives in the timer unit; document it next to the knobs.
## Ground rules
- **Base on `origin/main`.** Branch `feat/no-progress-lift-a` in its own
worktree. Do not touch, stash or commit anything in the main checkout.
- **Port 8080 is production.** Integration checks run on 8081
(`scripts/sandbox.sh`), against a sqlite-backup copy of the live DB.
- **Never edit** `config/config.yaml`'s existing values or
`config/config.local.yaml`. New keys get defaults in `config.yaml`.
- **Do not install or enable** the timer on this machine; that is an operator
step in the done report.
- **Commit by explicit path; never `git add -A`.** The orchestrator checks each
worker's `git show --stat` against its claims and the plan's paths before
accepting it.
- **Read in targeted slices.** Grep for the anchor, read around it. Do NOT read
`plans/no-progress-detection.md` whole; this brief is the contract.
- **If a worker's context passes about 150k tokens, stop it** and start a fresh
worker with a smaller prompt, rather than letting it grind.
- **ASCII only; no middle dots.** No prose paragraphs in the admin UI.
- **Lint:** pinned `uvx ruff@0.16.9`, no NEW findings in touched files against
the base, never `ruff --fix` outside changed lines.
- **pytest:** record the `origin/main` baseline first; it stays green apart
from the baseline.
## Done means
- Every must-flag and must-not-flag label passes in tests, and the live
backtest prints the same verdicts.
- One real `notify-send` from the admin test button on 8081.
- The watchdog runs once by hand (`python -m watchdog --once`) against live
opencode and prints its verdicts.
- The done report lists each commit, the tests, and the operator steps
(install and enable the timer).
## Addendum, 2026-09-26: surfacing and one-click block (North Star #4)
The owner set a new priority: surface waste in the admin portal, and make
stopping it one click (`CLAUDE.md` North Star #4). This lift therefore also
ships:
### 5. Loops panel and one-click model block (admin portal)
- **A Loops panel** on the Home page (or the Controls page, whichever the
cockpit layout fits), reading `watchdog_verdicts`: each flagged session with
agent, verdict, turns and $ since the last landed change, the top repeated
target, and when it was flagged.
- **A per-model rollup:** stall verdicts per model over the last hour, next to
each model's share of traffic. It needs session-to-model attribution; see
item 0.
- **A Block model button** beside each model in the rollup. It writes
`admin_model_overrides` with `availability = 'blocked'` and a reason (default
"stalled N sessions", editable), through the existing
`POST /admin/api/models/{model_id}/{provider}/availability` path. Verify that
routing excludes `blocked` exactly like `deprecated` (`load_candidates`), and
add a test for it.
- **A Blocked models list** with reason, time and a one-click Unblock
(`DELETE .../availability`).
- **Desktop alerts link to the panel** (`http://127.0.0.1:8080/admin/` plus an
anchor).
### 0. First todo: why do the identity headers not arrive?
`router-link.js` (#99) is installed in `~/.config/opencode/plugins/`, is
identical to the repo copy, and passes its 29 tests. The router reads
`X-Router-Conversation` (`src/dispatcher.py`, `resolve_identity(request.headers,
...)`). Yet production `route_decisions` after the 04:02 opencode restart still
show fingerprint `session_key`s and `agent = NULL`. Find out why, and fix it if
it is a small change.
The per-model rollup depends on it. If it cannot be fixed in this lift, the
rollup ships disabled with a visible "needs conversation identity" note, and
everything else ships.

View File

@@ -790,7 +790,8 @@ def _get_db_path(cfg: Any) -> str:
return "router.db" return "router.db"
if __name__ == "__main__": def main(argv: list[str] | None = None) -> int:
"""Run one watchdog tick; return the process exit code (0 when it ran)."""
ap = argparse.ArgumentParser( ap = argparse.ArgumentParser(
description="Watchdog orchestrator for opencode loop detection", description="Watchdog orchestrator for opencode loop detection",
) )
@@ -800,7 +801,7 @@ if __name__ == "__main__":
ap.add_argument( ap.add_argument(
"--config", default="config/config.yaml", help="config.yaml path", "--config", default="config/config.yaml", help="config.yaml path",
) )
args = ap.parse_args() args = ap.parse_args(argv)
logging.basicConfig( logging.basicConfig(
level=logging.INFO, level=logging.INFO,
@@ -841,6 +842,14 @@ if __name__ == "__main__":
last_fired_cb = _make_last_fired(conn) last_fired_cb = _make_last_fired(conn)
notifier = Notifier(notifier_cfg, last_fired=last_fired_cb) notifier = Notifier(notifier_cfg, last_fired=last_fired_cb)
rc = tick(cfg, conn, notifier=notifier) alerts_fired = tick(cfg, conn, notifier=notifier)
conn.close() conn.close()
sys.exit(rc or 0) logger.info("watchdog tick done: %d alert(s) fired", alerts_fired)
# A completed tick exits 0 however many alerts it fired. Returning the
# alert count made systemd mark the oneshot unit FAILED on every real
# alert (and wrap past 255), so a working alert looked like a crash.
return 0
if __name__ == "__main__":
sys.exit(main())

View File

@@ -1180,3 +1180,21 @@ class TestDescendantLandedScope:
) )
finally: finally:
_close_conn(conn, path) _close_conn(conn, path)
class TestMainExitCode:
"""`python -m watchdog --once` runs as a systemd oneshot: a tick that ran
and fired alerts is a success, so the exit code must not carry the alert
count (systemd would mark the unit FAILED on every real alert)."""
def test_main_returns_zero_when_alerts_fire(self, tmp_path):
import watchdog
cfg = MagicMock()
cfg.database.path = str(tmp_path / "router.db")
cfg.notifications.channels = []
with patch("watchdog.load_config", return_value=cfg), \
patch("watchdog.tick", return_value=3) as fake_tick:
rc = watchdog.main(["--once", "--config", "unused.yaml"])
assert fake_tick.called
assert rc == 0