diff --git a/CLAUDE.md b/CLAUDE.md index b2bec0a..2d7ce33 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -8,7 +8,7 @@ next steps, and is the one to trust on what is currently true. ## NORTH STAR GUIDELINES -Three rules that outrank local cleverness. Each exists because it was broken +Four rules that outrank local cleverness. Each exists because it was broken first and the breakage was expensive to find. When a change conflicts with one of these, the change is wrong — not the rule. @@ -87,6 +87,36 @@ before measuring anything, and open it read-only: `sqlite3.connect("file:...?mode=ro", uri=True)` — the `sqlite3` CLI here does not accept `-uri`. +### 4. Waste is surfaced in the admin portal, and stopping it is one click + +When the router can see money being wasted, it shows the operator where they +already look, with the evidence and the lever to stop it side by side. A +detector that only writes a log line, or a fix that needs a config edit, a +restart or a long table scan, has not met this rule. + +"Waste" here means **spend with no concrete change landing**: an agent session +looping, re-reading, or retrying the same failure, and a model that keeps +producing such sessions. It does **not** mean steady spend. A healthy agent run +can burn for hours, and a spend-rate alarm cannot tell the two apart. + +In practice: +- **Stalled sessions are visible** in the portal with their evidence: turns and + $ since the last landed change, and the top repeated target. They also reach + the operator when nobody is watching (desktop alert). +- **A model that keeps producing them is visible** as a per-model rollup, with + **Block model** beside the evidence. The block records its reason, shows in a + Blocked list, and is one click to undo. +- **Automatic responses come after the visible one,** never instead of it, and + every automatic action shows up in the same place. + +**Why:** incident #8 (`docs/incidents.md`). Agent sessions looped for hours on +2026-09-25/26: one worker read the same file 61 times, a planner re-read a spec +68 times its length, and a model confabulated truncation that was not there. +Every existing check stayed green. It was caught only by a human, or a Claude +session, reading opencode's session store by hand. Pulling the model took a trip +through a dropdown whose vocabulary is catalog `deprecated`. The detection +existed nowhere, and the lever existed only for someone who already knew. + ## What this is A router that uses a local model (served via Ollama) to classify incoming @@ -319,7 +349,12 @@ rather than from months of history. - `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md). - `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below. - `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted. -- `tests/` — 1155 tests across 40+ files, offline, verified on Python 3.10 and 3.14. [README](README.md). +- `progress_detect.py` — loop-detection signals over a window of per-session probe calls: duplicate-bulk (`dup_min`), top-similarity (`top_min`/`top_min_ro`), slow-progress (`cum_min`) and coverage (`cover_min`) heuristics, gated on `min_calls`. See [watchdog](docs/watchdog.md). +- `watchdog.py` — the per-session watchdog loop: every ~5 minutes it judges each session with a tool call since the last tick, on its full history, and writes a quiet `no_opencode` tick when opencode is not running. See [watchdog](docs/watchdog.md). +- `notifier.py` — alert fan-out: desktop/`notify-send` plus per-channel `min_severity` and a rate limit. See [watchdog](docs/watchdog.md). +- `watchdog_store.py` — `watchdog_ticks`, `watchdog_verdicts`, `watchdog_alerts`, `watchdog_channel_settings` (four tables + indexes). See [watchdog](docs/watchdog.md). +- `router-link.js` — the opencode plugin; exported as a factory with `parentCache` as a property, because opencode 1.18.x rejects the whole plugin when any export is not a function. See [watchdog](docs/watchdog.md). +- `tests/` — 2414 tests across 108 files, offline, verified on Python 3.10 and 3.14. [README](README.md). ## #45 — OpenRouter is an opt-in allowlist provider @@ -1261,10 +1296,14 @@ be on. ### When the router goes unreachable, start at docs/incidents.md -Four incidents so far, all sharing one shape: a change that looked local to the -router silently degraded the agent depending on it, and none announced itself as -a router problem. **`docs/incidents.md` carries the full write-ups plus a -symptom -> one-line-check table**; read it rather than re-deriving a diagnosis. +Eight incidents so far, nearly all sharing one shape: a change that looked local +to the router silently degraded the agent depending on it, and none announced +itself as a router problem. **`docs/incidents.md` carries the full write-ups plus +a symptom -> one-line-check table**; read it rather than re-deriving a diagnosis. +#8 is the exception worth knowing before an unattended agent run: the router +worked perfectly while agent workers looped for hours with no progress, and no +check noticed, because every check watched spend or availability rather than +whether changes landed (`plans/no-progress-detection.md`). Two conventions from those incidents that bind every session, and so stay here: @@ -1283,6 +1322,14 @@ Recovery for an unreachable-but-`active` service is `systemctl --user restart llm-router.service` -- a hung process was never in a tracked stop job, so this issues a fresh cycle systemd does enforce a timeout on. +The watchdog runs as its own systemd **timer** (`llm-router-watchdog.timer`), +separate from the dispatcher, so a stuck router cannot silence the thing meant +to notice it is stuck. Install and enable it with `systemctl --user enable +--now llm-router-watchdog.timer` (the timer's unit file ships pointing at a +placeholder home path and must be `sed`-repointed to the real one first), or +run it once by hand with the `--once` flag. See [watchdog](docs/watchdog.md) +for the signals it fires on, the alert lifecycle, and the known limits. + ## Pointing a coding agent at it The `/v1` endpoints are OpenAI-compatible, so any normal client works — diff --git a/docs/admin-portal.md b/docs/admin-portal.md index 2a4f632..b83d691 100644 --- a/docs/admin-portal.md +++ b/docs/admin-portal.md @@ -35,11 +35,30 @@ for the full model. Admin dashboard +**Loops** — the dashboard carries a Loops card (`admin/frontend/index.html#loops`) +showing the watchdog's open alerts, one row per alert. It pulls from +`GET /admin/api/watchdog/loops` (`watchdog_alerts` rows where `resolved_at` is +null) via `loadLoops()`. Each row shows exactly three things: a severity badge +(`critical` / `warning` / else `info`), the raw dedup key +(`opencode-loop:`), and `opened_at`. That is all the panel +shows today — it does **not** show the agent, model, top repeated target, calls, +or $ since the last landed change; those live in `watchdog_verdicts` and have +not shipped to this panel yet. + **Models** — `GET /admin/models` is the model-availability table. Each row shows the serving class, tier, status, and an override dropdown (`active` / -`deprecated` / `stale`) that writes through to the routing hard filters. By -default the table hides deprecated and stale rows; enable the **Show deprecated -/ stale** toggle to include them. +`blocked` / `deprecated` / `stale`) that writes through to the routing hard +filters. By default the table hides deprecated and stale rows; enable the +**Show deprecated / stale** toggle to include them. + +`blocked` is a fourth value in that same per-model override dropdown, distinct +from the catalog meaning of `deprecated`. It is an operator stop: a model an +operator has switched off, not one the catalog flagged. Routing and pinned +requests both exclude a blocked model (`_admin_excluded_models` in +`dispatcher.py` and `metrics.py`). Undo is setting the dropdown back to +`active` (clearing the override via `DELETE` on the availability path), not a +separate unlock — there is no per-model Block button beside evidence and no +Blocked list with reasons; those belong to the "not built yet" item below.
Admin models page @@ -126,8 +145,8 @@ land in `config/config.local.yaml`; see [config-local-overlay](config-local-overlay.md)). Changes are marked as dirty and written only on save. -It carries two dedicated cards, neither a row in the generic runtime-knob -list, because both have cross-field structure a flat scalar/boolean input +It carries three dedicated cards, none a row in the generic runtime-knob +list, because each has cross-field structure a flat scalar/boolean input can't safely represent: - **Local Compute** — "gaming mode". See [Gaming mode](#gaming-mode) below. @@ -145,6 +164,14 @@ can't safely represent: anything reaches `config.local.yaml`, and mode plus its companion block are written as one atomic change so an in-between invalid state is never even written transiently. +- **Watchdog** — the watchdog's live state (`controls.html#watchdog-card`, + `loadWatchdog()`): the last tick time (and a `no tick` badge before the + first one), how many sessions the last tick saw, how many verdicts it + flagged, and how many alerts are currently open. Below those, a set of + **notification channel** toggles (each channel with an enable/disable and a + minimum severity, persisted via `POST /admin/api/watchdog/channels`), a + **Refresh** button, and a **Send test alert** button (`POST + /admin/api/watchdog/test-alert`).
Admin controls page @@ -444,7 +471,15 @@ intermediate invalid state could land on disk. The GET response includes **Model availability overrides** `POST /admin/api/models/{model_id:path}/{provider}/availability` marks a model -as `active`, `deprecated`, or `stale`. `DELETE` on the same path removes the -override. The `{model_id:path}` converter accepts model ids that contain -slashes, so providers like OpenRouter with slash-bearing ids are handled the -same way as plain ids. Deprecation feeds into the routing hard filters. +as `active`, `blocked`, `deprecated`, or `stale`. `DELETE` on the same path +removes the override. The `{model_id:path}` converter accepts model ids that +contain slashes, so providers like OpenRouter with slash-bearing ids are +handled the same way as plain ids. Deprecation feeds into the routing hard +filters. + +**Not built yet.** The Loops panel does not yet show the evidence columns an +operator would act on — agent, model, top repeated target, calls, and $ since +the last landed change — and there is no per-model stall rollup, no **Block +model** button beside the evidence, and no Blocked list with reasons. Those are +the remaining surface of [North Star #4](../CLAUDE.md) in CLAUDE.md, whose first +half (conversation identity and the watchdog) is shipped. diff --git a/docs/api.md b/docs/api.md index 6309f58..207f6b8 100644 --- a/docs/api.md +++ b/docs/api.md @@ -173,6 +173,12 @@ These ids are what make `/outcome` attribution exact (below): an `X-Router-Conversation` on the request sets the `session_key` row that a later report can address by `conversation_id` without guessing `source` or ambiguity. +**Plugin loader contract:** the opencode plugin that stamps these headers +must export only functions — the opencode 1.18.x plugin loader rejects a +plugin outright if any export is not a function (so the plugin's +`parentCache` lives as a property on the factory rather than a named export). +See [clients.md](clients.md) for the full contract. + **Capability 422s**: When no model survives the hard filters, the 422 names the active constraints. That now includes "vision-capable model" or "json-mode-capable model" when the request carried images or a JSON-mode @@ -225,6 +231,65 @@ but a `true` report counts too. It's also the only quality signal that survives streaming, since a retry can't reach a response whose bytes are already gone, while a report arrives afterward and works either way. +## Admin watchdog and availability API + +The admin portal's watchdog and model-availability operations sit under +`/admin` (loopback-only, like the rest of the admin API). The watchdog +endpoints back the dashboard's **Loops** card and the **Watchdog** control +card; full behaviour is described in [watchdog.md](watchdog.md) and +[admin-portal.md](admin-portal.md). + +| Method | Path | Description | +|---|---|---| +| `GET` | `/admin/api/watchdog/status` | Latest watchdog tick plus open-alert count | +| `GET` | `/admin/api/watchdog/loops` | Open alert rows (one per looping session) | +| `GET` | `/admin/api/watchdog/channels` | Notification-channel settings | +| `POST` | `/admin/api/watchdog/channels` | Create or update a channel's enabled/min-severity | +| `POST` | `/admin/api/watchdog/test-alert` | Deliver a test alert through the enabled channels | + +`GET /admin/api/watchdog/status` returns a single object: + +- `last_tick` — the most recent `watchdog_ticks` row, or `null` before the + first tick +- `verdicts` — `{total, flagged}` for that tick's `watchdog_verdicts`, or + `null` when there is no tick yet +- `open_alerts` — count of `watchdog_alerts` rows with `resolved_at` null + +`GET /admin/api/watchdog/loops` returns an array of `watchdog_alerts` rows +where `resolved_at` is null, ordered by `opened_at` descending. Each element +is the full alert row. + +`GET /admin/api/watchdog/channels` returns an array of +`watchdog_channel_settings` rows ordered by `channel_name`. + +`POST /admin/api/watchdog/channels` upserts one channel. Body: + +- `channel_name` (required) — the channel to create or update; absence is a + 422 +- `enabled` (optional boolean) — when omitted the existing value is kept +- `min_severity` (optional: `info`, `warning`, or `critical`) — any other + value is a 422 + +Returns `{ok: true}`. + +`POST /admin/api/watchdog/test-alert` delivers a test alert through the +currently enabled channels. Body: + +- `severity` (optional, default `warning`; `info`, `warning`, or `critical`) +- `channel_name` (optional) — restrict to a single channel + +Returns `{sent: , severity: }`, or `{sent: 0, message: "no +enabled channels"}` when nothing is enabled. An invalid `severity` is a 422. + +**Model availability** — `POST +/admin/api/models/{model_id:path}/{provider}/availability` upserts an admin +override for a model's availability. `availability` must be one of `active`, +`blocked`, `deprecated`, or `stale` (anything else is a 422); `blocked` is an +operator stop that routing and pinned requests exclude +(`_admin_excluded_models`), distinct from the catalog meaning of `deprecated`. +`DELETE` on the same path removes the override. The `{model_id:path}` +converter accepts slash-bearing model ids (e.g. OpenRouter). + ## Pinning and auto behavior for local models `model: "auto"` will route eligible tasks to `qwen2.5-coder-router:14b` when the diff --git a/docs/clients.md b/docs/clients.md index dff01d9..7942d18 100644 --- a/docs/clients.md +++ b/docs/clients.md @@ -37,3 +37,11 @@ This resolved a real incident where reading a dependency's source 24 times outweighed writing to the project dir 22 times: the write-like 3× weighting puts the project directory that actually received the edits ahead of a dependency's source that was merely read. + +**Plugin loader contract:** opencode 1.18.x (the observed version) rejects the +entire plugin if *any* export is not a function. The `parentCache` map on +`router-link.js` is therefore set as a property on the `RouterLink` factory +function rather than as a named export (L221-224 of +`deploy/opencode-plugin/router-link.js`). Confirm the plugin loaded by +searching the log for `failed to load plugin`: `grep 'failed to load plugin' +~/.local/share/opencode/log/opencode.log`. diff --git a/docs/data-model.md b/docs/data-model.md index 5634090..438f149 100644 --- a/docs/data-model.md +++ b/docs/data-model.md @@ -31,7 +31,7 @@ Three data tables plus one observability table, `PRAGMA foreign_keys = ON`: | `access_level` | TEXT | `public` \| `preview` \| `canary` | | `pricing_tbd` | INTEGER | | | `deprecated` | INTEGER | | -| `availability` | TEXT | `active` \| `deprecated` \| `stale` | +| `availability` | TEXT | `active` \| `deprecated` \| `stale` \| `blocked` — the `blocked` value is an operator stop set via the admin override (see [admin-portal.md](admin-portal.md)); it is excluded from routing and pinned requests | | `last_updated` | TEXT | ISO8601 | **Serving class:** Neuralwatt ships ~6 base models as 19 catalog rows. The id @@ -326,3 +326,71 @@ prompts, and answers are deliberately excluded. A test enforces that the write path does not store prompt or answer text. The write is gated by `logging.log_route_decisions` and is **best-effort**: a failed write is logged at warning and swallowed so monitoring cannot slow or fail a request. + +## Watchdog Schema (SQLite) + +Four tables written by `watchdog.py` via `src/watchdog_store.py` (the code-side +inline-create mirrors this schema for live databases). See +[watchdog.md](watchdog.md) for how the loop reads them. + +### `watchdog_ticks` — one row per detection run + +| Column | Type | Notes | +|---|---|---| +| `id` | INTEGER | Autoincrement | +| `ticked_at` | TEXT | ISO8601 run time | +| `sessions_seen` | INTEGER | Count of sessions examined this run, default 0 | +| `outcome` | TEXT | | + +### `watchdog_verdicts` — one row per session per tick + +| Column | Type | Notes | +|---|---|---| +| `id` | INTEGER | Autoincrement | +| `tick_id` | INTEGER | FK → `watchdog_ticks.id`, NOT NULL | +| `session_id` | TEXT | NOT NULL | +| `session_root` | TEXT | | +| `agent` | TEXT | | +| `model_id` | TEXT | | +| `provider` | TEXT | | +| `flagged` | INTEGER | 0/1, default 0 | +| `dup` | REAL | Duplicate-fraction signal | +| `top` | INTEGER | | +| `top_what` | TEXT | | +| `landed` | INTEGER | | +| `slow` | INTEGER | | +| `coverage` | REAL | | +| `calls_since_landed` | INTEGER | | +| `cost_since_landed_usd` | REAL | | +| `llm_second_opinion` | TEXT | | +| `created_at` | TEXT | NOT NULL | + +Four indexes (`config/schema.sql` L482-485): + +| Index | On | +|---|---| +| `idx_watchdog_verdicts_model_id` | `watchdog_verdicts(model_id)` | +| `idx_watchdog_verdicts_created_at` | `watchdog_verdicts(created_at)` | +| `idx_watchdog_verdicts_flagged` | `watchdog_verdicts(flagged)` | +| `idx_watchdog_verdicts_session_root` | `watchdog_verdicts(session_root)` | + +### `watchdog_alerts` — one row per open alert + +| Column | Type | Notes | +|---|---|---| +| `dedup_key` | TEXT | PRIMARY KEY; e.g. `opencode-loop:` | +| `state` | TEXT | NOT NULL | +| `severity` | TEXT | NOT NULL | +| `flagged_ticks` | INTEGER | Default 0 | +| `opened_at` | TEXT | NOT NULL | +| `last_fired_at` | TEXT | NOT NULL | +| `resolved_at` | TEXT | | + +### `watchdog_channel_settings` — per-channel notification config + +| Column | Type | Notes | +|---|---|---| +| `id` | INTEGER | Autoincrement | +| `channel_name` | TEXT | NOT NULL, UNIQUE | +| `enabled` | INTEGER | Default 1 | +| `min_severity` | TEXT | Default `warning` | diff --git a/docs/incidents.md b/docs/incidents.md index 4ceea34..a4a0463 100644 --- a/docs/incidents.md +++ b/docs/incidents.md @@ -1,14 +1,18 @@ # Incidents: how this router has broken, and how to tell which one it is -Seven times now, a change that looked local to the router has silently -degraded either the agent depending on it or the operator trying to see it -clearly. They share a shape worth naming: **none of them announce themselves -as router problems.** Five of the seven presented as an opaque client-side -error — a connection refused, an "Unprocessable Content", an "internal server -error" — and diagnosing each meant knowing which log or table to look in. The -other two didn't error toward the client at all: #5 destroyed data outright, -and #6 just went quiet in one corner of the admin UI while `/health` stayed -green the whole time. +Eight times now, something around the router has silently degraded either the +agent depending on it or the operator trying to see it clearly. They share a +shape worth naming: **none of them announce themselves as router problems.** +Five of the eight presented as an opaque client-side error — a connection +refused, an "Unprocessable Content", an "internal server error" — and +diagnosing each meant knowing which log or table to look in. The other three +didn't error toward the client at all: #5 destroyed data outright, #6 just went +quiet in one corner of the admin UI while `/health` stayed green the whole time, +and #8 spent money for hours while every check stayed green. + +**#8 is the only one so far that cost money rather than capability**, and the +router behaved as configured throughout. Read it before leaving +an agent run unattended. **#5 is the only one so far that destroyed data**, and the only one where recovery depended on luck rather than design. Read it before running any @@ -21,6 +25,15 @@ or removing any config key. This page exists so the next one takes minutes rather than hours. Start with the symptom table, then read only the relevant section. +**How to write an entry (from #8 on).** +- Separate **Evidence** (measured, with where it was measured) from **Theory** + (inferred). Anything not established by a measurement or a reproduction goes + under Theory. It states what supports it and what would confirm or refute + it, and stays there until that test is done. +- Include a **Human response** timeline: what the operator noticed, decided and + did. Mark actions an assistant took, and at whose direction. +- Entries #1 to #7 predate this convention. + `opencode.json` points opencode's own model traffic at `http://127.0.0.1:8080/v1`, so on this machine the router *is* the coding agent's inference supply. Breaking it breaks the thing you would use to fix it. That is why these keep happening, @@ -38,6 +51,8 @@ and why the diagnostics below are worth having to hand. | 500s + `no such table`, venv/.env gone | #5 `git clean -fdx` | `ls -la router.db .env .venv` — a 0-byte db and a missing `.venv` is conclusive | | A newly-shipped admin knob just isn't there, `/health` is fine | #6 backend running stale code | `systemctl --user status llm-router` uptime vs. `git log -1 --format=%cd` on the commit that added the knob | | `activating (auto-restart)`, `pydantic_core...ValidationError: ...extra_forbidden` in the journal | #7 renamed key vs. un-migrated overlay | `journalctl --user -u llm-router --since '5 min ago' \| grep extra_forbidden` then `grep -n config/config.local.yaml` | +| Spend is high after an agent run, nothing errored, no warning fired | #8 agent looping with no progress | `sqlite3 router.db "select session_key, strftime('%H',observed_at,'localtime') h, round(sum(prompt_tokens)/1e6) mtok, round(sum(cost_usd),2) usd from energy_observations where observed_at > datetime('now','-12 hours') group by 1,2 order by usd desc limit 8;"` then check whether that run's commits actually landed | +| opencode plugin does nothing | #8 plugin loader rejects non-function export | `grep 'failed to load plugin' ~/.local/share/opencode/log/opencode.log` | That last row is not an incident yet — it is the silent-staleness failure described in the "Run as a service" section of `CLAUDE.md`. An unpolled catalog @@ -347,6 +362,253 @@ file is correct. --- +## #8 — Agent sessions looped for hours with no progress, and nothing noticed (2026-09-25) + +**Symptom.** Nothing errored. The router stayed healthy, `/metrics` showed zero +warnings, and the quota alarm read `kind: none`. The operator noticed early on +2026-09-26 that an opencode run had "burned a bunch of tokens over the hours". + +### Evidence + +Measured read-only against the live `router.db` and opencode's own session +store. The window is 2026-09-25 19:00 to 2026-09-26 02:00 local unless noted. + +**Spend.** + +| | | +|---|---| +| prompt tokens | 323M, 95% cached | +| billed | about $7.81, roughly 3x the busiest of the previous ten days ($2.66 on 09-20) | +| largest share | OpenRouter `z-ai/glm-5.3-flash`: 1,368 calls averaging 203k prompt tokens, $6.16 | +| largest contexts | NeuralWatt `glm-5.3`: 47 calls averaging 673k, max 682,679, the only model whose window still fit | +| 120k+ prompt calls | 921 of them, $6.57 of the total | +| routing | every decision `classification_source = classifier`, default profile; nothing pinned | + +**What the agents did.** An Atlas run (`.omo/plans/cockpit-quick-wins.md`) +delegated each item to a sub-agent worker: + +| worker session | tool calls | exact-duplicate calls | worst repeat | landed | +|---|---|---|---|---| +| item 2, 2nd attempt | 423 | 30% | `read navbar.js` x61 | nothing | +| item 3, 1st attempt | 132 | 23% | `read test_admin_js_units.py` x16 | nothing | +| item 4 | 134 | 19% | `read` of the plan x8 | code only | +| item 2, 1st attempt | 307 | 17% | `read navbar.js` x21 | nothing | +| healthy workers | 30 to 96 | 0 to 6% | 1 to 3 | yes | + +- Item 4's commit message claimed four test files it never wrote, and it + committed with `git add -A`. +- Atlas later said its own turns were "just re-reading the plan and state over + and over instead of doing work". +- A planning session described its reads as "heavily elided". The stored tool + outputs in opencode were intact. + +**Pinch was rewriting agent context.** +- `pinch.budget_tokens: 50000` pruned 1,915 of 1,992 decisions over nine hours, + keeping 48% of the prompt on average. +- The planning session's 527k-token prompts went out at about 125k. +- `src/context_prune.py` replaces older tool results with + `[read: result omitted]`, or keeps a 1,500-character head and tail around a + `[N chars trimmed...]` marker. Results inside the protected window get the + same cut above `protected_max_chars` (20,000). +- Pinch's config had not changed since 2026-09-06. It had pruned 90 to 100% of + requests daily for two weeks without incident. The input changed: + +| | 09-10 to 09-23 | 09-25 | 09-26 | +|---|---|---|---| +| avg prompt into pinch | 100k to 150k | 236k | 545k | +| p90 prompt | 170k to 260k | 528k | 1.31M | + +- `opencode.json` advertised `auto` with `limit.context: 782324`, so opencode did + not compact until near that size. + +**Routing drifted with size.** +- `deepseek/deepseek-v4-flash` has 384k effective context; + `z-ai/glm-5.3-flash` has 655k. +- Their proficiency was tied (`diff_checking` 0.933 vs 0.935) and unchanged + since 09-10. +- Deepseek went from 290 picks on 09-23 to 0 on 09-25, while appearing 851 + times as runner-up. + +**A clean reproduction, 2026-09-26 04:04 to 04:14.** A fresh planning +session, on the Lift A brief (188 lines): +- **Conditions:** context about 88k tokens, pinch off, 0 of 29 tool parts + marked `compacted` by opencode, every stored read intact. +- **What it did:** it still reported its reads "elided with `[...]`", wrote + fake "Earlier tool responses received ... (counts, not proof of task + completion)" blocks into its own replies (assistant text parts, not tool + output), invented session IDs, and called the brief "393 lines". +- **Where the phrase is not:** in opencode, oh-my-openagent (JS bundle and + native binary), or this repo. +- **Model:** `z-ai/glm-5.3-flash` served 27 of 34 chat decisions in the 15 + minutes before it was stopped. + +**Found along the way.** +- `deploy/opencode-plugin/router-outcome.js` has never read a real exit code. + On opencode 1.18, bash's exit is `output.metadata.exit`; the plugin reads + `output.exitCode`. +- PR #99 (conversation identity) was merged to `origin/main` but not running. + Production ran from a local checkout 29 commits behind. +- **Plugin loader rejected `router-link.js`.** The plugin exported `parentCache` + as a module-level const. Opencode 1.18's plugin loader discards any plugin + with a non-function export — the entire file is silently dropped. The fix was + to make `parentCache` a property on the factory function instead (L224, + `router-link.js` in PR #102, a24de1b). Confirmed by grepping + `~/.local/share/opencode/log/opencode.log` for `'failed to load plugin'`. +- How much of the $7.81 was waste cannot be computed. The run did land six + reviewed commits, and the deployed router cannot tell conversations apart. + +### Theory + +These are inferences, not measurements. Each lists what would confirm or +refute it. + +1. **Pinch cutting oversized sessions drove, or worsened, the re-read loops.** + Weakened by the clean reproduction, which looped with pinch off. As + sessions grew past about 250k, the fixed 50k budget removed most of each + agent's working memory, recent reads included. The agent saw its reads cut, + re-read them, grew the context, and got cut harder. + - Supported by: the agents' own descriptions ("elided", "drowning in + tool-output truncation"), intact stored outputs, the prune rates, and the + timing of the size jump. + - Not tested: no run has been observed with pinch off. + - Confirm: with pinch off and `auto` at 200k, the duplicate-read signature + should not recur on comparable runs. The interim watcher + (`plans/no-progress-detection-prototype.py`) and Phase 0 data can show it. +2. **Sessions got that large because long unattended orchestration ran without + compaction.** Atlas made 668 tool calls, one worker 423, under a 782k limit, + and the re-read loop in theory 1 fed the growth. The limit is a fact; that it + explains the growth is theory. +3. **Leading theory: `z-ai/glm-5.3-flash`, or its OpenRouter path, + confabulates truncation and then loops.** The clean reproduction above + removed pinch, context size and input corruption, and the behaviour stayed. + Theories 1 and 2 remain plausible aggravators, not the cause. + - Confirm: with glm excluded, comparable agent runs show no elision claims + and no duplicate-read signature. A direct A/B on one prompt would settle + model versus provider path. + +### Why nothing caught it + +Each existing guard answers a different question: +- `circuit_breaker.py` trips on provider 5xx. There were none. +- `runway_low_warning` fires when balance / burn drops under 6 hours. Burn is + averaged over a 24 h balance window: OpenRouter read $24.73 at $0.25/h, which + is 97 hours of runway. It asks "will I run out", not "is this being wasted". +- Spend rate was not abnormal. The worst hour ($2.42) and worst 3-hour window + ($4.13) sit inside the prior 30 days' range (hourly p90 $1.46, max $4.83; + worst 3 h $9.65). +- The router cannot see progress, and the deployed build could not tell + conversations apart. + +### Recurrences + +- **02:21 to about 02:56: Atlas looped after its plan was complete.** Every + checkbox was ticked, but `.omo/boulder.json` still carried `status: completed` + and `pr_url: .../pulls/77` from the previous plan (evidence: the file). The + continuation hook injected "continue" turns, 3 of 3 user turns in the window + (evidence: the session). Atlas re-derived its state each time: 54 turns, + about 19M cached-read tokens, `boulder.json` read 28 times, checkboxes grepped + 16 times. Cause: stale state, not pruning. The branch was never pushed, so PR + 77 was not touched. +- **About 03:05 to 03:17: a read-only `explore` subagent** re-read + `router-outcome.js` 20 times (19% duplicate calls across 322). The production + sync had just deleted that file. Read-only agents never land changes, so the + detector needs a separate signal for them. + +### Human response + +Times are local, 2026-09-26. +- **About 02:10.** Noticed the spend and asked for a mechanism in 6krrt to catch + it. Rejected a spend-rate alarm, since steady agent usage is legitimate, and + framed the target as "are concrete changes landing". +- **About 02:20 to 02:40.** Decided the design for `plans/no-progress-detection.md`: + - response modes (`warn`, `auto_compact`, `auto_limit_context`, a two-stage + `auto_recover`), with every mode warning + - `warn` as the default + - at most 2 recoveries per session tree + - `notify-send` now, with pluggable SMS, RingCentral and PagerDuty channels + later + - a local-model watchdog timer, "so I don't have to rely on Claude" + - `auto`'s context limit lowered to 200,000 in the repo and global + `opencode.json` +- **About 02:50.** Caught the post-completion Atlas loop by watching. At the + operator's direction, Claude aborted the session at 02:56 and cleared the + stale `pr_url` in `.omo/boulder.json`. +- **About 03:10.** Merged PR #101. A first `git pull` aborted on local changes, + and the restart on that line reloaded the old code. Then ran the prepared sync + (back up, drop the changes already upstream, `reset --keep origin/main`, + re-apply the rest) and restarted onto `827c408` at 03:15. +- **About 03:30.** Asked Claude to watch opencode for loops until the detector + ships (the interim watcher). +- **About 03:45.** Reported the planning session's "truncation hell". At the + operator's direction ("hot fix that"), Claude set `pinch.enabled` false at + runtime and persisted it to `config/config.local.yaml` (backup + `config.local.yaml.bak-20260926-pinch`). The operator restarted the service + at 03:52 and chose not to open a PR for the switch flip. +- **About 04:02.** Installed `router-link.js` in place of `router-outcome.js` + and restarted opencode (200k `auto` limit now active). Kicked off a fresh, + small Lift A planning session (`plans/no-progress-lift-a.md`). +- **About 04:14.** After the fresh session reproduced the confabulation, + excluded `z-ai/glm-5.3-flash` from routing. It had no picks after 04:13:57, + and traffic moved to deepseek and mimo. At the operator's direction, Claude + aborted the confabulating session. +- **~15:18.** Watchdog timer enabled (`systemctl --user enable --now + llm-router-watchdog.timer`). +- **15:19.** Plugin-loader root cause found in `opencode.log`; plugin fix + installed and opencode restarted. +- **15:19:57.** Router restarted. +- **After 15:19.** Conversation identity confirmed arriving: 17 of 17 requests + after the restart carry `c:` keys and an agent. + +### Status + +**Partly mitigated; detection not built; leading theory (3) unconfirmed.** + +**Lift A shipped.** PR #102 (a24de1b) landed the surfacing half of North Star #4 +(conversation identity and watchdog), live ~15:18 EDT. **Surfacing half of North +Star #4 is only partly shipped** — the identity half is confirmed (17 of 17 +requests after the restart carry `c:` keys), but the detection half still waits +on Lift B (`plans/no-progress-detection.md`). + +**Confounded by design, and say so.** Three mitigations landed within about +25 minutes: +- pinch off at 03:50 +- `auto` capped at 200k at 04:02 +- `z-ai/glm-5.3-flash` excluded at about 04:14 + +So "sessions look healthier since" cannot say which one mattered, or whether +it is several interacting: model degradation near a full window, oversized +sessions, pruning, and stale orchestration state. The one clean data point is +the 04:04 reproduction (glm, pinch off, small context, still confabulated), +which is why theory 3 leads. + +To attribute properly, change one variable at a time and watch the watchdog's +stall rate. For example, re-admit glm with everything else held, or re-enable +pinch with a budget above 200k. Until then, treat every theory here as +unproven. +Also missing: a **behavior breaker**. The availability breaker trips on 5xx and +PR #76 trips on malformed output; nothing trips on well-formed output that +confabulates or loops. Planned for Lift B (`plans/no-progress-detection.md`, +section 8). +- In place: + - pinch off, runtime and persisted + - `auto` limit 200,000, active for opencode sessions started after a restart + - #99 and #100 deployed at 03:15 +- Pending, operator: + - install `router-link.js` in place of `router-outcome.js` + - restart opencode +- Pending, lift: `plans/no-progress-detection.md`, the watchdog first. +- If pinch is re-enabled, its budget must sit above the 200k cap, and it must + never cut recent results. + +**Rule:** steady spend is not the failure; **spend with no concrete change +landing is.** Until the detector ships: +- Check an unattended agent run by whether commits are landing, not by the + quota page. +- Verify each worker commit's `git show --stat` against its message; a worker's + "done" is a claim. + +--- + ## The pattern The first five are defensible local decisions — free a port, stop a process, @@ -373,6 +635,14 @@ file by mistake; #7 was a *correct*, deliberate, reviewed code change that was still incomplete, because "every consumer" silently excluded a consumer that git cannot see. +#8 breaks the pattern from the other side: the router did what it was +configured to do, routing each request to a model that fit. It was the one +component that saw all the spend while being blind to whether any of it produced +anything, so it measured cost when the thing that mattered was progress. Whether +its own `pinch` pruning drove the loops is still a theory (see #8, Theory 1). If +confirmed, #8 joins the family of defensible settings, a 50k budget chosen for +~130k sessions, that turned harmful when the input changed shape. + The generalisable fix is the same each time: **compute the thing that is actually true, and surface it where the operator is already looking.** That is what the `/metrics` warnings and the inline admin warning are for, and it is the diff --git a/docs/watchdog.md b/docs/watchdog.md new file mode 100644 index 0000000..1f2665b --- /dev/null +++ b/docs/watchdog.md @@ -0,0 +1,408 @@ +> Deep dive into the watchdog subsystem — opencode session loop detection. Back to +> [README](../README.md). + +Watchdog is a periodic scanner that probes running opencode sessions for looping +patterns — repeated tool calls without progress. It runs as a systemd oneshot +every 5 minutes (`deploy/llm-router-watchdog.timer`) and, when a session flags, +writes into a local SQLite table and fires desktop notifications through the +configured channel pipeline. + +The detector itself is a pure module (`src/progress_detect.py`) that returns a +boolean verdict plus a reason dict. The orchestrator (`src/watchdog.py`) reads +opencode session data over HTTP, runs the detector, and manages the alert +state machine. The notifier (`src/notifier.py`) routes events through +configured channels to desktop alerts (`notify-send`). + +## What it catches and doesn't + +Watchdog looks for **looping** — sessions that make repeated tool calls against +the same targets without landing meaningful changes. It detects four signal +types: + +- **Dup signal**: a call appears more than 25% of the time in the sliding + window. The call matcher first normalises `bash` commands (strips env var + prefixes and `cd` prefixes, joins lines, collapses comments), then checks + tool name + canonicalised JSON args (sorted keys) for equality. + +- **Top signal**: a single target is called 12+ times in the window (or 8+ + for read-only agents like explore/librarian/oracle). Targets merge calls on + the same `(tool, basename)` for file reads, `(bash, first-two-words)` for + commands, and `(tool, pattern[:60])` for grep/glob. + +- **Slow signal**: a single tool call accounts for 15+ of the session's total + calls, AND no call in the entire session history has landed (tree write or + git commit). This catches sessions stuck on a single read/analysis. + +- **Coverage signal**: the session rereads one file at 4x its line count + within the window. The coverage score divides total bytes read by file length, + capped at the file's actual lines. A file read 4x its length means the model + is rereading without making progress. + +**NOT steady spend.** A session that steadily calls different files on a +difficult refactor is not flagged — each call lands a unique `(tool, args)` +and the top target never crosses the threshold. The detector only fires when +repetition outpaces the call budget, not when a session is quietly burning +tokens across many distinct tools. + +### Landed calls + +A call counts as **landed** when it: +- is an `edit`, `write`, or `patch` with a non-empty `metadata.diff` field, or +- is a `bash` command containing `git commit` with `metadata.exit == 0` + +Landed calls break both the slow and coverage signals. A session that reads a +file 4x, then edits it once, clears all signals immediately — the landed check +exits early and the verdict returns `False`. + +### Ancestral trees + +Sessions can be parented: the opencode API returns `parentID` on child sessions. +Watchdog resolves root sessions and evaluates EACH session in the tree on its +own calls, NOT merged into the root. A child's flagged verdict bubbles up to the +root level, which is what the alert dedup key uses (`opencode-loop:{root_session_id}`). + +For the landed-time check, a session sees NOT only its own landed calls but also +its transitive descendants' landed calls. This prevents a parent session from +being flagged just because its child landed — even though the parent may still +be looping independently. + +## Signals and thresholds + +Thresholds live in `DetectConfig` (`src/progress_detect.py:38-53`) and are +configurable under `watchdog.detector.*` in `config/config.yaml`. + +| Key | Default | Meaning | +|---|---|---| +| `watchdog.detector.dup_min` | 0.25 | Fraction: calls must repeat at this rate to flag | +| `watchdog.detector.top_min` | 12 | Minimum calls to a single target (normal agents) | +| `watchdog.detector.top_min_ro` | 8 | Minimum calls to a single target (read-only agents) | +| `watchdog.detector.cum_min` | 15 | Same call must account for this many total calls | +| `watchdog.detector.cover_min` | 4.0 | File must be reread this many times its length | +| `watchdog.detector.window` | 60 | Sliding window in number of tool calls (not seconds) | +| `watchdog.detector.min_calls` | 40 | Minimum total calls before evaluation triggers | +| `watchdog.detector.read_only_agent_keywords` | `["explore", "librarian", "oracle"]` | Title keywords that make an agent read-only (lowered `top_min` from 12 to 8) | + +Read-only agents get a lowered top threshold because agents whose job is +exploration or library work are expected to call the same targets repeatedly — +the floor is 8 instead of 12. + +## Alert lifecycle + +The alert state machine is per-root-session and tracked in the +`watchdog_alerts` table. + +### Trigger (initial detection) + +When the detector flags a root session's tree for the first time: + +1. The watchdog writes a `watchdog_alerts` row with `state='open'`. +2. The `local_llm_enabled` path asks a local Ollama model (defaulting to + `verification.model` or `qwen2.5-coder-router:14b`) for a second opinion. + It receives the agent name and the reason dict as context, and the model + must respond with only "yes" or "no" — no explanation. +3. If the local LLM answers "yes", severity is set to `critical`. If it + answers "no" or times out (returns `None`), severity is set to `warning`. +4. The notifier dispatches an `AlertEvent` through every matching channel. +5. The `flagged_ticks` counter starts at 1. + +A cap of 3 simultaneous LLM second-opinion calls (`_MAX_LLM = 3`) prevents a +burst of flagged sessions from flooding the local model. + +### Escalate (persistent looping) + +On every subsequent tick where the session is still flagged: + +1. `flagged_ticks` is incremented. +2. If `flagged_ticks >= 3` (15 minutes at 5-minute ticks) OR the LLM answers + "yes" again, the alert escalates to `severity='critical'` and the escalation + event is dispatched. +3. If the alert is already critical, ticking continues but no duplicate + escalation fires. + +### Resolve (session is gone or fixed) + +A root session alert resolves in two conditions: + +- **Session disappeared**: the root session ID is no longer returned by the + opencode `/session` API. The watchdog writes `state='resolved'`, + `resolved_at=`, and `severity='info'`. +- **Session fixed**: the root session's tree no longer has any flagged sessions + in the detector's verdict. The alert is resolved the same way and a resolve + event is dispatched. + +Resolve events always bypass the rate limit — they must clear the previous +notification so the operator knows the alert is over. + +### No-opencode + +When no `rc-servers.json` or no opencode server answers, the tick writes an +`outcome='no_opencode'` record and returns immediately with zero alerts. This +uses the filesystem lock (`router.db/.watchdog.lock`) so two watchdog instances +cannot run simultaneously. + +## Admin surfaces + +Watchdog data surfaces through three admin pages, each serving a different +operational question. + +### Loops panel — index.html#loops + +On the main dashboard (`admin/frontend/index.html`), the **Loops** card +(id="loops") shows all currently open alerts: + +- **Severity badge**: the alert's current severity (`critical` in red, + `warning` in orange-yellow, `info` in grey) +- **Dedup key**: the raw `opencode-loop:{root_session_id}` string, so the + operator can cross-reference with `journalctl` or `router.db` +- **Opened at**: `opened_at` ISO timestamp from the first trigger +- State: `open` or `resolved`; the panel filters to `resolved_at IS NULL` + +The panel is client-side only — it polls `/api/watchdog/loops` which queries +`watchdog_alerts` for all unresolved rows. + +### Watchdog card — controls.html + +The Controls page (`admin/frontend/controls.html`) includes a dedicated +**Watchdog** card with operational status: + +- **Last tick**: timestamp of the most recent `watchdog_ticks` row +- **Sessions seen**: how many opencode sessions were enumerated on that tick +- **Flagged verdicts**: count of flagged + total from the latest tick's + `watchdog_verdicts` rows +- **Open alerts**: count from `watchdog_alerts WHERE resolved_at IS NULL` +- **Notification channels**: list of configured channels with toggle controls + (enabled/min_severity), editable via `POST /api/watchdog/channels` +- **Refresh** button: reloads the watchdog status from the API +- **Send test alert** button: fires a test `AlertEvent` through enabled channels + +The status endpoints are: +- `GET /api/watchdog/status` — last tick, verdict counts, open alert count +- `GET /api/watchdog/channels` — per-channel settings from `watchdog_channel_settings` +- `POST /api/watchdog/channels` — upsert channel settings (enabled, min_severity) +- `POST /api/watchdog/test-alert` — deliver a test event (severity + optional + channel_name filter) to a live Notifier instance + +### Per-model stall rollup — models.html + +The `watchdog_verdicts` table has `model_id`/`provider` columns and indexes on +them, but no per-model stall rollup, no Block button beside evidence, and no +Blocked list are built yet. `blocked` is only a value in the Models override +dropdown (see docs/admin-portal.md). + +## Database schema + +Four tables live in `router.db`, created idempotently by both +`config/schema.sql` and `watchdog_store.py`: + +```sql +watchdog_ticks -- one row per watchdog tick +watchdog_verdicts -- per-session verdict at each tick (joined to ticks) +watchdog_alerts -- alert lifecycle state machine (dedup_key PK) +watchdog_channel_settings -- per-channel toggle + severity gate +``` + +Indexes on `watchdog_verdicts` cover `model_id`, `created_at`, `flagged`, and +`session_root` — the columns the admin endpoint queries most frequently. + +The ticks table carries `ticked_at`, `sessions_seen`, and `outcome` +(`'ok'`, `'flagged'`, or `'no_opencode'`). The verdicts table joins to ticks +via `tick_id` and carries the full reason dict fields: `dup`, `top`, +`top_what`, `landed`, `slow`, `coverage`, plus `calls_since_landed`, +`cost_since_landed_usd`, and the LLM second-opinion answer. + +Alerts are keyed by `dedup_key` (deduplicated per root session) and carry +`opened_at`, `last_fired_at`, and `resolved_at`. The transition logic lives in +the `_fire_alert()` helper: trigger inserts only if no open row exists, +escalate increments `flagged_ticks` and bumps severity, and resolve writes +`resolved_at` and downgrades to `severity='info'`. + +## Knobs + +All watchdog configuration lives under `watchdog:` in `config/config.yaml`: + +| Key | Default | Required | +|---|---|---| +| `watchdog.enabled` | `true` | Gates the entire subsystem | +| `watchdog.local_llm_enabled` | `true` | Whether to call Ollama for second opinions | +| `watchdog.model` | `null` (uses `verification.model`) | Local model name | +| `watchdog.detector.window` | 60 | Sliding window size in calls | +| `watchdog.detector.dup_min` | 0.25 | Dup threshold (fraction) | +| `watchdog.detector.top_min` | 12 | Top target threshold (normal agents) | +| `watchdog.detector.top_min_ro` | 8 | Top target threshold (read-only agents) | +| `watchdog.detector.cum_min` | 15 | Slow-signal threshold | +| `watchdog.detector.cover_min` | 4.0 | Coverage-signal multiplier | +| `watchdog.detector.min_calls` | 40 | Minimum eval calls | +| `watchdog.read_only_agents` | `["explore", "librarian", "oracle"]` | Agent title keywords for reduced threshold | +| `watchdog.dashboard_base_url` | `"http://127.0.0.1:8080/admin"` | Base URL for alert links | + +All seven `watchdog.detector.*` keys plus `watchdog.enabled` and +`watchdog.local_llm_enabled` (nine allowlisted keys total) are allowlisted for +live editing through the +admin portal (`POST /admin/api/config/{key}`), which writes to +`config.local.yaml` (the gitignored machine-local overlay). The changes are +validated by `RouterConfig` before reaching disk, and take effect at the next +tick — no service restart required since `detect_config_from_pydantic()` reads +cfg fresh each invocation. + +## Install + +The watchdog ships as two systemd user units: + +- `llm-router-watchdog.service` — a `Type=oneshot` that runs + `PYTHONPATH=%h/llm-router/src .venv/bin/python -m watchdog --once` +- `llm-router-watchdog.timer` — fires 5 min after boot, then every 5 min + +Installation follows the same pattern as the other deploy units. The shipped +unit files use `%h/llm-router` placeholders that need rewriting to your actual +repo path: + +```bash +REPO=$(pwd) +for u in deploy/llm-router-watchdog.{service,timer}; do + sed "s|%h/llm-router|${REPO}|g" "$u" \ + > ~/.config/systemd/user/"$(basename "$u")" +done +systemctl --user daemon-reload +systemctl --user enable --now llm-router-watchdog.timer +``` + +The service unit has **no `EnvironmentFile`** — watchdog needs no provider API +keys because it only queries the local opencode session APIs and the local +Ollama instance. + +Check installation: + +```bash +systemctl --user status llm-router-watchdog.timer +journalctl --user -u llm-router-watchdog.service --no-pager +``` + +## Run --once + +For ad-hoc debugging or pre-deploy smoke test: + +```bash +PYTHONPATH=src .venv/bin/python -m watchdog --once +``` + +The `--once` flag runs a single tick and exits 0 once the tick has run, +however many alerts it fired (the count is logged as `watchdog tick done: N +alert(s) fired`); systemd would otherwise mark the oneshot unit failed on every +real alert. It connects to the same `router.db` that the timed +service uses, acquires the filesystem lock, reads `rc-servers.json`, probes +opencode servers, and follows the full evaluate-alert-resolve pipeline. + +Add `--config /path/to/config.yaml` to point at a non-default config file. + +### Debug output + +The `--once` run emits structured `watchdog=` log lines for every state +transition: + +``` +2026-01-15 14:30:01 INFO watchdog: watchdog=2026-01-15T14:30:01+00:00 dedup='opencode-loop:ses_xxxxx' state=trigger severity=warning title='opencode loop: explore' +``` + +The `title` field carries the agent name derived from the session title. The +`dedup` field is the root session's dedup key, matching `journalctl` traces +and admin panel rows. + +## Backtest + +`scripts/progress_backtest.py` replays labelled sessions through the detector +to validate threshold calibration. It reads from a fixture JSON file and runs +a sliding-window evaluation with step 5. + +```bash +PYTHONPATH=src .venv/bin/python -m scripts.progress_backtest --fixture \ + tests/fixtures/progress/fixture.json +``` + +The fixture ships in `tests/fixtures/progress/fixture.json` and contains +15 labelled sessions: 8 marked `must_flag` and 7 marked `must_not_flag`. +Expected output after a successful run: + +``` +# backtest complete: 8/15 flagged +``` + +At least 8 sessions should trigger a flag. The fixture is generated from live +opencode sessions (see [export fixture](#re-export-fixture) below) and scrubbed +to protect file contents — content is SHA-1 hashed except for the few keys the +detector actually inspects (`filePath`, `command`, `pattern`, etc.). + +## Re-export fixture + +The fixture generator pulls sessions from the opencode SQLite database +(`~/.local/share/opencode/opencode.db`) and writes the scrubbed test fixture: + +```bash +python scripts/export_progress_fixture.py +``` + +Output goes to `tests/fixtures/progress/fixture.json`. The script takes a +session ID to label map (`LABELS` dict), extracts call history, resolves file +line counts (via `git show` for deleted worktree paths), and scrubs all strings +to SHA-1 prefixes, preserving only the detector's target keys. + +The fixture generation script hardcodes paths to the operator's home directory +and a specific worktree commit (`WORKTREE_COMMIT = "827c408"`). When +regenerating, update those paths to match the current environment. + +## Notifier pipeline + +Alert events flow through `src/notifier.py` before reaching the operator. Each +alert is an `AlertEvent` dataclass mirroring PagerDuty Events v2 shape: +`dedup_key`, `severity` (info/warning/critical), `state` (trigger/escalate/ +resolve), `title`, `summary`, `details`, and `source` (hardcoded to +`"6krrt-watchdog"`). + +Each configured channel has its own `min_severity` gate — a critical event +passes through a channel with `min_severity=warning`, but a warning event does +not pass a channel with `min_severity=critical`. Each channel also has a rate +limit window of 300 seconds (5 minutes). Resolve events always bypass the rate +limit to ensure clear-the-alert notifications are delivered. + +Currently only `type=desktop` channels are implemented; they invoke +`notify-send -u {urgency} {title} {summary}`. Unknown channel types are +logged as warnings and skipped. Missing `notify-send` (no `DISPLAY` or +`libnotify` installed) is logged at warning level — the notifier never raises. + +Channel settings (`watchdog_channel_settings`) are editable at runtime through +the admin portal and stored per-channel in the database, separate from the +base config. + +## Known limits + +- **Under 40 calls: invisible.** `min_calls=40` means the detector does not + evaluate sessions with fewer tool calls. Short troubleshooting sessions that + loop on 10 calls pass through undetected. + +- **Atlas at cover_min 4.0 exactly.** A session whose last 60 calls read a + single file exactly 4x its line count hits the coverage threshold at the + boundary. There is no margin: `>=` comparison means exactly 4.0x flags. + This matters for sessions that genuinely need to re-read a reference file + repeatedly — the coverage signal cannot be tuned per-file. + +- **Stale localhost:4096 probes.** Watchdog reads `rc-servers.json` from + `~/.local/share/opencode/rc-servers.json`, which may contain stale entries + for opencode sessions that have since ended. These produce empty call lists + and are silently skipped — but they add network round-trip latency to each + tick. Only the first answering server's sessions are evaluated (the watchdog + picks `answering[0]` from the list of servers that successfully return a + session list). + +- **No streaming path.** `POST /outcome` is the only ground truth signal for + streaming traffic; the watchdog has no equivalent endpoint to report streaming + session outcomes. It can only see tool-call repetition, not whether the answer + was useful. + +- **Local LLM gate.** Second-opinion calls are gated by both + `local_compute.enabled` (global) and `watchdog.local_llm_enabled` + (per-subsystem). If either is false, no LLM calls are made and severity stays + as `warning` on initial trigger regardless of model quality. The cap of 3 + simultaneous LLM calls means that if 5 sessions flag, 2 wait without opinion. + +- **No provider calls.** Watchdog does not touch the dispatch path. It only + reads opencode session data and runs local heuristics. No provider calls, no + quota consume, no cost incurred. diff --git a/opencode.json b/opencode.json index 91db3f1..eb67d2d 100644 --- a/opencode.json +++ b/opencode.json @@ -13,7 +13,7 @@ "auto": { "name": "auto (router picks, interactive)", "limit": { - "context": 782324, + "context": 200000, "output": 16384 }, "modalities": { diff --git a/plans/cockpit-brainstorm.md b/plans/cockpit-brainstorm.md new file mode 100644 index 0000000..93d321f --- /dev/null +++ b/plans/cockpit-brainstorm.md @@ -0,0 +1,397 @@ +# Cockpit brainstorm: the router as an operator's instrument panel + +Status: reference -- brainstorm, not a queue item; its six quick wins shipped in PR #101 + +**Date:** 2026-09-23 +**Read against:** the repo tarball as of this date (code, config, plans, +screenshots). **Not** read against the live `router.db`, so every "gap" below +is a gap in what the code surfaces, not a claim about what the data shows. +Section 5 is the exception: it was read against the live portal. + +## The framing + +6krrt as a local LLM router with a cockpit. The operator: + +1. adjusts dials and knobs in flight, +2. sees where tokens are wasted, +3. watches proficiency, and +4. sees poor results from a carrier or model surface, so they can decide + whether to preclude it (with the circuit breaker as the automatic version + of that decision for outages). + +The router makes the per-request decision. The cockpit is where the operator +makes the slower decisions: which models and carriers are trusted, for what, +and at what settings. Under this framing, the question for every feature is +whether it helps the operator notice something and act on it. + +--- + +## What the cockpit already has + +| job | what exists | where | +|---|---|---| +| in-flight dials | ~11 runtime knobs (bool/float/int tables in `admin.py`), persisted config writes via `/api/config/{key}`, classifier card, profiles CRUD, gaming mode | Controls, Profiles | +| knob coverage | North Star rule #1 + `test_admin_knob_coverage.py` | tests | +| token waste | pinch savings card; `cache_rate_series` + warnings; incumbent cache pricing dial; `cost_estimate_calibration` in `/metrics` | Dashboard, Controls, `/metrics` | +| proficiency | model × category matrix with evidence status (measured / inherited / thin, n per cell); client outcome log with attributable + applied flags | Proficiency | +| poor results | `content_fault_warnings` (malformed output rate); `rejection_warnings`; availability override (active / deprecated / stale); allowlist editor | warnings bell, Models, Providers | +| breaker | passive per-(model, provider) breaker on 5xx; on/off knob | Controls (toggle only) | + +The foundation is good. What's mostly missing is the link between a signal +and the action it should prompt, plus any record of which actions were taken. + +--- + +## Cross-cutting ideas (these help all four jobs) + +### X1. Flight recorder: a change log + +**Gap.** `admin_schema.sql` has only `admin_model_overrides`. Nothing records +when a knob, override, allowlist entry, profile or feedback fold changed, what +it changed from, or why. Runtime knobs are "NEVER persisted", so a restart +silently reverts them, and nothing notes that the revert happened. + +**Idea.** +- An `admin_changes` table: `(at, surface, key, old, new, reason, source)`, + where `source` is `runtime | persisted | override | allowlist | profile | + feedback_fold | restart`. Every admin write path appends a row. +- An `operator_epoch` id stamped on each `route_decisions` row, incremented + on each **operator action** (a row in `admin_changes`). Every metric can then + be split before and after a change without joining on timestamps. Scope it + to operator actions on purpose: "any change that affects routing" would also + cover catalog price polls and every feedback fold, so the epoch would tick + constantly and split nothing. +- Change markers drawn on the dashboard history charts. +- A drift badge when a runtime value differs from its persisted twin + ("this will revert on restart"). +- Restart reverts are loggable without persisting runtime knobs: the change + log already holds each knob's last runtime value, so at startup, compare it + with the loaded config and write one `restart` row per value discarded. + +**Why first.** Adjusting knobs in flight produces no learning if you can't +later tell what a change did. The other ideas here lean on this one. + +### X2. Structured warnings with actions attached + +**Gap.** `/metrics` warnings are plain strings. The dashboard decides which +page fixes a warning, and how severe it is, by regex on the text +(`warningTarget` and `warningSeverity` in `index.html`). Rewording a warning +silently breaks its link and its severity. + +**Idea.** Emit `{class, severity, subject: {model, provider, category?}, +evidence_url, actions: [...]}`. The warning registry in +`test_tui_warnings.py` already lists every class, so it can serve as the +schema. The bell can then offer an action in place: "exclude from +`tool_use_agentic`", "quarantine 24h", "open decisions filtered to this +model". + +### X3. Preview before apply (replay) + +**Gap.** The reviews replayed thousands of real decisions to test a setting, +but only as one-off scripts pasted into plan documents. + +**Idea.** For any change that affects ranking (`quality_tolerance`, the +incumbent dial, profile edits, availability overrides, graded exclusions +from P2 below), the portal replays the last N decisions through +`select_candidates` and `rank_candidates` with the proposed value, then shows: +- a winner-shift matrix: old winner → new winner, with counts, +- estimated cost delta, +- decisions that would become unroutable (the 2026-09-01 and 2026-09-04 + incidents were both deprecations that emptied a candidate set). + +This is cheap because the ranking modules are pure. Its output is an estimate +of routing change, not quality change, and should be labeled that way. + +--- + +## 1. Dials in flight + +- **D1. Timed changes.** "Apply for 2h, then revert." Turns a knob change into + a bounded experiment and lowers the risk of forgetting a test setting. + Needs X1 to record both the apply and the revert. +- **D2. Blast radius on hover.** How many of the last 24h of decisions this + knob would have touched. It's the X3 replay reduced to one number. +- **D3. Session-scoped A/B.** Assign new sessions to setting A or B and + compare cost and outcome rate per arm. It has to be per session, not per + turn, because a mid-session switch costs cache (token-waste Wave 2 + measured 0.919 → 0.348). Larger build; only worth it once X1 exists and + outcome volume can support a comparison. +- **D4. Grouping by effect.** Group controls by what they move (cost, + quality, latency, safety, measurement-only) rather than by config section, + and mark the few that actually change routing. Knob coverage guarantees + every control exists; this makes them usable. + +## 2. Token waste + +**Gap.** Waste signals are spread across pinch, cache rate, calibration and +the verifications table, and none is in dollars in one place. +`cost_estimate_calibration` is in `/metrics` but not in the portal. + +- **W1. Waste ledger.** One panel with dollars per waste class over a window, + each with a trend: + | class | how it's computed | + |---|---| + | switch cache loss | billed on switch turns − same prompt at the session's same-model cache rate | + | retries | re-billed prompt tokens from `iteration.py` attempts | + | paid for failure | billed cost of responses later reported `ok:false` via `/outcome` | + | empty-200 / malformed | billed cost of responses with a structural `malformed` verdict | + | truncation | responses stopped by `finish_reason: length` with no client cap | + | pinch (negative) | dollars saved, from `pinch_summary` | + + "Paid for failure" counts only decisions that received an outcome, so it + shows n and outcome coverage (the share of billed decisions with a report) + next to the dollar figure. +- **W2. Why did it switch?** Record a switch reason on each decision: category + changed, breaker open, exploration, override removed the incumbent, profile + change. Then rank sessions by switch cost and drill into the turns. Without + a reason, a switch is only a cost; with one, it points to a knob. +- **W3. Calibration panel.** Surface `cost_estimate_calibration` per model, + headlining the **spread** (as its docstring argues), and flag pairs where + the estimator and the bill disagree on order. Where that happens, the cost + tiebreak picks the wrong model. +- **W4. Cost per successful outcome.** Per (model, provider, category): + billed dollars ÷ `ok:true` outcomes. A cheap model that fails often isn't + cheap. This is the one waste figure that includes quality. + + Show n and outcome coverage next to it. Models that carry more traffic + collect more outcomes (the same exposure bias the feedback fold had to + correct), and if coverage stays hidden, a thinly reported cheap model will + look better than it is. + +## 3. Proficiency monitoring + +**Gap.** The matrix shows current state well. It doesn't show movement, and +it doesn't show which thin cells actually matter. + +- **P1. Trend and drift per cell.** A sparkline of the outcome rate over time. + Compare a recent window against the long-run rate with an interval, and + flag drops that fall outside it. Carriers change serving setups (quant, + engine, hardware) without notice, and a drift flag is how that shows up. +- **P2. Evidence priority.** Rank cells by `traffic share × uncertainty`. A + thin cell that routing never consults doesn't matter; a thin cell deciding + 30% of traffic does. The ranking tells you where to spend `eval_proficiency` + runs or exploration budget. +- **P3. The tolerance band per category.** Show which models sit within + `quality_tolerance` of the leader, so it's clear where cost is deciding and + where quality is. +- **P4. Label provenance per cell.** The share of each cell's outcomes whose + category came from a fresh classification versus a cached or borrowed one. + North Star rule #2 exists because borrowed labels trained the matrix; + this makes it visible cell by cell. + +## 4. Poor results → operator preclusion (and the breaker) + +**Gap.** The operator's only preclusion tools are binary: an availability +override (active / deprecated / stale) or removing a model from the +allowlist. The breaker is in-memory, trips only on 5xx, forgets everything +on restart, and shows nothing but an on/off toggle. + +- **Q1. Carrier/model scorecard.** One sortable table per (model, provider) + over a window: outcome fail rate (with n), malformed / empty-200 rate, + 5xx and timeout count, breaker trips, p95 latency, estimate-vs-bill error, + cache rate, cost per success (W4). This is the "who's misbehaving" view the + preclusion decision needs. +- **Q2. Same model, different carriers.** Group scorecard rows by base model, + e.g. `qwen3.6-35b` on NeuralWatt vs OpenRouter. This separates "the model is + bad at this" from "this carrier serves it badly", which calls for a + different fix: drop the carrier's row, or restrict the model. +- **Q3. Graded actions** in place of deprecate-or-not. Each action records a + reason and links its evidence in X1: + 1. exclude from category X. This **cannot** reuse `eligible_categories`: + that field is restrict-only (`admin.py`, the probe docstring), and NULL + means "every category", so a deny needs its own column, + 2. exclude when the request carries tools (the `deepseek-v4-flash` case). + This may be the way out of the frozen `tool_use_agentic` data: + `min_tool_proficiency` can only read scores that stopped moving on + 2026-09-15, while a per-model manual rule lets the operator decide from + Q1 scorecard evidence instead, + 3. restrict to batch / `-flex` only, + 4. quarantine: exploration-only, so it keeps collecting evidence without + carrying traffic (see the open question below; deferred until Q1 shows + it is needed), + 5. timed ban, e.g. 24h, auto-expiring and logged, + 6. deprecate (what exists today). + + Each action goes through at least the minimal X3 check (would this empty a + candidate set for any recent decision?) before committing. That check alone + would have caught both the 2026-09-01 and 2026-09-04 incidents; the full + winner-shift replay can come later. +- **Q4. Breaker visibility.** Show open circuits with `down_until`, current + cooldown and trip count; keep a trip history (persist it, since `_store` + is process-lifetime); add manual force-open (drain a model on purpose) and + force-close (reset after a known fix). +- **Q5. A quality breaker, separate from availability.** Trip on content + failures (empty-200, mangled output, a burst of `ok:false`) using the + novelty-or-rate rule already in `rejection_warnings`. This is roughly what + `plans/mangled-output-detection.md` specs (status: planned). Keep it a + separate breaker and panel so a quality trip doesn't read as an outage. + **One doc owns it:** fold Q5 into `mangled-output-detection.md` rather than + specifying it twice, and let this file point there. +- **Q6. Carrier-level breaker.** When several models on one provider trip + inside a short window, open the provider rather than walking its models one + cooldown at a time. Account exhaustion already has its own path; this is for + partial carrier outages. + +--- + +## 5. Quality of life: ergonomics and readouts + +Unlike the sections above, this one **was** read against the live portal +(8080, view-only, 2026-09-23), plus the frontend source. Each item names the +thing observed. + +### Readouts that mislead today + +- **R1. Home tiles use different windows.** Quota is "this period", + Decisions 7d, Models and Pinch 30d. Busiest models shows `kimi-k2.7-code` + at 10,028 while the Decisions tile says 2,887 for everything, and both are + correct. Label each tile's window on its face, or drive all of them from + the Activity card's 24h / 7d / 30d selector. +- **R2. The Decisions tile counts verifications, and mixes diagnostics with + ground truth.** It reads "2,887 verified in 7 days: 132 ok, 93 failed", but + `index.html` sums `verdict_mix`, which counts `verifications` rows, not + decisions. Checked read-only against the live DB on 2026-09-23: + | | tile | actually | + |---|---|---| + | headline | 2,887 | 2,678 route decisions; 2,887 is verification rows, 2,660 of them `unverifiable` | + | "ok" | 132 | 96 client `succeeded` + 29 `local_llm` ok + 7 structural ok | + | "failed" | 93 | 82 client `failed` + 11 `malformed` (local_llm + structural) | + The tile adds structural and `local_llm` verdicts, which CLAUDE.md calls + diagnostics only, to `/outcome` reports, the only ground truth. Show route + decisions as the headline, then client outcomes on their own: "178 client + reports (6.6%): 96 ok, 82 failed (46%)". Diagnostics, if shown at all, go on + a separate line. +- **R3. The Activity chart smooths across gaps.** The decisions series is a + spline through sparse points, so it draws continuous traffic across hours + that had none. Break the line at empty buckets or use bars. Separately, + billed requests exceed decisions at several peaks; the chart should say what + the excess is (cloud classifier calls? retries? unrouted pins?) rather than + leave two lines that disagree unexplained. +- **R4. Sparklines have no values.** The tile sparklines carry no axis and no + hover. Add a hover value, or min and max labels. +- **R5. The Classifier card flashes a wrong value.** On first paint the mode + select reads `local_llm` and only switches to `local_encoder` once config + loads. For a second, the card shows a value that isn't configured. Render + a loading state instead of the default. +- **R6. Config echo vs live state.** The Classifier card shows what the + overlay says (`local_encoder`, device `cpu`), not what the process actually + loaded. Show the resolved model, device and recent p50 latency from the + running classifier. (It also exposed doc drift: CLAUDE.md says the live + deployment runs `device: cuda`; the overlay says `cpu`.) +- **R7. Proficiency colour encodes evidence, not score.** Green / amber mean + measured / thin, so the best model in a column isn't visible without reading + every number. Mark each column's leader, outline the `quality_tolerance` + band (P3), make columns sortable, and add a "routable only" toggle so rows + that are deprecated or not allowlisted stop padding the matrix. + +### Decisions page ergonomics + +- **E1. Collapse runs.** The top 40 rows are one session: same category, + profile, tier, source and model, and only ctx and cost move. Fold + consecutive same-session, same-model rows into one expandable row: "38 turns, + ctx 75k to 98k, $0.27 total". Switches then stand out as row boundaries, + which is what W2 wants to surface. +- **E2. Session view.** A per-session rollup: turns, total cost, context + growth, switches, outcome count. Context climbing about 1k per turn is + visible in the raw table and invisible everywhere else. +- **E3. Filters in the URL.** `decisions.html` never reads `URLSearchParams`, + so no other page can link to "decisions for this model" or "this session". + X2 actions, the Q1 scorecard and the model modal all need that link. +- **E4. Row drill-down.** A row click does nothing. Open a panel with the + decision's rejected candidates, verification verdict, `/outcome` report, + and estimated vs billed cost from its energy observation. +- **E5. Server-side filtering.** The page loads 1,000 of 32,540 rows and + filters and searches only those, with a footnote saying so. Push filters to + the API so that "All" means all. +- **E6. Model, provider and source as filters.** Today they're reachable only + through free-text search. + +### Controls page + +- **C1. One knob table, not two.** Runtime Knobs and Persisted Config list + largely the same knobs in two columns, in different orders, with persisted + labels truncated (`objective.incumbent_cache_pri…`). Checking drift means + matching rows by eye. Use one table: knob, live value, persisted value, + layer (base / overlay), drift badge. That table is also X1's natural home. +- **C2. Separate Restart Service.** It sits in the same button row as Refresh + Catalog and Seed Energy. Move it apart, and have its confirmation list which + runtime values differ from persisted and will revert. +- **C3. Prose in cards.** Local Compute carries three paragraphs; the portal + style rule is no paragraphs. Keep one line plus a docs link or tooltip. +- **C4. Default and last change per knob.** Show the code default beside + each value, plus "changed 2h ago from 0.3" once X1 exists. + +### Portal-wide + +- **G1. Nav lives in eight files.** Each page hardcodes the same `
    ` of + nav links. `navbar.js` already notes that adding Quota was an eight-file + edit; have it render the links too. +- **G2. Width.** Most pages sit in `container-xl` and use well under half of + a wide screen, while Decisions, the densest table, is the most cramped. + Proficiency already goes full-width. Go fluid on the data-heavy pages. +- **G3. Dismissals don't stick on rate warnings.** Dismissal is keyed on the + warning's exact text, on purpose (a changed condition resurfaces it). But + warnings that embed a live rate ("3.2x pace") change text every poll, so a + dismissal never holds. X2's `class` + `subject` is the right key; resurface + on a severity change, not a digit change. +- **G4. Keyboard.** `/` to focus search, `Esc` to close modals, and `j`/`k` + on the Decisions table. Cheap, and this is a page someone lives in. + +--- + +## One thing the cockpit framing makes more urgent + +The portal is loopback-only with no auth, and it already includes +`/api/restart-service`, `/api/apply-feedback` and provider deletes. The more +it becomes the place decisions are made, the stronger the pull to reach it +from other machines, and moving the router onto a Proxmox box on the LAN is +exactly that move. Add auth before binding anything but `127.0.0.1`. Same +warning CLAUDE.md already gives for the API, with more at stake. + +--- + +## A possible order + +| # | item | why here | +|---|---|---| +| 0 | portal auth | **only if** the router is leaving `127.0.0.1` (the Proxmox move); blocks that move, not this list | +| 1 | X1 change log + operator epoch | everything else reads it | +| 2 | Q1 scorecard + Q4 breaker visibility | the preclusion decision needs one view first | +| 3 | Q3 graded actions (via X1) + minimal X3 (empties-a-candidate-set check) | turns the scorecard into decisions without repeating 09-01 / 09-04 | +| 4 | X2 structured warnings | lets warnings trigger Q3 actions directly | +| 5 | W1 waste ledger + W2 switch reasons | token waste in dollars, with causes | +| 6 | X3 full replay preview (winner shift, cost delta) | makes knob changes safe to try | +| 7 | P1 drift, P2 evidence priority | proficiency monitoring over time | +| 8 | Q5 quality breaker (owned by `mangled-output-detection.md`), D1 timed changes | automate what the operator has been doing by hand | +| 9 | D3 session A/B, Q6 carrier breaker | only after the above prove out | + +Section 5 sits outside this order: most of it is small and independent. +Quick wins that also unblock the numbered items: **E3** (URL filters, needed +by X2 and Q1), **R2** (outcome coverage, the same honesty W1 and W4 need), +**C1** (one knob table, X1's home), **G3** (fixed by X2's class keys). R5 and +G1 are fixes of a few lines each. + +**North Star #1 applies to every item here.** Q5 thresholds, D1 durations, +timed-ban lengths and any quarantine settings are new config knobs, so each +ships with its admin control or `test_admin_knob_coverage.py` fails. Scope the +control into the item, not a follow-up. + +### Open questions, with proposed answers + +- **Is the change log append-only history, or also an undo stack?** + Proposed: append-only, plus a per-row "revert to old value" that writes one + key and appends its own row. That is not the shape CLAUDE.md rejected for + gaming mode: that concern was one flag writing five keys and drifting apart. + A one-key revert can't drift, and the history stays intact. +- **Should graded exclusions live in the DB or in `config.local.yaml`?** + Proposed: the DB, next to `admin_model_overrides` (extend it or add a + sibling table). Timed bans need an expiry, and an expiry belongs in a DB + column. The overlay is per-machine and not in git, so it's the wrong home + for judgments about a carrier. And availability overrides already live in + the DB. The category exclusion still needs a new deny column (see Q3.1). +- **Does quarantine conflict with session-scoped exploration?** Yes, and the + weight goes up: a quarantined model would carry whole sessions, not single + turns. There's a second problem too. `exploration.py` only chooses among + models that already passed the hard filters and were ranked, so quarantine + needs a new state, "eligible to explore, barred from winning", rather than + reusing anything existing. Defer it until Q1 shows it's needed. diff --git a/plans/cockpit-quick-wins.md b/plans/cockpit-quick-wins.md new file mode 100644 index 0000000..f1bd843 --- /dev/null +++ b/plans/cockpit-quick-wins.md @@ -0,0 +1,217 @@ +# Cockpit quick wins: six small admin-portal fixes + +Status: done -- shipped in PR #101 +Date: 2026-09-23 +Source: `plans/cockpit-brainstorm.md` section 5 (items R2, R5, E3, C1, G1, G3). +Each item below was observed on the live portal or checked read-only against +the live `router.db` on 2026-09-23. + +## Ground rules (all items) + +- **Worktree, not the main checkout.** Create a worktree from `main` HEAD + (`c87e675` or later) on branch `feat/cockpit-quick-wins`. The main checkout + at `/home/alee/Sources/6krrt` has UNCOMMITTED edits from other work, + including `admin/frontend/controls.html` (a `confidence_threshold` -> + `confidence_min` rename around lines 1026 and 1150) and + `tests/test_admin_frontend.py`. Do not touch, stash, or commit them. C1 edits + a different region of `controls.html`; keep it that way so the later merge + is clean. +- Commit by explicit path. Never `git add -A` or `git add .`. +- **Never touch port 8080** (production). For visual checks, run a throwaway + instance on **8081** from the worktree, with a temp copy of the DB, never the + live `router.db` for writes. +- Never edit `config/config.yaml` or `config/config.local.yaml`. +- ASCII only in new code, comments, and UI strings. No middle-dot separators + (U+00B7); relate facts with layout, or with `:` or a rephrase. +- No paragraphs of prose in the UI. One line per hint at most. +- One commit per item, in the order listed. `pytest` (offline) and + `ruff check` stay green after each commit. +- For every UI change, take a screenshot on 8081 and check the changed area + zoomed in, not just the item's checklist. + +--- + +## 1. R5: Classifier card shows a wrong value on first paint + +**Problem.** `admin/frontend/controls.html:264-268`: the +`#classifier-mode-select` ``s are populated (~lines 405-420), read + `kind`, `category`, `profile`, `tier`, `q` (search) and `size` from + `location.search` and apply them. If a value is not among a select's + options, add it as an option rather than dropping it (a link may name a + category that has no rows in the loaded window). +- On any filter change (`resetPageAndRender`, the Clear button, page size), + write the current filters with `history.replaceState`. Omit empty values. + Don't push history entries. +- `q` also accepts a model id or a session id, since search already matches + both (`search` at ~line 433). + +**Test.** A static test that `decisions.html` reads `URLSearchParams` and calls +`history.replaceState`. On 8081, open +`/admin/decisions?category=diff_checking&q=deepseek`, confirm the filters and +the rows match, and screenshot it. + +## 4. R2: The Decisions tile counts verifications and mixes diagnostics with ground truth + +**Problem.** `admin/frontend/index.html:597-610`: the Home "Decisions" tile +sums `_snap.verdict_mix`, which is `metrics.verdict_mix()` and counts +`verifications` rows by verdict. Live DB, last 7 days: + +| verdict | kind | n | +|---|---|---| +| unverifiable | structural | 2,660 | +| succeeded | client_outcome | 96 | +| failed | client_outcome | 82 | +| ok | local_llm | 29 | +| ok | structural | 7 | +| malformed | local_llm / structural | 4 / 7 | +| truncated | structural | 2 | + +The tile shows "2,887 verified in 7 days: 132 ok, 93 failed". Actual route +decisions in 7 days: **2,678**. "ok" adds client `succeeded` to structural and +`local_llm` ok. "failed" adds client `failed` to `malformed`. Per CLAUDE.md, +structural and `local_llm` verdicts are diagnostics only; `/outcome` +(`kind = 'client_outcome'`) is the only ground truth. + +**Fix.** +- Backend: in `src/metrics.py`, add `decision_outcome_summary(conn, + since_days=7)` returning + `{decisions, client_reports, client_ok, client_failed}`: + `decisions` = `route_decisions` rows with `observed_at` inside the window; + `client_*` = `verifications` rows with `kind = 'client_outcome'` inside the + window (`succeeded` = ok, `failed` = failed). Include it in the admin + snapshot (`src/admin.py` near line 2131) as `decision_outcomes`. Leave + `verdict_mix` unchanged; the TUI uses it. +- Frontend: the tile headline becomes `decisions`. Qualifier lines: + - `routed in 7 days` + - `178 client reports (6.6%): 96 ok, 82 failed` (coverage = + `client_reports / decisions`; the failed count uses the warn colour when + the fail rate among reports is above 10%, matching the existing Signal + quality card threshold). + - With zero reports: `no client reports in 7 days`. +- Drop the word "verified". +- Use `metrics.py`'s existing window idiom: + `julianday(observed_at) > julianday('now', '-' || ? || ' days')`. + +**Test.** `tests/`: a metrics unit test with a seeded in-memory DB that +includes structural, `local_llm` and `client_outcome` rows, asserting only +`client_outcome` rows are counted and that `decisions` comes from +`route_decisions`. A frontend static test that the tile no longer reads +`verdict_mix`. + +## 5. G3: Warning dismissals don't stick on rate warnings + +**Problem.** `admin/frontend/index.html:1010-1040`: a dismissal is keyed on +the warning's exact text. That's deliberate ("when the underlying numbers move +the text moves with them... the warning comes back"). But warnings that embed +a live rate or count ("3.2x pace", "n=4") change text on nearly every poll, so +a dismissal never holds. + +**Fix (interim, until structured warnings exist).** Key each dismissal on +`digit-normalized text + severity`: replace every run of digits (and decimal +points inside a number) with `#`, then append `|` + `warningSeverity(text)`. +The warning comes back when its wording or its severity changes, not when a +digit moves. This is the same normalization `rejection_warnings` already uses +for grouping (CLAUDE.md, the reactive detector section). +- Migrate the existing stored keys on load: normalize each and dedupe, same + pattern as the existing one-time carry-over from `dismissedWarnings`. +- Update the block comment to state the new rule and why it changed. + +**Test.** A static or JS-extracted unit test: two texts differing only in +digits produce the same key; the same text at different severity produces a +different key. + +## 6. C1: One knob table instead of two columns + +**Problem.** `admin/frontend/controls.html`: "Runtime Knobs" (`#runtime-list`, +`renderRuntime` ~line 687) and "Persisted Config" (`#config-list`, +`renderConfig` ~line 821) list largely the same knobs in two side-by-side +cards, in different orders, and persisted labels truncate +(`objective.incumbent_cache_pri...`). Checking whether a runtime value has +drifted from its persisted value means matching rows by eye. + +**Fix.** +- One card, **Knobs**, full width, as a table: + `knob | live | persisted | layer | ` with the columns meaning: + - `knob`: the dotted config key, never truncated (wrap if needed). + - `live`: the runtime control, if the knob has one, else empty. + - `persisted`: the persisted-config control, if the key is in the persisted + allowlist, else empty. + - `layer`: the existing `base` / `overlay` badge. + - A **drift** badge in the row when live differs from persisted + ("reverts on restart"). +- Pair rows through the runtime-to-config mapping that already exists in + `src/admin.py` (`_BOOL_KNOBS` at ~line 348, plus the non-bool runtime knob + tables beside it). Expose that mapping to the frontend: add the dotted + config key to each item in the `GET /admin/api/runtime` response, rather + than duplicating the table in JS. +- Keep the existing save semantics exactly: runtime controls POST to + `api/runtime/{knob}` immediately, as today; persisted controls stay dirty + until the Save button, as today. Keep the existing hint badges (for example + "enables the reuse window"). +- Keep the help lines ("Live on the next request...", "Applied on + restart...") as one line each, placed as column-header tooltips or a single + line under the table. +- Don't touch the classifier-card region (lines ~1000-1160); another branch + edits it. + +**Test.** `test_admin_knob_coverage.py` must still pass unchanged. Add a +test that every runtime knob in the API response carries a config key, and +that the key resolves to a real `RouterConfig` path. Visual check on 8081: +set one runtime knob so it differs from persisted (on the 8081 instance +only), then confirm the drift badge shows. Screenshot the whole table and a +zoomed crop of one drifted row. + +--- + +## Done means + +- Six commits on `feat/cockpit-quick-wins`, one per item, in the order above. +- `pytest` and `ruff check` green on the worktree. +- Screenshots from 8081 for items 1, 3, 4 and 6. +- The 8081 instance stopped, by the PID you started. +- A short report listing each commit, the tests added, and anything + deferred, with the reason. diff --git a/plans/deploy-separation.md b/plans/deploy-separation.md new file mode 100644 index 0000000..caff23e --- /dev/null +++ b/plans/deploy-separation.md @@ -0,0 +1,167 @@ +# Deploy separation: production runs from its own clone, not the dev checkout + +Status: planned -- spec and runbook ready, dry run done (see "Dry run"); the cutover is an operator step not yet run +Date: 2026-09-26 +Why: incident #8, and #4, #5 and #7 before it. Production has been one `git pull`, +one uncommitted edit or one agent session away from the dev checkout. + +## Current state (measured 2026-09-26) + +- **Every installed user unit runs from `~/Sources/6krrt`**, the dev checkout: + - `llm-router.service` + - `llm-router-poller`, `-seed`, `-backup`, `-offsite` + - `llm-router-baseline-report` +- **What they take from it:** `WorkingDirectory`, `PYTHONPATH`, + `EnvironmentFile=.env`, `.venv/bin/*`, and `HF_HOME=.hf-cache`. + `ReadWritePaths` covers the whole source tree, so the live service can write + anywhere in the repo. +- **Everything else follows the working directory,** because the code resolves + paths against it: `config/config.yaml`, the `config/config.local.yaml` + overlay, and `database.path: "router.db"`. Moving the working directory moves + all of it, with no code change. +- **The repo's own unit templates** (`deploy/*.service`) already target a + separate `%h/llm-router` directory, except `-backup` and `-offsite`. The split + was the original design. The installed units drifted onto the dev checkout, + and `~/llm-router` does not exist. +- **Sizes:** `router.db` 48 MB, `.hf-cache` 2.8 GB, `.venv` 7.2 GB. `.env` + holds `NEURALWATT_API_KEY` and `OPENROUTER_API_KEY`. + +## Target layout + +``` +~/.local/share/6krrt/app production clone of the Gitea repo, detached at a + merged origin/main commit. Nobody edits here; no + agent works here; nothing here ever runs `git clean`. + router.db live DB (gitignored; checkout never touches it) + config/config.local.yaml live overlay (gitignored) + .env live keys, mode 600 (gitignored) + .venv/ production venv, built from the clone's requirements +~/.cache/6krrt-hf shared HF model cache (moved once from + ~/Sources/6krrt/.hf-cache; production and the dev + sandbox both point HF_HOME here) +~/Sources/6krrt dev checkout: source only. No live DB, no + production .env. The 8081 sandbox uses DB copies. +``` + +Units: every `llm-router*` unit's paths change from `%h/Sources/6krrt` to +`%h/.local/share/6krrt/app`, and `ReadWritePaths` narrows to that directory. +The backup and offsite scripts already take `REPO` from the environment; set it +to the app directory. + +## `scripts/deploy.sh` (to build; not needed for the first cutover) + +The only way code reaches production after the cutover: +1. **Show the change.** `git -C $APP fetch origin`, then + `git log --oneline HEAD..origin/main`. +2. **Preflight, without touching the running service:** + - load `origin/main`'s config against the real overlay (catches a #7-style + rename) + - check the requirements are satisfied in `$APP/.venv` + - `import dispatcher` from a temp export of `origin/main` +3. **Cut over.** Record the current SHA, `git checkout --detach origin/main`, + `uv pip install` only if the requirements changed, restart, and poll + `/health` for 30 s. +4. **Roll back on failure.** Check out the recorded SHA, restart, and alert. +5. **Record the deployed SHA** in `$APP/.deployed`, and expose it on `/health` + if that endpoint gains a field, so "what is running" has one answer. + +## Cutover runbook + +Downtime is about 1 minute. The venv is built BEFORE stopping anything. Run +the commands from a shell, not from an agent session. + +```sh +APP=~/.local/share/6krrt/app +DEV=~/Sources/6krrt + +# 0. Build the clone and venv while production keeps running. +mkdir -p ~/.local/share/6krrt +git clone ssh://git@git.adlee.work:2222/alee/6krrt.git "$APP" +git -C "$APP" checkout --detach origin/main +(cd "$APP" && uv venv .venv && uv pip install --python .venv/bin/python \ + -r requirements.txt -r requirements-encoder.txt) + +# 1. Share the HF cache (one move; the dev sandbox follows in step 6). +mkdir -p ~/.cache && mv "$DEV/.hf-cache" ~/.cache/6krrt-hf + +# 2. Stop production and every timer that touches the DB. +systemctl --user stop llm-router-poller.timer llm-router-seed.timer \ + llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer +systemctl --user stop llm-router.service + +# 3. Copy state; originals stay put until step 7. +sqlite3 "$DEV/router.db" ".backup '$APP/router.db'" +command cp -p "$DEV/config/config.local.yaml" "$APP/config/config.local.yaml" +command cp -p "$DEV/.env" "$APP/.env" && chmod 600 "$APP/.env" +sqlite3 "$APP/router.db" 'PRAGMA integrity_check;' # must print: ok + +# 4. Repoint the units (backups first), then reload. +cd ~/.config/systemd/user +mkdir -p ~/.local/share/6krrt/unit-backup && command cp -p llm-router* ~/.local/share/6krrt/unit-backup/ -r +sed -i "s#%h/Sources/6krrt#%h/.local/share/6krrt/app#g; s#/home/alee/Sources/6krrt#/home/alee/.local/share/6krrt/app#g" llm-router*.service +sed -i "s#^Environment=HF_HOME=.*#Environment=HF_HOME=%h/.cache/6krrt-hf#" llm-router.service +systemctl --user daemon-reload + +# 5. Start and verify. +systemctl --user start llm-router.service +sleep 3; curl -s localhost:8080/health +systemctl --user show llm-router.service -p WorkingDirectory +curl -s localhost:8080/admin/api/runtime | head -c 200 # pinch persisted false = overlay loaded +systemctl --user start llm-router-poller.timer llm-router-seed.timer \ + llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer + +# 6. Point the dev sandbox at the shared HF cache and the new live DB. +# (edit scripts/sandbox.sh: HF_HOME and the live-DB path; see "Follow-ups") + +# 7. Only after a day of healthy running: retire the dev copies. +mv "$DEV/router.db" "$DEV/router.db.pre-cutover-$(date +%Y%m%d)" +``` + +**Rollback, before step 7:** restore the unit backups from +`~/.local/share/6krrt/unit-backup/`, run `daemon-reload`, start the service. +The dev copies of the DB, overlay and `.env` were never moved, so production is +back where it started. Any decisions logged after the cutover stay in the +app-directory DB. + +## Follow-ups that must land with or right after the cutover + +- **`CLAUDE.md` North Star #3** and the "Measure the live `router.db`" memory + name `/home/alee/Sources/6krrt/router.db` as the live DB. Update both to + `~/.local/share/6krrt/app/router.db`, or agents will measure a stale file. +- **`scripts/sandbox.sh`:** `HF_HOME` goes to `~/.cache/6krrt-hf`. Its DB copy + instructions change to the app directory's DB. +- **`deploy/*.service` templates:** make the repo's templates say + `%h/.local/share/6krrt/app` so reinstalling from the repo reproduces this + layout, and fix `-backup` and `-offsite` to match. +- **`deploy/README.md`:** replace the install section with this layout and + `deploy.sh`. +- **The admin portal's "Restart Service" trigger** keeps working (it calls + `systemctl`), but any trigger that runs a script by repo-relative path should + be checked against the new working directory. + +## Dry run + +See the result appended below. + +**Result, 2026-09-26 04:3x, against `origin/main` `827c408`:** +- **Clone from Gitea and detached checkout:** OK. +- **Fresh venv, `uv venv` plus both requirements files:** 2 s (all from `uv`'s + cache); `torch 2.9.1+cu128` and `transformers` import. Building the venv + does not affect downtime, and step 0 can run hours ahead. +- **State:** the DB copied by sqlite backup with the source opened `mode=ro`, + and `integrity_check` = ok; overlay copied; `.env` replaced with dummy keys + for the dry run. +- **Served on 8081** from the clone. `/proc//cwd` confirmed the clone + directory, with `HF_HOME` on the dev cache read-only: + - `/health` 200; `/admin/controls` 200 + - `/admin/api/runtime` showed `pinch_enabled` `persisted: False`, so the + copied overlay loaded + - the snapshot read the copied DB: 3,994 decisions, 200 client reports over 7 + days + - `/route` considered 41 candidates and picked `deepseek/deepseek-v4-flash` +- **Stopped** by the recorded PID; 8081 free; the dummy `.env` deleted. + +**Not exercised by the dry run:** the systemd unit edits (step 4); the timers +under the new `ReadWritePaths`; the HF cache move (step 1); and real provider +keys. The unit edit is the step to watch during the cutover. `systemctl --user +show llm-router.service -p WorkingDirectory` in step 5 confirms it took. diff --git a/plans/docs-lift-a-pass.md b/plans/docs-lift-a-pass.md new file mode 100644 index 0000000..d0a3111 --- /dev/null +++ b/plans/docs-lift-a-pass.md @@ -0,0 +1,105 @@ +# Docs pass after Lift A (PR #102) + +Status: done -- executed on branch docs/lift-a-docs-pass +Date: 2026-09-26 +This brief is the contract. Keep the plan small: one todo per doc, each with +exact section anchors. Do NOT read `plans/no-progress-detection.md` (reference +only). + +## Goal + +Bring the docs in line with what PR #102 shipped and what is running now. The +docs describe the code as it is on this branch; where a doc and the code +disagree, the code wins, and the todo says which line it checked. + +## Where the work happens + +- Worktree: `/home/alee/Sources/6krrt-worktrees/docs-lift-a-pass`, branch + `docs/lift-a-docs-pass`, based on `origin/main` `a24de1b` plus three commits + that already carry North Star #4 in `CLAUDE.md`, the incident #8 rewrite in + `docs/incidents.md`, and the plans. Build on them; do not rewrite them. +- **Docs only.** No code, config, test or unit-file changes. +- **Do not touch** `README.md`, `AGENTS.md`, `deploy/README.md`: the owner has + uncommitted edits to them in the main checkout. Anything that belongs in + `deploy/README.md` goes in `docs/watchdog.md` instead, and the done report + lists it for the owner to fold in. + +## What shipped (verify each against the code before writing it) + +- `src/progress_detect.py`: pure detector. Signals dup, top, slow, coverage; + landed gating; read-only agents; canonical (sorted-key) args; bash command + normalization (comment lines, `cd dir &&`, `VAR=x` prefixes dropped). +- `src/watchdog.py` + `deploy/llm-router-watchdog.{service,timer}`: 5-minute + oneshot. Judges only sessions with a tool call since the last tick, on their + full history, per session; landed counts the session's own descendants only. + Alert lifecycle in `watchdog_alerts`: trigger once, escalate once (LLM yes or + 3 ticks), resolve once. Quiet `no_opencode` tick when opencode is closed. + Optional local-model second opinion (does not veto). +- `src/notifier.py`: `desktop` channel (`notify-send`), PagerDuty-Events-v2 + shaped `AlertEvent`, per-channel `min_severity`, rate limit backed by the DB. +- `src/watchdog_store.py`: tables `watchdog_ticks`, `watchdog_verdicts`, + `watchdog_alerts`, `watchdog_channel_settings`. +- Admin: Loops panel (`index.html#loops`), Watchdog card on Controls (last + tick, verdicts, channels, Send test alert), per-model stall rollup and + Block/Unblock on Models. New availability value `blocked`, excluded from + routing and from pinned requests via `_admin_excluded_models`. +- Endpoints: `GET /admin/api/watchdog/status`, `POST /admin/api/watchdog/channels`, + `POST /admin/api/watchdog/test-alert`; availability endpoint accepts `blocked`. +- Config: `watchdog.*` (incl. `detector.*`, `dashboard_base_url`), + `notifications.channels`. Watchdog knobs are persisted-only (separate process). +- `deploy/opencode-plugin/router-link.js`: the `parentCache` export made + opencode 1.18 reject the whole plugin ("Plugin export is not a function" in + `~/.local/share/opencode/log/opencode.log`), so identity headers never + reached the router from #99 until this fix. A loader-contract test now asserts + every export is a function. +- `scripts/progress_backtest.py` (`--fixture` replays the 15 labelled sessions; + expect 8/15 flagged) and `scripts/export_progress_fixture.py`. +- Operator state 2026-09-26: timer enabled ~15:18 EDT with unit paths repointed + from `%h/llm-router` to `%h/Sources/6krrt` (production still runs from the + dev checkout; see `plans/deploy-separation.md`); plugin fix installed and + opencode restarted 15:19; router restarted 15:19:57. + +## Todos (one per doc) + +1. **New `docs/watchdog.md`**: what it catches and does not (steady spend is not + waste), signals and thresholds, alert lifecycle, the admin surfaces, knobs, + install (the `sed` repoint + `enable --now`), how to run `--once`, backtest, + re-exporting the fixture, and known limits: sessions under 40 calls are + invisible; the Atlas must-flag case sits exactly at `cover_min` 4.0; the + stale `localhost:4096` entry in `rc-servers.json` logs a harmless probe + warning. +2. **`CLAUDE.md`**: "What's built and working" (anchor `## What's built and + working`): add the modules above in the existing one-bullet style, linking + `docs/watchdog.md`; update the `tests/` bullet count from the real + `pytest --collect-only -q` total. "Run as a service" (anchor `## Run as a + service`): one short paragraph on the watchdog timer. Do not edit the North + Star section. +3. **`docs/incidents.md` #8**: under Status, add that Lift A shipped (PR #102, + `a24de1b`) and the timer is live; add the plugin-loader finding as Evidence + (with the log line) and to the Human response timeline; add a symptom-table + row "opencode plugin does nothing -> grep opencode.log for `failed to load + plugin`". Keep the Evidence / labelled Theory convention in the file header. +4. **`docs/admin-portal.md`**: sections for the Loops panel, the Watchdog card, + and the Models rollup with Block/Unblock (where `blocked` differs from + `deprecated`: an operator's reasoned stop, one click to undo). +5. **`docs/api.md`**: the three watchdog endpoints and `blocked` on the + availability endpoint (request/response shapes read from `src/admin.py`). +6. **`docs/data-model.md`**: the four watchdog tables (from `config/schema.sql`), + and `blocked` in the `availability` row at the `models` section (line ~34) + and wherever `admin_model_overrides` is described. +7. **`docs/clients.md`** (anchor `## Session-directory attribution (opencode + plugin)`) and `docs/api.md` near the `X-Router-*` table (line ~151): the + plugin loader contract (every export must be a function) and how to confirm + the plugin loaded. + +## Rules + +- ASCII only; no middle dots. Match each file's existing style and heading + depth. Admin-portal docs describe the UI, they do not add prose to it. +- Numbers and names come from the code on this branch, never from this brief. +- Commit per todo, by explicit path. Never `git add -A`. +- Do not push and do not open a PR. The done report lists commits and the + items for `deploy/README.md`. +- If a worker's context passes ~150k tokens, stop it and start a fresh one. + +Write `.omo/plans/docs-lift-a-pass.md`, then stop and report. diff --git a/plans/no-progress-detection-prototype.py b/plans/no-progress-detection-prototype.py new file mode 100644 index 0000000..9268b8b --- /dev/null +++ b/plans/no-progress-detection-prototype.py @@ -0,0 +1,241 @@ +#!/usr/bin/env python3 +"""Interim opencode loop watcher (stand-in until the no-progress lift ships). + +Signals per session, over a sliding window of its last WINDOW tool calls: + dup share of calls that repeat an earlier (tool, args) in the window + top count of the most repeated (tool, args) in the window + landed any edit/write with a non-empty diff, or a `git commit` with exit 0, + by the session OR any descendant within the window's time span + +Flag when the window is full enough and nothing landed and + dup >= DUP_MIN or top >= TOP_MIN. +Read-only agents (explore/librarian/oracle) never land, so for them only +top >= TOP_MIN_RO counts. + +Modes: + backtest : replay every session's history, report worst window per session + watch : poll live sessions; on a new flag print it, notify-send, exit 0 +""" +import json, sys, time, subprocess, urllib.request, collections, os + +U = os.environ.get("OC_URL", "http://127.0.0.1:4097") +WINDOW, MIN_CALLS = 60, 40 +DUP_MIN, TOP_MIN, TOP_MIN_RO = 0.25, 12, 8 +CUM_MIN = 15 # one exact call repeated this often over the whole session: a slow loop +COVER_MIN = 4.0 # lines read from one file in the window / its length: re-reading, whatever the offsets +READ_ONLY = ("@explore", "@librarian", "@oracle", "explore subagent") + + +def get(path): + return json.load(urllib.request.urlopen(U + path, timeout=30)) + + +def calls_of(sid): + out = [] + for m in get(f"/session/{sid}/message"): + for p in m["parts"]: + if p.get("type") != "tool": + continue + st = p.get("state") or {} + inp = st.get("input") or {} + md = st.get("metadata") or {} + t = (st.get("time") or {}).get("start") or m["info"].get("time", {}).get("created", 0) + tool = p.get("tool") + landed = (tool in ("edit", "write", "patch") and bool(md.get("diff"))) or ( + tool == "bash" and "git commit" in (inp.get("command") or "") and md.get("exit") == 0) + out.append((t, tool, json.dumps(inp, sort_keys=True), landed)) + return out + + +import re +_PATHISH = re.compile(r"[\w./-]+\.(?:json|md|py|js|html|yaml|yml|toml|txt|sql|sh|mjs|css)\b") + + +_CD_PREFIX = re.compile(r"^\s*cd\s+\S+\s*(?:&&|;)\s*") +_ENV_PREFIX = re.compile(r"^\s*(?:[A-Z_][A-Z0-9_]*=\S+\s+)+") + + +def _normalize_bash(cmd): + """Drop comment lines and leading `cd dir &&` / `VAR=x` prefixes, so + differently-worded commands under one prefix do not collapse to one key.""" + lines = [ln for ln in cmd.splitlines() if not ln.lstrip().startswith("#")] + c = " ".join(lines).strip() + for _ in range(3): + c2 = _ENV_PREFIX.sub("", _CD_PREFIX.sub("", c)) + if c2 == c: + break + c = c2 + return c + + +def target_of(tool, args_json): + """Coarse target: the file a call is about, ignoring offsets and the exact + command wording, so 28 differently-worded `cat boulder.json` count as one.""" + try: + a = json.loads(args_json) + except Exception: + return (tool, args_json[:80]) + if a.get("filePath"): + # A read keys on its slice: walking a big file in pieces is progress, + # re-reading the same slice is not. + return ("file", os.path.basename(a["filePath"]), a.get("offset"), a.get("limit")) if tool == "read" \ + else ("file", os.path.basename(a["filePath"])) + if tool == "bash": + cmd = _normalize_bash(a.get("command") or "") + files = sorted({os.path.basename(f) for f in _PATHISH.findall(cmd)}) + if files: + # Numbers in the command are the slice (sed -n 100,160p, head -40): + # walking a file is progress; the same slice again is not. + nums = tuple(sorted(set(re.findall(r"\b\d+\b", cmd)))) + return ("bash-files", ",".join(files[:3]), nums) + return ("bash", " ".join(cmd.split()[:2])) + if tool in ("grep", "glob"): + return (tool, str(a.get("pattern"))[:60]) + return (tool, args_json[:80]) + + +def window_stats(win): + fp = collections.Counter((c[1], c[2]) for c in win) + dup = sum(v - 1 for v in fp.values() if v > 1) / max(1, len(win)) + tg = collections.Counter(target_of(c[1], c[2]) for c in win) + (top_k, top_n) = tg.most_common(1)[0] if tg else (("", ""), 0) + return dup, top_n, top_k + + +_LEN = {} + + +def file_lines(path): + if path not in _LEN: + try: + with open(path, "rb") as fh: + _LEN[path] = max(1, fh.read().count(b"\n")) + except OSError: + _LEN[path] = None + return _LEN[path] + + +def coverage(win): + """Max over files of (lines requested in the window / file length).""" + req = collections.Counter() + for c in win: + if c[1] != "read": + continue + a = json.loads(c[2]) + fp = a.get("filePath") + n = file_lines(fp) if fp else None + if not n: + continue + req[fp] += min(a.get("limit") or n, n) + best = max(((v / file_lines(k), k) for k, v in req.items()), default=(0.0, "")) + return round(best[0], 1), os.path.basename(best[1]) + + +def is_ro(title): + t = (title or "").lower() + return any(k in t for k in READ_ONLY) + + +def evaluate(sess, calls, landed_times_tree, end_idx=None): + calls = calls if end_idx is None else calls[:end_idx] + win = calls[-WINDOW:] + if len(win) < MIN_CALLS: + return None + t0, t1 = win[0][0], win[-1][0] + landed = any(t0 <= lt <= t1 for lt in landed_times_tree) + dup, top_n, top_k = window_stats(win) + ro = is_ro(sess.get("title")) + cum_k, cum_n = collections.Counter((c[1], c[2]) for c in calls).most_common(1)[0] + slow = (not landed) and cum_n >= CUM_MIN + cov, cov_file = coverage(win) + reread = (not landed) and cov >= COVER_MIN + if reread and not (top_n >= TOP_MIN): + top_n, top_k = int(cov), ("coverage", f"{cov_file} read {cov}x its length") + if ro: + flag = top_n >= TOP_MIN_RO or slow or reread + else: + flag = ((not landed) and (dup >= DUP_MIN or top_n >= TOP_MIN)) or slow or reread + if slow and not (top_n >= TOP_MIN): + top_n, top_k = cum_n, ("exact-total", cum_k[0], cum_k[1][:80]) + return dict(flag=flag, dup=round(dup, 2), top=top_n, top_what=" ".join(str(x) for x in top_k)[:120], + landed=landed, ro=ro, n=len(calls), t1=t1) + + +def tree_landed(sessions, calls_by): + kids = collections.defaultdict(list) + for s in sessions: + if s.get("parentID"): + kids[s["parentID"]].append(s["id"]) + memo = {} + def lt(sid): + if sid in memo: + return memo[sid] + own = [c[0] for c in calls_by.get(sid, []) if c[3]] + for k in kids.get(sid, []): + own += lt(k) + memo[sid] = own + return own + return lt + + +def backtest(since_ms): + sessions = [s for s in get("/session") if s.get("time", {}).get("updated", 0) >= since_ms] + calls_by = {s["id"]: calls_of(s["id"]) for s in sessions} + lt = tree_landed(sessions, calls_by) + for s in sessions: + cs = calls_by[s["id"]] + worst, first_flag = None, None + for i in range(MIN_CALLS, len(cs) + 1, 5): + r = evaluate(s, cs, lt(s["id"]), i) + if r and (worst is None or (r["flag"], r["dup"], r["top"]) > (worst["flag"], worst["dup"], worst["top"])): + worst = r + if r and r["flag"] and first_flag is None: + first_flag = i + if worst: + print(f"{'FLAG' if worst['flag'] else 'ok '} calls={len(cs):4d} first_flag_at={first_flag} " + f"dup={worst['dup']} top={worst['top']} landed={worst['landed']} ro={worst['ro']} | {s.get('title','')[:50]}" + + (f"\n top: {worst['top_what'][:110]}" if worst['flag'] else "")) + + +def watch(minutes, state_path): + seen = set() + if os.path.exists(state_path): + seen = set(json.load(open(state_path))) + deadline = time.time() + minutes * 60 + while time.time() < deadline: + try: + now = time.time() * 1000 + sessions = get("/session") + status = get("/session/status") + recent = [s for s in sessions if s["id"] in status or now - s.get("time", {}).get("updated", 0) < 5 * 60e3] + calls_by = {s["id"]: calls_of(s["id"]) for s in recent} + # Judge only sessions doing work now: busy, or a tool call in the + # last 5 min. A session's "updated" stamp moves on abort, restart + # and title edits, so it cannot stand in for activity. + active = [s for s in recent if s["id"] in status + or (calls_by[s["id"]] and now - calls_by[s["id"]][-1][0] < 5 * 60e3)] + lt = tree_landed(sessions, calls_by) + for s in active: + r = evaluate(s, calls_by[s["id"]], lt(s["id"])) + key = f"{s['id']}:{len(calls_by[s['id']]) // 60}" + if r and r["flag"] and key not in seen: + seen.add(key) + json.dump(sorted(seen), open(state_path, "w")) + msg = (f"{s.get('title','')[:60]} ({s['id']}): {r['n']} calls, dup {r['dup']}, " + f"top x{r['top']}: {r['top_what']}, landed={r['landed']}, busy={s['id'] in status}") + print("LOOP SUSPECT:", msg) + subprocess.run(["notify-send", "-u", "critical", "opencode loop suspected", msg[:250]], + check=False, timeout=5) + return 0 + except Exception as e: # opencode down or restarting: keep watching + print(f"[{time.strftime('%H:%M:%S')}] poll error: {e}", file=sys.stderr) + time.sleep(120) + print("no loops seen in", minutes, "min") + return 0 + + +if __name__ == "__main__": + if sys.argv[1] == "backtest": + backtest(int(sys.argv[2])) + else: + sys.exit(watch(int(sys.argv[2]), sys.argv[3])) diff --git a/plans/no-progress-detection.md b/plans/no-progress-detection.md new file mode 100644 index 0000000..2a98575 --- /dev/null +++ b/plans/no-progress-detection.md @@ -0,0 +1,572 @@ +# No-progress detection: catch an agent that is spinning, not one that is working + +Status: reference -- parent spec; Lift A shipped in PR #102, Lift B (router-side judge, modes, behavior breaker) not started +Date: 2026-09-26 +Trigger: the `cockpit-quick-wins` Atlas run, 2026-09-25 19:00 to 2026-09-26 02:00. +Incident write-up: `docs/incidents.md` #8. + +## What happened, measured + +Read-only against the live `router.db` and opencode's own session store +(`http://127.0.0.1:4097`), 2026-09-26. + +Spend, 19:00 to 02:00 local: +- 323M prompt tokens (95% cached) and about $7.81, over roughly 8 hours. +- 79% of the money went to OpenRouter `z-ai/glm-5.3-flash`: 1,368 calls + averaging 203k prompt tokens. +- 921 calls carried 120k+ prompt tokens and cost $6.57. NeuralWatt `glm-5.3` + took 47 calls averaging 673k context, max 682,679, because nothing else fit. +- Every decision was `classification_source = classifier` on the default + profile. The router did what it was told; nothing was pinned. + +**Spend rate is the wrong signal.** Tonight's worst hour ($2.42) and worst +3-hour window ($4.13) sit inside the prior 30 days' normal range (hourly p90 +$1.46, max $4.83; worst 3h $9.65). A rate alarm quiet enough for a healthy +all-day agent run would have missed this entirely, and one tight enough to catch +it would fire on good work. The operator's framing is the right one: steady +usage is fine. What is not fine is spend with **no concrete change landing**. + +**The sessions that wasted money are visibly different in their tool calls:** + +| session (opencode) | tool calls | exact-duplicate calls | worst repeat | landed a change? | +|---|---|---|---|---| +| Item 2 worker, 2nd attempt | 423 | 30% | `read navbar.js` x61 | no (worktree unchanged) | +| Item 3 worker, 1st attempt | 132 | 23% | `read test_admin_js_units.py` x16 | no | +| Item 4 worker | 134 | 19% | `read .omo/plans/...` x8 | partly (code, no tests) | +| Item 2 worker, 1st attempt | 307 | 17% | `read navbar.js` x21 | no | +| Atlas orchestrator | 576 | 11% | playwright console check x10 | yes (via workers) | +| healthy workers (items 1, 5, 6, 3-retry) | 30-96 | 0-6% | 1-3 | yes | + +**A second loop shape, 2026-09-26 02:21 to about 02:56:** the Atlas orchestrator +itself, after its plan was complete: +- `.omo/boulder.json` carried a stale `status: completed` and `pr_url` from the + previous plan. +- oh-my-openagent's continuation hook injected "continue" turns; all 3 user + turns in the window were injected. +- Atlas, working from a trimmed context, re-derived its state every time: + 54 turns, about 19M cached-read tokens, `boulder.json` read 28 times, the + plan's checkboxes grepped 16 times, nothing landed. + +The operator caught it only by watching. The judge must treat this as +`stalling`, and Phase 0 must include it as a calibration case. The plugin can +also see injected continuation turns (`chat.message`), which is a useful extra +signal: "continuation turns keep arriving and nothing lands" is this shape +exactly. + +**A third shape, 2026-09-26 around 03:05 to 03:17: a read-only agent.** +Prometheus's `explore` subagent reached 322 tool calls, 19% exact duplicates, +and read `deploy/opencode-plugin/router-outcome.js` 20 times. The file had been +deleted from under it: the production checkout was synced to `origin/main`, +where #99 retired it. The subagent was aborted. + +Read-only agents (`explore`, `librarian`, `oracle`, and the like) never land a +change by design, so "no landed change" cannot be their signal. For them the +judge uses: +- target novelty: new files or queries per N calls, and +- a failure streak on the same target, such as repeated reads of a path that + no longer exists. + +Identify them by the `X-Router-Agent` value, via a configurable list of +read-only agent names. + +"Exact duplicate" means the same tool with byte-identical arguments. Duplicate +share alone does not separate the orchestrator (11%, productive) from a stuck +worker (17%). Duplicates **plus** no landed change does. + +## Why nothing caught it + +- `circuit_breaker.py` is availability-only: it trips on provider 5xx. +- The only spend alarm is `runway_low_warning`: balance / burn < 6 h, with burn + averaged over a 24 h balance window. OpenRouter read $24.73 at $0.25/h, so 97 h + of runway. It answers "will I run out", not "is this being wasted". +- The router cannot see progress at all. It sees prompts, not whether an edit + landed or a command keeps failing. The opencode plugin sees both. +- The DEPLOYED router cannot tell conversations apart either. `session_key` + merges concurrent conversations under one multi-agent client, and `/outcome` + attribution falls back to a directory match that 409s when several sessions + share a cwd, which is exactly the Atlas-plus-workers shape. **The fix already + exists on `origin/main` (PR #99, conversation-identity) but is not deployed:** + production runs from a local checkout 29 commits behind `origin/main`, the + live `route_decisions` has no `agent` / `parent_key` columns, and the + installed plugin is still `router-outcome.js` rather than #99's + `router-link.js`. + +## Design + +Split the job where the information lives. **The plugin is the sensor; the +router is the judge.** No raw task text leaves opencode, and the router never +stores any; the "never store raw task text" rule is unchanged. + +### 1. Exact session identity: ALREADY BUILT upstream (#99), build on it + +PR #99 (`conversation-identity`, merged to `origin/main` as `fe8c813`) did this. +Reuse it; do not rebuild it: +- **Plugin.** `deploy/opencode-plugin/router-link.js` replaces + `router-outcome.js`. Its `chat.headers` hook sends `X-Router-Conversation` + (the opencode session id), `X-Router-Agent` (slugged to the router's charset) + and `X-Router-Parent`, with a bounded parent-lookup cache. +- **Router.** `src/conversation_identity.py` resolves identity from those + headers. `session_key` becomes `'c:' + conversation id`. `route_decisions` + gained `agent` and `parent_key` columns. `/outcome` attributes to the exact + conversation. + +What remains for this lift: +- **Deploy #99.** It is merged but not running. Syncing the production checkout + to `origin/main` and installing `router-link.js` are operator steps: the main + checkout has uncommitted work, and part of it, the `confidence_min` rename, + also landed upstream via #100. +- **Fix the exit code in `router-link.js`.** It kept `router-outcome.js`'s + `output?.exitCode ?? output?.exit_code` (line 83 on `origin/main`); see "Fixes + that ride along". +- **Build the progress events (section 2) into `router-link.js`,** keyed by the + same conversation and parent ids, so the judge rolls children up to parents + through `parent_key`. + +Everywhere below, "session" means #99's conversation. Where this spec says +`X-Opencode-Session`, read `X-Router-Conversation`. + +### 2. Progress events (plugin `tool.execute.after` -> router `POST /progress`) + +The plugin sends one small event per tool call, fire-and-forget with a short +timeout, never blocking the session: + +``` +{session_id, parent_session_id, agent, tool, call_id, + fingerprint, # sha256 of tool + normalized args; never the args themselves + target, # coarse, normalized: the file path for read/edit/write, the + # command's first two tokens for bash; no contents + exit, # bash: output.metadata.exit + landed} # bool, see below +``` + +`landed` is the load-bearing definition, and it is deliberately narrow: +- `edit` / `write` / `patch` with a non-empty `metadata.diff` +- `bash` whose command is `git commit` (or `git commit --amend`) with exit 0 +- A test or lint command (the plugin's existing `TEST_COMMAND` list) that PASSES + after the session's previous run of the same fingerprint FAILED, i.e. a fix + landed. + +Reads, greps, todo writes, sub-agent dispatches and passing-again tests are not +progress. An orchestrator's progress is its children's: roll child `landed` +events up to the parent via `parent_session_id`. + +Storage: a new `progress_events` table, pruned to a rolling window (knob). It +holds fingerprints, targets and booleans, no content. + +### 3. The judge (router, per client session) + +Signals per session, over a rolling window of its last N tool events: +- `turns_since_landed`: LLM requests since the last landed event (self plus + children for a parent). +- `usd_since_landed`: billed spend since then, joined through `request_id`. +- `dup_share`: exact-duplicate fingerprint share in the window. +- `fail_streak`: consecutive nonzero exits on the same fingerprint (the + "sandbox is broken and the agent keeps retrying" case). +- `context_tokens`: the latest request's size, to report alongside, since a + stall at 600k context costs far more per turn than one at 30k. + +Verdicts: +- `ok`: something landed recently. +- `stalling` (warn): no landed change for `stall_turns` turns or + `stall_usd` dollars, AND (`dup_share >= dup_warn` OR `fail_streak >= fail_warn`). +- `looping` (act, if enabled): the stall condition at the higher `loop_*` + thresholds. + +Both conditions have to hold. Long-running healthy work with no duplicates (a +big read-heavy exploration) stays `ok` until it hits the plain turns/spend +ceiling, which is a separate, looser knob. + +**Calibrate before trusting it.** Phase 0 replays the judge offline over +opencode's session store (the 11 sessions above plus older history), and reports +the verdict each session would have received and when. Thresholds are picked +from that, not guessed. The Atlas orchestrator (11% duplicates, productive +through children) is the key false-positive case, and it must stay `ok`. + +### 4. Optional: a local model as a second opinion + +The deterministic signals should carry this. Where they are ambiguous, for +example near-duplicates (the same file read at shifting offsets, or the same +failing test with a different `-k`), there are two cheaper steps before any +model: +- normalize `read` fingerprints to the path, and bash to command plus target +- the local encoder already loaded for `classifier.mode: local_encoder` can + embed the `target` strings to cluster "similar-ish" events + +A local LLM judge ("does this digest look like a loop?") runs only on a session +already flagged `stalling`, on a digest of tool names, targets and exit codes, +never content. It is gated by `local_compute.enabled`, since the operator pays +the local power bill. Treat this as a later phase, justified by Phase 0 showing +ambiguous cases the deterministic signals miss. + +### 5. What happens on a detection: modes + +One knob, `progress.mode`, picks the response. **Every mode warns:** +- A `/metrics` warning (admin bell and TUI). It names the session, agent, turns + and $ since the last landed change, and the top repeated target. +- An opencode toast via `POST /tui/show-toast` in the window the operator is + watching. +- An **alert** through the router's notifier (section 6), which reaches the + operator when nobody is looking at either screen. + +| mode | on detection | +|---|---| +| `warn` | Warnings only. The neutral default: today's behaviour plus a warning. | +| `auto_compact` | Compact the stuck session with the incident report injected, then continue. No refusal, no orchestrator stop. | +| `auto_limit_context` | Cap the stuck session's context. No stop. | +| `auto_recover` | Two stages, below. | + +**`auto_recover`, stage 1 (first detection in a session tree):** +1. The router refuses the stuck session's chat requests (429, scoped by + `X-Router-Conversation`), so nothing more is spent while the plugin acts. +2. The plugin walks `parentID` to the tree's root, usually Atlas, and calls + `POST /session/{id}/abort` on the stuck session and on the root. +3. The plugin compacts the root: `POST /session/{root}/summarize`. Its + `experimental.session.compacting` hook appends the **incident report** to + the compaction prompt, so the report survives into the compacted context. +4. `experimental.compaction.autocontinue` stays enabled, so the root restarts + from the compacted context with the report in view. The router lifts the + refusal when the compaction completes. +5. Toast: "recovered : ". + +**`auto_recover`, stage 2 (detected again in the same tree within +`recover_window`):** apply the `auto_limit_context` cap to the tree, and toast +again. + +**Stage 3, a hard stop.** Recovery can itself loop: recover, loop, recover. After +`max_recoveries` in the window, the router refuses the tree and does not +restart it. It toasts and raises a high-severity warning, and the tree waits for +an admin Resume. Without this cap, auto-recovery is an unattended retry loop, +which is the failure it exists to stop. + +**The incident report** is built deterministically from `progress_events` +(fingerprints, targets and exit codes, never content), so it is free and cannot +hallucinate. Example: + +``` +[router] A sub-agent stalled and was stopped. + session: Item 2 G1 nav refactor (Sisyphus-Junior), child of this session + 423 tool calls, 0 landed changes (no edit with a diff, no commit) + most repeated: read admin/frontend/navbar.js x61, read tests/... x16 + spend since last landed change: $1.84, 70M prompt tokens +Avoid: re-reading whole files, since reads repeat without edits. Give the +worker the exact edit to make, and require an on-disk diff or commit before it +reports done. +``` + +The "Avoid:" line maps from the dominant signal: duplicates, a failure streak, +or context size. A local model may rephrase it later (section 4); it is not +needed to produce it. + +**"Limit context"** has two candidate mechanisms. Phase 0 must verify which works +before stage 2 is built: +- **(a) Router-side, per-session tighter `pinch` budget.** The machinery + exists (`context_prune.py`), but pruning rewrites the prefix, and + `plans/token-waste-waves.md` Wave 3 measured rewrites costing the cache. So + this trades cache hits for size. Measure it. +- **(b) The router answers an over-cap request with a context-overflow-shaped + error**, so opencode runs its own overflow compaction. The autocontinue hook + receives `overflow: boolean`, so opencode does have such a path. Unverified: + whether it recognizes a router error as an overflow. Test this on 8081 before + relying on it. + +Also, the Controls or Home page gets a live "Sessions" readout: session, agent, +parent, verdict, mode stage, turns and $ since landed, and the top repeated +target. This fits `plans/cockpit-brainstorm.md`. Each tree with a refusal gets +a Resume button. + +Every threshold and the mode are knobs with admin controls (North Star #1). +Per the "build for re-tuning" rule, the neutral default is `warn`, with Phase +0-calibrated thresholds. + +Without the plugin (another client, or a stale install), only the router half +works: warnings, alerts and the scoped 429. Abort, compact and restart need the +plugin. The Sessions readout says which sessions have a live plugin (from the +headers in section 1), so a mode that silently cannot act is visible. + +### 6. Alerting: a notifier with pluggable channels + +Toasts and the admin bell only help someone who is looking. Incident #8 ran +overnight. Alerts come from the **router**, not the plugin: the router is the +judge and is always running, while the plugin dies with the opencode process +it lives in. + +**Ship now:** one channel, `desktop` (`notify-send`). **Design for later:** SMS +(Twilio or similar), RingCentral, PagerDuty, a generic webhook (ntfy, Slack). +Each of those should be a small adapter added without touching the detector. + +Shape: +- **An alert is an event with a lifecycle**, not a message. Fields: + - `dedup_key`: for this detector, the tree root session id + - `severity`: `info`, `warning` or `critical` + - `state`: `trigger`, `escalate` or `resolve` + - `title` and a one-line `summary` + - `details`, the incident report from section 5 + - `source`: `no-progress`, so other detectors can reuse the notifier later + + This maps one-to-one onto PagerDuty Events v2 (`dedup_key`, + `event_action: trigger/resolve`), so that adapter is thin. Channels without + a lifecycle (SMS, desktop) render `trigger` and `escalate` as a message and + `resolve` as a short "recovered" message, or skip it (per-channel knob). +- **Transitions, not repeats.** An alert fires when a verdict *changes*, e.g. + `ok -> stalling`, `stalling -> looping`, a recovery, the hard stop, or + `-> ok`. It never fires per request. Per-channel rate limit and a quiet + window stop a flapping session from paging anyone at 3am every minute. +- **Severity routing per channel:** each channel has `min_severity`. The + intended setup: desktop gets `warning` and up; an SMS or PagerDuty channel + added later gets only `critical`, i.e. the hard stop, or `looping` in `warn` + mode. +- **Delivery never blocks routing.** Alerts go through a small queue with a + timeout and one retry. A failed delivery is logged and shown in the admin + UI; it never slows or fails a chat request. +- **Secrets stay out of config files.** Channel credentials (API keys, routing + keys, phone numbers) are read from environment variables named in config, + the same `api_key_env` pattern `dispatch_providers` already uses, and loaded + from `.env` by the systemd unit. `config.yaml` and `config.local.yaml` never + hold a secret. +- **Config:** a `notifications.channels` list, each entry + `{name, type, enabled, min_severity, resolve_messages, ...type-specific}`. + Only `type: desktop` exists in this lift. An unknown `type` fails config load + (StrictModel), so a typo cannot silently disable alerting. +- **Admin (North Star #1):** per channel, an enabled toggle, a `min_severity` + select, and a **Send test alert** button. The test button is the only way to + know a channel works before the night it matters. Recent deliveries and + failures show under the Sessions readout. + +**`desktop` specifics, to verify in Phase 0.** The router runs as a systemd +*user* service, so `notify-send` needs the session D-Bus +(`DBUS_SESSION_BUS_ADDRESS`, normally `unix:path=/run/user//bus`). Check +that the unit's environment has it and that its sandboxing (`ProtectHome` and +friends) allows the socket. If not, say so in the plan and fix it in +`deploy/`, not by weakening the sandbox wholesale. Use urgency `critical` for +`critical` alerts so they persist on screen. + +### 7. The watchdog timer: detection that needs nobody watching + +**Why.** All three loops in incident #8 were caught by a human or by a Claude +Code session reading opencode's session store by hand. The operator does not +want to depend on either. The watchdog is an independent, local, scheduled +sensor. It runs whether or not anyone is at the screen, whether or not a Claude +session is open, and whether or not the plugin loaded. **It calls no cloud +model.** + +**What.** A systemd user timer, `deploy/llm-router-watchdog.{service,timer}`, +shaped like the existing `llm-router-*` timers, every `watchdog.interval` +(proposed 5 min). Each tick: +1. **Find opencode.** Read `~/.local/share/opencode/rc-servers.json`, written + by the `session-registry` plugin. Probe each recorded `serverUrl`. If none + answers, exit quietly; no opencode means nothing to watch. +2. **List sessions active since the last tick** via `GET /session`, and read + each one's recent messages via `GET /session/{id}/message`. +3. **Deterministic check, free, on every active session.** The same signals as + the judge (section 3): exact-duplicate share, the most-repeated target, the + failure streak, landed changes (for read-only agents, target novelty + instead), and injected continuation turns. This is Phase 0's replay code + run live. One implementation, imported by both, not a copy. +4. **Local LLM second opinion, only on sessions step 3 flags.** Skipped when + `local_compute.enabled` is off, which keeps gaming mode honest. The model is + whatever is configured: a `watchdog.model` knob that defaults to + `verification.model`, on that section's base URL. Today that is + `qwen2.5-coder-router:14b` on the local Ollama (`localhost:11434`), already + kept warm by local verification, so a call pays no model-load delay. It + runs on a digest of tool names, targets and exit codes, never content. + Phase 0 measures this model's verdicts on the three incident #8 cases + (must flag: the stuck workers and the post-completion Atlas loop; must not + flag: productive Atlas) before its answer is trusted. It asks "is this agent looping without progress? yes or no, + with a one-line reason". A healthy tick costs zero inference. +5. **Report, do not act.** POST the verdict to the router (the `/progress` + judge or a sibling endpoint), tagged `source: watchdog`. The **router** + applies the configured mode, alerts through the notifier, and enforces + `max_recoveries` and Resume. The watchdog never aborts or compacts anything + itself, so there is exactly one recovery path and one place that decides. + +**If the router is down,** the watchdog still sends a `critical` desktop +alert directly with `notify-send`, and nothing else. A dead router is exactly +when an unattended run needs a human. + +**Failure modes to design for:** +- A tick that overruns the interval: use a lock file and skip the tick, do not + stack them. +- An opencode API shape change. The v2 port breaks this the same way it breaks + `oc_work_start.py`. Detect the unexpected shape and raise one + `warning`-level alert saying the watchdog is blind, rather than silently + reporting "all healthy". +- The watchdog's own health. Record `last_tick_at` and the result (via the + router, or a state file), and show it in the admin Sessions readout. A + watchdog that stopped running must be visible. + +**Knobs (North Star #1):** +- `watchdog.enabled` +- `watchdog.interval` +- `watchdog.local_llm_enabled` (default on, still gated by `local_compute`) +- `watchdog.model` (default: `verification.model`) +- `watchdog.read_only_agents`, the list from the read-only case above + +The timer ships enabled, since detection plus a warning is the neutral, +low-risk default; the actions follow `progress.mode`. + +### 8. The behavior breaker (Lift B): trip a model that stalls sessions + +**Gap.** The breaker family covers availability (`circuit_breaker.py`, 5xx) and +malformed output (PR #76: mojibake, empty 200s). Nothing trips on +**well-formed output that does not do the job**. In incident #8, +`z-ai/glm-5.3-flash` produced fluent text that confabulated truncation and +re-read in loops, even with pinch off and a small context. It was only +excluded by hand. + +**Trigger.** Verdicts from this detector (the watchdog, and the router judge +once built), attributed to the model(s) that served each flagged session in its +window. Trip when: +- stall verdicts on one model span at least `min_sessions` distinct sessions + (proposed 3) within `window` (proposed 60 min), AND +- that model's stall rate is well above the other models' on the same window's + traffic, so one hard task cannot blame a model. + +**Response.** The same shape as the existing breakers: a passive routing skip +with cooldown and backoff, plus a critical alert through the notifier and an +admin Resume. Key it on the model, not the provider, per +`plans/mangled-output-detection.md`'s scope reasoning: here one model misbehaved +while its siblings on the same provider did not. Record every trip with its +evidence (sessions, verdicts, rates) so a human can review it. + +**Blocked on identity.** Joining "this session stalled" to "these models served +it" needs #99's conversation headers on each request. As of 2026-09-26 they do +not arrive: `router-link.js` is installed and passes its 29 tests, and the router +reads the headers, but production decisions still carry fingerprint +`session_key`s with no `agent`. Lift B starts by finding out why. + +## Fixes that ride along + +- **The plugin never reads a real exit code.** Both `router-outcome.js` + (deployed) and its replacement `router-link.js` (#99, line 83) have this. In + 1.18, + `tool.execute.after` gives `output = {title, output, metadata}`, and bash's + exit is `output.metadata.exit` (verified on a live session). `looksFailed` + checks `output.exitCode` / `output.exit_code`, which do not exist, so every + verdict has come from the failure-text regex alone. Read `metadata.exit` + first. +- **The installed plugin differs from the repo copy.** The main checkout has an + uncommitted comment block. It is harmless, but the lift should add a check + (install script or a `/health` field reporting the plugin's version string + from a header) so a stale install is visible. +- **`opencode.json` advertised `auto` with `limit.context: 782324`**, the + largest window in the catalog, so opencode only compacted near 782k. Tonight's + contexts reached 682k and were resent every turn. Lowered to 200,000 on + 2026-09-26 by owner decision (see the end). It caps what `auto` is asked to + carry, so Phase 0 should also measure how often a real agent turn now + compacts, and whether any task needed more. + +## Ground rules for the executor + +- **Base on `origin/main`, NOT the local `main`.** The local `main` checkout + (`c87e675`) is 29 commits behind `origin/main` (`f554fc1` as of 2026-09-26) + and 2 ahead with local-only commits. It lacks #99 (conversation identity, + `router-link.js`) and #100 (local-encoder rebuild). `git fetch origin` first, + then branch `feat/no-progress-detection` from `origin/main` in its own + worktree. Read code from that worktree or `git show origin/main:`, never + from the main checkout's files. Do not touch, stash or commit anything in the + main checkout; it has uncommitted work from other efforts. +- The plugin to extend is `deploy/opencode-plugin/router-link.js` (with its + `router-link.test.mjs`). `router-outcome.js` was replaced by #99. +- **Port 8080 is production; never touch it.** Integration checks run on 8081 + via `scripts/sandbox.sh`, against a copy of the live DB made with the sqlite + backup API (source opened `mode=ro`). +- Never edit `config/config.yaml` or `config/config.local.yaml`. New knobs get + defaults in `config.yaml` only as part of this lift's own tracked changes. +- **The installed plugin at `~/.config/opencode/plugins/` is live for every + opencode session on this machine,** including the one running the lift. + Develop and test against the repo copy. Installing a new version is an + operator step, stated in the done report, not something the executor does. +- **Commit by explicit path. Never `git add -A` or `git add .`** Incident #8's + item 4 swept an `.omc/` state file into a commit that way. +- **Every worker claim is verified by the orchestrator** against + `git show --stat` and the actual test files before acceptance. Incident #8 + had a commit message listing tests that did not exist. +- **Read large files in targeted slices** (grep for the anchor, then read + around it). Incident #8's stuck workers re-read whole files dozens of times. +- ASCII only in new code, comments and UI strings; no middle-dot separators. + No prose paragraphs in the admin UI. +- Lint gate: pinned `uvx ruff@0.16.9`, no NEW findings in touched files + compared with the base commit (same recipe as + `.omo/plans/cockpit-quick-wins.md` runbook step 5). Never `ruff --fix` + outside the lines the item changes. +- `pytest` offline stays green after every commit. Record the baseline on + `origin/main` first. On `c87e675`, one test, + `tests/test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity`, + failed before any change; check whether it still does on `origin/main` rather + than assuming. +- New `route_decisions` columns are registered in + `tests/test_tui_schema_drift.py`; every new knob gets an admin control or a + recorded `DELIBERATELY_NOT_IN_ADMIN` reason (`tests/test_admin_knob_coverage.py`). + +## Phases + +0. **Calibrate offline.** Replay over opencode session history; produce the + verdict timeline per session and a proposed threshold table. No router or + plugin changes. **Start from `plans/no-progress-detection-prototype.py`,** + an interim watcher tuned on 2026-09-26 against the 15 sessions of incident + #8. On that set it flags every known loop (both item-2 workers, the item-3 + first attempt, item 4, the post-completion Atlas stretch, the explore + helper stuck on a deleted file, a helper that printed the spec 12 times) + and none of the healthy sessions. Four findings from tuning it: + - Exact `(tool, args)` fingerprints alone miss reworded commands: 28 + differently-worded `cat boulder.json` never matched. + - Collapsing to the bare file over-flags ordinary sliced reading, which is + what the ground rules tell workers to do. Key a `read` on file plus + offset/limit, and a `bash` target on its referenced files plus the + numbers in the command (the slice). + - Slow loops spread across 300+ calls escape a 60-call window. Add a + whole-session rule: one exact call repeated 15+ times with nothing landed + recently. + - Read-only agents need the separate rule (repeats of one target), since + "nothing landed" is always true for them. + Its thresholds (`WINDOW=60`, `DUP_MIN=0.25`, `TOP_MIN=12`, `TOP_MIN_RO=8`, + `CUM_MIN=15`) are a starting point fitted on one night, not a result. +1. **Identity: build on #99, do not rebuild it.** The `metadata.exit` fix in + `router-link.js` (and its `.test.mjs`), plus whatever #99 lacks for + progress roll-up, if anything. Deploying #99 itself (syncing the + production checkout, installing `router-link.js`) is an operator step. It + goes in the done report as a prerequisite for Phase 2 to see real traffic, + not an executor task. +2. **Progress events, the judge and the watchdog, `warn` mode.** `/progress`, + `progress_events`, verdicts, the `/metrics` warning, toast, Sessions readout, + and the knobs. Also the notifier (section 6) with the `desktop` channel and + the admin Send test alert button, and the watchdog timer (section 7) with + its local-LLM second opinion. **Order within the phase: build the watchdog + and the notifier first.** Together with Phase 0's detection code, they + protect unattended runs before the plugin half is even deployed. +3. **`auto_compact` and `auto_recover` stage 1.** Abort, compact with the + injected report, restart; the scoped 429; the `max_recoveries` hard stop and + admin Resume. The incident report builder lands here. +4. **`auto_limit_context` and `auto_recover` stage 2**, using whichever + mechanism Phase 0 verified. +5. **Local-model second opinion inside the router's live judge** (section 4), + only if earlier phases show ambiguous cases. The watchdog already has one + (section 7); this would add it to the per-request path. + +The v2 port (`project_opencode_v2_pinned`) changes hook registration. Keep the +plugin's logic in plain functions so the v1 and v2 shells stay thin. + +## Owner decisions + +Settled 2026-09-26: the response is a mode (`warn`, `auto_compact`, +`auto_limit_context`, `auto_recover` with two stages), and every mode warns. +`auto_recover` stops and compacts the tree's root (Atlas), not only the stuck +worker. + +Also settled 2026-09-26: +- **The default mode is `warn`.** +- **`max_recoveries` is 2** per tree per window, then the hard stop. +- **The `auto` context limit is 200,000.** Applied the same day, ahead of this + lift, in the repo `opencode.json` and the global + `~/.config/opencode/opencode.json` (backup at + `opencode.json.bak-20260926-auto-context`). The global file pins every agent + (atlas, sisyphus-junior, prometheus, ...) to `llm-router/auto`. opencode + reads config at startup, so it takes effect for sessions started after a + restart. `auto:batch` stays at 782,324 on purpose: batch is where long + contexts are expected. + +- **Alerting ships with `desktop` (`notify-send`) only**, behind a notifier + built for more channels. SMS, RingCentral, PagerDuty and webhooks are later + adapters (section 6), not part of this lift. + +No owner decisions remain open. Thresholds come from Phase 0. diff --git a/plans/no-progress-lift-a.md b/plans/no-progress-lift-a.md new file mode 100644 index 0000000..fdfa95c --- /dev/null +++ b/plans/no-progress-lift-a.md @@ -0,0 +1,227 @@ +# Lift A: loop watchdog, desktop alerts, and the shared detector + +Status: done -- shipped in PR #102; the admin surfacing (component 5) shipped only partly +Date: 2026-09-26 +Parent spec: `plans/no-progress-detection.md` (reference only; do NOT read it +end to end). This brief is the contract for Lift A. +Incident: `docs/incidents.md` #8. + +## Goal + +Catch an opencode agent session that is spinning without progress, and alert +the operator on the desktop, with nobody watching and no cloud model involved. + +## Why this slice first, and why it is small + +Incident #8's loops were caught only by a human or a Claude session reading +opencode's session store by hand. Lift A replaces that with a local timer. + +It is deliberately narrow because the last planning attempt failed on size: +- a 540-line spec +- a planning context that grew to about 600k tokens +- the model reporting "elided" reads that were stored intact, and re-reading + the spec 84 times (68x its length) + +Keep this lift's inputs, todos and contexts small. + +**Out of scope (Lift B):** +- router-side progress events, the live judge, and `/progress` +- recovery modes (`auto_compact`, `auto_recover`, `auto_limit_context`) +- scoped 429s +- the `router-link.js` changes +- SMS, PagerDuty and other alert channels + +## Components + +### 1. Shared detector: `src/progress_detect.py` (pure, stdlib only) + +Port the rules from `plans/no-progress-detection-prototype.py`. It is tuned on +the incident #8 sessions; keep its behaviour and make it a clean, tested +module. Input: a session's tool calls as `(time, tool, args_json, landed)`. +Output: a verdict with the reason. Signals, over a window of the last +`window` calls: +- `dup`: exact `(tool, args)` duplicate share +- `top`: the most-repeated target. A `read` target is the file plus + offset/limit. A `bash` target is the files it names plus the numbers in the + command (the slice). +- `slow`: one exact call repeated `cum_min`+ times over the whole session +- `coverage`: lines requested from one file in the window divided by its + length. This catches re-reading a file through shifting offsets, which the + other signals miss. +- `landed`: an edit or write with a non-empty diff, or a `git commit` with + exit 0, by the session or any descendant (`parentID`) within the window. + +Verdict: +- Normal agents are flagged when nothing landed and any of dup, top, slow or + coverage crosses its threshold. +- Read-only agents (explore, librarian, oracle; a configurable list) never + land, so for them top, slow or coverage alone decides. +- Thresholds start at the prototype's values: `window` 60, `dup_min` 0.25, + `top_min` 12, `top_min_ro` 8, `cum_min` 15, `cover_min` 4.0. + +### 2. Calibration backtest: `scripts/progress_backtest.py` + +It replays sessions from opencode's HTTP API through `progress_detect` and +prints a per-session verdict with the first flag point. Commit a small +**fixture** of the labelled sessions below, exported as tool-call fingerprints +(tool name, args, landed; never message text), under `tests/fixtures/`. A test +asserts every label. Live opencode is not needed in tests. + +MUST FLAG: + +| session | what it was | +|---|---| +| `ses_f24cc0258ffepbJPoOzoiLZi9s` | item 2 worker, read navbar.js x61 | +| `ses_f24fcf651ffeOm9O4E14QSZeJ4` | item 2 worker, 1st attempt | +| `ses_f24890640ffewKnIPqjjAfFvOk` | item 3 worker, 1st attempt | +| `ses_f245607eaffeGxPH8F1tvvdvkG` | item 4 worker | +| `ses_f237e7e31ffeOzz4kUKu4CWzMu` | explore helper re-reading a deleted file | +| `ses_f235d755bffeo0uWtD1g2HwQUg` | helper that printed the spec 12x | +| `ses_f2506ac70ffebLUxBsq0aauOn9` | Atlas; must flag in its post-02:07 stretch | +| `ses_f2380dd4cffeVwcZn3Kv3nQN7v` | planner that re-read the spec 68x its length | + +MUST NOT FLAG: + +| session | what it was | +|---|---| +| `ses_f246e699fffeCHUPTLr5ajBhWC` | item 3 retry | +| `ses_f245f78f8ffeGMn3PgArpOwhFy` | item 3 compact | +| `ses_f243ea08effemMwXdMRE1gpnhK` | item 5 | +| `ses_f242e5e28ffehF97ZxxdbTeylQ` | item 6 | +| `ses_f236bbdaaffezwnvH10X2rgrxc` | explore, sliced reads of dispatcher.py | +| `ses_f236bc7baffeG5LQsr2K1c46xO` | explore | +| `ses_f2e81929fffeaVbWui0X9QjmRQ` | cockpit planner | + +The Atlas session also must NOT flag before its 02:07 stretch. + +Export the fixture FIRST, in the first todo. opencode sessions can be deleted, +and the fixture is the ground truth. + +### 3. Notifier: `src/notifier.py`, `desktop` channel only + +- **An alert is an event:** + - `dedup_key` (the session tree root) + - `severity` (`info`, `warning` or `critical`) + - `state` (`trigger`, `escalate` or `resolve`) + - `title`, `summary`, `details`, `source` + + The shape maps one-to-one onto PagerDuty Events v2, so later adapters are + thin. Only the `desktop` channel ships. +- **`desktop` runs `notify-send`,** with urgency `critical` for critical + alerts. +- **Fire on transitions only,** with a per-channel rate limit and a quiet + window. +- **A delivery failure is logged, never raised.** +- **Config:** a `notifications.channels` list of + `{name, type, enabled, min_severity}`. An unknown `type` fails config load + (StrictModel). Secrets would come from env vars named in config; desktop has + none. + +### 4. Watchdog: `src/watchdog.py` + `deploy/llm-router-watchdog.{service,timer}` + +A systemd user timer, shaped like the existing `llm-router-*` timers, every 5 +min. Each tick: +1. **Find opencode.** Read `~/.local/share/opencode/rc-servers.json` and probe + each `serverUrl`. If none answers, exit 0 quietly. +2. **Read activity.** List sessions active since the last tick and read their + tool calls. +3. **Detect.** Run `progress_detect` on each active session. +4. **Second opinion, flagged sessions only.** Ask the configured local model + (`watchdog.model`, default `verification.model`, today + `qwen2.5-coder-router:14b` on the local Ollama) "looping without progress? + yes/no + one line", on a digest of tool names, targets and exit codes, never + content. Skip this when `local_compute.enabled` is false or + `watchdog.local_llm_enabled` is false. Report the model's answer alongside + the deterministic verdict; it does not veto it in Lift A. +5. **Alert.** Alert through the notifier directly, and write rows to the + router DB: a `watchdog_ticks` table (time, sessions seen, outcome) and a + `watchdog_verdicts` table. The router need not be up; SQLite WAL handles the + second writer. + +Also: +- A lock file prevents overlapping ticks. +- An unexpected opencode API shape raises one `warning` alert ("watchdog is + blind"), never a silent all-clear. +- Admin (North Star #1): a small Watchdog card showing the last tick, recent + verdicts, the channel list with enabled and `min_severity` controls, and a + **Send test alert** button. Every new knob gets a control or a recorded + reason; `tests/test_admin_knob_coverage.py` enforces it. +- Knobs: + - `watchdog.enabled`, `watchdog.local_llm_enabled`, `watchdog.model`, + `watchdog.read_only_agents` + - the detector thresholds from section 1 + - `notifications.channels` + +The tick interval lives in the timer unit; document it next to the knobs. + +## Ground rules + +- **Base on `origin/main`.** Branch `feat/no-progress-lift-a` in its own + worktree. Do not touch, stash or commit anything in the main checkout. +- **Port 8080 is production.** Integration checks run on 8081 + (`scripts/sandbox.sh`), against a sqlite-backup copy of the live DB. +- **Never edit** `config/config.yaml`'s existing values or + `config/config.local.yaml`. New keys get defaults in `config.yaml`. +- **Do not install or enable** the timer on this machine; that is an operator + step in the done report. +- **Commit by explicit path; never `git add -A`.** The orchestrator checks each + worker's `git show --stat` against its claims and the plan's paths before + accepting it. +- **Read in targeted slices.** Grep for the anchor, read around it. Do NOT read + `plans/no-progress-detection.md` whole; this brief is the contract. +- **If a worker's context passes about 150k tokens, stop it** and start a fresh + worker with a smaller prompt, rather than letting it grind. +- **ASCII only; no middle dots.** No prose paragraphs in the admin UI. +- **Lint:** pinned `uvx ruff@0.16.9`, no NEW findings in touched files against + the base, never `ruff --fix` outside changed lines. +- **pytest:** record the `origin/main` baseline first; it stays green apart + from the baseline. + +## Done means + +- Every must-flag and must-not-flag label passes in tests, and the live + backtest prints the same verdicts. +- One real `notify-send` from the admin test button on 8081. +- The watchdog runs once by hand (`python -m watchdog --once`) against live + opencode and prints its verdicts. +- The done report lists each commit, the tests, and the operator steps + (install and enable the timer). + +## Addendum, 2026-09-26: surfacing and one-click block (North Star #4) + +The owner set a new priority: surface waste in the admin portal, and make +stopping it one click (`CLAUDE.md` North Star #4). This lift therefore also +ships: + +### 5. Loops panel and one-click model block (admin portal) + +- **A Loops panel** on the Home page (or the Controls page, whichever the + cockpit layout fits), reading `watchdog_verdicts`: each flagged session with + agent, verdict, turns and $ since the last landed change, the top repeated + target, and when it was flagged. +- **A per-model rollup:** stall verdicts per model over the last hour, next to + each model's share of traffic. It needs session-to-model attribution; see + item 0. +- **A Block model button** beside each model in the rollup. It writes + `admin_model_overrides` with `availability = 'blocked'` and a reason (default + "stalled N sessions", editable), through the existing + `POST /admin/api/models/{model_id}/{provider}/availability` path. Verify that + routing excludes `blocked` exactly like `deprecated` (`load_candidates`), and + add a test for it. +- **A Blocked models list** with reason, time and a one-click Unblock + (`DELETE .../availability`). +- **Desktop alerts link to the panel** (`http://127.0.0.1:8080/admin/` plus an + anchor). + +### 0. First todo: why do the identity headers not arrive? + +`router-link.js` (#99) is installed in `~/.config/opencode/plugins/`, is +identical to the repo copy, and passes its 29 tests. The router reads +`X-Router-Conversation` (`src/dispatcher.py`, `resolve_identity(request.headers, +...)`). Yet production `route_decisions` after the 04:02 opencode restart still +show fingerprint `session_key`s and `agent = NULL`. Find out why, and fix it if +it is a small change. + +The per-model rollup depends on it. If it cannot be fixed in this lift, the +rollup ships disabled with a visible "needs conversation identity" note, and +everything else ships. diff --git a/src/watchdog.py b/src/watchdog.py index cdce6d8..d0e70a7 100644 --- a/src/watchdog.py +++ b/src/watchdog.py @@ -790,7 +790,8 @@ def _get_db_path(cfg: Any) -> str: return "router.db" -if __name__ == "__main__": +def main(argv: list[str] | None = None) -> int: + """Run one watchdog tick; return the process exit code (0 when it ran).""" ap = argparse.ArgumentParser( description="Watchdog orchestrator for opencode loop detection", ) @@ -800,7 +801,7 @@ if __name__ == "__main__": ap.add_argument( "--config", default="config/config.yaml", help="config.yaml path", ) - args = ap.parse_args() + args = ap.parse_args(argv) logging.basicConfig( level=logging.INFO, @@ -841,6 +842,14 @@ if __name__ == "__main__": last_fired_cb = _make_last_fired(conn) notifier = Notifier(notifier_cfg, last_fired=last_fired_cb) - rc = tick(cfg, conn, notifier=notifier) + alerts_fired = tick(cfg, conn, notifier=notifier) conn.close() - sys.exit(rc or 0) \ No newline at end of file + logger.info("watchdog tick done: %d alert(s) fired", alerts_fired) + # A completed tick exits 0 however many alerts it fired. Returning the + # alert count made systemd mark the oneshot unit FAILED on every real + # alert (and wrap past 255), so a working alert looked like a crash. + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tests/test_watchdog.py b/tests/test_watchdog.py index 04766d6..51c983b 100644 --- a/tests/test_watchdog.py +++ b/tests/test_watchdog.py @@ -1180,3 +1180,21 @@ class TestDescendantLandedScope: ) finally: _close_conn(conn, path) + + +class TestMainExitCode: + """`python -m watchdog --once` runs as a systemd oneshot: a tick that ran + and fired alerts is a success, so the exit code must not carry the alert + count (systemd would mark the unit FAILED on every real alert).""" + + def test_main_returns_zero_when_alerts_fire(self, tmp_path): + import watchdog + + cfg = MagicMock() + cfg.database.path = str(tmp_path / "router.db") + cfg.notifications.channels = [] + with patch("watchdog.load_config", return_value=cfg), \ + patch("watchdog.tick", return_value=3) as fake_tick: + rc = watchdog.main(["--once", "--config", "unused.yaml"]) + assert fake_tick.called + assert rc == 0