# No-progress detection: catch an agent that is spinning, not one that is working Status: reference -- parent spec; Lift A shipped in PR #102, Lift B (router-side judge, modes, behavior breaker) not started Date: 2026-09-26 Trigger: the `cockpit-quick-wins` Atlas run, 2026-09-25 19:00 to 2026-09-26 02:00. Incident write-up: `docs/incidents.md` #8. ## What happened, measured Read-only against the live `router.db` and opencode's own session store (`http://127.0.0.1:4097`), 2026-09-26. Spend, 19:00 to 02:00 local: - 323M prompt tokens (95% cached) and about $7.81, over roughly 8 hours. - 79% of the money went to OpenRouter `z-ai/glm-5.3-flash`: 1,368 calls averaging 203k prompt tokens. - 921 calls carried 120k+ prompt tokens and cost $6.57. NeuralWatt `glm-5.3` took 47 calls averaging 673k context, max 682,679, because nothing else fit. - Every decision was `classification_source = classifier` on the default profile. The router did what it was told; nothing was pinned. **Spend rate is the wrong signal.** Tonight's worst hour ($2.42) and worst 3-hour window ($4.13) sit inside the prior 30 days' normal range (hourly p90 $1.46, max $4.83; worst 3h $9.65). A rate alarm quiet enough for a healthy all-day agent run would have missed this entirely, and one tight enough to catch it would fire on good work. The operator's framing is the right one: steady usage is fine. What is not fine is spend with **no concrete change landing**. **The sessions that wasted money are visibly different in their tool calls:** | session (opencode) | tool calls | exact-duplicate calls | worst repeat | landed a change? | |---|---|---|---|---| | Item 2 worker, 2nd attempt | 423 | 30% | `read navbar.js` x61 | no (worktree unchanged) | | Item 3 worker, 1st attempt | 132 | 23% | `read test_admin_js_units.py` x16 | no | | Item 4 worker | 134 | 19% | `read .omo/plans/...` x8 | partly (code, no tests) | | Item 2 worker, 1st attempt | 307 | 17% | `read navbar.js` x21 | no | | Atlas orchestrator | 576 | 11% | playwright console check x10 | yes (via workers) | | healthy workers (items 1, 5, 6, 3-retry) | 30-96 | 0-6% | 1-3 | yes | **A second loop shape, 2026-09-26 02:21 to about 02:56:** the Atlas orchestrator itself, after its plan was complete: - `.omo/boulder.json` carried a stale `status: completed` and `pr_url` from the previous plan. - oh-my-openagent's continuation hook injected "continue" turns; all 3 user turns in the window were injected. - Atlas, working from a trimmed context, re-derived its state every time: 54 turns, about 19M cached-read tokens, `boulder.json` read 28 times, the plan's checkboxes grepped 16 times, nothing landed. The operator caught it only by watching. The judge must treat this as `stalling`, and Phase 0 must include it as a calibration case. The plugin can also see injected continuation turns (`chat.message`), which is a useful extra signal: "continuation turns keep arriving and nothing lands" is this shape exactly. **A third shape, 2026-09-26 around 03:05 to 03:17: a read-only agent.** Prometheus's `explore` subagent reached 322 tool calls, 19% exact duplicates, and read `deploy/opencode-plugin/router-outcome.js` 20 times. The file had been deleted from under it: the production checkout was synced to `origin/main`, where #99 retired it. The subagent was aborted. Read-only agents (`explore`, `librarian`, `oracle`, and the like) never land a change by design, so "no landed change" cannot be their signal. For them the judge uses: - target novelty: new files or queries per N calls, and - a failure streak on the same target, such as repeated reads of a path that no longer exists. Identify them by the `X-Router-Agent` value, via a configurable list of read-only agent names. "Exact duplicate" means the same tool with byte-identical arguments. Duplicate share alone does not separate the orchestrator (11%, productive) from a stuck worker (17%). Duplicates **plus** no landed change does. ## Why nothing caught it - `circuit_breaker.py` is availability-only: it trips on provider 5xx. - The only spend alarm is `runway_low_warning`: balance / burn < 6 h, with burn averaged over a 24 h balance window. OpenRouter read $24.73 at $0.25/h, so 97 h of runway. It answers "will I run out", not "is this being wasted". - The router cannot see progress at all. It sees prompts, not whether an edit landed or a command keeps failing. The opencode plugin sees both. - The DEPLOYED router cannot tell conversations apart either. `session_key` merges concurrent conversations under one multi-agent client, and `/outcome` attribution falls back to a directory match that 409s when several sessions share a cwd, which is exactly the Atlas-plus-workers shape. **The fix already exists on `origin/main` (PR #99, conversation-identity) but is not deployed:** production runs from a local checkout 29 commits behind `origin/main`, the live `route_decisions` has no `agent` / `parent_key` columns, and the installed plugin is still `router-outcome.js` rather than #99's `router-link.js`. ## Design Split the job where the information lives. **The plugin is the sensor; the router is the judge.** No raw task text leaves opencode, and the router never stores any; the "never store raw task text" rule is unchanged. ### 1. Exact session identity: ALREADY BUILT upstream (#99), build on it PR #99 (`conversation-identity`, merged to `origin/main` as `fe8c813`) did this. Reuse it; do not rebuild it: - **Plugin.** `deploy/opencode-plugin/router-link.js` replaces `router-outcome.js`. Its `chat.headers` hook sends `X-Router-Conversation` (the opencode session id), `X-Router-Agent` (slugged to the router's charset) and `X-Router-Parent`, with a bounded parent-lookup cache. - **Router.** `src/conversation_identity.py` resolves identity from those headers. `session_key` becomes `'c:' + conversation id`. `route_decisions` gained `agent` and `parent_key` columns. `/outcome` attributes to the exact conversation. What remains for this lift: - **Deploy #99.** It is merged but not running. Syncing the production checkout to `origin/main` and installing `router-link.js` are operator steps: the main checkout has uncommitted work, and part of it, the `confidence_min` rename, also landed upstream via #100. - **Fix the exit code in `router-link.js`.** It kept `router-outcome.js`'s `output?.exitCode ?? output?.exit_code` (line 83 on `origin/main`); see "Fixes that ride along". - **Build the progress events (section 2) into `router-link.js`,** keyed by the same conversation and parent ids, so the judge rolls children up to parents through `parent_key`. Everywhere below, "session" means #99's conversation. Where this spec says `X-Opencode-Session`, read `X-Router-Conversation`. ### 2. Progress events (plugin `tool.execute.after` -> router `POST /progress`) The plugin sends one small event per tool call, fire-and-forget with a short timeout, never blocking the session: ``` {session_id, parent_session_id, agent, tool, call_id, fingerprint, # sha256 of tool + normalized args; never the args themselves target, # coarse, normalized: the file path for read/edit/write, the # command's first two tokens for bash; no contents exit, # bash: output.metadata.exit landed} # bool, see below ``` `landed` is the load-bearing definition, and it is deliberately narrow: - `edit` / `write` / `patch` with a non-empty `metadata.diff` - `bash` whose command is `git commit` (or `git commit --amend`) with exit 0 - A test or lint command (the plugin's existing `TEST_COMMAND` list) that PASSES after the session's previous run of the same fingerprint FAILED, i.e. a fix landed. Reads, greps, todo writes, sub-agent dispatches and passing-again tests are not progress. An orchestrator's progress is its children's: roll child `landed` events up to the parent via `parent_session_id`. Storage: a new `progress_events` table, pruned to a rolling window (knob). It holds fingerprints, targets and booleans, no content. ### 3. The judge (router, per client session) Signals per session, over a rolling window of its last N tool events: - `turns_since_landed`: LLM requests since the last landed event (self plus children for a parent). - `usd_since_landed`: billed spend since then, joined through `request_id`. - `dup_share`: exact-duplicate fingerprint share in the window. - `fail_streak`: consecutive nonzero exits on the same fingerprint (the "sandbox is broken and the agent keeps retrying" case). - `context_tokens`: the latest request's size, to report alongside, since a stall at 600k context costs far more per turn than one at 30k. Verdicts: - `ok`: something landed recently. - `stalling` (warn): no landed change for `stall_turns` turns or `stall_usd` dollars, AND (`dup_share >= dup_warn` OR `fail_streak >= fail_warn`). - `looping` (act, if enabled): the stall condition at the higher `loop_*` thresholds. Both conditions have to hold. Long-running healthy work with no duplicates (a big read-heavy exploration) stays `ok` until it hits the plain turns/spend ceiling, which is a separate, looser knob. **Calibrate before trusting it.** Phase 0 replays the judge offline over opencode's session store (the 11 sessions above plus older history), and reports the verdict each session would have received and when. Thresholds are picked from that, not guessed. The Atlas orchestrator (11% duplicates, productive through children) is the key false-positive case, and it must stay `ok`. ### 4. Optional: a local model as a second opinion The deterministic signals should carry this. Where they are ambiguous, for example near-duplicates (the same file read at shifting offsets, or the same failing test with a different `-k`), there are two cheaper steps before any model: - normalize `read` fingerprints to the path, and bash to command plus target - the local encoder already loaded for `classifier.mode: local_encoder` can embed the `target` strings to cluster "similar-ish" events A local LLM judge ("does this digest look like a loop?") runs only on a session already flagged `stalling`, on a digest of tool names, targets and exit codes, never content. It is gated by `local_compute.enabled`, since the operator pays the local power bill. Treat this as a later phase, justified by Phase 0 showing ambiguous cases the deterministic signals miss. ### 5. What happens on a detection: modes One knob, `progress.mode`, picks the response. **Every mode warns:** - A `/metrics` warning (admin bell and TUI). It names the session, agent, turns and $ since the last landed change, and the top repeated target. - An opencode toast via `POST /tui/show-toast` in the window the operator is watching. - An **alert** through the router's notifier (section 6), which reaches the operator when nobody is looking at either screen. | mode | on detection | |---|---| | `warn` | Warnings only. The neutral default: today's behaviour plus a warning. | | `auto_compact` | Compact the stuck session with the incident report injected, then continue. No refusal, no orchestrator stop. | | `auto_limit_context` | Cap the stuck session's context. No stop. | | `auto_recover` | Two stages, below. | **`auto_recover`, stage 1 (first detection in a session tree):** 1. The router refuses the stuck session's chat requests (429, scoped by `X-Router-Conversation`), so nothing more is spent while the plugin acts. 2. The plugin walks `parentID` to the tree's root, usually Atlas, and calls `POST /session/{id}/abort` on the stuck session and on the root. 3. The plugin compacts the root: `POST /session/{root}/summarize`. Its `experimental.session.compacting` hook appends the **incident report** to the compaction prompt, so the report survives into the compacted context. 4. `experimental.compaction.autocontinue` stays enabled, so the root restarts from the compacted context with the report in view. The router lifts the refusal when the compaction completes. 5. Toast: "recovered : ". **`auto_recover`, stage 2 (detected again in the same tree within `recover_window`):** apply the `auto_limit_context` cap to the tree, and toast again. **Stage 3, a hard stop.** Recovery can itself loop: recover, loop, recover. After `max_recoveries` in the window, the router refuses the tree and does not restart it. It toasts and raises a high-severity warning, and the tree waits for an admin Resume. Without this cap, auto-recovery is an unattended retry loop, which is the failure it exists to stop. **The incident report** is built deterministically from `progress_events` (fingerprints, targets and exit codes, never content), so it is free and cannot hallucinate. Example: ``` [router] A sub-agent stalled and was stopped. session: Item 2 G1 nav refactor (Sisyphus-Junior), child of this session 423 tool calls, 0 landed changes (no edit with a diff, no commit) most repeated: read admin/frontend/navbar.js x61, read tests/... x16 spend since last landed change: $1.84, 70M prompt tokens Avoid: re-reading whole files, since reads repeat without edits. Give the worker the exact edit to make, and require an on-disk diff or commit before it reports done. ``` The "Avoid:" line maps from the dominant signal: duplicates, a failure streak, or context size. A local model may rephrase it later (section 4); it is not needed to produce it. **"Limit context"** has two candidate mechanisms. Phase 0 must verify which works before stage 2 is built: - **(a) Router-side, per-session tighter `pinch` budget.** The machinery exists (`context_prune.py`), but pruning rewrites the prefix, and `plans/token-waste-waves.md` Wave 3 measured rewrites costing the cache. So this trades cache hits for size. Measure it. - **(b) The router answers an over-cap request with a context-overflow-shaped error**, so opencode runs its own overflow compaction. The autocontinue hook receives `overflow: boolean`, so opencode does have such a path. Unverified: whether it recognizes a router error as an overflow. Test this on 8081 before relying on it. Also, the Controls or Home page gets a live "Sessions" readout: session, agent, parent, verdict, mode stage, turns and $ since landed, and the top repeated target. This fits `plans/cockpit-brainstorm.md`. Each tree with a refusal gets a Resume button. Every threshold and the mode are knobs with admin controls (North Star #1). Per the "build for re-tuning" rule, the neutral default is `warn`, with Phase 0-calibrated thresholds. Without the plugin (another client, or a stale install), only the router half works: warnings, alerts and the scoped 429. Abort, compact and restart need the plugin. The Sessions readout says which sessions have a live plugin (from the headers in section 1), so a mode that silently cannot act is visible. ### 6. Alerting: a notifier with pluggable channels Toasts and the admin bell only help someone who is looking. Incident #8 ran overnight. Alerts come from the **router**, not the plugin: the router is the judge and is always running, while the plugin dies with the opencode process it lives in. **Ship now:** one channel, `desktop` (`notify-send`). **Design for later:** SMS (Twilio or similar), RingCentral, PagerDuty, a generic webhook (ntfy, Slack). Each of those should be a small adapter added without touching the detector. Shape: - **An alert is an event with a lifecycle**, not a message. Fields: - `dedup_key`: for this detector, the tree root session id - `severity`: `info`, `warning` or `critical` - `state`: `trigger`, `escalate` or `resolve` - `title` and a one-line `summary` - `details`, the incident report from section 5 - `source`: `no-progress`, so other detectors can reuse the notifier later This maps one-to-one onto PagerDuty Events v2 (`dedup_key`, `event_action: trigger/resolve`), so that adapter is thin. Channels without a lifecycle (SMS, desktop) render `trigger` and `escalate` as a message and `resolve` as a short "recovered" message, or skip it (per-channel knob). - **Transitions, not repeats.** An alert fires when a verdict *changes*, e.g. `ok -> stalling`, `stalling -> looping`, a recovery, the hard stop, or `-> ok`. It never fires per request. Per-channel rate limit and a quiet window stop a flapping session from paging anyone at 3am every minute. - **Severity routing per channel:** each channel has `min_severity`. The intended setup: desktop gets `warning` and up; an SMS or PagerDuty channel added later gets only `critical`, i.e. the hard stop, or `looping` in `warn` mode. - **Delivery never blocks routing.** Alerts go through a small queue with a timeout and one retry. A failed delivery is logged and shown in the admin UI; it never slows or fails a chat request. - **Secrets stay out of config files.** Channel credentials (API keys, routing keys, phone numbers) are read from environment variables named in config, the same `api_key_env` pattern `dispatch_providers` already uses, and loaded from `.env` by the systemd unit. `config.yaml` and `config.local.yaml` never hold a secret. - **Config:** a `notifications.channels` list, each entry `{name, type, enabled, min_severity, resolve_messages, ...type-specific}`. Only `type: desktop` exists in this lift. An unknown `type` fails config load (StrictModel), so a typo cannot silently disable alerting. - **Admin (North Star #1):** per channel, an enabled toggle, a `min_severity` select, and a **Send test alert** button. The test button is the only way to know a channel works before the night it matters. Recent deliveries and failures show under the Sessions readout. **`desktop` specifics, to verify in Phase 0.** The router runs as a systemd *user* service, so `notify-send` needs the session D-Bus (`DBUS_SESSION_BUS_ADDRESS`, normally `unix:path=/run/user//bus`). Check that the unit's environment has it and that its sandboxing (`ProtectHome` and friends) allows the socket. If not, say so in the plan and fix it in `deploy/`, not by weakening the sandbox wholesale. Use urgency `critical` for `critical` alerts so they persist on screen. ### 7. The watchdog timer: detection that needs nobody watching **Why.** All three loops in incident #8 were caught by a human or by a Claude Code session reading opencode's session store by hand. The operator does not want to depend on either. The watchdog is an independent, local, scheduled sensor. It runs whether or not anyone is at the screen, whether or not a Claude session is open, and whether or not the plugin loaded. **It calls no cloud model.** **What.** A systemd user timer, `deploy/llm-router-watchdog.{service,timer}`, shaped like the existing `llm-router-*` timers, every `watchdog.interval` (proposed 5 min). Each tick: 1. **Find opencode.** Read `~/.local/share/opencode/rc-servers.json`, written by the `session-registry` plugin. Probe each recorded `serverUrl`. If none answers, exit quietly; no opencode means nothing to watch. 2. **List sessions active since the last tick** via `GET /session`, and read each one's recent messages via `GET /session/{id}/message`. 3. **Deterministic check, free, on every active session.** The same signals as the judge (section 3): exact-duplicate share, the most-repeated target, the failure streak, landed changes (for read-only agents, target novelty instead), and injected continuation turns. This is Phase 0's replay code run live. One implementation, imported by both, not a copy. 4. **Local LLM second opinion, only on sessions step 3 flags.** Skipped when `local_compute.enabled` is off, which keeps gaming mode honest. The model is whatever is configured: a `watchdog.model` knob that defaults to `verification.model`, on that section's base URL. Today that is `qwen2.5-coder-router:14b` on the local Ollama (`localhost:11434`), already kept warm by local verification, so a call pays no model-load delay. It runs on a digest of tool names, targets and exit codes, never content. Phase 0 measures this model's verdicts on the three incident #8 cases (must flag: the stuck workers and the post-completion Atlas loop; must not flag: productive Atlas) before its answer is trusted. It asks "is this agent looping without progress? yes or no, with a one-line reason". A healthy tick costs zero inference. 5. **Report, do not act.** POST the verdict to the router (the `/progress` judge or a sibling endpoint), tagged `source: watchdog`. The **router** applies the configured mode, alerts through the notifier, and enforces `max_recoveries` and Resume. The watchdog never aborts or compacts anything itself, so there is exactly one recovery path and one place that decides. **If the router is down,** the watchdog still sends a `critical` desktop alert directly with `notify-send`, and nothing else. A dead router is exactly when an unattended run needs a human. **Failure modes to design for:** - A tick that overruns the interval: use a lock file and skip the tick, do not stack them. - An opencode API shape change. The v2 port breaks this the same way it breaks `oc_work_start.py`. Detect the unexpected shape and raise one `warning`-level alert saying the watchdog is blind, rather than silently reporting "all healthy". - The watchdog's own health. Record `last_tick_at` and the result (via the router, or a state file), and show it in the admin Sessions readout. A watchdog that stopped running must be visible. **Knobs (North Star #1):** - `watchdog.enabled` - `watchdog.interval` - `watchdog.local_llm_enabled` (default on, still gated by `local_compute`) - `watchdog.model` (default: `verification.model`) - `watchdog.read_only_agents`, the list from the read-only case above The timer ships enabled, since detection plus a warning is the neutral, low-risk default; the actions follow `progress.mode`. ### 8. The behavior breaker (Lift B): trip a model that stalls sessions **Gap.** The breaker family covers availability (`circuit_breaker.py`, 5xx) and malformed output (PR #76: mojibake, empty 200s). Nothing trips on **well-formed output that does not do the job**. In incident #8, `z-ai/glm-5.3-flash` produced fluent text that confabulated truncation and re-read in loops, even with pinch off and a small context. It was only excluded by hand. **Trigger.** Verdicts from this detector (the watchdog, and the router judge once built), attributed to the model(s) that served each flagged session in its window. Trip when: - stall verdicts on one model span at least `min_sessions` distinct sessions (proposed 3) within `window` (proposed 60 min), AND - that model's stall rate is well above the other models' on the same window's traffic, so one hard task cannot blame a model. **Response.** The same shape as the existing breakers: a passive routing skip with cooldown and backoff, plus a critical alert through the notifier and an admin Resume. Key it on the model, not the provider, per `plans/mangled-output-detection.md`'s scope reasoning: here one model misbehaved while its siblings on the same provider did not. Record every trip with its evidence (sessions, verdicts, rates) so a human can review it. **Blocked on identity.** Joining "this session stalled" to "these models served it" needs #99's conversation headers on each request. As of 2026-09-26 they do not arrive: `router-link.js` is installed and passes its 29 tests, and the router reads the headers, but production decisions still carry fingerprint `session_key`s with no `agent`. Lift B starts by finding out why. ## Fixes that ride along - **The plugin never reads a real exit code.** Both `router-outcome.js` (deployed) and its replacement `router-link.js` (#99, line 83) have this. In 1.18, `tool.execute.after` gives `output = {title, output, metadata}`, and bash's exit is `output.metadata.exit` (verified on a live session). `looksFailed` checks `output.exitCode` / `output.exit_code`, which do not exist, so every verdict has come from the failure-text regex alone. Read `metadata.exit` first. - **The installed plugin differs from the repo copy.** The main checkout has an uncommitted comment block. It is harmless, but the lift should add a check (install script or a `/health` field reporting the plugin's version string from a header) so a stale install is visible. - **`opencode.json` advertised `auto` with `limit.context: 782324`**, the largest window in the catalog, so opencode only compacted near 782k. Tonight's contexts reached 682k and were resent every turn. Lowered to 200,000 on 2026-09-26 by owner decision (see the end). It caps what `auto` is asked to carry, so Phase 0 should also measure how often a real agent turn now compacts, and whether any task needed more. ## Ground rules for the executor - **Base on `origin/main`, NOT the local `main`.** The local `main` checkout (`c87e675`) is 29 commits behind `origin/main` (`f554fc1` as of 2026-09-26) and 2 ahead with local-only commits. It lacks #99 (conversation identity, `router-link.js`) and #100 (local-encoder rebuild). `git fetch origin` first, then branch `feat/no-progress-detection` from `origin/main` in its own worktree. Read code from that worktree or `git show origin/main:`, never from the main checkout's files. Do not touch, stash or commit anything in the main checkout; it has uncommitted work from other efforts. - The plugin to extend is `deploy/opencode-plugin/router-link.js` (with its `router-link.test.mjs`). `router-outcome.js` was replaced by #99. - **Port 8080 is production; never touch it.** Integration checks run on 8081 via `scripts/sandbox.sh`, against a copy of the live DB made with the sqlite backup API (source opened `mode=ro`). - Never edit `config/config.yaml` or `config/config.local.yaml`. New knobs get defaults in `config.yaml` only as part of this lift's own tracked changes. - **The installed plugin at `~/.config/opencode/plugins/` is live for every opencode session on this machine,** including the one running the lift. Develop and test against the repo copy. Installing a new version is an operator step, stated in the done report, not something the executor does. - **Commit by explicit path. Never `git add -A` or `git add .`** Incident #8's item 4 swept an `.omc/` state file into a commit that way. - **Every worker claim is verified by the orchestrator** against `git show --stat` and the actual test files before acceptance. Incident #8 had a commit message listing tests that did not exist. - **Read large files in targeted slices** (grep for the anchor, then read around it). Incident #8's stuck workers re-read whole files dozens of times. - ASCII only in new code, comments and UI strings; no middle-dot separators. No prose paragraphs in the admin UI. - Lint gate: pinned `uvx ruff@0.16.9`, no NEW findings in touched files compared with the base commit (same recipe as `.omo/plans/cockpit-quick-wins.md` runbook step 5). Never `ruff --fix` outside the lines the item changes. - `pytest` offline stays green after every commit. Record the baseline on `origin/main` first. On `c87e675`, one test, `tests/test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity`, failed before any change; check whether it still does on `origin/main` rather than assuming. - New `route_decisions` columns are registered in `tests/test_tui_schema_drift.py`; every new knob gets an admin control or a recorded `DELIBERATELY_NOT_IN_ADMIN` reason (`tests/test_admin_knob_coverage.py`). ## Phases 0. **Calibrate offline.** Replay over opencode session history; produce the verdict timeline per session and a proposed threshold table. No router or plugin changes. **Start from `plans/no-progress-detection-prototype.py`,** an interim watcher tuned on 2026-09-26 against the 15 sessions of incident #8. On that set it flags every known loop (both item-2 workers, the item-3 first attempt, item 4, the post-completion Atlas stretch, the explore helper stuck on a deleted file, a helper that printed the spec 12 times) and none of the healthy sessions. Four findings from tuning it: - Exact `(tool, args)` fingerprints alone miss reworded commands: 28 differently-worded `cat boulder.json` never matched. - Collapsing to the bare file over-flags ordinary sliced reading, which is what the ground rules tell workers to do. Key a `read` on file plus offset/limit, and a `bash` target on its referenced files plus the numbers in the command (the slice). - Slow loops spread across 300+ calls escape a 60-call window. Add a whole-session rule: one exact call repeated 15+ times with nothing landed recently. - Read-only agents need the separate rule (repeats of one target), since "nothing landed" is always true for them. Its thresholds (`WINDOW=60`, `DUP_MIN=0.25`, `TOP_MIN=12`, `TOP_MIN_RO=8`, `CUM_MIN=15`) are a starting point fitted on one night, not a result. 1. **Identity: build on #99, do not rebuild it.** The `metadata.exit` fix in `router-link.js` (and its `.test.mjs`), plus whatever #99 lacks for progress roll-up, if anything. Deploying #99 itself (syncing the production checkout, installing `router-link.js`) is an operator step. It goes in the done report as a prerequisite for Phase 2 to see real traffic, not an executor task. 2. **Progress events, the judge and the watchdog, `warn` mode.** `/progress`, `progress_events`, verdicts, the `/metrics` warning, toast, Sessions readout, and the knobs. Also the notifier (section 6) with the `desktop` channel and the admin Send test alert button, and the watchdog timer (section 7) with its local-LLM second opinion. **Order within the phase: build the watchdog and the notifier first.** Together with Phase 0's detection code, they protect unattended runs before the plugin half is even deployed. 3. **`auto_compact` and `auto_recover` stage 1.** Abort, compact with the injected report, restart; the scoped 429; the `max_recoveries` hard stop and admin Resume. The incident report builder lands here. 4. **`auto_limit_context` and `auto_recover` stage 2**, using whichever mechanism Phase 0 verified. 5. **Local-model second opinion inside the router's live judge** (section 4), only if earlier phases show ambiguous cases. The watchdog already has one (section 7); this would add it to the per-request path. The v2 port (`project_opencode_v2_pinned`) changes hook registration. Keep the plugin's logic in plain functions so the v1 and v2 shells stay thin. ## Owner decisions Settled 2026-09-26: the response is a mode (`warn`, `auto_compact`, `auto_limit_context`, `auto_recover` with two stages), and every mode warns. `auto_recover` stops and compacts the tree's root (Atlas), not only the stuck worker. Also settled 2026-09-26: - **The default mode is `warn`.** - **`max_recoveries` is 2** per tree per window, then the hard stop. - **The `auto` context limit is 200,000.** Applied the same day, ahead of this lift, in the repo `opencode.json` and the global `~/.config/opencode/opencode.json` (backup at `opencode.json.bak-20260926-auto-context`). The global file pins every agent (atlas, sisyphus-junior, prometheus, ...) to `llm-router/auto`. opencode reads config at startup, so it takes effect for sessions started after a restart. `auto:batch` stays at 782,324 on purpose: batch is where long contexts are expected. - **Alerting ships with `desktop` (`notify-send`) only**, behind a notifier built for more channels. SMS, RingCentral, PagerDuty and webhooks are later adapters (section 6), not part of this lift. No owner decisions remain open. Thresholds come from Phase 0.