Files
6krrt/plans/no-progress-detection.md
adlee-was-taken 69c969d104 plans: declare a valid Status on the six plans this branch adds
tests/test_plans_declare_status.py requires 'Status: <done|planned|in
progress|parked|reference> -- <reason>' in the first 8 lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-26 21:46:15 -04:00

31 KiB

No-progress detection: catch an agent that is spinning, not one that is working

Status: reference -- parent spec; Lift A shipped in PR #102, Lift B (router-side judge, modes, behavior breaker) not started Date: 2026-09-26 Trigger: the cockpit-quick-wins Atlas run, 2026-09-25 19:00 to 2026-09-26 02:00. Incident write-up: docs/incidents.md #8.

What happened, measured

Read-only against the live router.db and opencode's own session store (http://127.0.0.1:4097), 2026-09-26.

Spend, 19:00 to 02:00 local:

  • 323M prompt tokens (95% cached) and about $7.81, over roughly 8 hours.
  • 79% of the money went to OpenRouter z-ai/glm-5.3-flash: 1,368 calls averaging 203k prompt tokens.
  • 921 calls carried 120k+ prompt tokens and cost $6.57. NeuralWatt glm-5.3 took 47 calls averaging 673k context, max 682,679, because nothing else fit.
  • Every decision was classification_source = classifier on the default profile. The router did what it was told; nothing was pinned.

Spend rate is the wrong signal. Tonight's worst hour ($2.42) and worst 3-hour window ($4.13) sit inside the prior 30 days' normal range (hourly p90 $1.46, max $4.83; worst 3h $9.65). A rate alarm quiet enough for a healthy all-day agent run would have missed this entirely, and one tight enough to catch it would fire on good work. The operator's framing is the right one: steady usage is fine. What is not fine is spend with no concrete change landing.

The sessions that wasted money are visibly different in their tool calls:

session (opencode) tool calls exact-duplicate calls worst repeat landed a change?
Item 2 worker, 2nd attempt 423 30% read navbar.js x61 no (worktree unchanged)
Item 3 worker, 1st attempt 132 23% read test_admin_js_units.py x16 no
Item 4 worker 134 19% read .omo/plans/... x8 partly (code, no tests)
Item 2 worker, 1st attempt 307 17% read navbar.js x21 no
Atlas orchestrator 576 11% playwright console check x10 yes (via workers)
healthy workers (items 1, 5, 6, 3-retry) 30-96 0-6% 1-3 yes

A second loop shape, 2026-09-26 02:21 to about 02:56: the Atlas orchestrator itself, after its plan was complete:

  • .omo/boulder.json carried a stale status: completed and pr_url from the previous plan.
  • oh-my-openagent's continuation hook injected "continue" turns; all 3 user turns in the window were injected.
  • Atlas, working from a trimmed context, re-derived its state every time: 54 turns, about 19M cached-read tokens, boulder.json read 28 times, the plan's checkboxes grepped 16 times, nothing landed.

The operator caught it only by watching. The judge must treat this as stalling, and Phase 0 must include it as a calibration case. The plugin can also see injected continuation turns (chat.message), which is a useful extra signal: "continuation turns keep arriving and nothing lands" is this shape exactly.

A third shape, 2026-09-26 around 03:05 to 03:17: a read-only agent. Prometheus's explore subagent reached 322 tool calls, 19% exact duplicates, and read deploy/opencode-plugin/router-outcome.js 20 times. The file had been deleted from under it: the production checkout was synced to origin/main, where #99 retired it. The subagent was aborted.

Read-only agents (explore, librarian, oracle, and the like) never land a change by design, so "no landed change" cannot be their signal. For them the judge uses:

  • target novelty: new files or queries per N calls, and
  • a failure streak on the same target, such as repeated reads of a path that no longer exists.

Identify them by the X-Router-Agent value, via a configurable list of read-only agent names.

"Exact duplicate" means the same tool with byte-identical arguments. Duplicate share alone does not separate the orchestrator (11%, productive) from a stuck worker (17%). Duplicates plus no landed change does.

Why nothing caught it

  • circuit_breaker.py is availability-only: it trips on provider 5xx.
  • The only spend alarm is runway_low_warning: balance / burn < 6 h, with burn averaged over a 24 h balance window. OpenRouter read $24.73 at $0.25/h, so 97 h of runway. It answers "will I run out", not "is this being wasted".
  • The router cannot see progress at all. It sees prompts, not whether an edit landed or a command keeps failing. The opencode plugin sees both.
  • The DEPLOYED router cannot tell conversations apart either. session_key merges concurrent conversations under one multi-agent client, and /outcome attribution falls back to a directory match that 409s when several sessions share a cwd, which is exactly the Atlas-plus-workers shape. The fix already exists on origin/main (PR #99, conversation-identity) but is not deployed: production runs from a local checkout 29 commits behind origin/main, the live route_decisions has no agent / parent_key columns, and the installed plugin is still router-outcome.js rather than #99's router-link.js.

Design

Split the job where the information lives. The plugin is the sensor; the router is the judge. No raw task text leaves opencode, and the router never stores any; the "never store raw task text" rule is unchanged.

1. Exact session identity: ALREADY BUILT upstream (#99), build on it

PR #99 (conversation-identity, merged to origin/main as fe8c813) did this. Reuse it; do not rebuild it:

  • Plugin. deploy/opencode-plugin/router-link.js replaces router-outcome.js. Its chat.headers hook sends X-Router-Conversation (the opencode session id), X-Router-Agent (slugged to the router's charset) and X-Router-Parent, with a bounded parent-lookup cache.
  • Router. src/conversation_identity.py resolves identity from those headers. session_key becomes 'c:' + conversation id. route_decisions gained agent and parent_key columns. /outcome attributes to the exact conversation.

What remains for this lift:

  • Deploy #99. It is merged but not running. Syncing the production checkout to origin/main and installing router-link.js are operator steps: the main checkout has uncommitted work, and part of it, the confidence_min rename, also landed upstream via #100.
  • Fix the exit code in router-link.js. It kept router-outcome.js's output?.exitCode ?? output?.exit_code (line 83 on origin/main); see "Fixes that ride along".
  • Build the progress events (section 2) into router-link.js, keyed by the same conversation and parent ids, so the judge rolls children up to parents through parent_key.

Everywhere below, "session" means #99's conversation. Where this spec says X-Opencode-Session, read X-Router-Conversation.

2. Progress events (plugin tool.execute.after -> router POST /progress)

The plugin sends one small event per tool call, fire-and-forget with a short timeout, never blocking the session:

{session_id, parent_session_id, agent, tool, call_id,
 fingerprint,   # sha256 of tool + normalized args; never the args themselves
 target,        # coarse, normalized: the file path for read/edit/write, the
                #   command's first two tokens for bash; no contents
 exit,          # bash: output.metadata.exit
 landed}        # bool, see below

landed is the load-bearing definition, and it is deliberately narrow:

  • edit / write / patch with a non-empty metadata.diff
  • bash whose command is git commit (or git commit --amend) with exit 0
  • A test or lint command (the plugin's existing TEST_COMMAND list) that PASSES after the session's previous run of the same fingerprint FAILED, i.e. a fix landed.

Reads, greps, todo writes, sub-agent dispatches and passing-again tests are not progress. An orchestrator's progress is its children's: roll child landed events up to the parent via parent_session_id.

Storage: a new progress_events table, pruned to a rolling window (knob). It holds fingerprints, targets and booleans, no content.

3. The judge (router, per client session)

Signals per session, over a rolling window of its last N tool events:

  • turns_since_landed: LLM requests since the last landed event (self plus children for a parent).
  • usd_since_landed: billed spend since then, joined through request_id.
  • dup_share: exact-duplicate fingerprint share in the window.
  • fail_streak: consecutive nonzero exits on the same fingerprint (the "sandbox is broken and the agent keeps retrying" case).
  • context_tokens: the latest request's size, to report alongside, since a stall at 600k context costs far more per turn than one at 30k.

Verdicts:

  • ok: something landed recently.
  • stalling (warn): no landed change for stall_turns turns or stall_usd dollars, AND (dup_share >= dup_warn OR fail_streak >= fail_warn).
  • looping (act, if enabled): the stall condition at the higher loop_* thresholds.

Both conditions have to hold. Long-running healthy work with no duplicates (a big read-heavy exploration) stays ok until it hits the plain turns/spend ceiling, which is a separate, looser knob.

Calibrate before trusting it. Phase 0 replays the judge offline over opencode's session store (the 11 sessions above plus older history), and reports the verdict each session would have received and when. Thresholds are picked from that, not guessed. The Atlas orchestrator (11% duplicates, productive through children) is the key false-positive case, and it must stay ok.

4. Optional: a local model as a second opinion

The deterministic signals should carry this. Where they are ambiguous, for example near-duplicates (the same file read at shifting offsets, or the same failing test with a different -k), there are two cheaper steps before any model:

  • normalize read fingerprints to the path, and bash to command plus target
  • the local encoder already loaded for classifier.mode: local_encoder can embed the target strings to cluster "similar-ish" events

A local LLM judge ("does this digest look like a loop?") runs only on a session already flagged stalling, on a digest of tool names, targets and exit codes, never content. It is gated by local_compute.enabled, since the operator pays the local power bill. Treat this as a later phase, justified by Phase 0 showing ambiguous cases the deterministic signals miss.

5. What happens on a detection: modes

One knob, progress.mode, picks the response. Every mode warns:

  • A /metrics warning (admin bell and TUI). It names the session, agent, turns and $ since the last landed change, and the top repeated target.
  • An opencode toast via POST /tui/show-toast in the window the operator is watching.
  • An alert through the router's notifier (section 6), which reaches the operator when nobody is looking at either screen.
mode on detection
warn Warnings only. The neutral default: today's behaviour plus a warning.
auto_compact Compact the stuck session with the incident report injected, then continue. No refusal, no orchestrator stop.
auto_limit_context Cap the stuck session's context. No stop.
auto_recover Two stages, below.

auto_recover, stage 1 (first detection in a session tree):

  1. The router refuses the stuck session's chat requests (429, scoped by X-Router-Conversation), so nothing more is spent while the plugin acts.
  2. The plugin walks parentID to the tree's root, usually Atlas, and calls POST /session/{id}/abort on the stuck session and on the root.
  3. The plugin compacts the root: POST /session/{root}/summarize. Its experimental.session.compacting hook appends the incident report to the compaction prompt, so the report survives into the compacted context.
  4. experimental.compaction.autocontinue stays enabled, so the root restarts from the compacted context with the report in view. The router lifts the refusal when the compaction completes.
  5. Toast: "recovered : ".

auto_recover, stage 2 (detected again in the same tree within recover_window): apply the auto_limit_context cap to the tree, and toast again.

Stage 3, a hard stop. Recovery can itself loop: recover, loop, recover. After max_recoveries in the window, the router refuses the tree and does not restart it. It toasts and raises a high-severity warning, and the tree waits for an admin Resume. Without this cap, auto-recovery is an unattended retry loop, which is the failure it exists to stop.

The incident report is built deterministically from progress_events (fingerprints, targets and exit codes, never content), so it is free and cannot hallucinate. Example:

[router] A sub-agent stalled and was stopped.
  session: Item 2 G1 nav refactor (Sisyphus-Junior), child of this session
  423 tool calls, 0 landed changes (no edit with a diff, no commit)
  most repeated: read admin/frontend/navbar.js x61, read tests/... x16
  spend since last landed change: $1.84, 70M prompt tokens
Avoid: re-reading whole files, since reads repeat without edits. Give the
worker the exact edit to make, and require an on-disk diff or commit before it
reports done.

The "Avoid:" line maps from the dominant signal: duplicates, a failure streak, or context size. A local model may rephrase it later (section 4); it is not needed to produce it.

"Limit context" has two candidate mechanisms. Phase 0 must verify which works before stage 2 is built:

  • (a) Router-side, per-session tighter pinch budget. The machinery exists (context_prune.py), but pruning rewrites the prefix, and plans/token-waste-waves.md Wave 3 measured rewrites costing the cache. So this trades cache hits for size. Measure it.
  • (b) The router answers an over-cap request with a context-overflow-shaped error, so opencode runs its own overflow compaction. The autocontinue hook receives overflow: boolean, so opencode does have such a path. Unverified: whether it recognizes a router error as an overflow. Test this on 8081 before relying on it.

Also, the Controls or Home page gets a live "Sessions" readout: session, agent, parent, verdict, mode stage, turns and $ since landed, and the top repeated target. This fits plans/cockpit-brainstorm.md. Each tree with a refusal gets a Resume button.

Every threshold and the mode are knobs with admin controls (North Star #1). Per the "build for re-tuning" rule, the neutral default is warn, with Phase 0-calibrated thresholds.

Without the plugin (another client, or a stale install), only the router half works: warnings, alerts and the scoped 429. Abort, compact and restart need the plugin. The Sessions readout says which sessions have a live plugin (from the headers in section 1), so a mode that silently cannot act is visible.

6. Alerting: a notifier with pluggable channels

Toasts and the admin bell only help someone who is looking. Incident #8 ran overnight. Alerts come from the router, not the plugin: the router is the judge and is always running, while the plugin dies with the opencode process it lives in.

Ship now: one channel, desktop (notify-send). Design for later: SMS (Twilio or similar), RingCentral, PagerDuty, a generic webhook (ntfy, Slack). Each of those should be a small adapter added without touching the detector.

Shape:

  • An alert is an event with a lifecycle, not a message. Fields:

    • dedup_key: for this detector, the tree root session id
    • severity: info, warning or critical
    • state: trigger, escalate or resolve
    • title and a one-line summary
    • details, the incident report from section 5
    • source: no-progress, so other detectors can reuse the notifier later

    This maps one-to-one onto PagerDuty Events v2 (dedup_key, event_action: trigger/resolve), so that adapter is thin. Channels without a lifecycle (SMS, desktop) render trigger and escalate as a message and resolve as a short "recovered" message, or skip it (per-channel knob).

  • Transitions, not repeats. An alert fires when a verdict changes, e.g. ok -> stalling, stalling -> looping, a recovery, the hard stop, or -> ok. It never fires per request. Per-channel rate limit and a quiet window stop a flapping session from paging anyone at 3am every minute.

  • Severity routing per channel: each channel has min_severity. The intended setup: desktop gets warning and up; an SMS or PagerDuty channel added later gets only critical, i.e. the hard stop, or looping in warn mode.

  • Delivery never blocks routing. Alerts go through a small queue with a timeout and one retry. A failed delivery is logged and shown in the admin UI; it never slows or fails a chat request.

  • Secrets stay out of config files. Channel credentials (API keys, routing keys, phone numbers) are read from environment variables named in config, the same api_key_env pattern dispatch_providers already uses, and loaded from .env by the systemd unit. config.yaml and config.local.yaml never hold a secret.

  • Config: a notifications.channels list, each entry {name, type, enabled, min_severity, resolve_messages, ...type-specific}. Only type: desktop exists in this lift. An unknown type fails config load (StrictModel), so a typo cannot silently disable alerting.

  • Admin (North Star #1): per channel, an enabled toggle, a min_severity select, and a Send test alert button. The test button is the only way to know a channel works before the night it matters. Recent deliveries and failures show under the Sessions readout.

desktop specifics, to verify in Phase 0. The router runs as a systemd user service, so notify-send needs the session D-Bus (DBUS_SESSION_BUS_ADDRESS, normally unix:path=/run/user/<uid>/bus). Check that the unit's environment has it and that its sandboxing (ProtectHome and friends) allows the socket. If not, say so in the plan and fix it in deploy/, not by weakening the sandbox wholesale. Use urgency critical for critical alerts so they persist on screen.

7. The watchdog timer: detection that needs nobody watching

Why. All three loops in incident #8 were caught by a human or by a Claude Code session reading opencode's session store by hand. The operator does not want to depend on either. The watchdog is an independent, local, scheduled sensor. It runs whether or not anyone is at the screen, whether or not a Claude session is open, and whether or not the plugin loaded. It calls no cloud model.

What. A systemd user timer, deploy/llm-router-watchdog.{service,timer}, shaped like the existing llm-router-* timers, every watchdog.interval (proposed 5 min). Each tick:

  1. Find opencode. Read ~/.local/share/opencode/rc-servers.json, written by the session-registry plugin. Probe each recorded serverUrl. If none answers, exit quietly; no opencode means nothing to watch.
  2. List sessions active since the last tick via GET /session, and read each one's recent messages via GET /session/{id}/message.
  3. Deterministic check, free, on every active session. The same signals as the judge (section 3): exact-duplicate share, the most-repeated target, the failure streak, landed changes (for read-only agents, target novelty instead), and injected continuation turns. This is Phase 0's replay code run live. One implementation, imported by both, not a copy.
  4. Local LLM second opinion, only on sessions step 3 flags. Skipped when local_compute.enabled is off, which keeps gaming mode honest. The model is whatever is configured: a watchdog.model knob that defaults to verification.model, on that section's base URL. Today that is qwen2.5-coder-router:14b on the local Ollama (localhost:11434), already kept warm by local verification, so a call pays no model-load delay. It runs on a digest of tool names, targets and exit codes, never content. Phase 0 measures this model's verdicts on the three incident #8 cases (must flag: the stuck workers and the post-completion Atlas loop; must not flag: productive Atlas) before its answer is trusted. It asks "is this agent looping without progress? yes or no, with a one-line reason". A healthy tick costs zero inference.
  5. Report, do not act. POST the verdict to the router (the /progress judge or a sibling endpoint), tagged source: watchdog. The router applies the configured mode, alerts through the notifier, and enforces max_recoveries and Resume. The watchdog never aborts or compacts anything itself, so there is exactly one recovery path and one place that decides.

If the router is down, the watchdog still sends a critical desktop alert directly with notify-send, and nothing else. A dead router is exactly when an unattended run needs a human.

Failure modes to design for:

  • A tick that overruns the interval: use a lock file and skip the tick, do not stack them.
  • An opencode API shape change. The v2 port breaks this the same way it breaks oc_work_start.py. Detect the unexpected shape and raise one warning-level alert saying the watchdog is blind, rather than silently reporting "all healthy".
  • The watchdog's own health. Record last_tick_at and the result (via the router, or a state file), and show it in the admin Sessions readout. A watchdog that stopped running must be visible.

Knobs (North Star #1):

  • watchdog.enabled
  • watchdog.interval
  • watchdog.local_llm_enabled (default on, still gated by local_compute)
  • watchdog.model (default: verification.model)
  • watchdog.read_only_agents, the list from the read-only case above

The timer ships enabled, since detection plus a warning is the neutral, low-risk default; the actions follow progress.mode.

8. The behavior breaker (Lift B): trip a model that stalls sessions

Gap. The breaker family covers availability (circuit_breaker.py, 5xx) and malformed output (PR #76: mojibake, empty 200s). Nothing trips on well-formed output that does not do the job. In incident #8, z-ai/glm-5.3-flash produced fluent text that confabulated truncation and re-read in loops, even with pinch off and a small context. It was only excluded by hand.

Trigger. Verdicts from this detector (the watchdog, and the router judge once built), attributed to the model(s) that served each flagged session in its window. Trip when:

  • stall verdicts on one model span at least min_sessions distinct sessions (proposed 3) within window (proposed 60 min), AND
  • that model's stall rate is well above the other models' on the same window's traffic, so one hard task cannot blame a model.

Response. The same shape as the existing breakers: a passive routing skip with cooldown and backoff, plus a critical alert through the notifier and an admin Resume. Key it on the model, not the provider, per plans/mangled-output-detection.md's scope reasoning: here one model misbehaved while its siblings on the same provider did not. Record every trip with its evidence (sessions, verdicts, rates) so a human can review it.

Blocked on identity. Joining "this session stalled" to "these models served it" needs #99's conversation headers on each request. As of 2026-09-26 they do not arrive: router-link.js is installed and passes its 29 tests, and the router reads the headers, but production decisions still carry fingerprint session_keys with no agent. Lift B starts by finding out why.

Fixes that ride along

  • The plugin never reads a real exit code. Both router-outcome.js (deployed) and its replacement router-link.js (#99, line 83) have this. In 1.18, tool.execute.after gives output = {title, output, metadata}, and bash's exit is output.metadata.exit (verified on a live session). looksFailed checks output.exitCode / output.exit_code, which do not exist, so every verdict has come from the failure-text regex alone. Read metadata.exit first.
  • The installed plugin differs from the repo copy. The main checkout has an uncommitted comment block. It is harmless, but the lift should add a check (install script or a /health field reporting the plugin's version string from a header) so a stale install is visible.
  • opencode.json advertised auto with limit.context: 782324, the largest window in the catalog, so opencode only compacted near 782k. Tonight's contexts reached 682k and were resent every turn. Lowered to 200,000 on 2026-09-26 by owner decision (see the end). It caps what auto is asked to carry, so Phase 0 should also measure how often a real agent turn now compacts, and whether any task needed more.

Ground rules for the executor

  • Base on origin/main, NOT the local main. The local main checkout (c87e675) is 29 commits behind origin/main (f554fc1 as of 2026-09-26) and 2 ahead with local-only commits. It lacks #99 (conversation identity, router-link.js) and #100 (local-encoder rebuild). git fetch origin first, then branch feat/no-progress-detection from origin/main in its own worktree. Read code from that worktree or git show origin/main:<path>, never from the main checkout's files. Do not touch, stash or commit anything in the main checkout; it has uncommitted work from other efforts.
  • The plugin to extend is deploy/opencode-plugin/router-link.js (with its router-link.test.mjs). router-outcome.js was replaced by #99.
  • Port 8080 is production; never touch it. Integration checks run on 8081 via scripts/sandbox.sh, against a copy of the live DB made with the sqlite backup API (source opened mode=ro).
  • Never edit config/config.yaml or config/config.local.yaml. New knobs get defaults in config.yaml only as part of this lift's own tracked changes.
  • The installed plugin at ~/.config/opencode/plugins/ is live for every opencode session on this machine, including the one running the lift. Develop and test against the repo copy. Installing a new version is an operator step, stated in the done report, not something the executor does.
  • Commit by explicit path. Never git add -A or git add . Incident #8's item 4 swept an .omc/ state file into a commit that way.
  • Every worker claim is verified by the orchestrator against git show --stat and the actual test files before acceptance. Incident #8 had a commit message listing tests that did not exist.
  • Read large files in targeted slices (grep for the anchor, then read around it). Incident #8's stuck workers re-read whole files dozens of times.
  • ASCII only in new code, comments and UI strings; no middle-dot separators. No prose paragraphs in the admin UI.
  • Lint gate: pinned uvx ruff@0.16.9, no NEW findings in touched files compared with the base commit (same recipe as .omo/plans/cockpit-quick-wins.md runbook step 5). Never ruff --fix outside the lines the item changes.
  • pytest offline stays green after every commit. Record the baseline on origin/main first. On c87e675, one test, tests/test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity, failed before any change; check whether it still does on origin/main rather than assuming.
  • New route_decisions columns are registered in tests/test_tui_schema_drift.py; every new knob gets an admin control or a recorded DELIBERATELY_NOT_IN_ADMIN reason (tests/test_admin_knob_coverage.py).

Phases

  1. Calibrate offline. Replay over opencode session history; produce the verdict timeline per session and a proposed threshold table. No router or plugin changes. Start from plans/no-progress-detection-prototype.py, an interim watcher tuned on 2026-09-26 against the 15 sessions of incident #8. On that set it flags every known loop (both item-2 workers, the item-3 first attempt, item 4, the post-completion Atlas stretch, the explore helper stuck on a deleted file, a helper that printed the spec 12 times) and none of the healthy sessions. Four findings from tuning it:
    • Exact (tool, args) fingerprints alone miss reworded commands: 28 differently-worded cat boulder.json never matched.
    • Collapsing to the bare file over-flags ordinary sliced reading, which is what the ground rules tell workers to do. Key a read on file plus offset/limit, and a bash target on its referenced files plus the numbers in the command (the slice).
    • Slow loops spread across 300+ calls escape a 60-call window. Add a whole-session rule: one exact call repeated 15+ times with nothing landed recently.
    • Read-only agents need the separate rule (repeats of one target), since "nothing landed" is always true for them. Its thresholds (WINDOW=60, DUP_MIN=0.25, TOP_MIN=12, TOP_MIN_RO=8, CUM_MIN=15) are a starting point fitted on one night, not a result.
  2. Identity: build on #99, do not rebuild it. The metadata.exit fix in router-link.js (and its .test.mjs), plus whatever #99 lacks for progress roll-up, if anything. Deploying #99 itself (syncing the production checkout, installing router-link.js) is an operator step. It goes in the done report as a prerequisite for Phase 2 to see real traffic, not an executor task.
  3. Progress events, the judge and the watchdog, warn mode. /progress, progress_events, verdicts, the /metrics warning, toast, Sessions readout, and the knobs. Also the notifier (section 6) with the desktop channel and the admin Send test alert button, and the watchdog timer (section 7) with its local-LLM second opinion. Order within the phase: build the watchdog and the notifier first. Together with Phase 0's detection code, they protect unattended runs before the plugin half is even deployed.
  4. auto_compact and auto_recover stage 1. Abort, compact with the injected report, restart; the scoped 429; the max_recoveries hard stop and admin Resume. The incident report builder lands here.
  5. auto_limit_context and auto_recover stage 2, using whichever mechanism Phase 0 verified.
  6. Local-model second opinion inside the router's live judge (section 4), only if earlier phases show ambiguous cases. The watchdog already has one (section 7); this would add it to the per-request path.

The v2 port (project_opencode_v2_pinned) changes hook registration. Keep the plugin's logic in plain functions so the v1 and v2 shells stay thin.

Owner decisions

Settled 2026-09-26: the response is a mode (warn, auto_compact, auto_limit_context, auto_recover with two stages), and every mode warns. auto_recover stops and compacts the tree's root (Atlas), not only the stuck worker.

Also settled 2026-09-26:

  • The default mode is warn.

  • max_recoveries is 2 per tree per window, then the hard stop.

  • The auto context limit is 200,000. Applied the same day, ahead of this lift, in the repo opencode.json and the global ~/.config/opencode/opencode.json (backup at opencode.json.bak-20260926-auto-context). The global file pins every agent (atlas, sisyphus-junior, prometheus, ...) to llm-router/auto. opencode reads config at startup, so it takes effect for sessions started after a restart. auto:batch stays at 782,324 on purpose: batch is where long contexts are expected.

  • Alerting ships with desktop (notify-send) only, behind a notifier built for more channels. SMS, RingCentral, PagerDuty and webhooks are later adapters (section 6), not part of this lift.

No owner decisions remain open. Thresholds come from Phase 0.