tests/test_plans_declare_status.py requires 'Status: <done|planned|in progress|parked|reference> -- <reason>' in the first 8 lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
31 KiB
No-progress detection: catch an agent that is spinning, not one that is working
Status: reference -- parent spec; Lift A shipped in PR #102, Lift B (router-side judge, modes, behavior breaker) not started
Date: 2026-09-26
Trigger: the cockpit-quick-wins Atlas run, 2026-09-25 19:00 to 2026-09-26 02:00.
Incident write-up: docs/incidents.md #8.
What happened, measured
Read-only against the live router.db and opencode's own session store
(http://127.0.0.1:4097), 2026-09-26.
Spend, 19:00 to 02:00 local:
- 323M prompt tokens (95% cached) and about $7.81, over roughly 8 hours.
- 79% of the money went to OpenRouter
z-ai/glm-5.3-flash: 1,368 calls averaging 203k prompt tokens. - 921 calls carried 120k+ prompt tokens and cost $6.57. NeuralWatt
glm-5.3took 47 calls averaging 673k context, max 682,679, because nothing else fit. - Every decision was
classification_source = classifieron the default profile. The router did what it was told; nothing was pinned.
Spend rate is the wrong signal. Tonight's worst hour ($2.42) and worst 3-hour window ($4.13) sit inside the prior 30 days' normal range (hourly p90 $1.46, max $4.83; worst 3h $9.65). A rate alarm quiet enough for a healthy all-day agent run would have missed this entirely, and one tight enough to catch it would fire on good work. The operator's framing is the right one: steady usage is fine. What is not fine is spend with no concrete change landing.
The sessions that wasted money are visibly different in their tool calls:
| session (opencode) | tool calls | exact-duplicate calls | worst repeat | landed a change? |
|---|---|---|---|---|
| Item 2 worker, 2nd attempt | 423 | 30% | read navbar.js x61 |
no (worktree unchanged) |
| Item 3 worker, 1st attempt | 132 | 23% | read test_admin_js_units.py x16 |
no |
| Item 4 worker | 134 | 19% | read .omo/plans/... x8 |
partly (code, no tests) |
| Item 2 worker, 1st attempt | 307 | 17% | read navbar.js x21 |
no |
| Atlas orchestrator | 576 | 11% | playwright console check x10 | yes (via workers) |
| healthy workers (items 1, 5, 6, 3-retry) | 30-96 | 0-6% | 1-3 | yes |
A second loop shape, 2026-09-26 02:21 to about 02:56: the Atlas orchestrator itself, after its plan was complete:
.omo/boulder.jsoncarried a stalestatus: completedandpr_urlfrom the previous plan.- oh-my-openagent's continuation hook injected "continue" turns; all 3 user turns in the window were injected.
- Atlas, working from a trimmed context, re-derived its state every time:
54 turns, about 19M cached-read tokens,
boulder.jsonread 28 times, the plan's checkboxes grepped 16 times, nothing landed.
The operator caught it only by watching. The judge must treat this as
stalling, and Phase 0 must include it as a calibration case. The plugin can
also see injected continuation turns (chat.message), which is a useful extra
signal: "continuation turns keep arriving and nothing lands" is this shape
exactly.
A third shape, 2026-09-26 around 03:05 to 03:17: a read-only agent.
Prometheus's explore subagent reached 322 tool calls, 19% exact duplicates,
and read deploy/opencode-plugin/router-outcome.js 20 times. The file had been
deleted from under it: the production checkout was synced to origin/main,
where #99 retired it. The subagent was aborted.
Read-only agents (explore, librarian, oracle, and the like) never land a
change by design, so "no landed change" cannot be their signal. For them the
judge uses:
- target novelty: new files or queries per N calls, and
- a failure streak on the same target, such as repeated reads of a path that no longer exists.
Identify them by the X-Router-Agent value, via a configurable list of
read-only agent names.
"Exact duplicate" means the same tool with byte-identical arguments. Duplicate share alone does not separate the orchestrator (11%, productive) from a stuck worker (17%). Duplicates plus no landed change does.
Why nothing caught it
circuit_breaker.pyis availability-only: it trips on provider 5xx.- The only spend alarm is
runway_low_warning: balance / burn < 6 h, with burn averaged over a 24 h balance window. OpenRouter read $24.73 at $0.25/h, so 97 h of runway. It answers "will I run out", not "is this being wasted". - The router cannot see progress at all. It sees prompts, not whether an edit landed or a command keeps failing. The opencode plugin sees both.
- The DEPLOYED router cannot tell conversations apart either.
session_keymerges concurrent conversations under one multi-agent client, and/outcomeattribution falls back to a directory match that 409s when several sessions share a cwd, which is exactly the Atlas-plus-workers shape. The fix already exists onorigin/main(PR #99, conversation-identity) but is not deployed: production runs from a local checkout 29 commits behindorigin/main, the liveroute_decisionshas noagent/parent_keycolumns, and the installed plugin is stillrouter-outcome.jsrather than #99'srouter-link.js.
Design
Split the job where the information lives. The plugin is the sensor; the router is the judge. No raw task text leaves opencode, and the router never stores any; the "never store raw task text" rule is unchanged.
1. Exact session identity: ALREADY BUILT upstream (#99), build on it
PR #99 (conversation-identity, merged to origin/main as fe8c813) did this.
Reuse it; do not rebuild it:
- Plugin.
deploy/opencode-plugin/router-link.jsreplacesrouter-outcome.js. Itschat.headershook sendsX-Router-Conversation(the opencode session id),X-Router-Agent(slugged to the router's charset) andX-Router-Parent, with a bounded parent-lookup cache. - Router.
src/conversation_identity.pyresolves identity from those headers.session_keybecomes'c:' + conversation id.route_decisionsgainedagentandparent_keycolumns./outcomeattributes to the exact conversation.
What remains for this lift:
- Deploy #99. It is merged but not running. Syncing the production checkout
to
origin/mainand installingrouter-link.jsare operator steps: the main checkout has uncommitted work, and part of it, theconfidence_minrename, also landed upstream via #100. - Fix the exit code in
router-link.js. It keptrouter-outcome.js'soutput?.exitCode ?? output?.exit_code(line 83 onorigin/main); see "Fixes that ride along". - Build the progress events (section 2) into
router-link.js, keyed by the same conversation and parent ids, so the judge rolls children up to parents throughparent_key.
Everywhere below, "session" means #99's conversation. Where this spec says
X-Opencode-Session, read X-Router-Conversation.
2. Progress events (plugin tool.execute.after -> router POST /progress)
The plugin sends one small event per tool call, fire-and-forget with a short timeout, never blocking the session:
{session_id, parent_session_id, agent, tool, call_id,
fingerprint, # sha256 of tool + normalized args; never the args themselves
target, # coarse, normalized: the file path for read/edit/write, the
# command's first two tokens for bash; no contents
exit, # bash: output.metadata.exit
landed} # bool, see below
landed is the load-bearing definition, and it is deliberately narrow:
edit/write/patchwith a non-emptymetadata.diffbashwhose command isgit commit(orgit commit --amend) with exit 0- A test or lint command (the plugin's existing
TEST_COMMANDlist) that PASSES after the session's previous run of the same fingerprint FAILED, i.e. a fix landed.
Reads, greps, todo writes, sub-agent dispatches and passing-again tests are not
progress. An orchestrator's progress is its children's: roll child landed
events up to the parent via parent_session_id.
Storage: a new progress_events table, pruned to a rolling window (knob). It
holds fingerprints, targets and booleans, no content.
3. The judge (router, per client session)
Signals per session, over a rolling window of its last N tool events:
turns_since_landed: LLM requests since the last landed event (self plus children for a parent).usd_since_landed: billed spend since then, joined throughrequest_id.dup_share: exact-duplicate fingerprint share in the window.fail_streak: consecutive nonzero exits on the same fingerprint (the "sandbox is broken and the agent keeps retrying" case).context_tokens: the latest request's size, to report alongside, since a stall at 600k context costs far more per turn than one at 30k.
Verdicts:
ok: something landed recently.stalling(warn): no landed change forstall_turnsturns orstall_usddollars, AND (dup_share >= dup_warnORfail_streak >= fail_warn).looping(act, if enabled): the stall condition at the higherloop_*thresholds.
Both conditions have to hold. Long-running healthy work with no duplicates (a
big read-heavy exploration) stays ok until it hits the plain turns/spend
ceiling, which is a separate, looser knob.
Calibrate before trusting it. Phase 0 replays the judge offline over
opencode's session store (the 11 sessions above plus older history), and reports
the verdict each session would have received and when. Thresholds are picked
from that, not guessed. The Atlas orchestrator (11% duplicates, productive
through children) is the key false-positive case, and it must stay ok.
4. Optional: a local model as a second opinion
The deterministic signals should carry this. Where they are ambiguous, for
example near-duplicates (the same file read at shifting offsets, or the same
failing test with a different -k), there are two cheaper steps before any
model:
- normalize
readfingerprints to the path, and bash to command plus target - the local encoder already loaded for
classifier.mode: local_encodercan embed thetargetstrings to cluster "similar-ish" events
A local LLM judge ("does this digest look like a loop?") runs only on a session
already flagged stalling, on a digest of tool names, targets and exit codes,
never content. It is gated by local_compute.enabled, since the operator pays
the local power bill. Treat this as a later phase, justified by Phase 0 showing
ambiguous cases the deterministic signals miss.
5. What happens on a detection: modes
One knob, progress.mode, picks the response. Every mode warns:
- A
/metricswarning (admin bell and TUI). It names the session, agent, turns and $ since the last landed change, and the top repeated target. - An opencode toast via
POST /tui/show-toastin the window the operator is watching. - An alert through the router's notifier (section 6), which reaches the operator when nobody is looking at either screen.
| mode | on detection |
|---|---|
warn |
Warnings only. The neutral default: today's behaviour plus a warning. |
auto_compact |
Compact the stuck session with the incident report injected, then continue. No refusal, no orchestrator stop. |
auto_limit_context |
Cap the stuck session's context. No stop. |
auto_recover |
Two stages, below. |
auto_recover, stage 1 (first detection in a session tree):
- The router refuses the stuck session's chat requests (429, scoped by
X-Router-Conversation), so nothing more is spent while the plugin acts. - The plugin walks
parentIDto the tree's root, usually Atlas, and callsPOST /session/{id}/aborton the stuck session and on the root. - The plugin compacts the root:
POST /session/{root}/summarize. Itsexperimental.session.compactinghook appends the incident report to the compaction prompt, so the report survives into the compacted context. experimental.compaction.autocontinuestays enabled, so the root restarts from the compacted context with the report in view. The router lifts the refusal when the compaction completes.- Toast: "recovered : ".
auto_recover, stage 2 (detected again in the same tree within
recover_window): apply the auto_limit_context cap to the tree, and toast
again.
Stage 3, a hard stop. Recovery can itself loop: recover, loop, recover. After
max_recoveries in the window, the router refuses the tree and does not
restart it. It toasts and raises a high-severity warning, and the tree waits for
an admin Resume. Without this cap, auto-recovery is an unattended retry loop,
which is the failure it exists to stop.
The incident report is built deterministically from progress_events
(fingerprints, targets and exit codes, never content), so it is free and cannot
hallucinate. Example:
[router] A sub-agent stalled and was stopped.
session: Item 2 G1 nav refactor (Sisyphus-Junior), child of this session
423 tool calls, 0 landed changes (no edit with a diff, no commit)
most repeated: read admin/frontend/navbar.js x61, read tests/... x16
spend since last landed change: $1.84, 70M prompt tokens
Avoid: re-reading whole files, since reads repeat without edits. Give the
worker the exact edit to make, and require an on-disk diff or commit before it
reports done.
The "Avoid:" line maps from the dominant signal: duplicates, a failure streak, or context size. A local model may rephrase it later (section 4); it is not needed to produce it.
"Limit context" has two candidate mechanisms. Phase 0 must verify which works before stage 2 is built:
- (a) Router-side, per-session tighter
pinchbudget. The machinery exists (context_prune.py), but pruning rewrites the prefix, andplans/token-waste-waves.mdWave 3 measured rewrites costing the cache. So this trades cache hits for size. Measure it. - (b) The router answers an over-cap request with a context-overflow-shaped
error, so opencode runs its own overflow compaction. The autocontinue hook
receives
overflow: boolean, so opencode does have such a path. Unverified: whether it recognizes a router error as an overflow. Test this on 8081 before relying on it.
Also, the Controls or Home page gets a live "Sessions" readout: session, agent,
parent, verdict, mode stage, turns and $ since landed, and the top repeated
target. This fits plans/cockpit-brainstorm.md. Each tree with a refusal gets
a Resume button.
Every threshold and the mode are knobs with admin controls (North Star #1).
Per the "build for re-tuning" rule, the neutral default is warn, with Phase
0-calibrated thresholds.
Without the plugin (another client, or a stale install), only the router half works: warnings, alerts and the scoped 429. Abort, compact and restart need the plugin. The Sessions readout says which sessions have a live plugin (from the headers in section 1), so a mode that silently cannot act is visible.
6. Alerting: a notifier with pluggable channels
Toasts and the admin bell only help someone who is looking. Incident #8 ran overnight. Alerts come from the router, not the plugin: the router is the judge and is always running, while the plugin dies with the opencode process it lives in.
Ship now: one channel, desktop (notify-send). Design for later: SMS
(Twilio or similar), RingCentral, PagerDuty, a generic webhook (ntfy, Slack).
Each of those should be a small adapter added without touching the detector.
Shape:
-
An alert is an event with a lifecycle, not a message. Fields:
dedup_key: for this detector, the tree root session idseverity:info,warningorcriticalstate:trigger,escalateorresolvetitleand a one-linesummarydetails, the incident report from section 5source:no-progress, so other detectors can reuse the notifier later
This maps one-to-one onto PagerDuty Events v2 (
dedup_key,event_action: trigger/resolve), so that adapter is thin. Channels without a lifecycle (SMS, desktop) rendertriggerandescalateas a message andresolveas a short "recovered" message, or skip it (per-channel knob). -
Transitions, not repeats. An alert fires when a verdict changes, e.g.
ok -> stalling,stalling -> looping, a recovery, the hard stop, or-> ok. It never fires per request. Per-channel rate limit and a quiet window stop a flapping session from paging anyone at 3am every minute. -
Severity routing per channel: each channel has
min_severity. The intended setup: desktop getswarningand up; an SMS or PagerDuty channel added later gets onlycritical, i.e. the hard stop, orloopinginwarnmode. -
Delivery never blocks routing. Alerts go through a small queue with a timeout and one retry. A failed delivery is logged and shown in the admin UI; it never slows or fails a chat request.
-
Secrets stay out of config files. Channel credentials (API keys, routing keys, phone numbers) are read from environment variables named in config, the same
api_key_envpatterndispatch_providersalready uses, and loaded from.envby the systemd unit.config.yamlandconfig.local.yamlnever hold a secret. -
Config: a
notifications.channelslist, each entry{name, type, enabled, min_severity, resolve_messages, ...type-specific}. Onlytype: desktopexists in this lift. An unknowntypefails config load (StrictModel), so a typo cannot silently disable alerting. -
Admin (North Star #1): per channel, an enabled toggle, a
min_severityselect, and a Send test alert button. The test button is the only way to know a channel works before the night it matters. Recent deliveries and failures show under the Sessions readout.
desktop specifics, to verify in Phase 0. The router runs as a systemd
user service, so notify-send needs the session D-Bus
(DBUS_SESSION_BUS_ADDRESS, normally unix:path=/run/user/<uid>/bus). Check
that the unit's environment has it and that its sandboxing (ProtectHome and
friends) allows the socket. If not, say so in the plan and fix it in
deploy/, not by weakening the sandbox wholesale. Use urgency critical for
critical alerts so they persist on screen.
7. The watchdog timer: detection that needs nobody watching
Why. All three loops in incident #8 were caught by a human or by a Claude Code session reading opencode's session store by hand. The operator does not want to depend on either. The watchdog is an independent, local, scheduled sensor. It runs whether or not anyone is at the screen, whether or not a Claude session is open, and whether or not the plugin loaded. It calls no cloud model.
What. A systemd user timer, deploy/llm-router-watchdog.{service,timer},
shaped like the existing llm-router-* timers, every watchdog.interval
(proposed 5 min). Each tick:
- Find opencode. Read
~/.local/share/opencode/rc-servers.json, written by thesession-registryplugin. Probe each recordedserverUrl. If none answers, exit quietly; no opencode means nothing to watch. - List sessions active since the last tick via
GET /session, and read each one's recent messages viaGET /session/{id}/message. - Deterministic check, free, on every active session. The same signals as the judge (section 3): exact-duplicate share, the most-repeated target, the failure streak, landed changes (for read-only agents, target novelty instead), and injected continuation turns. This is Phase 0's replay code run live. One implementation, imported by both, not a copy.
- Local LLM second opinion, only on sessions step 3 flags. Skipped when
local_compute.enabledis off, which keeps gaming mode honest. The model is whatever is configured: awatchdog.modelknob that defaults toverification.model, on that section's base URL. Today that isqwen2.5-coder-router:14bon the local Ollama (localhost:11434), already kept warm by local verification, so a call pays no model-load delay. It runs on a digest of tool names, targets and exit codes, never content. Phase 0 measures this model's verdicts on the three incident #8 cases (must flag: the stuck workers and the post-completion Atlas loop; must not flag: productive Atlas) before its answer is trusted. It asks "is this agent looping without progress? yes or no, with a one-line reason". A healthy tick costs zero inference. - Report, do not act. POST the verdict to the router (the
/progressjudge or a sibling endpoint), taggedsource: watchdog. The router applies the configured mode, alerts through the notifier, and enforcesmax_recoveriesand Resume. The watchdog never aborts or compacts anything itself, so there is exactly one recovery path and one place that decides.
If the router is down, the watchdog still sends a critical desktop
alert directly with notify-send, and nothing else. A dead router is exactly
when an unattended run needs a human.
Failure modes to design for:
- A tick that overruns the interval: use a lock file and skip the tick, do not stack them.
- An opencode API shape change. The v2 port breaks this the same way it breaks
oc_work_start.py. Detect the unexpected shape and raise onewarning-level alert saying the watchdog is blind, rather than silently reporting "all healthy". - The watchdog's own health. Record
last_tick_atand the result (via the router, or a state file), and show it in the admin Sessions readout. A watchdog that stopped running must be visible.
Knobs (North Star #1):
watchdog.enabledwatchdog.intervalwatchdog.local_llm_enabled(default on, still gated bylocal_compute)watchdog.model(default:verification.model)watchdog.read_only_agents, the list from the read-only case above
The timer ships enabled, since detection plus a warning is the neutral,
low-risk default; the actions follow progress.mode.
8. The behavior breaker (Lift B): trip a model that stalls sessions
Gap. The breaker family covers availability (circuit_breaker.py, 5xx) and
malformed output (PR #76: mojibake, empty 200s). Nothing trips on
well-formed output that does not do the job. In incident #8,
z-ai/glm-5.3-flash produced fluent text that confabulated truncation and
re-read in loops, even with pinch off and a small context. It was only
excluded by hand.
Trigger. Verdicts from this detector (the watchdog, and the router judge once built), attributed to the model(s) that served each flagged session in its window. Trip when:
- stall verdicts on one model span at least
min_sessionsdistinct sessions (proposed 3) withinwindow(proposed 60 min), AND - that model's stall rate is well above the other models' on the same window's traffic, so one hard task cannot blame a model.
Response. The same shape as the existing breakers: a passive routing skip
with cooldown and backoff, plus a critical alert through the notifier and an
admin Resume. Key it on the model, not the provider, per
plans/mangled-output-detection.md's scope reasoning: here one model misbehaved
while its siblings on the same provider did not. Record every trip with its
evidence (sessions, verdicts, rates) so a human can review it.
Blocked on identity. Joining "this session stalled" to "these models served
it" needs #99's conversation headers on each request. As of 2026-09-26 they do
not arrive: router-link.js is installed and passes its 29 tests, and the router
reads the headers, but production decisions still carry fingerprint
session_keys with no agent. Lift B starts by finding out why.
Fixes that ride along
- The plugin never reads a real exit code. Both
router-outcome.js(deployed) and its replacementrouter-link.js(#99, line 83) have this. In 1.18,tool.execute.aftergivesoutput = {title, output, metadata}, and bash's exit isoutput.metadata.exit(verified on a live session).looksFailedchecksoutput.exitCode/output.exit_code, which do not exist, so every verdict has come from the failure-text regex alone. Readmetadata.exitfirst. - The installed plugin differs from the repo copy. The main checkout has an
uncommitted comment block. It is harmless, but the lift should add a check
(install script or a
/healthfield reporting the plugin's version string from a header) so a stale install is visible. opencode.jsonadvertisedautowithlimit.context: 782324, the largest window in the catalog, so opencode only compacted near 782k. Tonight's contexts reached 682k and were resent every turn. Lowered to 200,000 on 2026-09-26 by owner decision (see the end). It caps whatautois asked to carry, so Phase 0 should also measure how often a real agent turn now compacts, and whether any task needed more.
Ground rules for the executor
- Base on
origin/main, NOT the localmain. The localmaincheckout (c87e675) is 29 commits behindorigin/main(f554fc1as of 2026-09-26) and 2 ahead with local-only commits. It lacks #99 (conversation identity,router-link.js) and #100 (local-encoder rebuild).git fetch originfirst, then branchfeat/no-progress-detectionfromorigin/mainin its own worktree. Read code from that worktree orgit show origin/main:<path>, never from the main checkout's files. Do not touch, stash or commit anything in the main checkout; it has uncommitted work from other efforts. - The plugin to extend is
deploy/opencode-plugin/router-link.js(with itsrouter-link.test.mjs).router-outcome.jswas replaced by #99. - Port 8080 is production; never touch it. Integration checks run on 8081
via
scripts/sandbox.sh, against a copy of the live DB made with the sqlite backup API (source openedmode=ro). - Never edit
config/config.yamlorconfig/config.local.yaml. New knobs get defaults inconfig.yamlonly as part of this lift's own tracked changes. - The installed plugin at
~/.config/opencode/plugins/is live for every opencode session on this machine, including the one running the lift. Develop and test against the repo copy. Installing a new version is an operator step, stated in the done report, not something the executor does. - Commit by explicit path. Never
git add -Aorgit add .Incident #8's item 4 swept an.omc/state file into a commit that way. - Every worker claim is verified by the orchestrator against
git show --statand the actual test files before acceptance. Incident #8 had a commit message listing tests that did not exist. - Read large files in targeted slices (grep for the anchor, then read around it). Incident #8's stuck workers re-read whole files dozens of times.
- ASCII only in new code, comments and UI strings; no middle-dot separators. No prose paragraphs in the admin UI.
- Lint gate: pinned
uvx ruff@0.16.9, no NEW findings in touched files compared with the base commit (same recipe as.omo/plans/cockpit-quick-wins.mdrunbook step 5). Neverruff --fixoutside the lines the item changes. pytestoffline stays green after every commit. Record the baseline onorigin/mainfirst. Onc87e675, one test,tests/test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity, failed before any change; check whether it still does onorigin/mainrather than assuming.- New
route_decisionscolumns are registered intests/test_tui_schema_drift.py; every new knob gets an admin control or a recordedDELIBERATELY_NOT_IN_ADMINreason (tests/test_admin_knob_coverage.py).
Phases
- Calibrate offline. Replay over opencode session history; produce the
verdict timeline per session and a proposed threshold table. No router or
plugin changes. Start from
plans/no-progress-detection-prototype.py, an interim watcher tuned on 2026-09-26 against the 15 sessions of incident #8. On that set it flags every known loop (both item-2 workers, the item-3 first attempt, item 4, the post-completion Atlas stretch, the explore helper stuck on a deleted file, a helper that printed the spec 12 times) and none of the healthy sessions. Four findings from tuning it:- Exact
(tool, args)fingerprints alone miss reworded commands: 28 differently-wordedcat boulder.jsonnever matched. - Collapsing to the bare file over-flags ordinary sliced reading, which is
what the ground rules tell workers to do. Key a
readon file plus offset/limit, and abashtarget on its referenced files plus the numbers in the command (the slice). - Slow loops spread across 300+ calls escape a 60-call window. Add a whole-session rule: one exact call repeated 15+ times with nothing landed recently.
- Read-only agents need the separate rule (repeats of one target), since
"nothing landed" is always true for them.
Its thresholds (
WINDOW=60,DUP_MIN=0.25,TOP_MIN=12,TOP_MIN_RO=8,CUM_MIN=15) are a starting point fitted on one night, not a result.
- Exact
- Identity: build on #99, do not rebuild it. The
metadata.exitfix inrouter-link.js(and its.test.mjs), plus whatever #99 lacks for progress roll-up, if anything. Deploying #99 itself (syncing the production checkout, installingrouter-link.js) is an operator step. It goes in the done report as a prerequisite for Phase 2 to see real traffic, not an executor task. - Progress events, the judge and the watchdog,
warnmode./progress,progress_events, verdicts, the/metricswarning, toast, Sessions readout, and the knobs. Also the notifier (section 6) with thedesktopchannel and the admin Send test alert button, and the watchdog timer (section 7) with its local-LLM second opinion. Order within the phase: build the watchdog and the notifier first. Together with Phase 0's detection code, they protect unattended runs before the plugin half is even deployed. auto_compactandauto_recoverstage 1. Abort, compact with the injected report, restart; the scoped 429; themax_recoverieshard stop and admin Resume. The incident report builder lands here.auto_limit_contextandauto_recoverstage 2, using whichever mechanism Phase 0 verified.- Local-model second opinion inside the router's live judge (section 4), only if earlier phases show ambiguous cases. The watchdog already has one (section 7); this would add it to the per-request path.
The v2 port (project_opencode_v2_pinned) changes hook registration. Keep the
plugin's logic in plain functions so the v1 and v2 shells stay thin.
Owner decisions
Settled 2026-09-26: the response is a mode (warn, auto_compact,
auto_limit_context, auto_recover with two stages), and every mode warns.
auto_recover stops and compacts the tree's root (Atlas), not only the stuck
worker.
Also settled 2026-09-26:
-
The default mode is
warn. -
max_recoveriesis 2 per tree per window, then the hard stop. -
The
autocontext limit is 200,000. Applied the same day, ahead of this lift, in the repoopencode.jsonand the global~/.config/opencode/opencode.json(backup atopencode.json.bak-20260926-auto-context). The global file pins every agent (atlas, sisyphus-junior, prometheus, ...) tollm-router/auto. opencode reads config at startup, so it takes effect for sessions started after a restart.auto:batchstays at 782,324 on purpose: batch is where long contexts are expected. -
Alerting ships with
desktop(notify-send) only, behind a notifier built for more channels. SMS, RingCentral, PagerDuty and webhooks are later adapters (section 6), not part of this lift.
No owner decisions remain open. Thresholds come from Phase 0.