Files
6krrt/plans/no-progress-detection.md
adlee-was-taken 69c969d104 plans: declare a valid Status on the six plans this branch adds
tests/test_plans_declare_status.py requires 'Status: <done|planned|in
progress|parked|reference> -- <reason>' in the first 8 lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-26 21:46:15 -04:00

573 lines
31 KiB
Markdown

# No-progress detection: catch an agent that is spinning, not one that is working
Status: reference -- parent spec; Lift A shipped in PR #102, Lift B (router-side judge, modes, behavior breaker) not started
Date: 2026-09-26
Trigger: the `cockpit-quick-wins` Atlas run, 2026-09-25 19:00 to 2026-09-26 02:00.
Incident write-up: `docs/incidents.md` #8.
## What happened, measured
Read-only against the live `router.db` and opencode's own session store
(`http://127.0.0.1:4097`), 2026-09-26.
Spend, 19:00 to 02:00 local:
- 323M prompt tokens (95% cached) and about $7.81, over roughly 8 hours.
- 79% of the money went to OpenRouter `z-ai/glm-5.3-flash`: 1,368 calls
averaging 203k prompt tokens.
- 921 calls carried 120k+ prompt tokens and cost $6.57. NeuralWatt `glm-5.3`
took 47 calls averaging 673k context, max 682,679, because nothing else fit.
- Every decision was `classification_source = classifier` on the default
profile. The router did what it was told; nothing was pinned.
**Spend rate is the wrong signal.** Tonight's worst hour ($2.42) and worst
3-hour window ($4.13) sit inside the prior 30 days' normal range (hourly p90
$1.46, max $4.83; worst 3h $9.65). A rate alarm quiet enough for a healthy
all-day agent run would have missed this entirely, and one tight enough to catch
it would fire on good work. The operator's framing is the right one: steady
usage is fine. What is not fine is spend with **no concrete change landing**.
**The sessions that wasted money are visibly different in their tool calls:**
| session (opencode) | tool calls | exact-duplicate calls | worst repeat | landed a change? |
|---|---|---|---|---|
| Item 2 worker, 2nd attempt | 423 | 30% | `read navbar.js` x61 | no (worktree unchanged) |
| Item 3 worker, 1st attempt | 132 | 23% | `read test_admin_js_units.py` x16 | no |
| Item 4 worker | 134 | 19% | `read .omo/plans/...` x8 | partly (code, no tests) |
| Item 2 worker, 1st attempt | 307 | 17% | `read navbar.js` x21 | no |
| Atlas orchestrator | 576 | 11% | playwright console check x10 | yes (via workers) |
| healthy workers (items 1, 5, 6, 3-retry) | 30-96 | 0-6% | 1-3 | yes |
**A second loop shape, 2026-09-26 02:21 to about 02:56:** the Atlas orchestrator
itself, after its plan was complete:
- `.omo/boulder.json` carried a stale `status: completed` and `pr_url` from the
previous plan.
- oh-my-openagent's continuation hook injected "continue" turns; all 3 user
turns in the window were injected.
- Atlas, working from a trimmed context, re-derived its state every time:
54 turns, about 19M cached-read tokens, `boulder.json` read 28 times, the
plan's checkboxes grepped 16 times, nothing landed.
The operator caught it only by watching. The judge must treat this as
`stalling`, and Phase 0 must include it as a calibration case. The plugin can
also see injected continuation turns (`chat.message`), which is a useful extra
signal: "continuation turns keep arriving and nothing lands" is this shape
exactly.
**A third shape, 2026-09-26 around 03:05 to 03:17: a read-only agent.**
Prometheus's `explore` subagent reached 322 tool calls, 19% exact duplicates,
and read `deploy/opencode-plugin/router-outcome.js` 20 times. The file had been
deleted from under it: the production checkout was synced to `origin/main`,
where #99 retired it. The subagent was aborted.
Read-only agents (`explore`, `librarian`, `oracle`, and the like) never land a
change by design, so "no landed change" cannot be their signal. For them the
judge uses:
- target novelty: new files or queries per N calls, and
- a failure streak on the same target, such as repeated reads of a path that
no longer exists.
Identify them by the `X-Router-Agent` value, via a configurable list of
read-only agent names.
"Exact duplicate" means the same tool with byte-identical arguments. Duplicate
share alone does not separate the orchestrator (11%, productive) from a stuck
worker (17%). Duplicates **plus** no landed change does.
## Why nothing caught it
- `circuit_breaker.py` is availability-only: it trips on provider 5xx.
- The only spend alarm is `runway_low_warning`: balance / burn < 6 h, with burn
averaged over a 24 h balance window. OpenRouter read $24.73 at $0.25/h, so 97 h
of runway. It answers "will I run out", not "is this being wasted".
- The router cannot see progress at all. It sees prompts, not whether an edit
landed or a command keeps failing. The opencode plugin sees both.
- The DEPLOYED router cannot tell conversations apart either. `session_key`
merges concurrent conversations under one multi-agent client, and `/outcome`
attribution falls back to a directory match that 409s when several sessions
share a cwd, which is exactly the Atlas-plus-workers shape. **The fix already
exists on `origin/main` (PR #99, conversation-identity) but is not deployed:**
production runs from a local checkout 29 commits behind `origin/main`, the
live `route_decisions` has no `agent` / `parent_key` columns, and the
installed plugin is still `router-outcome.js` rather than #99's
`router-link.js`.
## Design
Split the job where the information lives. **The plugin is the sensor; the
router is the judge.** No raw task text leaves opencode, and the router never
stores any; the "never store raw task text" rule is unchanged.
### 1. Exact session identity: ALREADY BUILT upstream (#99), build on it
PR #99 (`conversation-identity`, merged to `origin/main` as `fe8c813`) did this.
Reuse it; do not rebuild it:
- **Plugin.** `deploy/opencode-plugin/router-link.js` replaces
`router-outcome.js`. Its `chat.headers` hook sends `X-Router-Conversation`
(the opencode session id), `X-Router-Agent` (slugged to the router's charset)
and `X-Router-Parent`, with a bounded parent-lookup cache.
- **Router.** `src/conversation_identity.py` resolves identity from those
headers. `session_key` becomes `'c:' + conversation id`. `route_decisions`
gained `agent` and `parent_key` columns. `/outcome` attributes to the exact
conversation.
What remains for this lift:
- **Deploy #99.** It is merged but not running. Syncing the production checkout
to `origin/main` and installing `router-link.js` are operator steps: the main
checkout has uncommitted work, and part of it, the `confidence_min` rename,
also landed upstream via #100.
- **Fix the exit code in `router-link.js`.** It kept `router-outcome.js`'s
`output?.exitCode ?? output?.exit_code` (line 83 on `origin/main`); see "Fixes
that ride along".
- **Build the progress events (section 2) into `router-link.js`,** keyed by the
same conversation and parent ids, so the judge rolls children up to parents
through `parent_key`.
Everywhere below, "session" means #99's conversation. Where this spec says
`X-Opencode-Session`, read `X-Router-Conversation`.
### 2. Progress events (plugin `tool.execute.after` -> router `POST /progress`)
The plugin sends one small event per tool call, fire-and-forget with a short
timeout, never blocking the session:
```
{session_id, parent_session_id, agent, tool, call_id,
fingerprint, # sha256 of tool + normalized args; never the args themselves
target, # coarse, normalized: the file path for read/edit/write, the
# command's first two tokens for bash; no contents
exit, # bash: output.metadata.exit
landed} # bool, see below
```
`landed` is the load-bearing definition, and it is deliberately narrow:
- `edit` / `write` / `patch` with a non-empty `metadata.diff`
- `bash` whose command is `git commit` (or `git commit --amend`) with exit 0
- A test or lint command (the plugin's existing `TEST_COMMAND` list) that PASSES
after the session's previous run of the same fingerprint FAILED, i.e. a fix
landed.
Reads, greps, todo writes, sub-agent dispatches and passing-again tests are not
progress. An orchestrator's progress is its children's: roll child `landed`
events up to the parent via `parent_session_id`.
Storage: a new `progress_events` table, pruned to a rolling window (knob). It
holds fingerprints, targets and booleans, no content.
### 3. The judge (router, per client session)
Signals per session, over a rolling window of its last N tool events:
- `turns_since_landed`: LLM requests since the last landed event (self plus
children for a parent).
- `usd_since_landed`: billed spend since then, joined through `request_id`.
- `dup_share`: exact-duplicate fingerprint share in the window.
- `fail_streak`: consecutive nonzero exits on the same fingerprint (the
"sandbox is broken and the agent keeps retrying" case).
- `context_tokens`: the latest request's size, to report alongside, since a
stall at 600k context costs far more per turn than one at 30k.
Verdicts:
- `ok`: something landed recently.
- `stalling` (warn): no landed change for `stall_turns` turns or
`stall_usd` dollars, AND (`dup_share >= dup_warn` OR `fail_streak >= fail_warn`).
- `looping` (act, if enabled): the stall condition at the higher `loop_*`
thresholds.
Both conditions have to hold. Long-running healthy work with no duplicates (a
big read-heavy exploration) stays `ok` until it hits the plain turns/spend
ceiling, which is a separate, looser knob.
**Calibrate before trusting it.** Phase 0 replays the judge offline over
opencode's session store (the 11 sessions above plus older history), and reports
the verdict each session would have received and when. Thresholds are picked
from that, not guessed. The Atlas orchestrator (11% duplicates, productive
through children) is the key false-positive case, and it must stay `ok`.
### 4. Optional: a local model as a second opinion
The deterministic signals should carry this. Where they are ambiguous, for
example near-duplicates (the same file read at shifting offsets, or the same
failing test with a different `-k`), there are two cheaper steps before any
model:
- normalize `read` fingerprints to the path, and bash to command plus target
- the local encoder already loaded for `classifier.mode: local_encoder` can
embed the `target` strings to cluster "similar-ish" events
A local LLM judge ("does this digest look like a loop?") runs only on a session
already flagged `stalling`, on a digest of tool names, targets and exit codes,
never content. It is gated by `local_compute.enabled`, since the operator pays
the local power bill. Treat this as a later phase, justified by Phase 0 showing
ambiguous cases the deterministic signals miss.
### 5. What happens on a detection: modes
One knob, `progress.mode`, picks the response. **Every mode warns:**
- A `/metrics` warning (admin bell and TUI). It names the session, agent, turns
and $ since the last landed change, and the top repeated target.
- An opencode toast via `POST /tui/show-toast` in the window the operator is
watching.
- An **alert** through the router's notifier (section 6), which reaches the
operator when nobody is looking at either screen.
| mode | on detection |
|---|---|
| `warn` | Warnings only. The neutral default: today's behaviour plus a warning. |
| `auto_compact` | Compact the stuck session with the incident report injected, then continue. No refusal, no orchestrator stop. |
| `auto_limit_context` | Cap the stuck session's context. No stop. |
| `auto_recover` | Two stages, below. |
**`auto_recover`, stage 1 (first detection in a session tree):**
1. The router refuses the stuck session's chat requests (429, scoped by
`X-Router-Conversation`), so nothing more is spent while the plugin acts.
2. The plugin walks `parentID` to the tree's root, usually Atlas, and calls
`POST /session/{id}/abort` on the stuck session and on the root.
3. The plugin compacts the root: `POST /session/{root}/summarize`. Its
`experimental.session.compacting` hook appends the **incident report** to
the compaction prompt, so the report survives into the compacted context.
4. `experimental.compaction.autocontinue` stays enabled, so the root restarts
from the compacted context with the report in view. The router lifts the
refusal when the compaction completes.
5. Toast: "recovered <agent>: <one-line reason>".
**`auto_recover`, stage 2 (detected again in the same tree within
`recover_window`):** apply the `auto_limit_context` cap to the tree, and toast
again.
**Stage 3, a hard stop.** Recovery can itself loop: recover, loop, recover. After
`max_recoveries` in the window, the router refuses the tree and does not
restart it. It toasts and raises a high-severity warning, and the tree waits for
an admin Resume. Without this cap, auto-recovery is an unattended retry loop,
which is the failure it exists to stop.
**The incident report** is built deterministically from `progress_events`
(fingerprints, targets and exit codes, never content), so it is free and cannot
hallucinate. Example:
```
[router] A sub-agent stalled and was stopped.
session: Item 2 G1 nav refactor (Sisyphus-Junior), child of this session
423 tool calls, 0 landed changes (no edit with a diff, no commit)
most repeated: read admin/frontend/navbar.js x61, read tests/... x16
spend since last landed change: $1.84, 70M prompt tokens
Avoid: re-reading whole files, since reads repeat without edits. Give the
worker the exact edit to make, and require an on-disk diff or commit before it
reports done.
```
The "Avoid:" line maps from the dominant signal: duplicates, a failure streak,
or context size. A local model may rephrase it later (section 4); it is not
needed to produce it.
**"Limit context"** has two candidate mechanisms. Phase 0 must verify which works
before stage 2 is built:
- **(a) Router-side, per-session tighter `pinch` budget.** The machinery
exists (`context_prune.py`), but pruning rewrites the prefix, and
`plans/token-waste-waves.md` Wave 3 measured rewrites costing the cache. So
this trades cache hits for size. Measure it.
- **(b) The router answers an over-cap request with a context-overflow-shaped
error**, so opencode runs its own overflow compaction. The autocontinue hook
receives `overflow: boolean`, so opencode does have such a path. Unverified:
whether it recognizes a router error as an overflow. Test this on 8081 before
relying on it.
Also, the Controls or Home page gets a live "Sessions" readout: session, agent,
parent, verdict, mode stage, turns and $ since landed, and the top repeated
target. This fits `plans/cockpit-brainstorm.md`. Each tree with a refusal gets
a Resume button.
Every threshold and the mode are knobs with admin controls (North Star #1).
Per the "build for re-tuning" rule, the neutral default is `warn`, with Phase
0-calibrated thresholds.
Without the plugin (another client, or a stale install), only the router half
works: warnings, alerts and the scoped 429. Abort, compact and restart need the
plugin. The Sessions readout says which sessions have a live plugin (from the
headers in section 1), so a mode that silently cannot act is visible.
### 6. Alerting: a notifier with pluggable channels
Toasts and the admin bell only help someone who is looking. Incident #8 ran
overnight. Alerts come from the **router**, not the plugin: the router is the
judge and is always running, while the plugin dies with the opencode process
it lives in.
**Ship now:** one channel, `desktop` (`notify-send`). **Design for later:** SMS
(Twilio or similar), RingCentral, PagerDuty, a generic webhook (ntfy, Slack).
Each of those should be a small adapter added without touching the detector.
Shape:
- **An alert is an event with a lifecycle**, not a message. Fields:
- `dedup_key`: for this detector, the tree root session id
- `severity`: `info`, `warning` or `critical`
- `state`: `trigger`, `escalate` or `resolve`
- `title` and a one-line `summary`
- `details`, the incident report from section 5
- `source`: `no-progress`, so other detectors can reuse the notifier later
This maps one-to-one onto PagerDuty Events v2 (`dedup_key`,
`event_action: trigger/resolve`), so that adapter is thin. Channels without
a lifecycle (SMS, desktop) render `trigger` and `escalate` as a message and
`resolve` as a short "recovered" message, or skip it (per-channel knob).
- **Transitions, not repeats.** An alert fires when a verdict *changes*, e.g.
`ok -> stalling`, `stalling -> looping`, a recovery, the hard stop, or
`-> ok`. It never fires per request. Per-channel rate limit and a quiet
window stop a flapping session from paging anyone at 3am every minute.
- **Severity routing per channel:** each channel has `min_severity`. The
intended setup: desktop gets `warning` and up; an SMS or PagerDuty channel
added later gets only `critical`, i.e. the hard stop, or `looping` in `warn`
mode.
- **Delivery never blocks routing.** Alerts go through a small queue with a
timeout and one retry. A failed delivery is logged and shown in the admin
UI; it never slows or fails a chat request.
- **Secrets stay out of config files.** Channel credentials (API keys, routing
keys, phone numbers) are read from environment variables named in config,
the same `api_key_env` pattern `dispatch_providers` already uses, and loaded
from `.env` by the systemd unit. `config.yaml` and `config.local.yaml` never
hold a secret.
- **Config:** a `notifications.channels` list, each entry
`{name, type, enabled, min_severity, resolve_messages, ...type-specific}`.
Only `type: desktop` exists in this lift. An unknown `type` fails config load
(StrictModel), so a typo cannot silently disable alerting.
- **Admin (North Star #1):** per channel, an enabled toggle, a `min_severity`
select, and a **Send test alert** button. The test button is the only way to
know a channel works before the night it matters. Recent deliveries and
failures show under the Sessions readout.
**`desktop` specifics, to verify in Phase 0.** The router runs as a systemd
*user* service, so `notify-send` needs the session D-Bus
(`DBUS_SESSION_BUS_ADDRESS`, normally `unix:path=/run/user/<uid>/bus`). Check
that the unit's environment has it and that its sandboxing (`ProtectHome` and
friends) allows the socket. If not, say so in the plan and fix it in
`deploy/`, not by weakening the sandbox wholesale. Use urgency `critical` for
`critical` alerts so they persist on screen.
### 7. The watchdog timer: detection that needs nobody watching
**Why.** All three loops in incident #8 were caught by a human or by a Claude
Code session reading opencode's session store by hand. The operator does not
want to depend on either. The watchdog is an independent, local, scheduled
sensor. It runs whether or not anyone is at the screen, whether or not a Claude
session is open, and whether or not the plugin loaded. **It calls no cloud
model.**
**What.** A systemd user timer, `deploy/llm-router-watchdog.{service,timer}`,
shaped like the existing `llm-router-*` timers, every `watchdog.interval`
(proposed 5 min). Each tick:
1. **Find opencode.** Read `~/.local/share/opencode/rc-servers.json`, written
by the `session-registry` plugin. Probe each recorded `serverUrl`. If none
answers, exit quietly; no opencode means nothing to watch.
2. **List sessions active since the last tick** via `GET /session`, and read
each one's recent messages via `GET /session/{id}/message`.
3. **Deterministic check, free, on every active session.** The same signals as
the judge (section 3): exact-duplicate share, the most-repeated target, the
failure streak, landed changes (for read-only agents, target novelty
instead), and injected continuation turns. This is Phase 0's replay code
run live. One implementation, imported by both, not a copy.
4. **Local LLM second opinion, only on sessions step 3 flags.** Skipped when
`local_compute.enabled` is off, which keeps gaming mode honest. The model is
whatever is configured: a `watchdog.model` knob that defaults to
`verification.model`, on that section's base URL. Today that is
`qwen2.5-coder-router:14b` on the local Ollama (`localhost:11434`), already
kept warm by local verification, so a call pays no model-load delay. It
runs on a digest of tool names, targets and exit codes, never content.
Phase 0 measures this model's verdicts on the three incident #8 cases
(must flag: the stuck workers and the post-completion Atlas loop; must not
flag: productive Atlas) before its answer is trusted. It asks "is this agent looping without progress? yes or no,
with a one-line reason". A healthy tick costs zero inference.
5. **Report, do not act.** POST the verdict to the router (the `/progress`
judge or a sibling endpoint), tagged `source: watchdog`. The **router**
applies the configured mode, alerts through the notifier, and enforces
`max_recoveries` and Resume. The watchdog never aborts or compacts anything
itself, so there is exactly one recovery path and one place that decides.
**If the router is down,** the watchdog still sends a `critical` desktop
alert directly with `notify-send`, and nothing else. A dead router is exactly
when an unattended run needs a human.
**Failure modes to design for:**
- A tick that overruns the interval: use a lock file and skip the tick, do not
stack them.
- An opencode API shape change. The v2 port breaks this the same way it breaks
`oc_work_start.py`. Detect the unexpected shape and raise one
`warning`-level alert saying the watchdog is blind, rather than silently
reporting "all healthy".
- The watchdog's own health. Record `last_tick_at` and the result (via the
router, or a state file), and show it in the admin Sessions readout. A
watchdog that stopped running must be visible.
**Knobs (North Star #1):**
- `watchdog.enabled`
- `watchdog.interval`
- `watchdog.local_llm_enabled` (default on, still gated by `local_compute`)
- `watchdog.model` (default: `verification.model`)
- `watchdog.read_only_agents`, the list from the read-only case above
The timer ships enabled, since detection plus a warning is the neutral,
low-risk default; the actions follow `progress.mode`.
### 8. The behavior breaker (Lift B): trip a model that stalls sessions
**Gap.** The breaker family covers availability (`circuit_breaker.py`, 5xx) and
malformed output (PR #76: mojibake, empty 200s). Nothing trips on
**well-formed output that does not do the job**. In incident #8,
`z-ai/glm-5.3-flash` produced fluent text that confabulated truncation and
re-read in loops, even with pinch off and a small context. It was only
excluded by hand.
**Trigger.** Verdicts from this detector (the watchdog, and the router judge
once built), attributed to the model(s) that served each flagged session in its
window. Trip when:
- stall verdicts on one model span at least `min_sessions` distinct sessions
(proposed 3) within `window` (proposed 60 min), AND
- that model's stall rate is well above the other models' on the same window's
traffic, so one hard task cannot blame a model.
**Response.** The same shape as the existing breakers: a passive routing skip
with cooldown and backoff, plus a critical alert through the notifier and an
admin Resume. Key it on the model, not the provider, per
`plans/mangled-output-detection.md`'s scope reasoning: here one model misbehaved
while its siblings on the same provider did not. Record every trip with its
evidence (sessions, verdicts, rates) so a human can review it.
**Blocked on identity.** Joining "this session stalled" to "these models served
it" needs #99's conversation headers on each request. As of 2026-09-26 they do
not arrive: `router-link.js` is installed and passes its 29 tests, and the router
reads the headers, but production decisions still carry fingerprint
`session_key`s with no `agent`. Lift B starts by finding out why.
## Fixes that ride along
- **The plugin never reads a real exit code.** Both `router-outcome.js`
(deployed) and its replacement `router-link.js` (#99, line 83) have this. In
1.18,
`tool.execute.after` gives `output = {title, output, metadata}`, and bash's
exit is `output.metadata.exit` (verified on a live session). `looksFailed`
checks `output.exitCode` / `output.exit_code`, which do not exist, so every
verdict has come from the failure-text regex alone. Read `metadata.exit`
first.
- **The installed plugin differs from the repo copy.** The main checkout has an
uncommitted comment block. It is harmless, but the lift should add a check
(install script or a `/health` field reporting the plugin's version string
from a header) so a stale install is visible.
- **`opencode.json` advertised `auto` with `limit.context: 782324`**, the
largest window in the catalog, so opencode only compacted near 782k. Tonight's
contexts reached 682k and were resent every turn. Lowered to 200,000 on
2026-09-26 by owner decision (see the end). It caps what `auto` is asked to
carry, so Phase 0 should also measure how often a real agent turn now
compacts, and whether any task needed more.
## Ground rules for the executor
- **Base on `origin/main`, NOT the local `main`.** The local `main` checkout
(`c87e675`) is 29 commits behind `origin/main` (`f554fc1` as of 2026-09-26)
and 2 ahead with local-only commits. It lacks #99 (conversation identity,
`router-link.js`) and #100 (local-encoder rebuild). `git fetch origin` first,
then branch `feat/no-progress-detection` from `origin/main` in its own
worktree. Read code from that worktree or `git show origin/main:<path>`, never
from the main checkout's files. Do not touch, stash or commit anything in the
main checkout; it has uncommitted work from other efforts.
- The plugin to extend is `deploy/opencode-plugin/router-link.js` (with its
`router-link.test.mjs`). `router-outcome.js` was replaced by #99.
- **Port 8080 is production; never touch it.** Integration checks run on 8081
via `scripts/sandbox.sh`, against a copy of the live DB made with the sqlite
backup API (source opened `mode=ro`).
- Never edit `config/config.yaml` or `config/config.local.yaml`. New knobs get
defaults in `config.yaml` only as part of this lift's own tracked changes.
- **The installed plugin at `~/.config/opencode/plugins/` is live for every
opencode session on this machine,** including the one running the lift.
Develop and test against the repo copy. Installing a new version is an
operator step, stated in the done report, not something the executor does.
- **Commit by explicit path. Never `git add -A` or `git add .`** Incident #8's
item 4 swept an `.omc/` state file into a commit that way.
- **Every worker claim is verified by the orchestrator** against
`git show --stat` and the actual test files before acceptance. Incident #8
had a commit message listing tests that did not exist.
- **Read large files in targeted slices** (grep for the anchor, then read
around it). Incident #8's stuck workers re-read whole files dozens of times.
- ASCII only in new code, comments and UI strings; no middle-dot separators.
No prose paragraphs in the admin UI.
- Lint gate: pinned `uvx ruff@0.16.9`, no NEW findings in touched files
compared with the base commit (same recipe as
`.omo/plans/cockpit-quick-wins.md` runbook step 5). Never `ruff --fix`
outside the lines the item changes.
- `pytest` offline stays green after every commit. Record the baseline on
`origin/main` first. On `c87e675`, one test,
`tests/test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity`,
failed before any change; check whether it still does on `origin/main` rather
than assuming.
- New `route_decisions` columns are registered in
`tests/test_tui_schema_drift.py`; every new knob gets an admin control or a
recorded `DELIBERATELY_NOT_IN_ADMIN` reason (`tests/test_admin_knob_coverage.py`).
## Phases
0. **Calibrate offline.** Replay over opencode session history; produce the
verdict timeline per session and a proposed threshold table. No router or
plugin changes. **Start from `plans/no-progress-detection-prototype.py`,**
an interim watcher tuned on 2026-09-26 against the 15 sessions of incident
#8. On that set it flags every known loop (both item-2 workers, the item-3
first attempt, item 4, the post-completion Atlas stretch, the explore
helper stuck on a deleted file, a helper that printed the spec 12 times)
and none of the healthy sessions. Four findings from tuning it:
- Exact `(tool, args)` fingerprints alone miss reworded commands: 28
differently-worded `cat boulder.json` never matched.
- Collapsing to the bare file over-flags ordinary sliced reading, which is
what the ground rules tell workers to do. Key a `read` on file plus
offset/limit, and a `bash` target on its referenced files plus the
numbers in the command (the slice).
- Slow loops spread across 300+ calls escape a 60-call window. Add a
whole-session rule: one exact call repeated 15+ times with nothing landed
recently.
- Read-only agents need the separate rule (repeats of one target), since
"nothing landed" is always true for them.
Its thresholds (`WINDOW=60`, `DUP_MIN=0.25`, `TOP_MIN=12`, `TOP_MIN_RO=8`,
`CUM_MIN=15`) are a starting point fitted on one night, not a result.
1. **Identity: build on #99, do not rebuild it.** The `metadata.exit` fix in
`router-link.js` (and its `.test.mjs`), plus whatever #99 lacks for
progress roll-up, if anything. Deploying #99 itself (syncing the
production checkout, installing `router-link.js`) is an operator step. It
goes in the done report as a prerequisite for Phase 2 to see real traffic,
not an executor task.
2. **Progress events, the judge and the watchdog, `warn` mode.** `/progress`,
`progress_events`, verdicts, the `/metrics` warning, toast, Sessions readout,
and the knobs. Also the notifier (section 6) with the `desktop` channel and
the admin Send test alert button, and the watchdog timer (section 7) with
its local-LLM second opinion. **Order within the phase: build the watchdog
and the notifier first.** Together with Phase 0's detection code, they
protect unattended runs before the plugin half is even deployed.
3. **`auto_compact` and `auto_recover` stage 1.** Abort, compact with the
injected report, restart; the scoped 429; the `max_recoveries` hard stop and
admin Resume. The incident report builder lands here.
4. **`auto_limit_context` and `auto_recover` stage 2**, using whichever
mechanism Phase 0 verified.
5. **Local-model second opinion inside the router's live judge** (section 4),
only if earlier phases show ambiguous cases. The watchdog already has one
(section 7); this would add it to the per-request path.
The v2 port (`project_opencode_v2_pinned`) changes hook registration. Keep the
plugin's logic in plain functions so the v1 and v2 shells stay thin.
## Owner decisions
Settled 2026-09-26: the response is a mode (`warn`, `auto_compact`,
`auto_limit_context`, `auto_recover` with two stages), and every mode warns.
`auto_recover` stops and compacts the tree's root (Atlas), not only the stuck
worker.
Also settled 2026-09-26:
- **The default mode is `warn`.**
- **`max_recoveries` is 2** per tree per window, then the hard stop.
- **The `auto` context limit is 200,000.** Applied the same day, ahead of this
lift, in the repo `opencode.json` and the global
`~/.config/opencode/opencode.json` (backup at
`opencode.json.bak-20260926-auto-context`). The global file pins every agent
(atlas, sisyphus-junior, prometheus, ...) to `llm-router/auto`. opencode
reads config at startup, so it takes effect for sessions started after a
restart. `auto:batch` stays at 782,324 on purpose: batch is where long
contexts are expected.
- **Alerting ships with `desktop` (`notify-send`) only**, behind a notifier
built for more channels. SMS, RingCentral, PagerDuty and webhooks are later
adapters (section 6), not part of this lift.
No owner decisions remain open. Thresholds come from Phase 0.