tests/test_plans_declare_status.py requires 'Status: <done|planned|in progress|parked|reference> -- <reason>' in the first 8 lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
573 lines
31 KiB
Markdown
573 lines
31 KiB
Markdown
# No-progress detection: catch an agent that is spinning, not one that is working
|
|
|
|
Status: reference -- parent spec; Lift A shipped in PR #102, Lift B (router-side judge, modes, behavior breaker) not started
|
|
Date: 2026-09-26
|
|
Trigger: the `cockpit-quick-wins` Atlas run, 2026-09-25 19:00 to 2026-09-26 02:00.
|
|
Incident write-up: `docs/incidents.md` #8.
|
|
|
|
## What happened, measured
|
|
|
|
Read-only against the live `router.db` and opencode's own session store
|
|
(`http://127.0.0.1:4097`), 2026-09-26.
|
|
|
|
Spend, 19:00 to 02:00 local:
|
|
- 323M prompt tokens (95% cached) and about $7.81, over roughly 8 hours.
|
|
- 79% of the money went to OpenRouter `z-ai/glm-5.3-flash`: 1,368 calls
|
|
averaging 203k prompt tokens.
|
|
- 921 calls carried 120k+ prompt tokens and cost $6.57. NeuralWatt `glm-5.3`
|
|
took 47 calls averaging 673k context, max 682,679, because nothing else fit.
|
|
- Every decision was `classification_source = classifier` on the default
|
|
profile. The router did what it was told; nothing was pinned.
|
|
|
|
**Spend rate is the wrong signal.** Tonight's worst hour ($2.42) and worst
|
|
3-hour window ($4.13) sit inside the prior 30 days' normal range (hourly p90
|
|
$1.46, max $4.83; worst 3h $9.65). A rate alarm quiet enough for a healthy
|
|
all-day agent run would have missed this entirely, and one tight enough to catch
|
|
it would fire on good work. The operator's framing is the right one: steady
|
|
usage is fine. What is not fine is spend with **no concrete change landing**.
|
|
|
|
**The sessions that wasted money are visibly different in their tool calls:**
|
|
|
|
| session (opencode) | tool calls | exact-duplicate calls | worst repeat | landed a change? |
|
|
|---|---|---|---|---|
|
|
| Item 2 worker, 2nd attempt | 423 | 30% | `read navbar.js` x61 | no (worktree unchanged) |
|
|
| Item 3 worker, 1st attempt | 132 | 23% | `read test_admin_js_units.py` x16 | no |
|
|
| Item 4 worker | 134 | 19% | `read .omo/plans/...` x8 | partly (code, no tests) |
|
|
| Item 2 worker, 1st attempt | 307 | 17% | `read navbar.js` x21 | no |
|
|
| Atlas orchestrator | 576 | 11% | playwright console check x10 | yes (via workers) |
|
|
| healthy workers (items 1, 5, 6, 3-retry) | 30-96 | 0-6% | 1-3 | yes |
|
|
|
|
**A second loop shape, 2026-09-26 02:21 to about 02:56:** the Atlas orchestrator
|
|
itself, after its plan was complete:
|
|
- `.omo/boulder.json` carried a stale `status: completed` and `pr_url` from the
|
|
previous plan.
|
|
- oh-my-openagent's continuation hook injected "continue" turns; all 3 user
|
|
turns in the window were injected.
|
|
- Atlas, working from a trimmed context, re-derived its state every time:
|
|
54 turns, about 19M cached-read tokens, `boulder.json` read 28 times, the
|
|
plan's checkboxes grepped 16 times, nothing landed.
|
|
|
|
The operator caught it only by watching. The judge must treat this as
|
|
`stalling`, and Phase 0 must include it as a calibration case. The plugin can
|
|
also see injected continuation turns (`chat.message`), which is a useful extra
|
|
signal: "continuation turns keep arriving and nothing lands" is this shape
|
|
exactly.
|
|
|
|
**A third shape, 2026-09-26 around 03:05 to 03:17: a read-only agent.**
|
|
Prometheus's `explore` subagent reached 322 tool calls, 19% exact duplicates,
|
|
and read `deploy/opencode-plugin/router-outcome.js` 20 times. The file had been
|
|
deleted from under it: the production checkout was synced to `origin/main`,
|
|
where #99 retired it. The subagent was aborted.
|
|
|
|
Read-only agents (`explore`, `librarian`, `oracle`, and the like) never land a
|
|
change by design, so "no landed change" cannot be their signal. For them the
|
|
judge uses:
|
|
- target novelty: new files or queries per N calls, and
|
|
- a failure streak on the same target, such as repeated reads of a path that
|
|
no longer exists.
|
|
|
|
Identify them by the `X-Router-Agent` value, via a configurable list of
|
|
read-only agent names.
|
|
|
|
"Exact duplicate" means the same tool with byte-identical arguments. Duplicate
|
|
share alone does not separate the orchestrator (11%, productive) from a stuck
|
|
worker (17%). Duplicates **plus** no landed change does.
|
|
|
|
## Why nothing caught it
|
|
|
|
- `circuit_breaker.py` is availability-only: it trips on provider 5xx.
|
|
- The only spend alarm is `runway_low_warning`: balance / burn < 6 h, with burn
|
|
averaged over a 24 h balance window. OpenRouter read $24.73 at $0.25/h, so 97 h
|
|
of runway. It answers "will I run out", not "is this being wasted".
|
|
- The router cannot see progress at all. It sees prompts, not whether an edit
|
|
landed or a command keeps failing. The opencode plugin sees both.
|
|
- The DEPLOYED router cannot tell conversations apart either. `session_key`
|
|
merges concurrent conversations under one multi-agent client, and `/outcome`
|
|
attribution falls back to a directory match that 409s when several sessions
|
|
share a cwd, which is exactly the Atlas-plus-workers shape. **The fix already
|
|
exists on `origin/main` (PR #99, conversation-identity) but is not deployed:**
|
|
production runs from a local checkout 29 commits behind `origin/main`, the
|
|
live `route_decisions` has no `agent` / `parent_key` columns, and the
|
|
installed plugin is still `router-outcome.js` rather than #99's
|
|
`router-link.js`.
|
|
|
|
## Design
|
|
|
|
Split the job where the information lives. **The plugin is the sensor; the
|
|
router is the judge.** No raw task text leaves opencode, and the router never
|
|
stores any; the "never store raw task text" rule is unchanged.
|
|
|
|
### 1. Exact session identity: ALREADY BUILT upstream (#99), build on it
|
|
|
|
PR #99 (`conversation-identity`, merged to `origin/main` as `fe8c813`) did this.
|
|
Reuse it; do not rebuild it:
|
|
- **Plugin.** `deploy/opencode-plugin/router-link.js` replaces
|
|
`router-outcome.js`. Its `chat.headers` hook sends `X-Router-Conversation`
|
|
(the opencode session id), `X-Router-Agent` (slugged to the router's charset)
|
|
and `X-Router-Parent`, with a bounded parent-lookup cache.
|
|
- **Router.** `src/conversation_identity.py` resolves identity from those
|
|
headers. `session_key` becomes `'c:' + conversation id`. `route_decisions`
|
|
gained `agent` and `parent_key` columns. `/outcome` attributes to the exact
|
|
conversation.
|
|
|
|
What remains for this lift:
|
|
- **Deploy #99.** It is merged but not running. Syncing the production checkout
|
|
to `origin/main` and installing `router-link.js` are operator steps: the main
|
|
checkout has uncommitted work, and part of it, the `confidence_min` rename,
|
|
also landed upstream via #100.
|
|
- **Fix the exit code in `router-link.js`.** It kept `router-outcome.js`'s
|
|
`output?.exitCode ?? output?.exit_code` (line 83 on `origin/main`); see "Fixes
|
|
that ride along".
|
|
- **Build the progress events (section 2) into `router-link.js`,** keyed by the
|
|
same conversation and parent ids, so the judge rolls children up to parents
|
|
through `parent_key`.
|
|
|
|
Everywhere below, "session" means #99's conversation. Where this spec says
|
|
`X-Opencode-Session`, read `X-Router-Conversation`.
|
|
|
|
### 2. Progress events (plugin `tool.execute.after` -> router `POST /progress`)
|
|
|
|
The plugin sends one small event per tool call, fire-and-forget with a short
|
|
timeout, never blocking the session:
|
|
|
|
```
|
|
{session_id, parent_session_id, agent, tool, call_id,
|
|
fingerprint, # sha256 of tool + normalized args; never the args themselves
|
|
target, # coarse, normalized: the file path for read/edit/write, the
|
|
# command's first two tokens for bash; no contents
|
|
exit, # bash: output.metadata.exit
|
|
landed} # bool, see below
|
|
```
|
|
|
|
`landed` is the load-bearing definition, and it is deliberately narrow:
|
|
- `edit` / `write` / `patch` with a non-empty `metadata.diff`
|
|
- `bash` whose command is `git commit` (or `git commit --amend`) with exit 0
|
|
- A test or lint command (the plugin's existing `TEST_COMMAND` list) that PASSES
|
|
after the session's previous run of the same fingerprint FAILED, i.e. a fix
|
|
landed.
|
|
|
|
Reads, greps, todo writes, sub-agent dispatches and passing-again tests are not
|
|
progress. An orchestrator's progress is its children's: roll child `landed`
|
|
events up to the parent via `parent_session_id`.
|
|
|
|
Storage: a new `progress_events` table, pruned to a rolling window (knob). It
|
|
holds fingerprints, targets and booleans, no content.
|
|
|
|
### 3. The judge (router, per client session)
|
|
|
|
Signals per session, over a rolling window of its last N tool events:
|
|
- `turns_since_landed`: LLM requests since the last landed event (self plus
|
|
children for a parent).
|
|
- `usd_since_landed`: billed spend since then, joined through `request_id`.
|
|
- `dup_share`: exact-duplicate fingerprint share in the window.
|
|
- `fail_streak`: consecutive nonzero exits on the same fingerprint (the
|
|
"sandbox is broken and the agent keeps retrying" case).
|
|
- `context_tokens`: the latest request's size, to report alongside, since a
|
|
stall at 600k context costs far more per turn than one at 30k.
|
|
|
|
Verdicts:
|
|
- `ok`: something landed recently.
|
|
- `stalling` (warn): no landed change for `stall_turns` turns or
|
|
`stall_usd` dollars, AND (`dup_share >= dup_warn` OR `fail_streak >= fail_warn`).
|
|
- `looping` (act, if enabled): the stall condition at the higher `loop_*`
|
|
thresholds.
|
|
|
|
Both conditions have to hold. Long-running healthy work with no duplicates (a
|
|
big read-heavy exploration) stays `ok` until it hits the plain turns/spend
|
|
ceiling, which is a separate, looser knob.
|
|
|
|
**Calibrate before trusting it.** Phase 0 replays the judge offline over
|
|
opencode's session store (the 11 sessions above plus older history), and reports
|
|
the verdict each session would have received and when. Thresholds are picked
|
|
from that, not guessed. The Atlas orchestrator (11% duplicates, productive
|
|
through children) is the key false-positive case, and it must stay `ok`.
|
|
|
|
### 4. Optional: a local model as a second opinion
|
|
|
|
The deterministic signals should carry this. Where they are ambiguous, for
|
|
example near-duplicates (the same file read at shifting offsets, or the same
|
|
failing test with a different `-k`), there are two cheaper steps before any
|
|
model:
|
|
- normalize `read` fingerprints to the path, and bash to command plus target
|
|
- the local encoder already loaded for `classifier.mode: local_encoder` can
|
|
embed the `target` strings to cluster "similar-ish" events
|
|
|
|
A local LLM judge ("does this digest look like a loop?") runs only on a session
|
|
already flagged `stalling`, on a digest of tool names, targets and exit codes,
|
|
never content. It is gated by `local_compute.enabled`, since the operator pays
|
|
the local power bill. Treat this as a later phase, justified by Phase 0 showing
|
|
ambiguous cases the deterministic signals miss.
|
|
|
|
### 5. What happens on a detection: modes
|
|
|
|
One knob, `progress.mode`, picks the response. **Every mode warns:**
|
|
- A `/metrics` warning (admin bell and TUI). It names the session, agent, turns
|
|
and $ since the last landed change, and the top repeated target.
|
|
- An opencode toast via `POST /tui/show-toast` in the window the operator is
|
|
watching.
|
|
- An **alert** through the router's notifier (section 6), which reaches the
|
|
operator when nobody is looking at either screen.
|
|
|
|
| mode | on detection |
|
|
|---|---|
|
|
| `warn` | Warnings only. The neutral default: today's behaviour plus a warning. |
|
|
| `auto_compact` | Compact the stuck session with the incident report injected, then continue. No refusal, no orchestrator stop. |
|
|
| `auto_limit_context` | Cap the stuck session's context. No stop. |
|
|
| `auto_recover` | Two stages, below. |
|
|
|
|
**`auto_recover`, stage 1 (first detection in a session tree):**
|
|
1. The router refuses the stuck session's chat requests (429, scoped by
|
|
`X-Router-Conversation`), so nothing more is spent while the plugin acts.
|
|
2. The plugin walks `parentID` to the tree's root, usually Atlas, and calls
|
|
`POST /session/{id}/abort` on the stuck session and on the root.
|
|
3. The plugin compacts the root: `POST /session/{root}/summarize`. Its
|
|
`experimental.session.compacting` hook appends the **incident report** to
|
|
the compaction prompt, so the report survives into the compacted context.
|
|
4. `experimental.compaction.autocontinue` stays enabled, so the root restarts
|
|
from the compacted context with the report in view. The router lifts the
|
|
refusal when the compaction completes.
|
|
5. Toast: "recovered <agent>: <one-line reason>".
|
|
|
|
**`auto_recover`, stage 2 (detected again in the same tree within
|
|
`recover_window`):** apply the `auto_limit_context` cap to the tree, and toast
|
|
again.
|
|
|
|
**Stage 3, a hard stop.** Recovery can itself loop: recover, loop, recover. After
|
|
`max_recoveries` in the window, the router refuses the tree and does not
|
|
restart it. It toasts and raises a high-severity warning, and the tree waits for
|
|
an admin Resume. Without this cap, auto-recovery is an unattended retry loop,
|
|
which is the failure it exists to stop.
|
|
|
|
**The incident report** is built deterministically from `progress_events`
|
|
(fingerprints, targets and exit codes, never content), so it is free and cannot
|
|
hallucinate. Example:
|
|
|
|
```
|
|
[router] A sub-agent stalled and was stopped.
|
|
session: Item 2 G1 nav refactor (Sisyphus-Junior), child of this session
|
|
423 tool calls, 0 landed changes (no edit with a diff, no commit)
|
|
most repeated: read admin/frontend/navbar.js x61, read tests/... x16
|
|
spend since last landed change: $1.84, 70M prompt tokens
|
|
Avoid: re-reading whole files, since reads repeat without edits. Give the
|
|
worker the exact edit to make, and require an on-disk diff or commit before it
|
|
reports done.
|
|
```
|
|
|
|
The "Avoid:" line maps from the dominant signal: duplicates, a failure streak,
|
|
or context size. A local model may rephrase it later (section 4); it is not
|
|
needed to produce it.
|
|
|
|
**"Limit context"** has two candidate mechanisms. Phase 0 must verify which works
|
|
before stage 2 is built:
|
|
- **(a) Router-side, per-session tighter `pinch` budget.** The machinery
|
|
exists (`context_prune.py`), but pruning rewrites the prefix, and
|
|
`plans/token-waste-waves.md` Wave 3 measured rewrites costing the cache. So
|
|
this trades cache hits for size. Measure it.
|
|
- **(b) The router answers an over-cap request with a context-overflow-shaped
|
|
error**, so opencode runs its own overflow compaction. The autocontinue hook
|
|
receives `overflow: boolean`, so opencode does have such a path. Unverified:
|
|
whether it recognizes a router error as an overflow. Test this on 8081 before
|
|
relying on it.
|
|
|
|
Also, the Controls or Home page gets a live "Sessions" readout: session, agent,
|
|
parent, verdict, mode stage, turns and $ since landed, and the top repeated
|
|
target. This fits `plans/cockpit-brainstorm.md`. Each tree with a refusal gets
|
|
a Resume button.
|
|
|
|
Every threshold and the mode are knobs with admin controls (North Star #1).
|
|
Per the "build for re-tuning" rule, the neutral default is `warn`, with Phase
|
|
0-calibrated thresholds.
|
|
|
|
Without the plugin (another client, or a stale install), only the router half
|
|
works: warnings, alerts and the scoped 429. Abort, compact and restart need the
|
|
plugin. The Sessions readout says which sessions have a live plugin (from the
|
|
headers in section 1), so a mode that silently cannot act is visible.
|
|
|
|
### 6. Alerting: a notifier with pluggable channels
|
|
|
|
Toasts and the admin bell only help someone who is looking. Incident #8 ran
|
|
overnight. Alerts come from the **router**, not the plugin: the router is the
|
|
judge and is always running, while the plugin dies with the opencode process
|
|
it lives in.
|
|
|
|
**Ship now:** one channel, `desktop` (`notify-send`). **Design for later:** SMS
|
|
(Twilio or similar), RingCentral, PagerDuty, a generic webhook (ntfy, Slack).
|
|
Each of those should be a small adapter added without touching the detector.
|
|
|
|
Shape:
|
|
- **An alert is an event with a lifecycle**, not a message. Fields:
|
|
- `dedup_key`: for this detector, the tree root session id
|
|
- `severity`: `info`, `warning` or `critical`
|
|
- `state`: `trigger`, `escalate` or `resolve`
|
|
- `title` and a one-line `summary`
|
|
- `details`, the incident report from section 5
|
|
- `source`: `no-progress`, so other detectors can reuse the notifier later
|
|
|
|
This maps one-to-one onto PagerDuty Events v2 (`dedup_key`,
|
|
`event_action: trigger/resolve`), so that adapter is thin. Channels without
|
|
a lifecycle (SMS, desktop) render `trigger` and `escalate` as a message and
|
|
`resolve` as a short "recovered" message, or skip it (per-channel knob).
|
|
- **Transitions, not repeats.** An alert fires when a verdict *changes*, e.g.
|
|
`ok -> stalling`, `stalling -> looping`, a recovery, the hard stop, or
|
|
`-> ok`. It never fires per request. Per-channel rate limit and a quiet
|
|
window stop a flapping session from paging anyone at 3am every minute.
|
|
- **Severity routing per channel:** each channel has `min_severity`. The
|
|
intended setup: desktop gets `warning` and up; an SMS or PagerDuty channel
|
|
added later gets only `critical`, i.e. the hard stop, or `looping` in `warn`
|
|
mode.
|
|
- **Delivery never blocks routing.** Alerts go through a small queue with a
|
|
timeout and one retry. A failed delivery is logged and shown in the admin
|
|
UI; it never slows or fails a chat request.
|
|
- **Secrets stay out of config files.** Channel credentials (API keys, routing
|
|
keys, phone numbers) are read from environment variables named in config,
|
|
the same `api_key_env` pattern `dispatch_providers` already uses, and loaded
|
|
from `.env` by the systemd unit. `config.yaml` and `config.local.yaml` never
|
|
hold a secret.
|
|
- **Config:** a `notifications.channels` list, each entry
|
|
`{name, type, enabled, min_severity, resolve_messages, ...type-specific}`.
|
|
Only `type: desktop` exists in this lift. An unknown `type` fails config load
|
|
(StrictModel), so a typo cannot silently disable alerting.
|
|
- **Admin (North Star #1):** per channel, an enabled toggle, a `min_severity`
|
|
select, and a **Send test alert** button. The test button is the only way to
|
|
know a channel works before the night it matters. Recent deliveries and
|
|
failures show under the Sessions readout.
|
|
|
|
**`desktop` specifics, to verify in Phase 0.** The router runs as a systemd
|
|
*user* service, so `notify-send` needs the session D-Bus
|
|
(`DBUS_SESSION_BUS_ADDRESS`, normally `unix:path=/run/user/<uid>/bus`). Check
|
|
that the unit's environment has it and that its sandboxing (`ProtectHome` and
|
|
friends) allows the socket. If not, say so in the plan and fix it in
|
|
`deploy/`, not by weakening the sandbox wholesale. Use urgency `critical` for
|
|
`critical` alerts so they persist on screen.
|
|
|
|
### 7. The watchdog timer: detection that needs nobody watching
|
|
|
|
**Why.** All three loops in incident #8 were caught by a human or by a Claude
|
|
Code session reading opencode's session store by hand. The operator does not
|
|
want to depend on either. The watchdog is an independent, local, scheduled
|
|
sensor. It runs whether or not anyone is at the screen, whether or not a Claude
|
|
session is open, and whether or not the plugin loaded. **It calls no cloud
|
|
model.**
|
|
|
|
**What.** A systemd user timer, `deploy/llm-router-watchdog.{service,timer}`,
|
|
shaped like the existing `llm-router-*` timers, every `watchdog.interval`
|
|
(proposed 5 min). Each tick:
|
|
1. **Find opencode.** Read `~/.local/share/opencode/rc-servers.json`, written
|
|
by the `session-registry` plugin. Probe each recorded `serverUrl`. If none
|
|
answers, exit quietly; no opencode means nothing to watch.
|
|
2. **List sessions active since the last tick** via `GET /session`, and read
|
|
each one's recent messages via `GET /session/{id}/message`.
|
|
3. **Deterministic check, free, on every active session.** The same signals as
|
|
the judge (section 3): exact-duplicate share, the most-repeated target, the
|
|
failure streak, landed changes (for read-only agents, target novelty
|
|
instead), and injected continuation turns. This is Phase 0's replay code
|
|
run live. One implementation, imported by both, not a copy.
|
|
4. **Local LLM second opinion, only on sessions step 3 flags.** Skipped when
|
|
`local_compute.enabled` is off, which keeps gaming mode honest. The model is
|
|
whatever is configured: a `watchdog.model` knob that defaults to
|
|
`verification.model`, on that section's base URL. Today that is
|
|
`qwen2.5-coder-router:14b` on the local Ollama (`localhost:11434`), already
|
|
kept warm by local verification, so a call pays no model-load delay. It
|
|
runs on a digest of tool names, targets and exit codes, never content.
|
|
Phase 0 measures this model's verdicts on the three incident #8 cases
|
|
(must flag: the stuck workers and the post-completion Atlas loop; must not
|
|
flag: productive Atlas) before its answer is trusted. It asks "is this agent looping without progress? yes or no,
|
|
with a one-line reason". A healthy tick costs zero inference.
|
|
5. **Report, do not act.** POST the verdict to the router (the `/progress`
|
|
judge or a sibling endpoint), tagged `source: watchdog`. The **router**
|
|
applies the configured mode, alerts through the notifier, and enforces
|
|
`max_recoveries` and Resume. The watchdog never aborts or compacts anything
|
|
itself, so there is exactly one recovery path and one place that decides.
|
|
|
|
**If the router is down,** the watchdog still sends a `critical` desktop
|
|
alert directly with `notify-send`, and nothing else. A dead router is exactly
|
|
when an unattended run needs a human.
|
|
|
|
**Failure modes to design for:**
|
|
- A tick that overruns the interval: use a lock file and skip the tick, do not
|
|
stack them.
|
|
- An opencode API shape change. The v2 port breaks this the same way it breaks
|
|
`oc_work_start.py`. Detect the unexpected shape and raise one
|
|
`warning`-level alert saying the watchdog is blind, rather than silently
|
|
reporting "all healthy".
|
|
- The watchdog's own health. Record `last_tick_at` and the result (via the
|
|
router, or a state file), and show it in the admin Sessions readout. A
|
|
watchdog that stopped running must be visible.
|
|
|
|
**Knobs (North Star #1):**
|
|
- `watchdog.enabled`
|
|
- `watchdog.interval`
|
|
- `watchdog.local_llm_enabled` (default on, still gated by `local_compute`)
|
|
- `watchdog.model` (default: `verification.model`)
|
|
- `watchdog.read_only_agents`, the list from the read-only case above
|
|
|
|
The timer ships enabled, since detection plus a warning is the neutral,
|
|
low-risk default; the actions follow `progress.mode`.
|
|
|
|
### 8. The behavior breaker (Lift B): trip a model that stalls sessions
|
|
|
|
**Gap.** The breaker family covers availability (`circuit_breaker.py`, 5xx) and
|
|
malformed output (PR #76: mojibake, empty 200s). Nothing trips on
|
|
**well-formed output that does not do the job**. In incident #8,
|
|
`z-ai/glm-5.3-flash` produced fluent text that confabulated truncation and
|
|
re-read in loops, even with pinch off and a small context. It was only
|
|
excluded by hand.
|
|
|
|
**Trigger.** Verdicts from this detector (the watchdog, and the router judge
|
|
once built), attributed to the model(s) that served each flagged session in its
|
|
window. Trip when:
|
|
- stall verdicts on one model span at least `min_sessions` distinct sessions
|
|
(proposed 3) within `window` (proposed 60 min), AND
|
|
- that model's stall rate is well above the other models' on the same window's
|
|
traffic, so one hard task cannot blame a model.
|
|
|
|
**Response.** The same shape as the existing breakers: a passive routing skip
|
|
with cooldown and backoff, plus a critical alert through the notifier and an
|
|
admin Resume. Key it on the model, not the provider, per
|
|
`plans/mangled-output-detection.md`'s scope reasoning: here one model misbehaved
|
|
while its siblings on the same provider did not. Record every trip with its
|
|
evidence (sessions, verdicts, rates) so a human can review it.
|
|
|
|
**Blocked on identity.** Joining "this session stalled" to "these models served
|
|
it" needs #99's conversation headers on each request. As of 2026-09-26 they do
|
|
not arrive: `router-link.js` is installed and passes its 29 tests, and the router
|
|
reads the headers, but production decisions still carry fingerprint
|
|
`session_key`s with no `agent`. Lift B starts by finding out why.
|
|
|
|
## Fixes that ride along
|
|
|
|
- **The plugin never reads a real exit code.** Both `router-outcome.js`
|
|
(deployed) and its replacement `router-link.js` (#99, line 83) have this. In
|
|
1.18,
|
|
`tool.execute.after` gives `output = {title, output, metadata}`, and bash's
|
|
exit is `output.metadata.exit` (verified on a live session). `looksFailed`
|
|
checks `output.exitCode` / `output.exit_code`, which do not exist, so every
|
|
verdict has come from the failure-text regex alone. Read `metadata.exit`
|
|
first.
|
|
- **The installed plugin differs from the repo copy.** The main checkout has an
|
|
uncommitted comment block. It is harmless, but the lift should add a check
|
|
(install script or a `/health` field reporting the plugin's version string
|
|
from a header) so a stale install is visible.
|
|
- **`opencode.json` advertised `auto` with `limit.context: 782324`**, the
|
|
largest window in the catalog, so opencode only compacted near 782k. Tonight's
|
|
contexts reached 682k and were resent every turn. Lowered to 200,000 on
|
|
2026-09-26 by owner decision (see the end). It caps what `auto` is asked to
|
|
carry, so Phase 0 should also measure how often a real agent turn now
|
|
compacts, and whether any task needed more.
|
|
|
|
## Ground rules for the executor
|
|
|
|
- **Base on `origin/main`, NOT the local `main`.** The local `main` checkout
|
|
(`c87e675`) is 29 commits behind `origin/main` (`f554fc1` as of 2026-09-26)
|
|
and 2 ahead with local-only commits. It lacks #99 (conversation identity,
|
|
`router-link.js`) and #100 (local-encoder rebuild). `git fetch origin` first,
|
|
then branch `feat/no-progress-detection` from `origin/main` in its own
|
|
worktree. Read code from that worktree or `git show origin/main:<path>`, never
|
|
from the main checkout's files. Do not touch, stash or commit anything in the
|
|
main checkout; it has uncommitted work from other efforts.
|
|
- The plugin to extend is `deploy/opencode-plugin/router-link.js` (with its
|
|
`router-link.test.mjs`). `router-outcome.js` was replaced by #99.
|
|
- **Port 8080 is production; never touch it.** Integration checks run on 8081
|
|
via `scripts/sandbox.sh`, against a copy of the live DB made with the sqlite
|
|
backup API (source opened `mode=ro`).
|
|
- Never edit `config/config.yaml` or `config/config.local.yaml`. New knobs get
|
|
defaults in `config.yaml` only as part of this lift's own tracked changes.
|
|
- **The installed plugin at `~/.config/opencode/plugins/` is live for every
|
|
opencode session on this machine,** including the one running the lift.
|
|
Develop and test against the repo copy. Installing a new version is an
|
|
operator step, stated in the done report, not something the executor does.
|
|
- **Commit by explicit path. Never `git add -A` or `git add .`** Incident #8's
|
|
item 4 swept an `.omc/` state file into a commit that way.
|
|
- **Every worker claim is verified by the orchestrator** against
|
|
`git show --stat` and the actual test files before acceptance. Incident #8
|
|
had a commit message listing tests that did not exist.
|
|
- **Read large files in targeted slices** (grep for the anchor, then read
|
|
around it). Incident #8's stuck workers re-read whole files dozens of times.
|
|
- ASCII only in new code, comments and UI strings; no middle-dot separators.
|
|
No prose paragraphs in the admin UI.
|
|
- Lint gate: pinned `uvx ruff@0.16.9`, no NEW findings in touched files
|
|
compared with the base commit (same recipe as
|
|
`.omo/plans/cockpit-quick-wins.md` runbook step 5). Never `ruff --fix`
|
|
outside the lines the item changes.
|
|
- `pytest` offline stays green after every commit. Record the baseline on
|
|
`origin/main` first. On `c87e675`, one test,
|
|
`tests/test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity`,
|
|
failed before any change; check whether it still does on `origin/main` rather
|
|
than assuming.
|
|
- New `route_decisions` columns are registered in
|
|
`tests/test_tui_schema_drift.py`; every new knob gets an admin control or a
|
|
recorded `DELIBERATELY_NOT_IN_ADMIN` reason (`tests/test_admin_knob_coverage.py`).
|
|
|
|
## Phases
|
|
|
|
0. **Calibrate offline.** Replay over opencode session history; produce the
|
|
verdict timeline per session and a proposed threshold table. No router or
|
|
plugin changes. **Start from `plans/no-progress-detection-prototype.py`,**
|
|
an interim watcher tuned on 2026-09-26 against the 15 sessions of incident
|
|
#8. On that set it flags every known loop (both item-2 workers, the item-3
|
|
first attempt, item 4, the post-completion Atlas stretch, the explore
|
|
helper stuck on a deleted file, a helper that printed the spec 12 times)
|
|
and none of the healthy sessions. Four findings from tuning it:
|
|
- Exact `(tool, args)` fingerprints alone miss reworded commands: 28
|
|
differently-worded `cat boulder.json` never matched.
|
|
- Collapsing to the bare file over-flags ordinary sliced reading, which is
|
|
what the ground rules tell workers to do. Key a `read` on file plus
|
|
offset/limit, and a `bash` target on its referenced files plus the
|
|
numbers in the command (the slice).
|
|
- Slow loops spread across 300+ calls escape a 60-call window. Add a
|
|
whole-session rule: one exact call repeated 15+ times with nothing landed
|
|
recently.
|
|
- Read-only agents need the separate rule (repeats of one target), since
|
|
"nothing landed" is always true for them.
|
|
Its thresholds (`WINDOW=60`, `DUP_MIN=0.25`, `TOP_MIN=12`, `TOP_MIN_RO=8`,
|
|
`CUM_MIN=15`) are a starting point fitted on one night, not a result.
|
|
1. **Identity: build on #99, do not rebuild it.** The `metadata.exit` fix in
|
|
`router-link.js` (and its `.test.mjs`), plus whatever #99 lacks for
|
|
progress roll-up, if anything. Deploying #99 itself (syncing the
|
|
production checkout, installing `router-link.js`) is an operator step. It
|
|
goes in the done report as a prerequisite for Phase 2 to see real traffic,
|
|
not an executor task.
|
|
2. **Progress events, the judge and the watchdog, `warn` mode.** `/progress`,
|
|
`progress_events`, verdicts, the `/metrics` warning, toast, Sessions readout,
|
|
and the knobs. Also the notifier (section 6) with the `desktop` channel and
|
|
the admin Send test alert button, and the watchdog timer (section 7) with
|
|
its local-LLM second opinion. **Order within the phase: build the watchdog
|
|
and the notifier first.** Together with Phase 0's detection code, they
|
|
protect unattended runs before the plugin half is even deployed.
|
|
3. **`auto_compact` and `auto_recover` stage 1.** Abort, compact with the
|
|
injected report, restart; the scoped 429; the `max_recoveries` hard stop and
|
|
admin Resume. The incident report builder lands here.
|
|
4. **`auto_limit_context` and `auto_recover` stage 2**, using whichever
|
|
mechanism Phase 0 verified.
|
|
5. **Local-model second opinion inside the router's live judge** (section 4),
|
|
only if earlier phases show ambiguous cases. The watchdog already has one
|
|
(section 7); this would add it to the per-request path.
|
|
|
|
The v2 port (`project_opencode_v2_pinned`) changes hook registration. Keep the
|
|
plugin's logic in plain functions so the v1 and v2 shells stay thin.
|
|
|
|
## Owner decisions
|
|
|
|
Settled 2026-09-26: the response is a mode (`warn`, `auto_compact`,
|
|
`auto_limit_context`, `auto_recover` with two stages), and every mode warns.
|
|
`auto_recover` stops and compacts the tree's root (Atlas), not only the stuck
|
|
worker.
|
|
|
|
Also settled 2026-09-26:
|
|
- **The default mode is `warn`.**
|
|
- **`max_recoveries` is 2** per tree per window, then the hard stop.
|
|
- **The `auto` context limit is 200,000.** Applied the same day, ahead of this
|
|
lift, in the repo `opencode.json` and the global
|
|
`~/.config/opencode/opencode.json` (backup at
|
|
`opencode.json.bak-20260926-auto-context`). The global file pins every agent
|
|
(atlas, sisyphus-junior, prometheus, ...) to `llm-router/auto`. opencode
|
|
reads config at startup, so it takes effect for sessions started after a
|
|
restart. `auto:batch` stays at 782,324 on purpose: batch is where long
|
|
contexts are expected.
|
|
|
|
- **Alerting ships with `desktop` (`notify-send`) only**, behind a notifier
|
|
built for more channels. SMS, RingCentral, PagerDuty and webhooks are later
|
|
adapters (section 6), not part of this lift.
|
|
|
|
No owner decisions remain open. Thresholds come from Phase 0.
|