tests/test_plans_declare_status.py requires 'Status: <done|planned|in progress|parked|reference> -- <reason>' in the first 8 lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
228 lines
10 KiB
Markdown
228 lines
10 KiB
Markdown
# Lift A: loop watchdog, desktop alerts, and the shared detector
|
|
|
|
Status: done -- shipped in PR #102; the admin surfacing (component 5) shipped only partly
|
|
Date: 2026-09-26
|
|
Parent spec: `plans/no-progress-detection.md` (reference only; do NOT read it
|
|
end to end). This brief is the contract for Lift A.
|
|
Incident: `docs/incidents.md` #8.
|
|
|
|
## Goal
|
|
|
|
Catch an opencode agent session that is spinning without progress, and alert
|
|
the operator on the desktop, with nobody watching and no cloud model involved.
|
|
|
|
## Why this slice first, and why it is small
|
|
|
|
Incident #8's loops were caught only by a human or a Claude session reading
|
|
opencode's session store by hand. Lift A replaces that with a local timer.
|
|
|
|
It is deliberately narrow because the last planning attempt failed on size:
|
|
- a 540-line spec
|
|
- a planning context that grew to about 600k tokens
|
|
- the model reporting "elided" reads that were stored intact, and re-reading
|
|
the spec 84 times (68x its length)
|
|
|
|
Keep this lift's inputs, todos and contexts small.
|
|
|
|
**Out of scope (Lift B):**
|
|
- router-side progress events, the live judge, and `/progress`
|
|
- recovery modes (`auto_compact`, `auto_recover`, `auto_limit_context`)
|
|
- scoped 429s
|
|
- the `router-link.js` changes
|
|
- SMS, PagerDuty and other alert channels
|
|
|
|
## Components
|
|
|
|
### 1. Shared detector: `src/progress_detect.py` (pure, stdlib only)
|
|
|
|
Port the rules from `plans/no-progress-detection-prototype.py`. It is tuned on
|
|
the incident #8 sessions; keep its behaviour and make it a clean, tested
|
|
module. Input: a session's tool calls as `(time, tool, args_json, landed)`.
|
|
Output: a verdict with the reason. Signals, over a window of the last
|
|
`window` calls:
|
|
- `dup`: exact `(tool, args)` duplicate share
|
|
- `top`: the most-repeated target. A `read` target is the file plus
|
|
offset/limit. A `bash` target is the files it names plus the numbers in the
|
|
command (the slice).
|
|
- `slow`: one exact call repeated `cum_min`+ times over the whole session
|
|
- `coverage`: lines requested from one file in the window divided by its
|
|
length. This catches re-reading a file through shifting offsets, which the
|
|
other signals miss.
|
|
- `landed`: an edit or write with a non-empty diff, or a `git commit` with
|
|
exit 0, by the session or any descendant (`parentID`) within the window.
|
|
|
|
Verdict:
|
|
- Normal agents are flagged when nothing landed and any of dup, top, slow or
|
|
coverage crosses its threshold.
|
|
- Read-only agents (explore, librarian, oracle; a configurable list) never
|
|
land, so for them top, slow or coverage alone decides.
|
|
- Thresholds start at the prototype's values: `window` 60, `dup_min` 0.25,
|
|
`top_min` 12, `top_min_ro` 8, `cum_min` 15, `cover_min` 4.0.
|
|
|
|
### 2. Calibration backtest: `scripts/progress_backtest.py`
|
|
|
|
It replays sessions from opencode's HTTP API through `progress_detect` and
|
|
prints a per-session verdict with the first flag point. Commit a small
|
|
**fixture** of the labelled sessions below, exported as tool-call fingerprints
|
|
(tool name, args, landed; never message text), under `tests/fixtures/`. A test
|
|
asserts every label. Live opencode is not needed in tests.
|
|
|
|
MUST FLAG:
|
|
|
|
| session | what it was |
|
|
|---|---|
|
|
| `ses_f24cc0258ffepbJPoOzoiLZi9s` | item 2 worker, read navbar.js x61 |
|
|
| `ses_f24fcf651ffeOm9O4E14QSZeJ4` | item 2 worker, 1st attempt |
|
|
| `ses_f24890640ffewKnIPqjjAfFvOk` | item 3 worker, 1st attempt |
|
|
| `ses_f245607eaffeGxPH8F1tvvdvkG` | item 4 worker |
|
|
| `ses_f237e7e31ffeOzz4kUKu4CWzMu` | explore helper re-reading a deleted file |
|
|
| `ses_f235d755bffeo0uWtD1g2HwQUg` | helper that printed the spec 12x |
|
|
| `ses_f2506ac70ffebLUxBsq0aauOn9` | Atlas; must flag in its post-02:07 stretch |
|
|
| `ses_f2380dd4cffeVwcZn3Kv3nQN7v` | planner that re-read the spec 68x its length |
|
|
|
|
MUST NOT FLAG:
|
|
|
|
| session | what it was |
|
|
|---|---|
|
|
| `ses_f246e699fffeCHUPTLr5ajBhWC` | item 3 retry |
|
|
| `ses_f245f78f8ffeGMn3PgArpOwhFy` | item 3 compact |
|
|
| `ses_f243ea08effemMwXdMRE1gpnhK` | item 5 |
|
|
| `ses_f242e5e28ffehF97ZxxdbTeylQ` | item 6 |
|
|
| `ses_f236bbdaaffezwnvH10X2rgrxc` | explore, sliced reads of dispatcher.py |
|
|
| `ses_f236bc7baffeG5LQsr2K1c46xO` | explore |
|
|
| `ses_f2e81929fffeaVbWui0X9QjmRQ` | cockpit planner |
|
|
|
|
The Atlas session also must NOT flag before its 02:07 stretch.
|
|
|
|
Export the fixture FIRST, in the first todo. opencode sessions can be deleted,
|
|
and the fixture is the ground truth.
|
|
|
|
### 3. Notifier: `src/notifier.py`, `desktop` channel only
|
|
|
|
- **An alert is an event:**
|
|
- `dedup_key` (the session tree root)
|
|
- `severity` (`info`, `warning` or `critical`)
|
|
- `state` (`trigger`, `escalate` or `resolve`)
|
|
- `title`, `summary`, `details`, `source`
|
|
|
|
The shape maps one-to-one onto PagerDuty Events v2, so later adapters are
|
|
thin. Only the `desktop` channel ships.
|
|
- **`desktop` runs `notify-send`,** with urgency `critical` for critical
|
|
alerts.
|
|
- **Fire on transitions only,** with a per-channel rate limit and a quiet
|
|
window.
|
|
- **A delivery failure is logged, never raised.**
|
|
- **Config:** a `notifications.channels` list of
|
|
`{name, type, enabled, min_severity}`. An unknown `type` fails config load
|
|
(StrictModel). Secrets would come from env vars named in config; desktop has
|
|
none.
|
|
|
|
### 4. Watchdog: `src/watchdog.py` + `deploy/llm-router-watchdog.{service,timer}`
|
|
|
|
A systemd user timer, shaped like the existing `llm-router-*` timers, every 5
|
|
min. Each tick:
|
|
1. **Find opencode.** Read `~/.local/share/opencode/rc-servers.json` and probe
|
|
each `serverUrl`. If none answers, exit 0 quietly.
|
|
2. **Read activity.** List sessions active since the last tick and read their
|
|
tool calls.
|
|
3. **Detect.** Run `progress_detect` on each active session.
|
|
4. **Second opinion, flagged sessions only.** Ask the configured local model
|
|
(`watchdog.model`, default `verification.model`, today
|
|
`qwen2.5-coder-router:14b` on the local Ollama) "looping without progress?
|
|
yes/no + one line", on a digest of tool names, targets and exit codes, never
|
|
content. Skip this when `local_compute.enabled` is false or
|
|
`watchdog.local_llm_enabled` is false. Report the model's answer alongside
|
|
the deterministic verdict; it does not veto it in Lift A.
|
|
5. **Alert.** Alert through the notifier directly, and write rows to the
|
|
router DB: a `watchdog_ticks` table (time, sessions seen, outcome) and a
|
|
`watchdog_verdicts` table. The router need not be up; SQLite WAL handles the
|
|
second writer.
|
|
|
|
Also:
|
|
- A lock file prevents overlapping ticks.
|
|
- An unexpected opencode API shape raises one `warning` alert ("watchdog is
|
|
blind"), never a silent all-clear.
|
|
- Admin (North Star #1): a small Watchdog card showing the last tick, recent
|
|
verdicts, the channel list with enabled and `min_severity` controls, and a
|
|
**Send test alert** button. Every new knob gets a control or a recorded
|
|
reason; `tests/test_admin_knob_coverage.py` enforces it.
|
|
- Knobs:
|
|
- `watchdog.enabled`, `watchdog.local_llm_enabled`, `watchdog.model`,
|
|
`watchdog.read_only_agents`
|
|
- the detector thresholds from section 1
|
|
- `notifications.channels`
|
|
|
|
The tick interval lives in the timer unit; document it next to the knobs.
|
|
|
|
## Ground rules
|
|
|
|
- **Base on `origin/main`.** Branch `feat/no-progress-lift-a` in its own
|
|
worktree. Do not touch, stash or commit anything in the main checkout.
|
|
- **Port 8080 is production.** Integration checks run on 8081
|
|
(`scripts/sandbox.sh`), against a sqlite-backup copy of the live DB.
|
|
- **Never edit** `config/config.yaml`'s existing values or
|
|
`config/config.local.yaml`. New keys get defaults in `config.yaml`.
|
|
- **Do not install or enable** the timer on this machine; that is an operator
|
|
step in the done report.
|
|
- **Commit by explicit path; never `git add -A`.** The orchestrator checks each
|
|
worker's `git show --stat` against its claims and the plan's paths before
|
|
accepting it.
|
|
- **Read in targeted slices.** Grep for the anchor, read around it. Do NOT read
|
|
`plans/no-progress-detection.md` whole; this brief is the contract.
|
|
- **If a worker's context passes about 150k tokens, stop it** and start a fresh
|
|
worker with a smaller prompt, rather than letting it grind.
|
|
- **ASCII only; no middle dots.** No prose paragraphs in the admin UI.
|
|
- **Lint:** pinned `uvx ruff@0.16.9`, no NEW findings in touched files against
|
|
the base, never `ruff --fix` outside changed lines.
|
|
- **pytest:** record the `origin/main` baseline first; it stays green apart
|
|
from the baseline.
|
|
|
|
## Done means
|
|
|
|
- Every must-flag and must-not-flag label passes in tests, and the live
|
|
backtest prints the same verdicts.
|
|
- One real `notify-send` from the admin test button on 8081.
|
|
- The watchdog runs once by hand (`python -m watchdog --once`) against live
|
|
opencode and prints its verdicts.
|
|
- The done report lists each commit, the tests, and the operator steps
|
|
(install and enable the timer).
|
|
|
|
## Addendum, 2026-09-26: surfacing and one-click block (North Star #4)
|
|
|
|
The owner set a new priority: surface waste in the admin portal, and make
|
|
stopping it one click (`CLAUDE.md` North Star #4). This lift therefore also
|
|
ships:
|
|
|
|
### 5. Loops panel and one-click model block (admin portal)
|
|
|
|
- **A Loops panel** on the Home page (or the Controls page, whichever the
|
|
cockpit layout fits), reading `watchdog_verdicts`: each flagged session with
|
|
agent, verdict, turns and $ since the last landed change, the top repeated
|
|
target, and when it was flagged.
|
|
- **A per-model rollup:** stall verdicts per model over the last hour, next to
|
|
each model's share of traffic. It needs session-to-model attribution; see
|
|
item 0.
|
|
- **A Block model button** beside each model in the rollup. It writes
|
|
`admin_model_overrides` with `availability = 'blocked'` and a reason (default
|
|
"stalled N sessions", editable), through the existing
|
|
`POST /admin/api/models/{model_id}/{provider}/availability` path. Verify that
|
|
routing excludes `blocked` exactly like `deprecated` (`load_candidates`), and
|
|
add a test for it.
|
|
- **A Blocked models list** with reason, time and a one-click Unblock
|
|
(`DELETE .../availability`).
|
|
- **Desktop alerts link to the panel** (`http://127.0.0.1:8080/admin/` plus an
|
|
anchor).
|
|
|
|
### 0. First todo: why do the identity headers not arrive?
|
|
|
|
`router-link.js` (#99) is installed in `~/.config/opencode/plugins/`, is
|
|
identical to the repo copy, and passes its 29 tests. The router reads
|
|
`X-Router-Conversation` (`src/dispatcher.py`, `resolve_identity(request.headers,
|
|
...)`). Yet production `route_decisions` after the 04:02 opencode restart still
|
|
show fingerprint `session_key`s and `agent = NULL`. Find out why, and fix it if
|
|
it is a small change.
|
|
|
|
The per-model rollup depends on it. If it cannot be fixed in this lift, the
|
|
rollup ships disabled with a visible "needs conversation identity" note, and
|
|
everything else ships.
|