Files
6krrt/plans/no-progress-lift-a.md
adlee-was-taken 69c969d104 plans: declare a valid Status on the six plans this branch adds
tests/test_plans_declare_status.py requires 'Status: <done|planned|in
progress|parked|reference> -- <reason>' in the first 8 lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-26 21:46:15 -04:00

228 lines
10 KiB
Markdown

# Lift A: loop watchdog, desktop alerts, and the shared detector
Status: done -- shipped in PR #102; the admin surfacing (component 5) shipped only partly
Date: 2026-09-26
Parent spec: `plans/no-progress-detection.md` (reference only; do NOT read it
end to end). This brief is the contract for Lift A.
Incident: `docs/incidents.md` #8.
## Goal
Catch an opencode agent session that is spinning without progress, and alert
the operator on the desktop, with nobody watching and no cloud model involved.
## Why this slice first, and why it is small
Incident #8's loops were caught only by a human or a Claude session reading
opencode's session store by hand. Lift A replaces that with a local timer.
It is deliberately narrow because the last planning attempt failed on size:
- a 540-line spec
- a planning context that grew to about 600k tokens
- the model reporting "elided" reads that were stored intact, and re-reading
the spec 84 times (68x its length)
Keep this lift's inputs, todos and contexts small.
**Out of scope (Lift B):**
- router-side progress events, the live judge, and `/progress`
- recovery modes (`auto_compact`, `auto_recover`, `auto_limit_context`)
- scoped 429s
- the `router-link.js` changes
- SMS, PagerDuty and other alert channels
## Components
### 1. Shared detector: `src/progress_detect.py` (pure, stdlib only)
Port the rules from `plans/no-progress-detection-prototype.py`. It is tuned on
the incident #8 sessions; keep its behaviour and make it a clean, tested
module. Input: a session's tool calls as `(time, tool, args_json, landed)`.
Output: a verdict with the reason. Signals, over a window of the last
`window` calls:
- `dup`: exact `(tool, args)` duplicate share
- `top`: the most-repeated target. A `read` target is the file plus
offset/limit. A `bash` target is the files it names plus the numbers in the
command (the slice).
- `slow`: one exact call repeated `cum_min`+ times over the whole session
- `coverage`: lines requested from one file in the window divided by its
length. This catches re-reading a file through shifting offsets, which the
other signals miss.
- `landed`: an edit or write with a non-empty diff, or a `git commit` with
exit 0, by the session or any descendant (`parentID`) within the window.
Verdict:
- Normal agents are flagged when nothing landed and any of dup, top, slow or
coverage crosses its threshold.
- Read-only agents (explore, librarian, oracle; a configurable list) never
land, so for them top, slow or coverage alone decides.
- Thresholds start at the prototype's values: `window` 60, `dup_min` 0.25,
`top_min` 12, `top_min_ro` 8, `cum_min` 15, `cover_min` 4.0.
### 2. Calibration backtest: `scripts/progress_backtest.py`
It replays sessions from opencode's HTTP API through `progress_detect` and
prints a per-session verdict with the first flag point. Commit a small
**fixture** of the labelled sessions below, exported as tool-call fingerprints
(tool name, args, landed; never message text), under `tests/fixtures/`. A test
asserts every label. Live opencode is not needed in tests.
MUST FLAG:
| session | what it was |
|---|---|
| `ses_f24cc0258ffepbJPoOzoiLZi9s` | item 2 worker, read navbar.js x61 |
| `ses_f24fcf651ffeOm9O4E14QSZeJ4` | item 2 worker, 1st attempt |
| `ses_f24890640ffewKnIPqjjAfFvOk` | item 3 worker, 1st attempt |
| `ses_f245607eaffeGxPH8F1tvvdvkG` | item 4 worker |
| `ses_f237e7e31ffeOzz4kUKu4CWzMu` | explore helper re-reading a deleted file |
| `ses_f235d755bffeo0uWtD1g2HwQUg` | helper that printed the spec 12x |
| `ses_f2506ac70ffebLUxBsq0aauOn9` | Atlas; must flag in its post-02:07 stretch |
| `ses_f2380dd4cffeVwcZn3Kv3nQN7v` | planner that re-read the spec 68x its length |
MUST NOT FLAG:
| session | what it was |
|---|---|
| `ses_f246e699fffeCHUPTLr5ajBhWC` | item 3 retry |
| `ses_f245f78f8ffeGMn3PgArpOwhFy` | item 3 compact |
| `ses_f243ea08effemMwXdMRE1gpnhK` | item 5 |
| `ses_f242e5e28ffehF97ZxxdbTeylQ` | item 6 |
| `ses_f236bbdaaffezwnvH10X2rgrxc` | explore, sliced reads of dispatcher.py |
| `ses_f236bc7baffeG5LQsr2K1c46xO` | explore |
| `ses_f2e81929fffeaVbWui0X9QjmRQ` | cockpit planner |
The Atlas session also must NOT flag before its 02:07 stretch.
Export the fixture FIRST, in the first todo. opencode sessions can be deleted,
and the fixture is the ground truth.
### 3. Notifier: `src/notifier.py`, `desktop` channel only
- **An alert is an event:**
- `dedup_key` (the session tree root)
- `severity` (`info`, `warning` or `critical`)
- `state` (`trigger`, `escalate` or `resolve`)
- `title`, `summary`, `details`, `source`
The shape maps one-to-one onto PagerDuty Events v2, so later adapters are
thin. Only the `desktop` channel ships.
- **`desktop` runs `notify-send`,** with urgency `critical` for critical
alerts.
- **Fire on transitions only,** with a per-channel rate limit and a quiet
window.
- **A delivery failure is logged, never raised.**
- **Config:** a `notifications.channels` list of
`{name, type, enabled, min_severity}`. An unknown `type` fails config load
(StrictModel). Secrets would come from env vars named in config; desktop has
none.
### 4. Watchdog: `src/watchdog.py` + `deploy/llm-router-watchdog.{service,timer}`
A systemd user timer, shaped like the existing `llm-router-*` timers, every 5
min. Each tick:
1. **Find opencode.** Read `~/.local/share/opencode/rc-servers.json` and probe
each `serverUrl`. If none answers, exit 0 quietly.
2. **Read activity.** List sessions active since the last tick and read their
tool calls.
3. **Detect.** Run `progress_detect` on each active session.
4. **Second opinion, flagged sessions only.** Ask the configured local model
(`watchdog.model`, default `verification.model`, today
`qwen2.5-coder-router:14b` on the local Ollama) "looping without progress?
yes/no + one line", on a digest of tool names, targets and exit codes, never
content. Skip this when `local_compute.enabled` is false or
`watchdog.local_llm_enabled` is false. Report the model's answer alongside
the deterministic verdict; it does not veto it in Lift A.
5. **Alert.** Alert through the notifier directly, and write rows to the
router DB: a `watchdog_ticks` table (time, sessions seen, outcome) and a
`watchdog_verdicts` table. The router need not be up; SQLite WAL handles the
second writer.
Also:
- A lock file prevents overlapping ticks.
- An unexpected opencode API shape raises one `warning` alert ("watchdog is
blind"), never a silent all-clear.
- Admin (North Star #1): a small Watchdog card showing the last tick, recent
verdicts, the channel list with enabled and `min_severity` controls, and a
**Send test alert** button. Every new knob gets a control or a recorded
reason; `tests/test_admin_knob_coverage.py` enforces it.
- Knobs:
- `watchdog.enabled`, `watchdog.local_llm_enabled`, `watchdog.model`,
`watchdog.read_only_agents`
- the detector thresholds from section 1
- `notifications.channels`
The tick interval lives in the timer unit; document it next to the knobs.
## Ground rules
- **Base on `origin/main`.** Branch `feat/no-progress-lift-a` in its own
worktree. Do not touch, stash or commit anything in the main checkout.
- **Port 8080 is production.** Integration checks run on 8081
(`scripts/sandbox.sh`), against a sqlite-backup copy of the live DB.
- **Never edit** `config/config.yaml`'s existing values or
`config/config.local.yaml`. New keys get defaults in `config.yaml`.
- **Do not install or enable** the timer on this machine; that is an operator
step in the done report.
- **Commit by explicit path; never `git add -A`.** The orchestrator checks each
worker's `git show --stat` against its claims and the plan's paths before
accepting it.
- **Read in targeted slices.** Grep for the anchor, read around it. Do NOT read
`plans/no-progress-detection.md` whole; this brief is the contract.
- **If a worker's context passes about 150k tokens, stop it** and start a fresh
worker with a smaller prompt, rather than letting it grind.
- **ASCII only; no middle dots.** No prose paragraphs in the admin UI.
- **Lint:** pinned `uvx ruff@0.16.9`, no NEW findings in touched files against
the base, never `ruff --fix` outside changed lines.
- **pytest:** record the `origin/main` baseline first; it stays green apart
from the baseline.
## Done means
- Every must-flag and must-not-flag label passes in tests, and the live
backtest prints the same verdicts.
- One real `notify-send` from the admin test button on 8081.
- The watchdog runs once by hand (`python -m watchdog --once`) against live
opencode and prints its verdicts.
- The done report lists each commit, the tests, and the operator steps
(install and enable the timer).
## Addendum, 2026-09-26: surfacing and one-click block (North Star #4)
The owner set a new priority: surface waste in the admin portal, and make
stopping it one click (`CLAUDE.md` North Star #4). This lift therefore also
ships:
### 5. Loops panel and one-click model block (admin portal)
- **A Loops panel** on the Home page (or the Controls page, whichever the
cockpit layout fits), reading `watchdog_verdicts`: each flagged session with
agent, verdict, turns and $ since the last landed change, the top repeated
target, and when it was flagged.
- **A per-model rollup:** stall verdicts per model over the last hour, next to
each model's share of traffic. It needs session-to-model attribution; see
item 0.
- **A Block model button** beside each model in the rollup. It writes
`admin_model_overrides` with `availability = 'blocked'` and a reason (default
"stalled N sessions", editable), through the existing
`POST /admin/api/models/{model_id}/{provider}/availability` path. Verify that
routing excludes `blocked` exactly like `deprecated` (`load_candidates`), and
add a test for it.
- **A Blocked models list** with reason, time and a one-click Unblock
(`DELETE .../availability`).
- **Desktop alerts link to the panel** (`http://127.0.0.1:8080/admin/` plus an
anchor).
### 0. First todo: why do the identity headers not arrive?
`router-link.js` (#99) is installed in `~/.config/opencode/plugins/`, is
identical to the repo copy, and passes its 29 tests. The router reads
`X-Router-Conversation` (`src/dispatcher.py`, `resolve_identity(request.headers,
...)`). Yet production `route_decisions` after the 04:02 opencode restart still
show fingerprint `session_key`s and `agent = NULL`. Find out why, and fix it if
it is a small change.
The per-model rollup depends on it. If it cannot be fixed in this lift, the
rollup ships disabled with a visible "needs conversation identity" note, and
everything else ships.