Files
6krrt/plans/no-progress-lift-a.md
adlee-was-taken 69c969d104 plans: declare a valid Status on the six plans this branch adds
tests/test_plans_declare_status.py requires 'Status: <done|planned|in
progress|parked|reference> -- <reason>' in the first 8 lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-26 21:46:15 -04:00

10 KiB

Lift A: loop watchdog, desktop alerts, and the shared detector

Status: done -- shipped in PR #102; the admin surfacing (component 5) shipped only partly Date: 2026-09-26 Parent spec: plans/no-progress-detection.md (reference only; do NOT read it end to end). This brief is the contract for Lift A. Incident: docs/incidents.md #8.

Goal

Catch an opencode agent session that is spinning without progress, and alert the operator on the desktop, with nobody watching and no cloud model involved.

Why this slice first, and why it is small

Incident #8's loops were caught only by a human or a Claude session reading opencode's session store by hand. Lift A replaces that with a local timer.

It is deliberately narrow because the last planning attempt failed on size:

  • a 540-line spec
  • a planning context that grew to about 600k tokens
  • the model reporting "elided" reads that were stored intact, and re-reading the spec 84 times (68x its length)

Keep this lift's inputs, todos and contexts small.

Out of scope (Lift B):

  • router-side progress events, the live judge, and /progress
  • recovery modes (auto_compact, auto_recover, auto_limit_context)
  • scoped 429s
  • the router-link.js changes
  • SMS, PagerDuty and other alert channels

Components

1. Shared detector: src/progress_detect.py (pure, stdlib only)

Port the rules from plans/no-progress-detection-prototype.py. It is tuned on the incident #8 sessions; keep its behaviour and make it a clean, tested module. Input: a session's tool calls as (time, tool, args_json, landed). Output: a verdict with the reason. Signals, over a window of the last window calls:

  • dup: exact (tool, args) duplicate share
  • top: the most-repeated target. A read target is the file plus offset/limit. A bash target is the files it names plus the numbers in the command (the slice).
  • slow: one exact call repeated cum_min+ times over the whole session
  • coverage: lines requested from one file in the window divided by its length. This catches re-reading a file through shifting offsets, which the other signals miss.
  • landed: an edit or write with a non-empty diff, or a git commit with exit 0, by the session or any descendant (parentID) within the window.

Verdict:

  • Normal agents are flagged when nothing landed and any of dup, top, slow or coverage crosses its threshold.
  • Read-only agents (explore, librarian, oracle; a configurable list) never land, so for them top, slow or coverage alone decides.
  • Thresholds start at the prototype's values: window 60, dup_min 0.25, top_min 12, top_min_ro 8, cum_min 15, cover_min 4.0.

2. Calibration backtest: scripts/progress_backtest.py

It replays sessions from opencode's HTTP API through progress_detect and prints a per-session verdict with the first flag point. Commit a small fixture of the labelled sessions below, exported as tool-call fingerprints (tool name, args, landed; never message text), under tests/fixtures/. A test asserts every label. Live opencode is not needed in tests.

MUST FLAG:

session what it was
ses_f24cc0258ffepbJPoOzoiLZi9s item 2 worker, read navbar.js x61
ses_f24fcf651ffeOm9O4E14QSZeJ4 item 2 worker, 1st attempt
ses_f24890640ffewKnIPqjjAfFvOk item 3 worker, 1st attempt
ses_f245607eaffeGxPH8F1tvvdvkG item 4 worker
ses_f237e7e31ffeOzz4kUKu4CWzMu explore helper re-reading a deleted file
ses_f235d755bffeo0uWtD1g2HwQUg helper that printed the spec 12x
ses_f2506ac70ffebLUxBsq0aauOn9 Atlas; must flag in its post-02:07 stretch
ses_f2380dd4cffeVwcZn3Kv3nQN7v planner that re-read the spec 68x its length

MUST NOT FLAG:

session what it was
ses_f246e699fffeCHUPTLr5ajBhWC item 3 retry
ses_f245f78f8ffeGMn3PgArpOwhFy item 3 compact
ses_f243ea08effemMwXdMRE1gpnhK item 5
ses_f242e5e28ffehF97ZxxdbTeylQ item 6
ses_f236bbdaaffezwnvH10X2rgrxc explore, sliced reads of dispatcher.py
ses_f236bc7baffeG5LQsr2K1c46xO explore
ses_f2e81929fffeaVbWui0X9QjmRQ cockpit planner

The Atlas session also must NOT flag before its 02:07 stretch.

Export the fixture FIRST, in the first todo. opencode sessions can be deleted, and the fixture is the ground truth.

3. Notifier: src/notifier.py, desktop channel only

  • An alert is an event:

    • dedup_key (the session tree root)
    • severity (info, warning or critical)
    • state (trigger, escalate or resolve)
    • title, summary, details, source

    The shape maps one-to-one onto PagerDuty Events v2, so later adapters are thin. Only the desktop channel ships.

  • desktop runs notify-send, with urgency critical for critical alerts.

  • Fire on transitions only, with a per-channel rate limit and a quiet window.

  • A delivery failure is logged, never raised.

  • Config: a notifications.channels list of {name, type, enabled, min_severity}. An unknown type fails config load (StrictModel). Secrets would come from env vars named in config; desktop has none.

4. Watchdog: src/watchdog.py + deploy/llm-router-watchdog.{service,timer}

A systemd user timer, shaped like the existing llm-router-* timers, every 5 min. Each tick:

  1. Find opencode. Read ~/.local/share/opencode/rc-servers.json and probe each serverUrl. If none answers, exit 0 quietly.
  2. Read activity. List sessions active since the last tick and read their tool calls.
  3. Detect. Run progress_detect on each active session.
  4. Second opinion, flagged sessions only. Ask the configured local model (watchdog.model, default verification.model, today qwen2.5-coder-router:14b on the local Ollama) "looping without progress? yes/no + one line", on a digest of tool names, targets and exit codes, never content. Skip this when local_compute.enabled is false or watchdog.local_llm_enabled is false. Report the model's answer alongside the deterministic verdict; it does not veto it in Lift A.
  5. Alert. Alert through the notifier directly, and write rows to the router DB: a watchdog_ticks table (time, sessions seen, outcome) and a watchdog_verdicts table. The router need not be up; SQLite WAL handles the second writer.

Also:

  • A lock file prevents overlapping ticks.
  • An unexpected opencode API shape raises one warning alert ("watchdog is blind"), never a silent all-clear.
  • Admin (North Star #1): a small Watchdog card showing the last tick, recent verdicts, the channel list with enabled and min_severity controls, and a Send test alert button. Every new knob gets a control or a recorded reason; tests/test_admin_knob_coverage.py enforces it.
  • Knobs:
    • watchdog.enabled, watchdog.local_llm_enabled, watchdog.model, watchdog.read_only_agents
    • the detector thresholds from section 1
    • notifications.channels

The tick interval lives in the timer unit; document it next to the knobs.

Ground rules

  • Base on origin/main. Branch feat/no-progress-lift-a in its own worktree. Do not touch, stash or commit anything in the main checkout.
  • Port 8080 is production. Integration checks run on 8081 (scripts/sandbox.sh), against a sqlite-backup copy of the live DB.
  • Never edit config/config.yaml's existing values or config/config.local.yaml. New keys get defaults in config.yaml.
  • Do not install or enable the timer on this machine; that is an operator step in the done report.
  • Commit by explicit path; never git add -A. The orchestrator checks each worker's git show --stat against its claims and the plan's paths before accepting it.
  • Read in targeted slices. Grep for the anchor, read around it. Do NOT read plans/no-progress-detection.md whole; this brief is the contract.
  • If a worker's context passes about 150k tokens, stop it and start a fresh worker with a smaller prompt, rather than letting it grind.
  • ASCII only; no middle dots. No prose paragraphs in the admin UI.
  • Lint: pinned uvx ruff@0.16.9, no NEW findings in touched files against the base, never ruff --fix outside changed lines.
  • pytest: record the origin/main baseline first; it stays green apart from the baseline.

Done means

  • Every must-flag and must-not-flag label passes in tests, and the live backtest prints the same verdicts.
  • One real notify-send from the admin test button on 8081.
  • The watchdog runs once by hand (python -m watchdog --once) against live opencode and prints its verdicts.
  • The done report lists each commit, the tests, and the operator steps (install and enable the timer).

Addendum, 2026-09-26: surfacing and one-click block (North Star #4)

The owner set a new priority: surface waste in the admin portal, and make stopping it one click (CLAUDE.md North Star #4). This lift therefore also ships:

5. Loops panel and one-click model block (admin portal)

  • A Loops panel on the Home page (or the Controls page, whichever the cockpit layout fits), reading watchdog_verdicts: each flagged session with agent, verdict, turns and $ since the last landed change, the top repeated target, and when it was flagged.
  • A per-model rollup: stall verdicts per model over the last hour, next to each model's share of traffic. It needs session-to-model attribution; see item 0.
  • A Block model button beside each model in the rollup. It writes admin_model_overrides with availability = 'blocked' and a reason (default "stalled N sessions", editable), through the existing POST /admin/api/models/{model_id}/{provider}/availability path. Verify that routing excludes blocked exactly like deprecated (load_candidates), and add a test for it.
  • A Blocked models list with reason, time and a one-click Unblock (DELETE .../availability).
  • Desktop alerts link to the panel (http://127.0.0.1:8080/admin/ plus an anchor).

0. First todo: why do the identity headers not arrive?

router-link.js (#99) is installed in ~/.config/opencode/plugins/, is identical to the repo copy, and passes its 29 tests. The router reads X-Router-Conversation (src/dispatcher.py, resolve_identity(request.headers, ...)). Yet production route_decisions after the 04:02 opencode restart still show fingerprint session_keys and agent = NULL. Find out why, and fix it if it is a small change.

The per-model rollup depends on it. If it cannot be fixed in this lift, the rollup ships disabled with a visible "needs conversation identity" note, and everything else ships.