Files
6krrt/docs/watchdog.md
adlee-was-taken 4d2bbf0253 fix(watchdog): exit 0 after a tick that ran, whatever it alerted
`python -m watchdog --once` exited with the number of alerts fired. It runs
as a systemd oneshot, so the first real alert would have marked
llm-router-watchdog.service FAILED (and a count past 255 wraps), making a
working alert look like a crash. The __main__ block becomes main(argv) and
returns 0 once the tick has run; the count is logged instead. The new test
fails against the old exit path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-26 21:44:06 -04:00

409 lines
18 KiB
Markdown

> Deep dive into the watchdog subsystem — opencode session loop detection. Back to
> [README](../README.md).
Watchdog is a periodic scanner that probes running opencode sessions for looping
patterns — repeated tool calls without progress. It runs as a systemd oneshot
every 5 minutes (`deploy/llm-router-watchdog.timer`) and, when a session flags,
writes into a local SQLite table and fires desktop notifications through the
configured channel pipeline.
The detector itself is a pure module (`src/progress_detect.py`) that returns a
boolean verdict plus a reason dict. The orchestrator (`src/watchdog.py`) reads
opencode session data over HTTP, runs the detector, and manages the alert
state machine. The notifier (`src/notifier.py`) routes events through
configured channels to desktop alerts (`notify-send`).
## What it catches and doesn't
Watchdog looks for **looping** — sessions that make repeated tool calls against
the same targets without landing meaningful changes. It detects four signal
types:
- **Dup signal**: a call appears more than 25% of the time in the sliding
window. The call matcher first normalises `bash` commands (strips env var
prefixes and `cd` prefixes, joins lines, collapses comments), then checks
tool name + canonicalised JSON args (sorted keys) for equality.
- **Top signal**: a single target is called 12+ times in the window (or 8+
for read-only agents like explore/librarian/oracle). Targets merge calls on
the same `(tool, basename)` for file reads, `(bash, first-two-words)` for
commands, and `(tool, pattern[:60])` for grep/glob.
- **Slow signal**: a single tool call accounts for 15+ of the session's total
calls, AND no call in the entire session history has landed (tree write or
git commit). This catches sessions stuck on a single read/analysis.
- **Coverage signal**: the session rereads one file at 4x its line count
within the window. The coverage score divides total bytes read by file length,
capped at the file's actual lines. A file read 4x its length means the model
is rereading without making progress.
**NOT steady spend.** A session that steadily calls different files on a
difficult refactor is not flagged — each call lands a unique `(tool, args)`
and the top target never crosses the threshold. The detector only fires when
repetition outpaces the call budget, not when a session is quietly burning
tokens across many distinct tools.
### Landed calls
A call counts as **landed** when it:
- is an `edit`, `write`, or `patch` with a non-empty `metadata.diff` field, or
- is a `bash` command containing `git commit` with `metadata.exit == 0`
Landed calls break both the slow and coverage signals. A session that reads a
file 4x, then edits it once, clears all signals immediately — the landed check
exits early and the verdict returns `False`.
### Ancestral trees
Sessions can be parented: the opencode API returns `parentID` on child sessions.
Watchdog resolves root sessions and evaluates EACH session in the tree on its
own calls, NOT merged into the root. A child's flagged verdict bubbles up to the
root level, which is what the alert dedup key uses (`opencode-loop:{root_session_id}`).
For the landed-time check, a session sees NOT only its own landed calls but also
its transitive descendants' landed calls. This prevents a parent session from
being flagged just because its child landed — even though the parent may still
be looping independently.
## Signals and thresholds
Thresholds live in `DetectConfig` (`src/progress_detect.py:38-53`) and are
configurable under `watchdog.detector.*` in `config/config.yaml`.
| Key | Default | Meaning |
|---|---|---|
| `watchdog.detector.dup_min` | 0.25 | Fraction: calls must repeat at this rate to flag |
| `watchdog.detector.top_min` | 12 | Minimum calls to a single target (normal agents) |
| `watchdog.detector.top_min_ro` | 8 | Minimum calls to a single target (read-only agents) |
| `watchdog.detector.cum_min` | 15 | Same call must account for this many total calls |
| `watchdog.detector.cover_min` | 4.0 | File must be reread this many times its length |
| `watchdog.detector.window` | 60 | Sliding window in number of tool calls (not seconds) |
| `watchdog.detector.min_calls` | 40 | Minimum total calls before evaluation triggers |
| `watchdog.detector.read_only_agent_keywords` | `["explore", "librarian", "oracle"]` | Title keywords that make an agent read-only (lowered `top_min` from 12 to 8) |
Read-only agents get a lowered top threshold because agents whose job is
exploration or library work are expected to call the same targets repeatedly —
the floor is 8 instead of 12.
## Alert lifecycle
The alert state machine is per-root-session and tracked in the
`watchdog_alerts` table.
### Trigger (initial detection)
When the detector flags a root session's tree for the first time:
1. The watchdog writes a `watchdog_alerts` row with `state='open'`.
2. The `local_llm_enabled` path asks a local Ollama model (defaulting to
`verification.model` or `qwen2.5-coder-router:14b`) for a second opinion.
It receives the agent name and the reason dict as context, and the model
must respond with only "yes" or "no" — no explanation.
3. If the local LLM answers "yes", severity is set to `critical`. If it
answers "no" or times out (returns `None`), severity is set to `warning`.
4. The notifier dispatches an `AlertEvent` through every matching channel.
5. The `flagged_ticks` counter starts at 1.
A cap of 3 simultaneous LLM second-opinion calls (`_MAX_LLM = 3`) prevents a
burst of flagged sessions from flooding the local model.
### Escalate (persistent looping)
On every subsequent tick where the session is still flagged:
1. `flagged_ticks` is incremented.
2. If `flagged_ticks >= 3` (15 minutes at 5-minute ticks) OR the LLM answers
"yes" again, the alert escalates to `severity='critical'` and the escalation
event is dispatched.
3. If the alert is already critical, ticking continues but no duplicate
escalation fires.
### Resolve (session is gone or fixed)
A root session alert resolves in two conditions:
- **Session disappeared**: the root session ID is no longer returned by the
opencode `/session` API. The watchdog writes `state='resolved'`,
`resolved_at=<now>`, and `severity='info'`.
- **Session fixed**: the root session's tree no longer has any flagged sessions
in the detector's verdict. The alert is resolved the same way and a resolve
event is dispatched.
Resolve events always bypass the rate limit — they must clear the previous
notification so the operator knows the alert is over.
### No-opencode
When no `rc-servers.json` or no opencode server answers, the tick writes an
`outcome='no_opencode'` record and returns immediately with zero alerts. This
uses the filesystem lock (`router.db/.watchdog.lock`) so two watchdog instances
cannot run simultaneously.
## Admin surfaces
Watchdog data surfaces through three admin pages, each serving a different
operational question.
### Loops panel — index.html#loops
On the main dashboard (`admin/frontend/index.html`), the **Loops** card
(id="loops") shows all currently open alerts:
- **Severity badge**: the alert's current severity (`critical` in red,
`warning` in orange-yellow, `info` in grey)
- **Dedup key**: the raw `opencode-loop:{root_session_id}` string, so the
operator can cross-reference with `journalctl` or `router.db`
- **Opened at**: `opened_at` ISO timestamp from the first trigger
- State: `open` or `resolved`; the panel filters to `resolved_at IS NULL`
The panel is client-side only — it polls `/api/watchdog/loops` which queries
`watchdog_alerts` for all unresolved rows.
### Watchdog card — controls.html
The Controls page (`admin/frontend/controls.html`) includes a dedicated
**Watchdog** card with operational status:
- **Last tick**: timestamp of the most recent `watchdog_ticks` row
- **Sessions seen**: how many opencode sessions were enumerated on that tick
- **Flagged verdicts**: count of flagged + total from the latest tick's
`watchdog_verdicts` rows
- **Open alerts**: count from `watchdog_alerts WHERE resolved_at IS NULL`
- **Notification channels**: list of configured channels with toggle controls
(enabled/min_severity), editable via `POST /api/watchdog/channels`
- **Refresh** button: reloads the watchdog status from the API
- **Send test alert** button: fires a test `AlertEvent` through enabled channels
The status endpoints are:
- `GET /api/watchdog/status` — last tick, verdict counts, open alert count
- `GET /api/watchdog/channels` — per-channel settings from `watchdog_channel_settings`
- `POST /api/watchdog/channels` — upsert channel settings (enabled, min_severity)
- `POST /api/watchdog/test-alert` — deliver a test event (severity + optional
channel_name filter) to a live Notifier instance
### Per-model stall rollup — models.html
The `watchdog_verdicts` table has `model_id`/`provider` columns and indexes on
them, but no per-model stall rollup, no Block button beside evidence, and no
Blocked list are built yet. `blocked` is only a value in the Models override
dropdown (see docs/admin-portal.md).
## Database schema
Four tables live in `router.db`, created idempotently by both
`config/schema.sql` and `watchdog_store.py`:
```sql
watchdog_ticks -- one row per watchdog tick
watchdog_verdicts -- per-session verdict at each tick (joined to ticks)
watchdog_alerts -- alert lifecycle state machine (dedup_key PK)
watchdog_channel_settings -- per-channel toggle + severity gate
```
Indexes on `watchdog_verdicts` cover `model_id`, `created_at`, `flagged`, and
`session_root` — the columns the admin endpoint queries most frequently.
The ticks table carries `ticked_at`, `sessions_seen`, and `outcome`
(`'ok'`, `'flagged'`, or `'no_opencode'`). The verdicts table joins to ticks
via `tick_id` and carries the full reason dict fields: `dup`, `top`,
`top_what`, `landed`, `slow`, `coverage`, plus `calls_since_landed`,
`cost_since_landed_usd`, and the LLM second-opinion answer.
Alerts are keyed by `dedup_key` (deduplicated per root session) and carry
`opened_at`, `last_fired_at`, and `resolved_at`. The transition logic lives in
the `_fire_alert()` helper: trigger inserts only if no open row exists,
escalate increments `flagged_ticks` and bumps severity, and resolve writes
`resolved_at` and downgrades to `severity='info'`.
## Knobs
All watchdog configuration lives under `watchdog:` in `config/config.yaml`:
| Key | Default | Required |
|---|---|---|
| `watchdog.enabled` | `true` | Gates the entire subsystem |
| `watchdog.local_llm_enabled` | `true` | Whether to call Ollama for second opinions |
| `watchdog.model` | `null` (uses `verification.model`) | Local model name |
| `watchdog.detector.window` | 60 | Sliding window size in calls |
| `watchdog.detector.dup_min` | 0.25 | Dup threshold (fraction) |
| `watchdog.detector.top_min` | 12 | Top target threshold (normal agents) |
| `watchdog.detector.top_min_ro` | 8 | Top target threshold (read-only agents) |
| `watchdog.detector.cum_min` | 15 | Slow-signal threshold |
| `watchdog.detector.cover_min` | 4.0 | Coverage-signal multiplier |
| `watchdog.detector.min_calls` | 40 | Minimum eval calls |
| `watchdog.read_only_agents` | `["explore", "librarian", "oracle"]` | Agent title keywords for reduced threshold |
| `watchdog.dashboard_base_url` | `"http://127.0.0.1:8080/admin"` | Base URL for alert links |
All seven `watchdog.detector.*` keys plus `watchdog.enabled` and
`watchdog.local_llm_enabled` (nine allowlisted keys total) are allowlisted for
live editing through the
admin portal (`POST /admin/api/config/{key}`), which writes to
`config.local.yaml` (the gitignored machine-local overlay). The changes are
validated by `RouterConfig` before reaching disk, and take effect at the next
tick — no service restart required since `detect_config_from_pydantic()` reads
cfg fresh each invocation.
## Install
The watchdog ships as two systemd user units:
- `llm-router-watchdog.service` — a `Type=oneshot` that runs
`PYTHONPATH=%h/llm-router/src .venv/bin/python -m watchdog --once`
- `llm-router-watchdog.timer` — fires 5 min after boot, then every 5 min
Installation follows the same pattern as the other deploy units. The shipped
unit files use `%h/llm-router` placeholders that need rewriting to your actual
repo path:
```bash
REPO=$(pwd)
for u in deploy/llm-router-watchdog.{service,timer}; do
sed "s|%h/llm-router|${REPO}|g" "$u" \
> ~/.config/systemd/user/"$(basename "$u")"
done
systemctl --user daemon-reload
systemctl --user enable --now llm-router-watchdog.timer
```
The service unit has **no `EnvironmentFile`** — watchdog needs no provider API
keys because it only queries the local opencode session APIs and the local
Ollama instance.
Check installation:
```bash
systemctl --user status llm-router-watchdog.timer
journalctl --user -u llm-router-watchdog.service --no-pager
```
## Run --once
For ad-hoc debugging or pre-deploy smoke test:
```bash
PYTHONPATH=src .venv/bin/python -m watchdog --once
```
The `--once` flag runs a single tick and exits 0 once the tick has run,
however many alerts it fired (the count is logged as `watchdog tick done: N
alert(s) fired`); systemd would otherwise mark the oneshot unit failed on every
real alert. It connects to the same `router.db` that the timed
service uses, acquires the filesystem lock, reads `rc-servers.json`, probes
opencode servers, and follows the full evaluate-alert-resolve pipeline.
Add `--config /path/to/config.yaml` to point at a non-default config file.
### Debug output
The `--once` run emits structured `watchdog=` log lines for every state
transition:
```
2026-01-15 14:30:01 INFO watchdog: watchdog=2026-01-15T14:30:01+00:00 dedup='opencode-loop:ses_xxxxx' state=trigger severity=warning title='opencode loop: explore'
```
The `title` field carries the agent name derived from the session title. The
`dedup` field is the root session's dedup key, matching `journalctl` traces
and admin panel rows.
## Backtest
`scripts/progress_backtest.py` replays labelled sessions through the detector
to validate threshold calibration. It reads from a fixture JSON file and runs
a sliding-window evaluation with step 5.
```bash
PYTHONPATH=src .venv/bin/python -m scripts.progress_backtest --fixture \
tests/fixtures/progress/fixture.json
```
The fixture ships in `tests/fixtures/progress/fixture.json` and contains
15 labelled sessions: 8 marked `must_flag` and 7 marked `must_not_flag`.
Expected output after a successful run:
```
# backtest complete: 8/15 flagged
```
At least 8 sessions should trigger a flag. The fixture is generated from live
opencode sessions (see [export fixture](#re-export-fixture) below) and scrubbed
to protect file contents — content is SHA-1 hashed except for the few keys the
detector actually inspects (`filePath`, `command`, `pattern`, etc.).
## Re-export fixture
The fixture generator pulls sessions from the opencode SQLite database
(`~/.local/share/opencode/opencode.db`) and writes the scrubbed test fixture:
```bash
python scripts/export_progress_fixture.py
```
Output goes to `tests/fixtures/progress/fixture.json`. The script takes a
session ID to label map (`LABELS` dict), extracts call history, resolves file
line counts (via `git show` for deleted worktree paths), and scrubs all strings
to SHA-1 prefixes, preserving only the detector's target keys.
The fixture generation script hardcodes paths to the operator's home directory
and a specific worktree commit (`WORKTREE_COMMIT = "827c408"`). When
regenerating, update those paths to match the current environment.
## Notifier pipeline
Alert events flow through `src/notifier.py` before reaching the operator. Each
alert is an `AlertEvent` dataclass mirroring PagerDuty Events v2 shape:
`dedup_key`, `severity` (info/warning/critical), `state` (trigger/escalate/
resolve), `title`, `summary`, `details`, and `source` (hardcoded to
`"6krrt-watchdog"`).
Each configured channel has its own `min_severity` gate — a critical event
passes through a channel with `min_severity=warning`, but a warning event does
not pass a channel with `min_severity=critical`. Each channel also has a rate
limit window of 300 seconds (5 minutes). Resolve events always bypass the rate
limit to ensure clear-the-alert notifications are delivered.
Currently only `type=desktop` channels are implemented; they invoke
`notify-send -u {urgency} {title} {summary}`. Unknown channel types are
logged as warnings and skipped. Missing `notify-send` (no `DISPLAY` or
`libnotify` installed) is logged at warning level — the notifier never raises.
Channel settings (`watchdog_channel_settings`) are editable at runtime through
the admin portal and stored per-channel in the database, separate from the
base config.
## Known limits
- **Under 40 calls: invisible.** `min_calls=40` means the detector does not
evaluate sessions with fewer tool calls. Short troubleshooting sessions that
loop on 10 calls pass through undetected.
- **Atlas at cover_min 4.0 exactly.** A session whose last 60 calls read a
single file exactly 4x its line count hits the coverage threshold at the
boundary. There is no margin: `>=` comparison means exactly 4.0x flags.
This matters for sessions that genuinely need to re-read a reference file
repeatedly — the coverage signal cannot be tuned per-file.
- **Stale localhost:4096 probes.** Watchdog reads `rc-servers.json` from
`~/.local/share/opencode/rc-servers.json`, which may contain stale entries
for opencode sessions that have since ended. These produce empty call lists
and are silently skipped — but they add network round-trip latency to each
tick. Only the first answering server's sessions are evaluated (the watchdog
picks `answering[0]` from the list of servers that successfully return a
session list).
- **No streaming path.** `POST /outcome` is the only ground truth signal for
streaming traffic; the watchdog has no equivalent endpoint to report streaming
session outcomes. It can only see tool-call repetition, not whether the answer
was useful.
- **Local LLM gate.** Second-opinion calls are gated by both
`local_compute.enabled` (global) and `watchdog.local_llm_enabled`
(per-subsystem). If either is false, no LLM calls are made and severity stays
as `warning` on initial trigger regardless of model quality. The cap of 3
simultaneous LLM calls means that if 5 sessions flag, 2 wait without opinion.
- **No provider calls.** Watchdog does not touch the dispatch path. It only
reads opencode session data and runs local heuristics. No provider calls, no
quota consume, no cost incurred.