`python -m watchdog --once` exited with the number of alerts fired. It runs as a systemd oneshot, so the first real alert would have marked llm-router-watchdog.service FAILED (and a count past 255 wraps), making a working alert look like a crash. The __main__ block becomes main(argv) and returns 0 once the tick has run; the count is logged instead. The new test fails against the old exit path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
409 lines
18 KiB
Markdown
409 lines
18 KiB
Markdown
> Deep dive into the watchdog subsystem — opencode session loop detection. Back to
|
|
> [README](../README.md).
|
|
|
|
Watchdog is a periodic scanner that probes running opencode sessions for looping
|
|
patterns — repeated tool calls without progress. It runs as a systemd oneshot
|
|
every 5 minutes (`deploy/llm-router-watchdog.timer`) and, when a session flags,
|
|
writes into a local SQLite table and fires desktop notifications through the
|
|
configured channel pipeline.
|
|
|
|
The detector itself is a pure module (`src/progress_detect.py`) that returns a
|
|
boolean verdict plus a reason dict. The orchestrator (`src/watchdog.py`) reads
|
|
opencode session data over HTTP, runs the detector, and manages the alert
|
|
state machine. The notifier (`src/notifier.py`) routes events through
|
|
configured channels to desktop alerts (`notify-send`).
|
|
|
|
## What it catches and doesn't
|
|
|
|
Watchdog looks for **looping** — sessions that make repeated tool calls against
|
|
the same targets without landing meaningful changes. It detects four signal
|
|
types:
|
|
|
|
- **Dup signal**: a call appears more than 25% of the time in the sliding
|
|
window. The call matcher first normalises `bash` commands (strips env var
|
|
prefixes and `cd` prefixes, joins lines, collapses comments), then checks
|
|
tool name + canonicalised JSON args (sorted keys) for equality.
|
|
|
|
- **Top signal**: a single target is called 12+ times in the window (or 8+
|
|
for read-only agents like explore/librarian/oracle). Targets merge calls on
|
|
the same `(tool, basename)` for file reads, `(bash, first-two-words)` for
|
|
commands, and `(tool, pattern[:60])` for grep/glob.
|
|
|
|
- **Slow signal**: a single tool call accounts for 15+ of the session's total
|
|
calls, AND no call in the entire session history has landed (tree write or
|
|
git commit). This catches sessions stuck on a single read/analysis.
|
|
|
|
- **Coverage signal**: the session rereads one file at 4x its line count
|
|
within the window. The coverage score divides total bytes read by file length,
|
|
capped at the file's actual lines. A file read 4x its length means the model
|
|
is rereading without making progress.
|
|
|
|
**NOT steady spend.** A session that steadily calls different files on a
|
|
difficult refactor is not flagged — each call lands a unique `(tool, args)`
|
|
and the top target never crosses the threshold. The detector only fires when
|
|
repetition outpaces the call budget, not when a session is quietly burning
|
|
tokens across many distinct tools.
|
|
|
|
### Landed calls
|
|
|
|
A call counts as **landed** when it:
|
|
- is an `edit`, `write`, or `patch` with a non-empty `metadata.diff` field, or
|
|
- is a `bash` command containing `git commit` with `metadata.exit == 0`
|
|
|
|
Landed calls break both the slow and coverage signals. A session that reads a
|
|
file 4x, then edits it once, clears all signals immediately — the landed check
|
|
exits early and the verdict returns `False`.
|
|
|
|
### Ancestral trees
|
|
|
|
Sessions can be parented: the opencode API returns `parentID` on child sessions.
|
|
Watchdog resolves root sessions and evaluates EACH session in the tree on its
|
|
own calls, NOT merged into the root. A child's flagged verdict bubbles up to the
|
|
root level, which is what the alert dedup key uses (`opencode-loop:{root_session_id}`).
|
|
|
|
For the landed-time check, a session sees NOT only its own landed calls but also
|
|
its transitive descendants' landed calls. This prevents a parent session from
|
|
being flagged just because its child landed — even though the parent may still
|
|
be looping independently.
|
|
|
|
## Signals and thresholds
|
|
|
|
Thresholds live in `DetectConfig` (`src/progress_detect.py:38-53`) and are
|
|
configurable under `watchdog.detector.*` in `config/config.yaml`.
|
|
|
|
| Key | Default | Meaning |
|
|
|---|---|---|
|
|
| `watchdog.detector.dup_min` | 0.25 | Fraction: calls must repeat at this rate to flag |
|
|
| `watchdog.detector.top_min` | 12 | Minimum calls to a single target (normal agents) |
|
|
| `watchdog.detector.top_min_ro` | 8 | Minimum calls to a single target (read-only agents) |
|
|
| `watchdog.detector.cum_min` | 15 | Same call must account for this many total calls |
|
|
| `watchdog.detector.cover_min` | 4.0 | File must be reread this many times its length |
|
|
| `watchdog.detector.window` | 60 | Sliding window in number of tool calls (not seconds) |
|
|
| `watchdog.detector.min_calls` | 40 | Minimum total calls before evaluation triggers |
|
|
| `watchdog.detector.read_only_agent_keywords` | `["explore", "librarian", "oracle"]` | Title keywords that make an agent read-only (lowered `top_min` from 12 to 8) |
|
|
|
|
Read-only agents get a lowered top threshold because agents whose job is
|
|
exploration or library work are expected to call the same targets repeatedly —
|
|
the floor is 8 instead of 12.
|
|
|
|
## Alert lifecycle
|
|
|
|
The alert state machine is per-root-session and tracked in the
|
|
`watchdog_alerts` table.
|
|
|
|
### Trigger (initial detection)
|
|
|
|
When the detector flags a root session's tree for the first time:
|
|
|
|
1. The watchdog writes a `watchdog_alerts` row with `state='open'`.
|
|
2. The `local_llm_enabled` path asks a local Ollama model (defaulting to
|
|
`verification.model` or `qwen2.5-coder-router:14b`) for a second opinion.
|
|
It receives the agent name and the reason dict as context, and the model
|
|
must respond with only "yes" or "no" — no explanation.
|
|
3. If the local LLM answers "yes", severity is set to `critical`. If it
|
|
answers "no" or times out (returns `None`), severity is set to `warning`.
|
|
4. The notifier dispatches an `AlertEvent` through every matching channel.
|
|
5. The `flagged_ticks` counter starts at 1.
|
|
|
|
A cap of 3 simultaneous LLM second-opinion calls (`_MAX_LLM = 3`) prevents a
|
|
burst of flagged sessions from flooding the local model.
|
|
|
|
### Escalate (persistent looping)
|
|
|
|
On every subsequent tick where the session is still flagged:
|
|
|
|
1. `flagged_ticks` is incremented.
|
|
2. If `flagged_ticks >= 3` (15 minutes at 5-minute ticks) OR the LLM answers
|
|
"yes" again, the alert escalates to `severity='critical'` and the escalation
|
|
event is dispatched.
|
|
3. If the alert is already critical, ticking continues but no duplicate
|
|
escalation fires.
|
|
|
|
### Resolve (session is gone or fixed)
|
|
|
|
A root session alert resolves in two conditions:
|
|
|
|
- **Session disappeared**: the root session ID is no longer returned by the
|
|
opencode `/session` API. The watchdog writes `state='resolved'`,
|
|
`resolved_at=<now>`, and `severity='info'`.
|
|
- **Session fixed**: the root session's tree no longer has any flagged sessions
|
|
in the detector's verdict. The alert is resolved the same way and a resolve
|
|
event is dispatched.
|
|
|
|
Resolve events always bypass the rate limit — they must clear the previous
|
|
notification so the operator knows the alert is over.
|
|
|
|
### No-opencode
|
|
|
|
When no `rc-servers.json` or no opencode server answers, the tick writes an
|
|
`outcome='no_opencode'` record and returns immediately with zero alerts. This
|
|
uses the filesystem lock (`router.db/.watchdog.lock`) so two watchdog instances
|
|
cannot run simultaneously.
|
|
|
|
## Admin surfaces
|
|
|
|
Watchdog data surfaces through three admin pages, each serving a different
|
|
operational question.
|
|
|
|
### Loops panel — index.html#loops
|
|
|
|
On the main dashboard (`admin/frontend/index.html`), the **Loops** card
|
|
(id="loops") shows all currently open alerts:
|
|
|
|
- **Severity badge**: the alert's current severity (`critical` in red,
|
|
`warning` in orange-yellow, `info` in grey)
|
|
- **Dedup key**: the raw `opencode-loop:{root_session_id}` string, so the
|
|
operator can cross-reference with `journalctl` or `router.db`
|
|
- **Opened at**: `opened_at` ISO timestamp from the first trigger
|
|
- State: `open` or `resolved`; the panel filters to `resolved_at IS NULL`
|
|
|
|
The panel is client-side only — it polls `/api/watchdog/loops` which queries
|
|
`watchdog_alerts` for all unresolved rows.
|
|
|
|
### Watchdog card — controls.html
|
|
|
|
The Controls page (`admin/frontend/controls.html`) includes a dedicated
|
|
**Watchdog** card with operational status:
|
|
|
|
- **Last tick**: timestamp of the most recent `watchdog_ticks` row
|
|
- **Sessions seen**: how many opencode sessions were enumerated on that tick
|
|
- **Flagged verdicts**: count of flagged + total from the latest tick's
|
|
`watchdog_verdicts` rows
|
|
- **Open alerts**: count from `watchdog_alerts WHERE resolved_at IS NULL`
|
|
- **Notification channels**: list of configured channels with toggle controls
|
|
(enabled/min_severity), editable via `POST /api/watchdog/channels`
|
|
- **Refresh** button: reloads the watchdog status from the API
|
|
- **Send test alert** button: fires a test `AlertEvent` through enabled channels
|
|
|
|
The status endpoints are:
|
|
- `GET /api/watchdog/status` — last tick, verdict counts, open alert count
|
|
- `GET /api/watchdog/channels` — per-channel settings from `watchdog_channel_settings`
|
|
- `POST /api/watchdog/channels` — upsert channel settings (enabled, min_severity)
|
|
- `POST /api/watchdog/test-alert` — deliver a test event (severity + optional
|
|
channel_name filter) to a live Notifier instance
|
|
|
|
### Per-model stall rollup — models.html
|
|
|
|
The `watchdog_verdicts` table has `model_id`/`provider` columns and indexes on
|
|
them, but no per-model stall rollup, no Block button beside evidence, and no
|
|
Blocked list are built yet. `blocked` is only a value in the Models override
|
|
dropdown (see docs/admin-portal.md).
|
|
|
|
## Database schema
|
|
|
|
Four tables live in `router.db`, created idempotently by both
|
|
`config/schema.sql` and `watchdog_store.py`:
|
|
|
|
```sql
|
|
watchdog_ticks -- one row per watchdog tick
|
|
watchdog_verdicts -- per-session verdict at each tick (joined to ticks)
|
|
watchdog_alerts -- alert lifecycle state machine (dedup_key PK)
|
|
watchdog_channel_settings -- per-channel toggle + severity gate
|
|
```
|
|
|
|
Indexes on `watchdog_verdicts` cover `model_id`, `created_at`, `flagged`, and
|
|
`session_root` — the columns the admin endpoint queries most frequently.
|
|
|
|
The ticks table carries `ticked_at`, `sessions_seen`, and `outcome`
|
|
(`'ok'`, `'flagged'`, or `'no_opencode'`). The verdicts table joins to ticks
|
|
via `tick_id` and carries the full reason dict fields: `dup`, `top`,
|
|
`top_what`, `landed`, `slow`, `coverage`, plus `calls_since_landed`,
|
|
`cost_since_landed_usd`, and the LLM second-opinion answer.
|
|
|
|
Alerts are keyed by `dedup_key` (deduplicated per root session) and carry
|
|
`opened_at`, `last_fired_at`, and `resolved_at`. The transition logic lives in
|
|
the `_fire_alert()` helper: trigger inserts only if no open row exists,
|
|
escalate increments `flagged_ticks` and bumps severity, and resolve writes
|
|
`resolved_at` and downgrades to `severity='info'`.
|
|
|
|
## Knobs
|
|
|
|
All watchdog configuration lives under `watchdog:` in `config/config.yaml`:
|
|
|
|
| Key | Default | Required |
|
|
|---|---|---|
|
|
| `watchdog.enabled` | `true` | Gates the entire subsystem |
|
|
| `watchdog.local_llm_enabled` | `true` | Whether to call Ollama for second opinions |
|
|
| `watchdog.model` | `null` (uses `verification.model`) | Local model name |
|
|
| `watchdog.detector.window` | 60 | Sliding window size in calls |
|
|
| `watchdog.detector.dup_min` | 0.25 | Dup threshold (fraction) |
|
|
| `watchdog.detector.top_min` | 12 | Top target threshold (normal agents) |
|
|
| `watchdog.detector.top_min_ro` | 8 | Top target threshold (read-only agents) |
|
|
| `watchdog.detector.cum_min` | 15 | Slow-signal threshold |
|
|
| `watchdog.detector.cover_min` | 4.0 | Coverage-signal multiplier |
|
|
| `watchdog.detector.min_calls` | 40 | Minimum eval calls |
|
|
| `watchdog.read_only_agents` | `["explore", "librarian", "oracle"]` | Agent title keywords for reduced threshold |
|
|
| `watchdog.dashboard_base_url` | `"http://127.0.0.1:8080/admin"` | Base URL for alert links |
|
|
|
|
All seven `watchdog.detector.*` keys plus `watchdog.enabled` and
|
|
`watchdog.local_llm_enabled` (nine allowlisted keys total) are allowlisted for
|
|
live editing through the
|
|
admin portal (`POST /admin/api/config/{key}`), which writes to
|
|
`config.local.yaml` (the gitignored machine-local overlay). The changes are
|
|
validated by `RouterConfig` before reaching disk, and take effect at the next
|
|
tick — no service restart required since `detect_config_from_pydantic()` reads
|
|
cfg fresh each invocation.
|
|
|
|
## Install
|
|
|
|
The watchdog ships as two systemd user units:
|
|
|
|
- `llm-router-watchdog.service` — a `Type=oneshot` that runs
|
|
`PYTHONPATH=%h/llm-router/src .venv/bin/python -m watchdog --once`
|
|
- `llm-router-watchdog.timer` — fires 5 min after boot, then every 5 min
|
|
|
|
Installation follows the same pattern as the other deploy units. The shipped
|
|
unit files use `%h/llm-router` placeholders that need rewriting to your actual
|
|
repo path:
|
|
|
|
```bash
|
|
REPO=$(pwd)
|
|
for u in deploy/llm-router-watchdog.{service,timer}; do
|
|
sed "s|%h/llm-router|${REPO}|g" "$u" \
|
|
> ~/.config/systemd/user/"$(basename "$u")"
|
|
done
|
|
systemctl --user daemon-reload
|
|
systemctl --user enable --now llm-router-watchdog.timer
|
|
```
|
|
|
|
The service unit has **no `EnvironmentFile`** — watchdog needs no provider API
|
|
keys because it only queries the local opencode session APIs and the local
|
|
Ollama instance.
|
|
|
|
Check installation:
|
|
|
|
```bash
|
|
systemctl --user status llm-router-watchdog.timer
|
|
journalctl --user -u llm-router-watchdog.service --no-pager
|
|
```
|
|
|
|
## Run --once
|
|
|
|
For ad-hoc debugging or pre-deploy smoke test:
|
|
|
|
```bash
|
|
PYTHONPATH=src .venv/bin/python -m watchdog --once
|
|
```
|
|
|
|
The `--once` flag runs a single tick and exits 0 once the tick has run,
|
|
however many alerts it fired (the count is logged as `watchdog tick done: N
|
|
alert(s) fired`); systemd would otherwise mark the oneshot unit failed on every
|
|
real alert. It connects to the same `router.db` that the timed
|
|
service uses, acquires the filesystem lock, reads `rc-servers.json`, probes
|
|
opencode servers, and follows the full evaluate-alert-resolve pipeline.
|
|
|
|
Add `--config /path/to/config.yaml` to point at a non-default config file.
|
|
|
|
### Debug output
|
|
|
|
The `--once` run emits structured `watchdog=` log lines for every state
|
|
transition:
|
|
|
|
```
|
|
2026-01-15 14:30:01 INFO watchdog: watchdog=2026-01-15T14:30:01+00:00 dedup='opencode-loop:ses_xxxxx' state=trigger severity=warning title='opencode loop: explore'
|
|
```
|
|
|
|
The `title` field carries the agent name derived from the session title. The
|
|
`dedup` field is the root session's dedup key, matching `journalctl` traces
|
|
and admin panel rows.
|
|
|
|
## Backtest
|
|
|
|
`scripts/progress_backtest.py` replays labelled sessions through the detector
|
|
to validate threshold calibration. It reads from a fixture JSON file and runs
|
|
a sliding-window evaluation with step 5.
|
|
|
|
```bash
|
|
PYTHONPATH=src .venv/bin/python -m scripts.progress_backtest --fixture \
|
|
tests/fixtures/progress/fixture.json
|
|
```
|
|
|
|
The fixture ships in `tests/fixtures/progress/fixture.json` and contains
|
|
15 labelled sessions: 8 marked `must_flag` and 7 marked `must_not_flag`.
|
|
Expected output after a successful run:
|
|
|
|
```
|
|
# backtest complete: 8/15 flagged
|
|
```
|
|
|
|
At least 8 sessions should trigger a flag. The fixture is generated from live
|
|
opencode sessions (see [export fixture](#re-export-fixture) below) and scrubbed
|
|
to protect file contents — content is SHA-1 hashed except for the few keys the
|
|
detector actually inspects (`filePath`, `command`, `pattern`, etc.).
|
|
|
|
## Re-export fixture
|
|
|
|
The fixture generator pulls sessions from the opencode SQLite database
|
|
(`~/.local/share/opencode/opencode.db`) and writes the scrubbed test fixture:
|
|
|
|
```bash
|
|
python scripts/export_progress_fixture.py
|
|
```
|
|
|
|
Output goes to `tests/fixtures/progress/fixture.json`. The script takes a
|
|
session ID to label map (`LABELS` dict), extracts call history, resolves file
|
|
line counts (via `git show` for deleted worktree paths), and scrubs all strings
|
|
to SHA-1 prefixes, preserving only the detector's target keys.
|
|
|
|
The fixture generation script hardcodes paths to the operator's home directory
|
|
and a specific worktree commit (`WORKTREE_COMMIT = "827c408"`). When
|
|
regenerating, update those paths to match the current environment.
|
|
|
|
## Notifier pipeline
|
|
|
|
Alert events flow through `src/notifier.py` before reaching the operator. Each
|
|
alert is an `AlertEvent` dataclass mirroring PagerDuty Events v2 shape:
|
|
`dedup_key`, `severity` (info/warning/critical), `state` (trigger/escalate/
|
|
resolve), `title`, `summary`, `details`, and `source` (hardcoded to
|
|
`"6krrt-watchdog"`).
|
|
|
|
Each configured channel has its own `min_severity` gate — a critical event
|
|
passes through a channel with `min_severity=warning`, but a warning event does
|
|
not pass a channel with `min_severity=critical`. Each channel also has a rate
|
|
limit window of 300 seconds (5 minutes). Resolve events always bypass the rate
|
|
limit to ensure clear-the-alert notifications are delivered.
|
|
|
|
Currently only `type=desktop` channels are implemented; they invoke
|
|
`notify-send -u {urgency} {title} {summary}`. Unknown channel types are
|
|
logged as warnings and skipped. Missing `notify-send` (no `DISPLAY` or
|
|
`libnotify` installed) is logged at warning level — the notifier never raises.
|
|
|
|
Channel settings (`watchdog_channel_settings`) are editable at runtime through
|
|
the admin portal and stored per-channel in the database, separate from the
|
|
base config.
|
|
|
|
## Known limits
|
|
|
|
- **Under 40 calls: invisible.** `min_calls=40` means the detector does not
|
|
evaluate sessions with fewer tool calls. Short troubleshooting sessions that
|
|
loop on 10 calls pass through undetected.
|
|
|
|
- **Atlas at cover_min 4.0 exactly.** A session whose last 60 calls read a
|
|
single file exactly 4x its line count hits the coverage threshold at the
|
|
boundary. There is no margin: `>=` comparison means exactly 4.0x flags.
|
|
This matters for sessions that genuinely need to re-read a reference file
|
|
repeatedly — the coverage signal cannot be tuned per-file.
|
|
|
|
- **Stale localhost:4096 probes.** Watchdog reads `rc-servers.json` from
|
|
`~/.local/share/opencode/rc-servers.json`, which may contain stale entries
|
|
for opencode sessions that have since ended. These produce empty call lists
|
|
and are silently skipped — but they add network round-trip latency to each
|
|
tick. Only the first answering server's sessions are evaluated (the watchdog
|
|
picks `answering[0]` from the list of servers that successfully return a
|
|
session list).
|
|
|
|
- **No streaming path.** `POST /outcome` is the only ground truth signal for
|
|
streaming traffic; the watchdog has no equivalent endpoint to report streaming
|
|
session outcomes. It can only see tool-call repetition, not whether the answer
|
|
was useful.
|
|
|
|
- **Local LLM gate.** Second-opinion calls are gated by both
|
|
`local_compute.enabled` (global) and `watchdog.local_llm_enabled`
|
|
(per-subsystem). If either is false, no LLM calls are made and severity stays
|
|
as `warning` on initial trigger regardless of model quality. The cap of 3
|
|
simultaneous LLM calls means that if 5 sessions flag, 2 wait without opinion.
|
|
|
|
- **No provider calls.** Watchdog does not touch the dispatch path. It only
|
|
reads opencode session data and runs local heuristics. No provider calls, no
|
|
quota consume, no cost incurred.
|