Files
6krrt/docs/watchdog.md
adlee-was-taken 4d2bbf0253 fix(watchdog): exit 0 after a tick that ran, whatever it alerted
`python -m watchdog --once` exited with the number of alerts fired. It runs
as a systemd oneshot, so the first real alert would have marked
llm-router-watchdog.service FAILED (and a count past 255 wraps), making a
working alert look like a crash. The __main__ block becomes main(argv) and
returns 0 once the tick has run; the count is logged instead. The new test
fails against the old exit path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-26 21:44:06 -04:00

18 KiB

Deep dive into the watchdog subsystem — opencode session loop detection. Back to README.

Watchdog is a periodic scanner that probes running opencode sessions for looping patterns — repeated tool calls without progress. It runs as a systemd oneshot every 5 minutes (deploy/llm-router-watchdog.timer) and, when a session flags, writes into a local SQLite table and fires desktop notifications through the configured channel pipeline.

The detector itself is a pure module (src/progress_detect.py) that returns a boolean verdict plus a reason dict. The orchestrator (src/watchdog.py) reads opencode session data over HTTP, runs the detector, and manages the alert state machine. The notifier (src/notifier.py) routes events through configured channels to desktop alerts (notify-send).

What it catches and doesn't

Watchdog looks for looping — sessions that make repeated tool calls against the same targets without landing meaningful changes. It detects four signal types:

  • Dup signal: a call appears more than 25% of the time in the sliding window. The call matcher first normalises bash commands (strips env var prefixes and cd prefixes, joins lines, collapses comments), then checks tool name + canonicalised JSON args (sorted keys) for equality.

  • Top signal: a single target is called 12+ times in the window (or 8+ for read-only agents like explore/librarian/oracle). Targets merge calls on the same (tool, basename) for file reads, (bash, first-two-words) for commands, and (tool, pattern[:60]) for grep/glob.

  • Slow signal: a single tool call accounts for 15+ of the session's total calls, AND no call in the entire session history has landed (tree write or git commit). This catches sessions stuck on a single read/analysis.

  • Coverage signal: the session rereads one file at 4x its line count within the window. The coverage score divides total bytes read by file length, capped at the file's actual lines. A file read 4x its length means the model is rereading without making progress.

NOT steady spend. A session that steadily calls different files on a difficult refactor is not flagged — each call lands a unique (tool, args) and the top target never crosses the threshold. The detector only fires when repetition outpaces the call budget, not when a session is quietly burning tokens across many distinct tools.

Landed calls

A call counts as landed when it:

  • is an edit, write, or patch with a non-empty metadata.diff field, or
  • is a bash command containing git commit with metadata.exit == 0

Landed calls break both the slow and coverage signals. A session that reads a file 4x, then edits it once, clears all signals immediately — the landed check exits early and the verdict returns False.

Ancestral trees

Sessions can be parented: the opencode API returns parentID on child sessions. Watchdog resolves root sessions and evaluates EACH session in the tree on its own calls, NOT merged into the root. A child's flagged verdict bubbles up to the root level, which is what the alert dedup key uses (opencode-loop:{root_session_id}).

For the landed-time check, a session sees NOT only its own landed calls but also its transitive descendants' landed calls. This prevents a parent session from being flagged just because its child landed — even though the parent may still be looping independently.

Signals and thresholds

Thresholds live in DetectConfig (src/progress_detect.py:38-53) and are configurable under watchdog.detector.* in config/config.yaml.

Key Default Meaning
watchdog.detector.dup_min 0.25 Fraction: calls must repeat at this rate to flag
watchdog.detector.top_min 12 Minimum calls to a single target (normal agents)
watchdog.detector.top_min_ro 8 Minimum calls to a single target (read-only agents)
watchdog.detector.cum_min 15 Same call must account for this many total calls
watchdog.detector.cover_min 4.0 File must be reread this many times its length
watchdog.detector.window 60 Sliding window in number of tool calls (not seconds)
watchdog.detector.min_calls 40 Minimum total calls before evaluation triggers
watchdog.detector.read_only_agent_keywords ["explore", "librarian", "oracle"] Title keywords that make an agent read-only (lowered top_min from 12 to 8)

Read-only agents get a lowered top threshold because agents whose job is exploration or library work are expected to call the same targets repeatedly — the floor is 8 instead of 12.

Alert lifecycle

The alert state machine is per-root-session and tracked in the watchdog_alerts table.

Trigger (initial detection)

When the detector flags a root session's tree for the first time:

  1. The watchdog writes a watchdog_alerts row with state='open'.
  2. The local_llm_enabled path asks a local Ollama model (defaulting to verification.model or qwen2.5-coder-router:14b) for a second opinion. It receives the agent name and the reason dict as context, and the model must respond with only "yes" or "no" — no explanation.
  3. If the local LLM answers "yes", severity is set to critical. If it answers "no" or times out (returns None), severity is set to warning.
  4. The notifier dispatches an AlertEvent through every matching channel.
  5. The flagged_ticks counter starts at 1.

A cap of 3 simultaneous LLM second-opinion calls (_MAX_LLM = 3) prevents a burst of flagged sessions from flooding the local model.

Escalate (persistent looping)

On every subsequent tick where the session is still flagged:

  1. flagged_ticks is incremented.
  2. If flagged_ticks >= 3 (15 minutes at 5-minute ticks) OR the LLM answers "yes" again, the alert escalates to severity='critical' and the escalation event is dispatched.
  3. If the alert is already critical, ticking continues but no duplicate escalation fires.

Resolve (session is gone or fixed)

A root session alert resolves in two conditions:

  • Session disappeared: the root session ID is no longer returned by the opencode /session API. The watchdog writes state='resolved', resolved_at=<now>, and severity='info'.
  • Session fixed: the root session's tree no longer has any flagged sessions in the detector's verdict. The alert is resolved the same way and a resolve event is dispatched.

Resolve events always bypass the rate limit — they must clear the previous notification so the operator knows the alert is over.

No-opencode

When no rc-servers.json or no opencode server answers, the tick writes an outcome='no_opencode' record and returns immediately with zero alerts. This uses the filesystem lock (router.db/.watchdog.lock) so two watchdog instances cannot run simultaneously.

Admin surfaces

Watchdog data surfaces through three admin pages, each serving a different operational question.

Loops panel — index.html#loops

On the main dashboard (admin/frontend/index.html), the Loops card (id="loops") shows all currently open alerts:

  • Severity badge: the alert's current severity (critical in red, warning in orange-yellow, info in grey)
  • Dedup key: the raw opencode-loop:{root_session_id} string, so the operator can cross-reference with journalctl or router.db
  • Opened at: opened_at ISO timestamp from the first trigger
  • State: open or resolved; the panel filters to resolved_at IS NULL

The panel is client-side only — it polls /api/watchdog/loops which queries watchdog_alerts for all unresolved rows.

Watchdog card — controls.html

The Controls page (admin/frontend/controls.html) includes a dedicated Watchdog card with operational status:

  • Last tick: timestamp of the most recent watchdog_ticks row
  • Sessions seen: how many opencode sessions were enumerated on that tick
  • Flagged verdicts: count of flagged + total from the latest tick's watchdog_verdicts rows
  • Open alerts: count from watchdog_alerts WHERE resolved_at IS NULL
  • Notification channels: list of configured channels with toggle controls (enabled/min_severity), editable via POST /api/watchdog/channels
  • Refresh button: reloads the watchdog status from the API
  • Send test alert button: fires a test AlertEvent through enabled channels

The status endpoints are:

  • GET /api/watchdog/status — last tick, verdict counts, open alert count
  • GET /api/watchdog/channels — per-channel settings from watchdog_channel_settings
  • POST /api/watchdog/channels — upsert channel settings (enabled, min_severity)
  • POST /api/watchdog/test-alert — deliver a test event (severity + optional channel_name filter) to a live Notifier instance

Per-model stall rollup — models.html

The watchdog_verdicts table has model_id/provider columns and indexes on them, but no per-model stall rollup, no Block button beside evidence, and no Blocked list are built yet. blocked is only a value in the Models override dropdown (see docs/admin-portal.md).

Database schema

Four tables live in router.db, created idempotently by both config/schema.sql and watchdog_store.py:

watchdog_ticks           -- one row per watchdog tick
watchdog_verdicts        -- per-session verdict at each tick (joined to ticks)
watchdog_alerts          -- alert lifecycle state machine (dedup_key PK)
watchdog_channel_settings -- per-channel toggle + severity gate

Indexes on watchdog_verdicts cover model_id, created_at, flagged, and session_root — the columns the admin endpoint queries most frequently.

The ticks table carries ticked_at, sessions_seen, and outcome ('ok', 'flagged', or 'no_opencode'). The verdicts table joins to ticks via tick_id and carries the full reason dict fields: dup, top, top_what, landed, slow, coverage, plus calls_since_landed, cost_since_landed_usd, and the LLM second-opinion answer.

Alerts are keyed by dedup_key (deduplicated per root session) and carry opened_at, last_fired_at, and resolved_at. The transition logic lives in the _fire_alert() helper: trigger inserts only if no open row exists, escalate increments flagged_ticks and bumps severity, and resolve writes resolved_at and downgrades to severity='info'.

Knobs

All watchdog configuration lives under watchdog: in config/config.yaml:

Key Default Required
watchdog.enabled true Gates the entire subsystem
watchdog.local_llm_enabled true Whether to call Ollama for second opinions
watchdog.model null (uses verification.model) Local model name
watchdog.detector.window 60 Sliding window size in calls
watchdog.detector.dup_min 0.25 Dup threshold (fraction)
watchdog.detector.top_min 12 Top target threshold (normal agents)
watchdog.detector.top_min_ro 8 Top target threshold (read-only agents)
watchdog.detector.cum_min 15 Slow-signal threshold
watchdog.detector.cover_min 4.0 Coverage-signal multiplier
watchdog.detector.min_calls 40 Minimum eval calls
watchdog.read_only_agents ["explore", "librarian", "oracle"] Agent title keywords for reduced threshold
watchdog.dashboard_base_url "http://127.0.0.1:8080/admin" Base URL for alert links

All seven watchdog.detector.* keys plus watchdog.enabled and watchdog.local_llm_enabled (nine allowlisted keys total) are allowlisted for live editing through the admin portal (POST /admin/api/config/{key}), which writes to config.local.yaml (the gitignored machine-local overlay). The changes are validated by RouterConfig before reaching disk, and take effect at the next tick — no service restart required since detect_config_from_pydantic() reads cfg fresh each invocation.

Install

The watchdog ships as two systemd user units:

  • llm-router-watchdog.service — a Type=oneshot that runs PYTHONPATH=%h/llm-router/src .venv/bin/python -m watchdog --once
  • llm-router-watchdog.timer — fires 5 min after boot, then every 5 min

Installation follows the same pattern as the other deploy units. The shipped unit files use %h/llm-router placeholders that need rewriting to your actual repo path:

REPO=$(pwd)
for u in deploy/llm-router-watchdog.{service,timer}; do
  sed "s|%h/llm-router|${REPO}|g" "$u" \
    > ~/.config/systemd/user/"$(basename "$u")"
done
systemctl --user daemon-reload
systemctl --user enable --now llm-router-watchdog.timer

The service unit has no EnvironmentFile — watchdog needs no provider API keys because it only queries the local opencode session APIs and the local Ollama instance.

Check installation:

systemctl --user status llm-router-watchdog.timer
journalctl --user -u llm-router-watchdog.service --no-pager

Run --once

For ad-hoc debugging or pre-deploy smoke test:

PYTHONPATH=src .venv/bin/python -m watchdog --once

The --once flag runs a single tick and exits 0 once the tick has run, however many alerts it fired (the count is logged as watchdog tick done: N alert(s) fired); systemd would otherwise mark the oneshot unit failed on every real alert. It connects to the same router.db that the timed service uses, acquires the filesystem lock, reads rc-servers.json, probes opencode servers, and follows the full evaluate-alert-resolve pipeline.

Add --config /path/to/config.yaml to point at a non-default config file.

Debug output

The --once run emits structured watchdog= log lines for every state transition:

2026-01-15 14:30:01 INFO watchdog: watchdog=2026-01-15T14:30:01+00:00 dedup='opencode-loop:ses_xxxxx' state=trigger severity=warning title='opencode loop: explore'

The title field carries the agent name derived from the session title. The dedup field is the root session's dedup key, matching journalctl traces and admin panel rows.

Backtest

scripts/progress_backtest.py replays labelled sessions through the detector to validate threshold calibration. It reads from a fixture JSON file and runs a sliding-window evaluation with step 5.

PYTHONPATH=src .venv/bin/python -m scripts.progress_backtest --fixture \
  tests/fixtures/progress/fixture.json

The fixture ships in tests/fixtures/progress/fixture.json and contains 15 labelled sessions: 8 marked must_flag and 7 marked must_not_flag. Expected output after a successful run:

# backtest complete: 8/15 flagged

At least 8 sessions should trigger a flag. The fixture is generated from live opencode sessions (see export fixture below) and scrubbed to protect file contents — content is SHA-1 hashed except for the few keys the detector actually inspects (filePath, command, pattern, etc.).

Re-export fixture

The fixture generator pulls sessions from the opencode SQLite database (~/.local/share/opencode/opencode.db) and writes the scrubbed test fixture:

python scripts/export_progress_fixture.py

Output goes to tests/fixtures/progress/fixture.json. The script takes a session ID to label map (LABELS dict), extracts call history, resolves file line counts (via git show for deleted worktree paths), and scrubs all strings to SHA-1 prefixes, preserving only the detector's target keys.

The fixture generation script hardcodes paths to the operator's home directory and a specific worktree commit (WORKTREE_COMMIT = "827c408"). When regenerating, update those paths to match the current environment.

Notifier pipeline

Alert events flow through src/notifier.py before reaching the operator. Each alert is an AlertEvent dataclass mirroring PagerDuty Events v2 shape: dedup_key, severity (info/warning/critical), state (trigger/escalate/ resolve), title, summary, details, and source (hardcoded to "6krrt-watchdog").

Each configured channel has its own min_severity gate — a critical event passes through a channel with min_severity=warning, but a warning event does not pass a channel with min_severity=critical. Each channel also has a rate limit window of 300 seconds (5 minutes). Resolve events always bypass the rate limit to ensure clear-the-alert notifications are delivered.

Currently only type=desktop channels are implemented; they invoke notify-send -u {urgency} {title} {summary}. Unknown channel types are logged as warnings and skipped. Missing notify-send (no DISPLAY or libnotify installed) is logged at warning level — the notifier never raises.

Channel settings (watchdog_channel_settings) are editable at runtime through the admin portal and stored per-channel in the database, separate from the base config.

Known limits

  • Under 40 calls: invisible. min_calls=40 means the detector does not evaluate sessions with fewer tool calls. Short troubleshooting sessions that loop on 10 calls pass through undetected.

  • Atlas at cover_min 4.0 exactly. A session whose last 60 calls read a single file exactly 4x its line count hits the coverage threshold at the boundary. There is no margin: >= comparison means exactly 4.0x flags. This matters for sessions that genuinely need to re-read a reference file repeatedly — the coverage signal cannot be tuned per-file.

  • Stale localhost:4096 probes. Watchdog reads rc-servers.json from ~/.local/share/opencode/rc-servers.json, which may contain stale entries for opencode sessions that have since ended. These produce empty call lists and are silently skipped — but they add network round-trip latency to each tick. Only the first answering server's sessions are evaluated (the watchdog picks answering[0] from the list of servers that successfully return a session list).

  • No streaming path. POST /outcome is the only ground truth signal for streaming traffic; the watchdog has no equivalent endpoint to report streaming session outcomes. It can only see tool-call repetition, not whether the answer was useful.

  • Local LLM gate. Second-opinion calls are gated by both local_compute.enabled (global) and watchdog.local_llm_enabled (per-subsystem). If either is false, no LLM calls are made and severity stays as warning on initial trigger regardless of model quality. The cap of 3 simultaneous LLM calls means that if 5 sessions flag, 2 wait without opinion.

  • No provider calls. Watchdog does not touch the dispatch path. It only reads opencode session data and runs local heuristics. No provider calls, no quota consume, no cost incurred.