`python -m watchdog --once` exited with the number of alerts fired. It runs as a systemd oneshot, so the first real alert would have marked llm-router-watchdog.service FAILED (and a count past 255 wraps), making a working alert look like a crash. The __main__ block becomes main(argv) and returns 0 once the tick has run; the count is logged instead. The new test fails against the old exit path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
18 KiB
Deep dive into the watchdog subsystem — opencode session loop detection. Back to README.
Watchdog is a periodic scanner that probes running opencode sessions for looping
patterns — repeated tool calls without progress. It runs as a systemd oneshot
every 5 minutes (deploy/llm-router-watchdog.timer) and, when a session flags,
writes into a local SQLite table and fires desktop notifications through the
configured channel pipeline.
The detector itself is a pure module (src/progress_detect.py) that returns a
boolean verdict plus a reason dict. The orchestrator (src/watchdog.py) reads
opencode session data over HTTP, runs the detector, and manages the alert
state machine. The notifier (src/notifier.py) routes events through
configured channels to desktop alerts (notify-send).
What it catches and doesn't
Watchdog looks for looping — sessions that make repeated tool calls against the same targets without landing meaningful changes. It detects four signal types:
-
Dup signal: a call appears more than 25% of the time in the sliding window. The call matcher first normalises
bashcommands (strips env var prefixes andcdprefixes, joins lines, collapses comments), then checks tool name + canonicalised JSON args (sorted keys) for equality. -
Top signal: a single target is called 12+ times in the window (or 8+ for read-only agents like explore/librarian/oracle). Targets merge calls on the same
(tool, basename)for file reads,(bash, first-two-words)for commands, and(tool, pattern[:60])for grep/glob. -
Slow signal: a single tool call accounts for 15+ of the session's total calls, AND no call in the entire session history has landed (tree write or git commit). This catches sessions stuck on a single read/analysis.
-
Coverage signal: the session rereads one file at 4x its line count within the window. The coverage score divides total bytes read by file length, capped at the file's actual lines. A file read 4x its length means the model is rereading without making progress.
NOT steady spend. A session that steadily calls different files on a
difficult refactor is not flagged — each call lands a unique (tool, args)
and the top target never crosses the threshold. The detector only fires when
repetition outpaces the call budget, not when a session is quietly burning
tokens across many distinct tools.
Landed calls
A call counts as landed when it:
- is an
edit,write, orpatchwith a non-emptymetadata.difffield, or - is a
bashcommand containinggit commitwithmetadata.exit == 0
Landed calls break both the slow and coverage signals. A session that reads a
file 4x, then edits it once, clears all signals immediately — the landed check
exits early and the verdict returns False.
Ancestral trees
Sessions can be parented: the opencode API returns parentID on child sessions.
Watchdog resolves root sessions and evaluates EACH session in the tree on its
own calls, NOT merged into the root. A child's flagged verdict bubbles up to the
root level, which is what the alert dedup key uses (opencode-loop:{root_session_id}).
For the landed-time check, a session sees NOT only its own landed calls but also its transitive descendants' landed calls. This prevents a parent session from being flagged just because its child landed — even though the parent may still be looping independently.
Signals and thresholds
Thresholds live in DetectConfig (src/progress_detect.py:38-53) and are
configurable under watchdog.detector.* in config/config.yaml.
| Key | Default | Meaning |
|---|---|---|
watchdog.detector.dup_min |
0.25 | Fraction: calls must repeat at this rate to flag |
watchdog.detector.top_min |
12 | Minimum calls to a single target (normal agents) |
watchdog.detector.top_min_ro |
8 | Minimum calls to a single target (read-only agents) |
watchdog.detector.cum_min |
15 | Same call must account for this many total calls |
watchdog.detector.cover_min |
4.0 | File must be reread this many times its length |
watchdog.detector.window |
60 | Sliding window in number of tool calls (not seconds) |
watchdog.detector.min_calls |
40 | Minimum total calls before evaluation triggers |
watchdog.detector.read_only_agent_keywords |
["explore", "librarian", "oracle"] |
Title keywords that make an agent read-only (lowered top_min from 12 to 8) |
Read-only agents get a lowered top threshold because agents whose job is exploration or library work are expected to call the same targets repeatedly — the floor is 8 instead of 12.
Alert lifecycle
The alert state machine is per-root-session and tracked in the
watchdog_alerts table.
Trigger (initial detection)
When the detector flags a root session's tree for the first time:
- The watchdog writes a
watchdog_alertsrow withstate='open'. - The
local_llm_enabledpath asks a local Ollama model (defaulting toverification.modelorqwen2.5-coder-router:14b) for a second opinion. It receives the agent name and the reason dict as context, and the model must respond with only "yes" or "no" — no explanation. - If the local LLM answers "yes", severity is set to
critical. If it answers "no" or times out (returnsNone), severity is set towarning. - The notifier dispatches an
AlertEventthrough every matching channel. - The
flagged_tickscounter starts at 1.
A cap of 3 simultaneous LLM second-opinion calls (_MAX_LLM = 3) prevents a
burst of flagged sessions from flooding the local model.
Escalate (persistent looping)
On every subsequent tick where the session is still flagged:
flagged_ticksis incremented.- If
flagged_ticks >= 3(15 minutes at 5-minute ticks) OR the LLM answers "yes" again, the alert escalates toseverity='critical'and the escalation event is dispatched. - If the alert is already critical, ticking continues but no duplicate escalation fires.
Resolve (session is gone or fixed)
A root session alert resolves in two conditions:
- Session disappeared: the root session ID is no longer returned by the
opencode
/sessionAPI. The watchdog writesstate='resolved',resolved_at=<now>, andseverity='info'. - Session fixed: the root session's tree no longer has any flagged sessions in the detector's verdict. The alert is resolved the same way and a resolve event is dispatched.
Resolve events always bypass the rate limit — they must clear the previous notification so the operator knows the alert is over.
No-opencode
When no rc-servers.json or no opencode server answers, the tick writes an
outcome='no_opencode' record and returns immediately with zero alerts. This
uses the filesystem lock (router.db/.watchdog.lock) so two watchdog instances
cannot run simultaneously.
Admin surfaces
Watchdog data surfaces through three admin pages, each serving a different operational question.
Loops panel — index.html#loops
On the main dashboard (admin/frontend/index.html), the Loops card
(id="loops") shows all currently open alerts:
- Severity badge: the alert's current severity (
criticalin red,warningin orange-yellow,infoin grey) - Dedup key: the raw
opencode-loop:{root_session_id}string, so the operator can cross-reference withjournalctlorrouter.db - Opened at:
opened_atISO timestamp from the first trigger - State:
openorresolved; the panel filters toresolved_at IS NULL
The panel is client-side only — it polls /api/watchdog/loops which queries
watchdog_alerts for all unresolved rows.
Watchdog card — controls.html
The Controls page (admin/frontend/controls.html) includes a dedicated
Watchdog card with operational status:
- Last tick: timestamp of the most recent
watchdog_ticksrow - Sessions seen: how many opencode sessions were enumerated on that tick
- Flagged verdicts: count of flagged + total from the latest tick's
watchdog_verdictsrows - Open alerts: count from
watchdog_alerts WHERE resolved_at IS NULL - Notification channels: list of configured channels with toggle controls
(enabled/min_severity), editable via
POST /api/watchdog/channels - Refresh button: reloads the watchdog status from the API
- Send test alert button: fires a test
AlertEventthrough enabled channels
The status endpoints are:
GET /api/watchdog/status— last tick, verdict counts, open alert countGET /api/watchdog/channels— per-channel settings fromwatchdog_channel_settingsPOST /api/watchdog/channels— upsert channel settings (enabled, min_severity)POST /api/watchdog/test-alert— deliver a test event (severity + optional channel_name filter) to a live Notifier instance
Per-model stall rollup — models.html
The watchdog_verdicts table has model_id/provider columns and indexes on
them, but no per-model stall rollup, no Block button beside evidence, and no
Blocked list are built yet. blocked is only a value in the Models override
dropdown (see docs/admin-portal.md).
Database schema
Four tables live in router.db, created idempotently by both
config/schema.sql and watchdog_store.py:
watchdog_ticks -- one row per watchdog tick
watchdog_verdicts -- per-session verdict at each tick (joined to ticks)
watchdog_alerts -- alert lifecycle state machine (dedup_key PK)
watchdog_channel_settings -- per-channel toggle + severity gate
Indexes on watchdog_verdicts cover model_id, created_at, flagged, and
session_root — the columns the admin endpoint queries most frequently.
The ticks table carries ticked_at, sessions_seen, and outcome
('ok', 'flagged', or 'no_opencode'). The verdicts table joins to ticks
via tick_id and carries the full reason dict fields: dup, top,
top_what, landed, slow, coverage, plus calls_since_landed,
cost_since_landed_usd, and the LLM second-opinion answer.
Alerts are keyed by dedup_key (deduplicated per root session) and carry
opened_at, last_fired_at, and resolved_at. The transition logic lives in
the _fire_alert() helper: trigger inserts only if no open row exists,
escalate increments flagged_ticks and bumps severity, and resolve writes
resolved_at and downgrades to severity='info'.
Knobs
All watchdog configuration lives under watchdog: in config/config.yaml:
| Key | Default | Required |
|---|---|---|
watchdog.enabled |
true |
Gates the entire subsystem |
watchdog.local_llm_enabled |
true |
Whether to call Ollama for second opinions |
watchdog.model |
null (uses verification.model) |
Local model name |
watchdog.detector.window |
60 | Sliding window size in calls |
watchdog.detector.dup_min |
0.25 | Dup threshold (fraction) |
watchdog.detector.top_min |
12 | Top target threshold (normal agents) |
watchdog.detector.top_min_ro |
8 | Top target threshold (read-only agents) |
watchdog.detector.cum_min |
15 | Slow-signal threshold |
watchdog.detector.cover_min |
4.0 | Coverage-signal multiplier |
watchdog.detector.min_calls |
40 | Minimum eval calls |
watchdog.read_only_agents |
["explore", "librarian", "oracle"] |
Agent title keywords for reduced threshold |
watchdog.dashboard_base_url |
"http://127.0.0.1:8080/admin" |
Base URL for alert links |
All seven watchdog.detector.* keys plus watchdog.enabled and
watchdog.local_llm_enabled (nine allowlisted keys total) are allowlisted for
live editing through the
admin portal (POST /admin/api/config/{key}), which writes to
config.local.yaml (the gitignored machine-local overlay). The changes are
validated by RouterConfig before reaching disk, and take effect at the next
tick — no service restart required since detect_config_from_pydantic() reads
cfg fresh each invocation.
Install
The watchdog ships as two systemd user units:
llm-router-watchdog.service— aType=oneshotthat runsPYTHONPATH=%h/llm-router/src .venv/bin/python -m watchdog --oncellm-router-watchdog.timer— fires 5 min after boot, then every 5 min
Installation follows the same pattern as the other deploy units. The shipped
unit files use %h/llm-router placeholders that need rewriting to your actual
repo path:
REPO=$(pwd)
for u in deploy/llm-router-watchdog.{service,timer}; do
sed "s|%h/llm-router|${REPO}|g" "$u" \
> ~/.config/systemd/user/"$(basename "$u")"
done
systemctl --user daemon-reload
systemctl --user enable --now llm-router-watchdog.timer
The service unit has no EnvironmentFile — watchdog needs no provider API
keys because it only queries the local opencode session APIs and the local
Ollama instance.
Check installation:
systemctl --user status llm-router-watchdog.timer
journalctl --user -u llm-router-watchdog.service --no-pager
Run --once
For ad-hoc debugging or pre-deploy smoke test:
PYTHONPATH=src .venv/bin/python -m watchdog --once
The --once flag runs a single tick and exits 0 once the tick has run,
however many alerts it fired (the count is logged as watchdog tick done: N alert(s) fired); systemd would otherwise mark the oneshot unit failed on every
real alert. It connects to the same router.db that the timed
service uses, acquires the filesystem lock, reads rc-servers.json, probes
opencode servers, and follows the full evaluate-alert-resolve pipeline.
Add --config /path/to/config.yaml to point at a non-default config file.
Debug output
The --once run emits structured watchdog= log lines for every state
transition:
2026-01-15 14:30:01 INFO watchdog: watchdog=2026-01-15T14:30:01+00:00 dedup='opencode-loop:ses_xxxxx' state=trigger severity=warning title='opencode loop: explore'
The title field carries the agent name derived from the session title. The
dedup field is the root session's dedup key, matching journalctl traces
and admin panel rows.
Backtest
scripts/progress_backtest.py replays labelled sessions through the detector
to validate threshold calibration. It reads from a fixture JSON file and runs
a sliding-window evaluation with step 5.
PYTHONPATH=src .venv/bin/python -m scripts.progress_backtest --fixture \
tests/fixtures/progress/fixture.json
The fixture ships in tests/fixtures/progress/fixture.json and contains
15 labelled sessions: 8 marked must_flag and 7 marked must_not_flag.
Expected output after a successful run:
# backtest complete: 8/15 flagged
At least 8 sessions should trigger a flag. The fixture is generated from live
opencode sessions (see export fixture below) and scrubbed
to protect file contents — content is SHA-1 hashed except for the few keys the
detector actually inspects (filePath, command, pattern, etc.).
Re-export fixture
The fixture generator pulls sessions from the opencode SQLite database
(~/.local/share/opencode/opencode.db) and writes the scrubbed test fixture:
python scripts/export_progress_fixture.py
Output goes to tests/fixtures/progress/fixture.json. The script takes a
session ID to label map (LABELS dict), extracts call history, resolves file
line counts (via git show for deleted worktree paths), and scrubs all strings
to SHA-1 prefixes, preserving only the detector's target keys.
The fixture generation script hardcodes paths to the operator's home directory
and a specific worktree commit (WORKTREE_COMMIT = "827c408"). When
regenerating, update those paths to match the current environment.
Notifier pipeline
Alert events flow through src/notifier.py before reaching the operator. Each
alert is an AlertEvent dataclass mirroring PagerDuty Events v2 shape:
dedup_key, severity (info/warning/critical), state (trigger/escalate/
resolve), title, summary, details, and source (hardcoded to
"6krrt-watchdog").
Each configured channel has its own min_severity gate — a critical event
passes through a channel with min_severity=warning, but a warning event does
not pass a channel with min_severity=critical. Each channel also has a rate
limit window of 300 seconds (5 minutes). Resolve events always bypass the rate
limit to ensure clear-the-alert notifications are delivered.
Currently only type=desktop channels are implemented; they invoke
notify-send -u {urgency} {title} {summary}. Unknown channel types are
logged as warnings and skipped. Missing notify-send (no DISPLAY or
libnotify installed) is logged at warning level — the notifier never raises.
Channel settings (watchdog_channel_settings) are editable at runtime through
the admin portal and stored per-channel in the database, separate from the
base config.
Known limits
-
Under 40 calls: invisible.
min_calls=40means the detector does not evaluate sessions with fewer tool calls. Short troubleshooting sessions that loop on 10 calls pass through undetected. -
Atlas at cover_min 4.0 exactly. A session whose last 60 calls read a single file exactly 4x its line count hits the coverage threshold at the boundary. There is no margin:
>=comparison means exactly 4.0x flags. This matters for sessions that genuinely need to re-read a reference file repeatedly — the coverage signal cannot be tuned per-file. -
Stale localhost:4096 probes. Watchdog reads
rc-servers.jsonfrom~/.local/share/opencode/rc-servers.json, which may contain stale entries for opencode sessions that have since ended. These produce empty call lists and are silently skipped — but they add network round-trip latency to each tick. Only the first answering server's sessions are evaluated (the watchdog picksanswering[0]from the list of servers that successfully return a session list). -
No streaming path.
POST /outcomeis the only ground truth signal for streaming traffic; the watchdog has no equivalent endpoint to report streaming session outcomes. It can only see tool-call repetition, not whether the answer was useful. -
Local LLM gate. Second-opinion calls are gated by both
local_compute.enabled(global) andwatchdog.local_llm_enabled(per-subsystem). If either is false, no LLM calls are made and severity stays aswarningon initial trigger regardless of model quality. The cap of 3 simultaneous LLM calls means that if 5 sessions flag, 2 wait without opinion. -
No provider calls. Watchdog does not touch the dispatch path. It only reads opencode session data and runs local heuristics. No provider calls, no quota consume, no cost incurred.