docs: Lift A docs pass, plus watchdog exits 0 after a tick #103
59
CLAUDE.md
59
CLAUDE.md
@@ -8,7 +8,7 @@ next steps, and is the one to trust on what is currently true.
|
|||||||
|
|
||||||
## NORTH STAR GUIDELINES
|
## NORTH STAR GUIDELINES
|
||||||
|
|
||||||
Three rules that outrank local cleverness. Each exists because it was broken
|
Four rules that outrank local cleverness. Each exists because it was broken
|
||||||
first and the breakage was expensive to find. When a change conflicts with one
|
first and the breakage was expensive to find. When a change conflicts with one
|
||||||
of these, the change is wrong — not the rule.
|
of these, the change is wrong — not the rule.
|
||||||
|
|
||||||
@@ -87,6 +87,36 @@ before measuring anything, and open it read-only:
|
|||||||
`sqlite3.connect("file:...?mode=ro", uri=True)` — the `sqlite3` CLI here does
|
`sqlite3.connect("file:...?mode=ro", uri=True)` — the `sqlite3` CLI here does
|
||||||
not accept `-uri`.
|
not accept `-uri`.
|
||||||
|
|
||||||
|
### 4. Waste is surfaced in the admin portal, and stopping it is one click
|
||||||
|
|
||||||
|
When the router can see money being wasted, it shows the operator where they
|
||||||
|
already look, with the evidence and the lever to stop it side by side. A
|
||||||
|
detector that only writes a log line, or a fix that needs a config edit, a
|
||||||
|
restart or a long table scan, has not met this rule.
|
||||||
|
|
||||||
|
"Waste" here means **spend with no concrete change landing**: an agent session
|
||||||
|
looping, re-reading, or retrying the same failure, and a model that keeps
|
||||||
|
producing such sessions. It does **not** mean steady spend. A healthy agent run
|
||||||
|
can burn for hours, and a spend-rate alarm cannot tell the two apart.
|
||||||
|
|
||||||
|
In practice:
|
||||||
|
- **Stalled sessions are visible** in the portal with their evidence: turns and
|
||||||
|
$ since the last landed change, and the top repeated target. They also reach
|
||||||
|
the operator when nobody is watching (desktop alert).
|
||||||
|
- **A model that keeps producing them is visible** as a per-model rollup, with
|
||||||
|
**Block model** beside the evidence. The block records its reason, shows in a
|
||||||
|
Blocked list, and is one click to undo.
|
||||||
|
- **Automatic responses come after the visible one,** never instead of it, and
|
||||||
|
every automatic action shows up in the same place.
|
||||||
|
|
||||||
|
**Why:** incident #8 (`docs/incidents.md`). Agent sessions looped for hours on
|
||||||
|
2026-09-25/26: one worker read the same file 61 times, a planner re-read a spec
|
||||||
|
68 times its length, and a model confabulated truncation that was not there.
|
||||||
|
Every existing check stayed green. It was caught only by a human, or a Claude
|
||||||
|
session, reading opencode's session store by hand. Pulling the model took a trip
|
||||||
|
through a dropdown whose vocabulary is catalog `deprecated`. The detection
|
||||||
|
existed nowhere, and the lever existed only for someone who already knew.
|
||||||
|
|
||||||
## What this is
|
## What this is
|
||||||
|
|
||||||
A router that uses a local model (served via Ollama) to classify incoming
|
A router that uses a local model (served via Ollama) to classify incoming
|
||||||
@@ -319,7 +349,12 @@ rather than from months of history.
|
|||||||
- `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md).
|
- `local_encoder.py` — zero-shot category classification via a non-generative encoder, backing `classifier.mode: local_encoder`. `transformers`/`torch` imported lazily; a deployment that never selects the mode needs neither installed. [local-models](docs/local-models.md).
|
||||||
- `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below.
|
- `provider_model_allowlist` table — DB gate for opt-in providers. Models are not ingested unless explicitly allowlisted, and rows that leave the allowlist become `deprecated` on the next poll. Used by OpenRouter; Neuralwatt is unaffected. See the #45 OpenRouter opt-in allowlist section below.
|
||||||
- `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted.
|
- `config.py` / `DispatchProvider.require_allowlist` — Pydantic flag that switches a provider from ingest-everything to allowlist-gated. A missing allowlist is treated as empty: every active row for that provider is deprecated and no new rows are upserted.
|
||||||
- `tests/` — 1155 tests across 40+ files, offline, verified on Python 3.10 and 3.14. [README](README.md).
|
- `progress_detect.py` — loop-detection signals over a window of per-session probe calls: duplicate-bulk (`dup_min`), top-similarity (`top_min`/`top_min_ro`), slow-progress (`cum_min`) and coverage (`cover_min`) heuristics, gated on `min_calls`. See [watchdog](docs/watchdog.md).
|
||||||
|
- `watchdog.py` — the per-session watchdog loop: every ~5 minutes it judges each session with a tool call since the last tick, on its full history, and writes a quiet `no_opencode` tick when opencode is not running. See [watchdog](docs/watchdog.md).
|
||||||
|
- `notifier.py` — alert fan-out: desktop/`notify-send` plus per-channel `min_severity` and a rate limit. See [watchdog](docs/watchdog.md).
|
||||||
|
- `watchdog_store.py` — `watchdog_ticks`, `watchdog_verdicts`, `watchdog_alerts`, `watchdog_channel_settings` (four tables + indexes). See [watchdog](docs/watchdog.md).
|
||||||
|
- `router-link.js` — the opencode plugin; exported as a factory with `parentCache` as a property, because opencode 1.18.x rejects the whole plugin when any export is not a function. See [watchdog](docs/watchdog.md).
|
||||||
|
- `tests/` — 2414 tests across 108 files, offline, verified on Python 3.10 and 3.14. [README](README.md).
|
||||||
|
|
||||||
## #45 — OpenRouter is an opt-in allowlist provider
|
## #45 — OpenRouter is an opt-in allowlist provider
|
||||||
|
|
||||||
@@ -1261,10 +1296,14 @@ be on.
|
|||||||
|
|
||||||
### When the router goes unreachable, start at docs/incidents.md
|
### When the router goes unreachable, start at docs/incidents.md
|
||||||
|
|
||||||
Four incidents so far, all sharing one shape: a change that looked local to the
|
Eight incidents so far, nearly all sharing one shape: a change that looked local
|
||||||
router silently degraded the agent depending on it, and none announced itself as
|
to the router silently degraded the agent depending on it, and none announced
|
||||||
a router problem. **`docs/incidents.md` carries the full write-ups plus a
|
itself as a router problem. **`docs/incidents.md` carries the full write-ups plus
|
||||||
symptom -> one-line-check table**; read it rather than re-deriving a diagnosis.
|
a symptom -> one-line-check table**; read it rather than re-deriving a diagnosis.
|
||||||
|
#8 is the exception worth knowing before an unattended agent run: the router
|
||||||
|
worked perfectly while agent workers looped for hours with no progress, and no
|
||||||
|
check noticed, because every check watched spend or availability rather than
|
||||||
|
whether changes landed (`plans/no-progress-detection.md`).
|
||||||
|
|
||||||
Two conventions from those incidents that bind every session, and so stay here:
|
Two conventions from those incidents that bind every session, and so stay here:
|
||||||
|
|
||||||
@@ -1283,6 +1322,14 @@ Recovery for an unreachable-but-`active` service is
|
|||||||
`systemctl --user restart llm-router.service` -- a hung process was never in a
|
`systemctl --user restart llm-router.service` -- a hung process was never in a
|
||||||
tracked stop job, so this issues a fresh cycle systemd does enforce a timeout on.
|
tracked stop job, so this issues a fresh cycle systemd does enforce a timeout on.
|
||||||
|
|
||||||
|
The watchdog runs as its own systemd **timer** (`llm-router-watchdog.timer`),
|
||||||
|
separate from the dispatcher, so a stuck router cannot silence the thing meant
|
||||||
|
to notice it is stuck. Install and enable it with `systemctl --user enable
|
||||||
|
--now llm-router-watchdog.timer` (the timer's unit file ships pointing at a
|
||||||
|
placeholder home path and must be `sed`-repointed to the real one first), or
|
||||||
|
run it once by hand with the `--once` flag. See [watchdog](docs/watchdog.md)
|
||||||
|
for the signals it fires on, the alert lifecycle, and the known limits.
|
||||||
|
|
||||||
## Pointing a coding agent at it
|
## Pointing a coding agent at it
|
||||||
|
|
||||||
The `/v1` endpoints are OpenAI-compatible, so any normal client works —
|
The `/v1` endpoints are OpenAI-compatible, so any normal client works —
|
||||||
|
|||||||
@@ -35,11 +35,30 @@ for the full model.
|
|||||||
<img src="../assets/screenshots/dashboard.png" alt="Admin dashboard" width="820">
|
<img src="../assets/screenshots/dashboard.png" alt="Admin dashboard" width="820">
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
|
**Loops** — the dashboard carries a Loops card (`admin/frontend/index.html#loops`)
|
||||||
|
showing the watchdog's open alerts, one row per alert. It pulls from
|
||||||
|
`GET /admin/api/watchdog/loops` (`watchdog_alerts` rows where `resolved_at` is
|
||||||
|
null) via `loadLoops()`. Each row shows exactly three things: a severity badge
|
||||||
|
(`critical` / `warning` / else `info`), the raw dedup key
|
||||||
|
(`opencode-loop:<root session id>`), and `opened_at`. That is all the panel
|
||||||
|
shows today — it does **not** show the agent, model, top repeated target, calls,
|
||||||
|
or $ since the last landed change; those live in `watchdog_verdicts` and have
|
||||||
|
not shipped to this panel yet.
|
||||||
|
|
||||||
**Models** — `GET /admin/models` is the model-availability table. Each row shows
|
**Models** — `GET /admin/models` is the model-availability table. Each row shows
|
||||||
the serving class, tier, status, and an override dropdown (`active` /
|
the serving class, tier, status, and an override dropdown (`active` /
|
||||||
`deprecated` / `stale`) that writes through to the routing hard filters. By
|
`blocked` / `deprecated` / `stale`) that writes through to the routing hard
|
||||||
default the table hides deprecated and stale rows; enable the **Show deprecated
|
filters. By default the table hides deprecated and stale rows; enable the
|
||||||
/ stale** toggle to include them.
|
**Show deprecated / stale** toggle to include them.
|
||||||
|
|
||||||
|
`blocked` is a fourth value in that same per-model override dropdown, distinct
|
||||||
|
from the catalog meaning of `deprecated`. It is an operator stop: a model an
|
||||||
|
operator has switched off, not one the catalog flagged. Routing and pinned
|
||||||
|
requests both exclude a blocked model (`_admin_excluded_models` in
|
||||||
|
`dispatcher.py` and `metrics.py`). Undo is setting the dropdown back to
|
||||||
|
`active` (clearing the override via `DELETE` on the availability path), not a
|
||||||
|
separate unlock — there is no per-model Block button beside evidence and no
|
||||||
|
Blocked list with reasons; those belong to the "not built yet" item below.
|
||||||
|
|
||||||
<div align="center">
|
<div align="center">
|
||||||
<img src="../assets/screenshots/models.png" alt="Admin models page" width="820">
|
<img src="../assets/screenshots/models.png" alt="Admin models page" width="820">
|
||||||
@@ -126,8 +145,8 @@ land in `config/config.local.yaml`; see
|
|||||||
[config-local-overlay](config-local-overlay.md)). Changes are marked as dirty
|
[config-local-overlay](config-local-overlay.md)). Changes are marked as dirty
|
||||||
and written only on save.
|
and written only on save.
|
||||||
|
|
||||||
It carries two dedicated cards, neither a row in the generic runtime-knob
|
It carries three dedicated cards, none a row in the generic runtime-knob
|
||||||
list, because both have cross-field structure a flat scalar/boolean input
|
list, because each has cross-field structure a flat scalar/boolean input
|
||||||
can't safely represent:
|
can't safely represent:
|
||||||
|
|
||||||
- **Local Compute** — "gaming mode". See [Gaming mode](#gaming-mode) below.
|
- **Local Compute** — "gaming mode". See [Gaming mode](#gaming-mode) below.
|
||||||
@@ -145,6 +164,14 @@ can't safely represent:
|
|||||||
anything reaches `config.local.yaml`, and mode plus its companion block are
|
anything reaches `config.local.yaml`, and mode plus its companion block are
|
||||||
written as one atomic change so an in-between invalid state is never even
|
written as one atomic change so an in-between invalid state is never even
|
||||||
written transiently.
|
written transiently.
|
||||||
|
- **Watchdog** — the watchdog's live state (`controls.html#watchdog-card`,
|
||||||
|
`loadWatchdog()`): the last tick time (and a `no tick` badge before the
|
||||||
|
first one), how many sessions the last tick saw, how many verdicts it
|
||||||
|
flagged, and how many alerts are currently open. Below those, a set of
|
||||||
|
**notification channel** toggles (each channel with an enable/disable and a
|
||||||
|
minimum severity, persisted via `POST /admin/api/watchdog/channels`), a
|
||||||
|
**Refresh** button, and a **Send test alert** button (`POST
|
||||||
|
/admin/api/watchdog/test-alert`).
|
||||||
|
|
||||||
<div align="center">
|
<div align="center">
|
||||||
<img src="../assets/screenshots/controls.png" alt="Admin controls page" width="820">
|
<img src="../assets/screenshots/controls.png" alt="Admin controls page" width="820">
|
||||||
@@ -444,7 +471,15 @@ intermediate invalid state could land on disk. The GET response includes
|
|||||||
**Model availability overrides**
|
**Model availability overrides**
|
||||||
|
|
||||||
`POST /admin/api/models/{model_id:path}/{provider}/availability` marks a model
|
`POST /admin/api/models/{model_id:path}/{provider}/availability` marks a model
|
||||||
as `active`, `deprecated`, or `stale`. `DELETE` on the same path removes the
|
as `active`, `blocked`, `deprecated`, or `stale`. `DELETE` on the same path
|
||||||
override. The `{model_id:path}` converter accepts model ids that contain
|
removes the override. The `{model_id:path}` converter accepts model ids that
|
||||||
slashes, so providers like OpenRouter with slash-bearing ids are handled the
|
contain slashes, so providers like OpenRouter with slash-bearing ids are
|
||||||
same way as plain ids. Deprecation feeds into the routing hard filters.
|
handled the same way as plain ids. Deprecation feeds into the routing hard
|
||||||
|
filters.
|
||||||
|
|
||||||
|
**Not built yet.** The Loops panel does not yet show the evidence columns an
|
||||||
|
operator would act on — agent, model, top repeated target, calls, and $ since
|
||||||
|
the last landed change — and there is no per-model stall rollup, no **Block
|
||||||
|
model** button beside the evidence, and no Blocked list with reasons. Those are
|
||||||
|
the remaining surface of [North Star #4](../CLAUDE.md) in CLAUDE.md, whose first
|
||||||
|
half (conversation identity and the watchdog) is shipped.
|
||||||
|
|||||||
65
docs/api.md
65
docs/api.md
@@ -173,6 +173,12 @@ These ids are what make `/outcome` attribution exact (below): an
|
|||||||
`X-Router-Conversation` on the request sets the `session_key` row that a later
|
`X-Router-Conversation` on the request sets the `session_key` row that a later
|
||||||
report can address by `conversation_id` without guessing `source` or ambiguity.
|
report can address by `conversation_id` without guessing `source` or ambiguity.
|
||||||
|
|
||||||
|
**Plugin loader contract:** the opencode plugin that stamps these headers
|
||||||
|
must export only functions — the opencode 1.18.x plugin loader rejects a
|
||||||
|
plugin outright if any export is not a function (so the plugin's
|
||||||
|
`parentCache` lives as a property on the factory rather than a named export).
|
||||||
|
See [clients.md](clients.md) for the full contract.
|
||||||
|
|
||||||
**Capability 422s**: When no model survives the hard filters, the 422 names the
|
**Capability 422s**: When no model survives the hard filters, the 422 names the
|
||||||
active constraints. That now includes "vision-capable model" or
|
active constraints. That now includes "vision-capable model" or
|
||||||
"json-mode-capable model" when the request carried images or a JSON-mode
|
"json-mode-capable model" when the request carried images or a JSON-mode
|
||||||
@@ -225,6 +231,65 @@ but a `true` report counts too. It's also the only quality signal that survives
|
|||||||
streaming, since a retry can't reach a response whose bytes are already gone,
|
streaming, since a retry can't reach a response whose bytes are already gone,
|
||||||
while a report arrives afterward and works either way.
|
while a report arrives afterward and works either way.
|
||||||
|
|
||||||
|
## Admin watchdog and availability API
|
||||||
|
|
||||||
|
The admin portal's watchdog and model-availability operations sit under
|
||||||
|
`/admin` (loopback-only, like the rest of the admin API). The watchdog
|
||||||
|
endpoints back the dashboard's **Loops** card and the **Watchdog** control
|
||||||
|
card; full behaviour is described in [watchdog.md](watchdog.md) and
|
||||||
|
[admin-portal.md](admin-portal.md).
|
||||||
|
|
||||||
|
| Method | Path | Description |
|
||||||
|
|---|---|---|
|
||||||
|
| `GET` | `/admin/api/watchdog/status` | Latest watchdog tick plus open-alert count |
|
||||||
|
| `GET` | `/admin/api/watchdog/loops` | Open alert rows (one per looping session) |
|
||||||
|
| `GET` | `/admin/api/watchdog/channels` | Notification-channel settings |
|
||||||
|
| `POST` | `/admin/api/watchdog/channels` | Create or update a channel's enabled/min-severity |
|
||||||
|
| `POST` | `/admin/api/watchdog/test-alert` | Deliver a test alert through the enabled channels |
|
||||||
|
|
||||||
|
`GET /admin/api/watchdog/status` returns a single object:
|
||||||
|
|
||||||
|
- `last_tick` — the most recent `watchdog_ticks` row, or `null` before the
|
||||||
|
first tick
|
||||||
|
- `verdicts` — `{total, flagged}` for that tick's `watchdog_verdicts`, or
|
||||||
|
`null` when there is no tick yet
|
||||||
|
- `open_alerts` — count of `watchdog_alerts` rows with `resolved_at` null
|
||||||
|
|
||||||
|
`GET /admin/api/watchdog/loops` returns an array of `watchdog_alerts` rows
|
||||||
|
where `resolved_at` is null, ordered by `opened_at` descending. Each element
|
||||||
|
is the full alert row.
|
||||||
|
|
||||||
|
`GET /admin/api/watchdog/channels` returns an array of
|
||||||
|
`watchdog_channel_settings` rows ordered by `channel_name`.
|
||||||
|
|
||||||
|
`POST /admin/api/watchdog/channels` upserts one channel. Body:
|
||||||
|
|
||||||
|
- `channel_name` (required) — the channel to create or update; absence is a
|
||||||
|
422
|
||||||
|
- `enabled` (optional boolean) — when omitted the existing value is kept
|
||||||
|
- `min_severity` (optional: `info`, `warning`, or `critical`) — any other
|
||||||
|
value is a 422
|
||||||
|
|
||||||
|
Returns `{ok: true}`.
|
||||||
|
|
||||||
|
`POST /admin/api/watchdog/test-alert` delivers a test alert through the
|
||||||
|
currently enabled channels. Body:
|
||||||
|
|
||||||
|
- `severity` (optional, default `warning`; `info`, `warning`, or `critical`)
|
||||||
|
- `channel_name` (optional) — restrict to a single channel
|
||||||
|
|
||||||
|
Returns `{sent: <n>, severity: <severity>}`, or `{sent: 0, message: "no
|
||||||
|
enabled channels"}` when nothing is enabled. An invalid `severity` is a 422.
|
||||||
|
|
||||||
|
**Model availability** — `POST
|
||||||
|
/admin/api/models/{model_id:path}/{provider}/availability` upserts an admin
|
||||||
|
override for a model's availability. `availability` must be one of `active`,
|
||||||
|
`blocked`, `deprecated`, or `stale` (anything else is a 422); `blocked` is an
|
||||||
|
operator stop that routing and pinned requests exclude
|
||||||
|
(`_admin_excluded_models`), distinct from the catalog meaning of `deprecated`.
|
||||||
|
`DELETE` on the same path removes the override. The `{model_id:path}`
|
||||||
|
converter accepts slash-bearing model ids (e.g. OpenRouter).
|
||||||
|
|
||||||
## Pinning and auto behavior for local models
|
## Pinning and auto behavior for local models
|
||||||
|
|
||||||
`model: "auto"` will route eligible tasks to `qwen2.5-coder-router:14b` when the
|
`model: "auto"` will route eligible tasks to `qwen2.5-coder-router:14b` when the
|
||||||
|
|||||||
@@ -37,3 +37,11 @@ This resolved a real incident where reading a dependency's source 24 times
|
|||||||
outweighed writing to the project dir 22 times: the write-like 3× weighting
|
outweighed writing to the project dir 22 times: the write-like 3× weighting
|
||||||
puts the project directory that actually received the edits ahead of a
|
puts the project directory that actually received the edits ahead of a
|
||||||
dependency's source that was merely read.
|
dependency's source that was merely read.
|
||||||
|
|
||||||
|
**Plugin loader contract:** opencode 1.18.x (the observed version) rejects the
|
||||||
|
entire plugin if *any* export is not a function. The `parentCache` map on
|
||||||
|
`router-link.js` is therefore set as a property on the `RouterLink` factory
|
||||||
|
function rather than as a named export (L221-224 of
|
||||||
|
`deploy/opencode-plugin/router-link.js`). Confirm the plugin loaded by
|
||||||
|
searching the log for `failed to load plugin`: `grep 'failed to load plugin'
|
||||||
|
~/.local/share/opencode/log/opencode.log`.
|
||||||
|
|||||||
@@ -31,7 +31,7 @@ Three data tables plus one observability table, `PRAGMA foreign_keys = ON`:
|
|||||||
| `access_level` | TEXT | `public` \| `preview` \| `canary` |
|
| `access_level` | TEXT | `public` \| `preview` \| `canary` |
|
||||||
| `pricing_tbd` | INTEGER | |
|
| `pricing_tbd` | INTEGER | |
|
||||||
| `deprecated` | INTEGER | |
|
| `deprecated` | INTEGER | |
|
||||||
| `availability` | TEXT | `active` \| `deprecated` \| `stale` |
|
| `availability` | TEXT | `active` \| `deprecated` \| `stale` \| `blocked` — the `blocked` value is an operator stop set via the admin override (see [admin-portal.md](admin-portal.md)); it is excluded from routing and pinned requests |
|
||||||
| `last_updated` | TEXT | ISO8601 |
|
| `last_updated` | TEXT | ISO8601 |
|
||||||
|
|
||||||
**Serving class:** Neuralwatt ships ~6 base models as 19 catalog rows. The id
|
**Serving class:** Neuralwatt ships ~6 base models as 19 catalog rows. The id
|
||||||
@@ -326,3 +326,71 @@ prompts, and answers are deliberately excluded. A test enforces that the write
|
|||||||
path does not store prompt or answer text. The write is gated by
|
path does not store prompt or answer text. The write is gated by
|
||||||
`logging.log_route_decisions` and is **best-effort**: a failed write is logged
|
`logging.log_route_decisions` and is **best-effort**: a failed write is logged
|
||||||
at warning and swallowed so monitoring cannot slow or fail a request.
|
at warning and swallowed so monitoring cannot slow or fail a request.
|
||||||
|
|
||||||
|
## Watchdog Schema (SQLite)
|
||||||
|
|
||||||
|
Four tables written by `watchdog.py` via `src/watchdog_store.py` (the code-side
|
||||||
|
inline-create mirrors this schema for live databases). See
|
||||||
|
[watchdog.md](watchdog.md) for how the loop reads them.
|
||||||
|
|
||||||
|
### `watchdog_ticks` — one row per detection run
|
||||||
|
|
||||||
|
| Column | Type | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| `id` | INTEGER | Autoincrement |
|
||||||
|
| `ticked_at` | TEXT | ISO8601 run time |
|
||||||
|
| `sessions_seen` | INTEGER | Count of sessions examined this run, default 0 |
|
||||||
|
| `outcome` | TEXT | |
|
||||||
|
|
||||||
|
### `watchdog_verdicts` — one row per session per tick
|
||||||
|
|
||||||
|
| Column | Type | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| `id` | INTEGER | Autoincrement |
|
||||||
|
| `tick_id` | INTEGER | FK → `watchdog_ticks.id`, NOT NULL |
|
||||||
|
| `session_id` | TEXT | NOT NULL |
|
||||||
|
| `session_root` | TEXT | |
|
||||||
|
| `agent` | TEXT | |
|
||||||
|
| `model_id` | TEXT | |
|
||||||
|
| `provider` | TEXT | |
|
||||||
|
| `flagged` | INTEGER | 0/1, default 0 |
|
||||||
|
| `dup` | REAL | Duplicate-fraction signal |
|
||||||
|
| `top` | INTEGER | |
|
||||||
|
| `top_what` | TEXT | |
|
||||||
|
| `landed` | INTEGER | |
|
||||||
|
| `slow` | INTEGER | |
|
||||||
|
| `coverage` | REAL | |
|
||||||
|
| `calls_since_landed` | INTEGER | |
|
||||||
|
| `cost_since_landed_usd` | REAL | |
|
||||||
|
| `llm_second_opinion` | TEXT | |
|
||||||
|
| `created_at` | TEXT | NOT NULL |
|
||||||
|
|
||||||
|
Four indexes (`config/schema.sql` L482-485):
|
||||||
|
|
||||||
|
| Index | On |
|
||||||
|
|---|---|
|
||||||
|
| `idx_watchdog_verdicts_model_id` | `watchdog_verdicts(model_id)` |
|
||||||
|
| `idx_watchdog_verdicts_created_at` | `watchdog_verdicts(created_at)` |
|
||||||
|
| `idx_watchdog_verdicts_flagged` | `watchdog_verdicts(flagged)` |
|
||||||
|
| `idx_watchdog_verdicts_session_root` | `watchdog_verdicts(session_root)` |
|
||||||
|
|
||||||
|
### `watchdog_alerts` — one row per open alert
|
||||||
|
|
||||||
|
| Column | Type | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| `dedup_key` | TEXT | PRIMARY KEY; e.g. `opencode-loop:<root session id>` |
|
||||||
|
| `state` | TEXT | NOT NULL |
|
||||||
|
| `severity` | TEXT | NOT NULL |
|
||||||
|
| `flagged_ticks` | INTEGER | Default 0 |
|
||||||
|
| `opened_at` | TEXT | NOT NULL |
|
||||||
|
| `last_fired_at` | TEXT | NOT NULL |
|
||||||
|
| `resolved_at` | TEXT | |
|
||||||
|
|
||||||
|
### `watchdog_channel_settings` — per-channel notification config
|
||||||
|
|
||||||
|
| Column | Type | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| `id` | INTEGER | Autoincrement |
|
||||||
|
| `channel_name` | TEXT | NOT NULL, UNIQUE |
|
||||||
|
| `enabled` | INTEGER | Default 1 |
|
||||||
|
| `min_severity` | TEXT | Default `warning` |
|
||||||
|
|||||||
@@ -1,14 +1,18 @@
|
|||||||
# Incidents: how this router has broken, and how to tell which one it is
|
# Incidents: how this router has broken, and how to tell which one it is
|
||||||
|
|
||||||
Seven times now, a change that looked local to the router has silently
|
Eight times now, something around the router has silently degraded either the
|
||||||
degraded either the agent depending on it or the operator trying to see it
|
agent depending on it or the operator trying to see it clearly. They share a
|
||||||
clearly. They share a shape worth naming: **none of them announce themselves
|
shape worth naming: **none of them announce themselves as router problems.**
|
||||||
as router problems.** Five of the seven presented as an opaque client-side
|
Five of the eight presented as an opaque client-side error — a connection
|
||||||
error — a connection refused, an "Unprocessable Content", an "internal server
|
refused, an "Unprocessable Content", an "internal server error" — and
|
||||||
error" — and diagnosing each meant knowing which log or table to look in. The
|
diagnosing each meant knowing which log or table to look in. The other three
|
||||||
other two didn't error toward the client at all: #5 destroyed data outright,
|
didn't error toward the client at all: #5 destroyed data outright, #6 just went
|
||||||
and #6 just went quiet in one corner of the admin UI while `/health` stayed
|
quiet in one corner of the admin UI while `/health` stayed green the whole time,
|
||||||
green the whole time.
|
and #8 spent money for hours while every check stayed green.
|
||||||
|
|
||||||
|
**#8 is the only one so far that cost money rather than capability**, and the
|
||||||
|
router behaved as configured throughout. Read it before leaving
|
||||||
|
an agent run unattended.
|
||||||
|
|
||||||
**#5 is the only one so far that destroyed data**, and the only one where
|
**#5 is the only one so far that destroyed data**, and the only one where
|
||||||
recovery depended on luck rather than design. Read it before running any
|
recovery depended on luck rather than design. Read it before running any
|
||||||
@@ -21,6 +25,15 @@ or removing any config key.
|
|||||||
This page exists so the next one takes minutes rather than hours. Start with the
|
This page exists so the next one takes minutes rather than hours. Start with the
|
||||||
symptom table, then read only the relevant section.
|
symptom table, then read only the relevant section.
|
||||||
|
|
||||||
|
**How to write an entry (from #8 on).**
|
||||||
|
- Separate **Evidence** (measured, with where it was measured) from **Theory**
|
||||||
|
(inferred). Anything not established by a measurement or a reproduction goes
|
||||||
|
under Theory. It states what supports it and what would confirm or refute
|
||||||
|
it, and stays there until that test is done.
|
||||||
|
- Include a **Human response** timeline: what the operator noticed, decided and
|
||||||
|
did. Mark actions an assistant took, and at whose direction.
|
||||||
|
- Entries #1 to #7 predate this convention.
|
||||||
|
|
||||||
`opencode.json` points opencode's own model traffic at `http://127.0.0.1:8080/v1`,
|
`opencode.json` points opencode's own model traffic at `http://127.0.0.1:8080/v1`,
|
||||||
so on this machine the router *is* the coding agent's inference supply. Breaking
|
so on this machine the router *is* the coding agent's inference supply. Breaking
|
||||||
it breaks the thing you would use to fix it. That is why these keep happening,
|
it breaks the thing you would use to fix it. That is why these keep happening,
|
||||||
@@ -38,6 +51,8 @@ and why the diagnostics below are worth having to hand.
|
|||||||
| 500s + `no such table`, venv/.env gone | #5 `git clean -fdx` | `ls -la router.db .env .venv` — a 0-byte db and a missing `.venv` is conclusive |
|
| 500s + `no such table`, venv/.env gone | #5 `git clean -fdx` | `ls -la router.db .env .venv` — a 0-byte db and a missing `.venv` is conclusive |
|
||||||
| A newly-shipped admin knob just isn't there, `/health` is fine | #6 backend running stale code | `systemctl --user status llm-router` uptime vs. `git log -1 --format=%cd` on the commit that added the knob |
|
| A newly-shipped admin knob just isn't there, `/health` is fine | #6 backend running stale code | `systemctl --user status llm-router` uptime vs. `git log -1 --format=%cd` on the commit that added the knob |
|
||||||
| `activating (auto-restart)`, `pydantic_core...ValidationError: ...extra_forbidden` in the journal | #7 renamed key vs. un-migrated overlay | `journalctl --user -u llm-router --since '5 min ago' \| grep extra_forbidden` then `grep -n <old key name> config/config.local.yaml` |
|
| `activating (auto-restart)`, `pydantic_core...ValidationError: ...extra_forbidden` in the journal | #7 renamed key vs. un-migrated overlay | `journalctl --user -u llm-router --since '5 min ago' \| grep extra_forbidden` then `grep -n <old key name> config/config.local.yaml` |
|
||||||
|
| Spend is high after an agent run, nothing errored, no warning fired | #8 agent looping with no progress | `sqlite3 router.db "select session_key, strftime('%H',observed_at,'localtime') h, round(sum(prompt_tokens)/1e6) mtok, round(sum(cost_usd),2) usd from energy_observations where observed_at > datetime('now','-12 hours') group by 1,2 order by usd desc limit 8;"` then check whether that run's commits actually landed |
|
||||||
|
| opencode plugin does nothing | #8 plugin loader rejects non-function export | `grep 'failed to load plugin' ~/.local/share/opencode/log/opencode.log` |
|
||||||
|
|
||||||
That last row is not an incident yet — it is the silent-staleness failure
|
That last row is not an incident yet — it is the silent-staleness failure
|
||||||
described in the "Run as a service" section of `CLAUDE.md`. An unpolled catalog
|
described in the "Run as a service" section of `CLAUDE.md`. An unpolled catalog
|
||||||
@@ -347,6 +362,253 @@ file is correct.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## #8 — Agent sessions looped for hours with no progress, and nothing noticed (2026-09-25)
|
||||||
|
|
||||||
|
**Symptom.** Nothing errored. The router stayed healthy, `/metrics` showed zero
|
||||||
|
warnings, and the quota alarm read `kind: none`. The operator noticed early on
|
||||||
|
2026-09-26 that an opencode run had "burned a bunch of tokens over the hours".
|
||||||
|
|
||||||
|
### Evidence
|
||||||
|
|
||||||
|
Measured read-only against the live `router.db` and opencode's own session
|
||||||
|
store. The window is 2026-09-25 19:00 to 2026-09-26 02:00 local unless noted.
|
||||||
|
|
||||||
|
**Spend.**
|
||||||
|
|
||||||
|
| | |
|
||||||
|
|---|---|
|
||||||
|
| prompt tokens | 323M, 95% cached |
|
||||||
|
| billed | about $7.81, roughly 3x the busiest of the previous ten days ($2.66 on 09-20) |
|
||||||
|
| largest share | OpenRouter `z-ai/glm-5.3-flash`: 1,368 calls averaging 203k prompt tokens, $6.16 |
|
||||||
|
| largest contexts | NeuralWatt `glm-5.3`: 47 calls averaging 673k, max 682,679, the only model whose window still fit |
|
||||||
|
| 120k+ prompt calls | 921 of them, $6.57 of the total |
|
||||||
|
| routing | every decision `classification_source = classifier`, default profile; nothing pinned |
|
||||||
|
|
||||||
|
**What the agents did.** An Atlas run (`.omo/plans/cockpit-quick-wins.md`)
|
||||||
|
delegated each item to a sub-agent worker:
|
||||||
|
|
||||||
|
| worker session | tool calls | exact-duplicate calls | worst repeat | landed |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| item 2, 2nd attempt | 423 | 30% | `read navbar.js` x61 | nothing |
|
||||||
|
| item 3, 1st attempt | 132 | 23% | `read test_admin_js_units.py` x16 | nothing |
|
||||||
|
| item 4 | 134 | 19% | `read` of the plan x8 | code only |
|
||||||
|
| item 2, 1st attempt | 307 | 17% | `read navbar.js` x21 | nothing |
|
||||||
|
| healthy workers | 30 to 96 | 0 to 6% | 1 to 3 | yes |
|
||||||
|
|
||||||
|
- Item 4's commit message claimed four test files it never wrote, and it
|
||||||
|
committed with `git add -A`.
|
||||||
|
- Atlas later said its own turns were "just re-reading the plan and state over
|
||||||
|
and over instead of doing work".
|
||||||
|
- A planning session described its reads as "heavily elided". The stored tool
|
||||||
|
outputs in opencode were intact.
|
||||||
|
|
||||||
|
**Pinch was rewriting agent context.**
|
||||||
|
- `pinch.budget_tokens: 50000` pruned 1,915 of 1,992 decisions over nine hours,
|
||||||
|
keeping 48% of the prompt on average.
|
||||||
|
- The planning session's 527k-token prompts went out at about 125k.
|
||||||
|
- `src/context_prune.py` replaces older tool results with
|
||||||
|
`[read: result omitted]`, or keeps a 1,500-character head and tail around a
|
||||||
|
`[N chars trimmed...]` marker. Results inside the protected window get the
|
||||||
|
same cut above `protected_max_chars` (20,000).
|
||||||
|
- Pinch's config had not changed since 2026-09-06. It had pruned 90 to 100% of
|
||||||
|
requests daily for two weeks without incident. The input changed:
|
||||||
|
|
||||||
|
| | 09-10 to 09-23 | 09-25 | 09-26 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| avg prompt into pinch | 100k to 150k | 236k | 545k |
|
||||||
|
| p90 prompt | 170k to 260k | 528k | 1.31M |
|
||||||
|
|
||||||
|
- `opencode.json` advertised `auto` with `limit.context: 782324`, so opencode did
|
||||||
|
not compact until near that size.
|
||||||
|
|
||||||
|
**Routing drifted with size.**
|
||||||
|
- `deepseek/deepseek-v4-flash` has 384k effective context;
|
||||||
|
`z-ai/glm-5.3-flash` has 655k.
|
||||||
|
- Their proficiency was tied (`diff_checking` 0.933 vs 0.935) and unchanged
|
||||||
|
since 09-10.
|
||||||
|
- Deepseek went from 290 picks on 09-23 to 0 on 09-25, while appearing 851
|
||||||
|
times as runner-up.
|
||||||
|
|
||||||
|
**A clean reproduction, 2026-09-26 04:04 to 04:14.** A fresh planning
|
||||||
|
session, on the Lift A brief (188 lines):
|
||||||
|
- **Conditions:** context about 88k tokens, pinch off, 0 of 29 tool parts
|
||||||
|
marked `compacted` by opencode, every stored read intact.
|
||||||
|
- **What it did:** it still reported its reads "elided with `[...]`", wrote
|
||||||
|
fake "Earlier tool responses received ... (counts, not proof of task
|
||||||
|
completion)" blocks into its own replies (assistant text parts, not tool
|
||||||
|
output), invented session IDs, and called the brief "393 lines".
|
||||||
|
- **Where the phrase is not:** in opencode, oh-my-openagent (JS bundle and
|
||||||
|
native binary), or this repo.
|
||||||
|
- **Model:** `z-ai/glm-5.3-flash` served 27 of 34 chat decisions in the 15
|
||||||
|
minutes before it was stopped.
|
||||||
|
|
||||||
|
**Found along the way.**
|
||||||
|
- `deploy/opencode-plugin/router-outcome.js` has never read a real exit code.
|
||||||
|
On opencode 1.18, bash's exit is `output.metadata.exit`; the plugin reads
|
||||||
|
`output.exitCode`.
|
||||||
|
- PR #99 (conversation identity) was merged to `origin/main` but not running.
|
||||||
|
Production ran from a local checkout 29 commits behind.
|
||||||
|
- **Plugin loader rejected `router-link.js`.** The plugin exported `parentCache`
|
||||||
|
as a module-level const. Opencode 1.18's plugin loader discards any plugin
|
||||||
|
with a non-function export — the entire file is silently dropped. The fix was
|
||||||
|
to make `parentCache` a property on the factory function instead (L224,
|
||||||
|
`router-link.js` in PR #102, a24de1b). Confirmed by grepping
|
||||||
|
`~/.local/share/opencode/log/opencode.log` for `'failed to load plugin'`.
|
||||||
|
- How much of the $7.81 was waste cannot be computed. The run did land six
|
||||||
|
reviewed commits, and the deployed router cannot tell conversations apart.
|
||||||
|
|
||||||
|
### Theory
|
||||||
|
|
||||||
|
These are inferences, not measurements. Each lists what would confirm or
|
||||||
|
refute it.
|
||||||
|
|
||||||
|
1. **Pinch cutting oversized sessions drove, or worsened, the re-read loops.**
|
||||||
|
Weakened by the clean reproduction, which looped with pinch off. As
|
||||||
|
sessions grew past about 250k, the fixed 50k budget removed most of each
|
||||||
|
agent's working memory, recent reads included. The agent saw its reads cut,
|
||||||
|
re-read them, grew the context, and got cut harder.
|
||||||
|
- Supported by: the agents' own descriptions ("elided", "drowning in
|
||||||
|
tool-output truncation"), intact stored outputs, the prune rates, and the
|
||||||
|
timing of the size jump.
|
||||||
|
- Not tested: no run has been observed with pinch off.
|
||||||
|
- Confirm: with pinch off and `auto` at 200k, the duplicate-read signature
|
||||||
|
should not recur on comparable runs. The interim watcher
|
||||||
|
(`plans/no-progress-detection-prototype.py`) and Phase 0 data can show it.
|
||||||
|
2. **Sessions got that large because long unattended orchestration ran without
|
||||||
|
compaction.** Atlas made 668 tool calls, one worker 423, under a 782k limit,
|
||||||
|
and the re-read loop in theory 1 fed the growth. The limit is a fact; that it
|
||||||
|
explains the growth is theory.
|
||||||
|
3. **Leading theory: `z-ai/glm-5.3-flash`, or its OpenRouter path,
|
||||||
|
confabulates truncation and then loops.** The clean reproduction above
|
||||||
|
removed pinch, context size and input corruption, and the behaviour stayed.
|
||||||
|
Theories 1 and 2 remain plausible aggravators, not the cause.
|
||||||
|
- Confirm: with glm excluded, comparable agent runs show no elision claims
|
||||||
|
and no duplicate-read signature. A direct A/B on one prompt would settle
|
||||||
|
model versus provider path.
|
||||||
|
|
||||||
|
### Why nothing caught it
|
||||||
|
|
||||||
|
Each existing guard answers a different question:
|
||||||
|
- `circuit_breaker.py` trips on provider 5xx. There were none.
|
||||||
|
- `runway_low_warning` fires when balance / burn drops under 6 hours. Burn is
|
||||||
|
averaged over a 24 h balance window: OpenRouter read $24.73 at $0.25/h, which
|
||||||
|
is 97 hours of runway. It asks "will I run out", not "is this being wasted".
|
||||||
|
- Spend rate was not abnormal. The worst hour ($2.42) and worst 3-hour window
|
||||||
|
($4.13) sit inside the prior 30 days' range (hourly p90 $1.46, max $4.83;
|
||||||
|
worst 3 h $9.65).
|
||||||
|
- The router cannot see progress, and the deployed build could not tell
|
||||||
|
conversations apart.
|
||||||
|
|
||||||
|
### Recurrences
|
||||||
|
|
||||||
|
- **02:21 to about 02:56: Atlas looped after its plan was complete.** Every
|
||||||
|
checkbox was ticked, but `.omo/boulder.json` still carried `status: completed`
|
||||||
|
and `pr_url: .../pulls/77` from the previous plan (evidence: the file). The
|
||||||
|
continuation hook injected "continue" turns, 3 of 3 user turns in the window
|
||||||
|
(evidence: the session). Atlas re-derived its state each time: 54 turns,
|
||||||
|
about 19M cached-read tokens, `boulder.json` read 28 times, checkboxes grepped
|
||||||
|
16 times. Cause: stale state, not pruning. The branch was never pushed, so PR
|
||||||
|
77 was not touched.
|
||||||
|
- **About 03:05 to 03:17: a read-only `explore` subagent** re-read
|
||||||
|
`router-outcome.js` 20 times (19% duplicate calls across 322). The production
|
||||||
|
sync had just deleted that file. Read-only agents never land changes, so the
|
||||||
|
detector needs a separate signal for them.
|
||||||
|
|
||||||
|
### Human response
|
||||||
|
|
||||||
|
Times are local, 2026-09-26.
|
||||||
|
- **About 02:10.** Noticed the spend and asked for a mechanism in 6krrt to catch
|
||||||
|
it. Rejected a spend-rate alarm, since steady agent usage is legitimate, and
|
||||||
|
framed the target as "are concrete changes landing".
|
||||||
|
- **About 02:20 to 02:40.** Decided the design for `plans/no-progress-detection.md`:
|
||||||
|
- response modes (`warn`, `auto_compact`, `auto_limit_context`, a two-stage
|
||||||
|
`auto_recover`), with every mode warning
|
||||||
|
- `warn` as the default
|
||||||
|
- at most 2 recoveries per session tree
|
||||||
|
- `notify-send` now, with pluggable SMS, RingCentral and PagerDuty channels
|
||||||
|
later
|
||||||
|
- a local-model watchdog timer, "so I don't have to rely on Claude"
|
||||||
|
- `auto`'s context limit lowered to 200,000 in the repo and global
|
||||||
|
`opencode.json`
|
||||||
|
- **About 02:50.** Caught the post-completion Atlas loop by watching. At the
|
||||||
|
operator's direction, Claude aborted the session at 02:56 and cleared the
|
||||||
|
stale `pr_url` in `.omo/boulder.json`.
|
||||||
|
- **About 03:10.** Merged PR #101. A first `git pull` aborted on local changes,
|
||||||
|
and the restart on that line reloaded the old code. Then ran the prepared sync
|
||||||
|
(back up, drop the changes already upstream, `reset --keep origin/main`,
|
||||||
|
re-apply the rest) and restarted onto `827c408` at 03:15.
|
||||||
|
- **About 03:30.** Asked Claude to watch opencode for loops until the detector
|
||||||
|
ships (the interim watcher).
|
||||||
|
- **About 03:45.** Reported the planning session's "truncation hell". At the
|
||||||
|
operator's direction ("hot fix that"), Claude set `pinch.enabled` false at
|
||||||
|
runtime and persisted it to `config/config.local.yaml` (backup
|
||||||
|
`config.local.yaml.bak-20260926-pinch`). The operator restarted the service
|
||||||
|
at 03:52 and chose not to open a PR for the switch flip.
|
||||||
|
- **About 04:02.** Installed `router-link.js` in place of `router-outcome.js`
|
||||||
|
and restarted opencode (200k `auto` limit now active). Kicked off a fresh,
|
||||||
|
small Lift A planning session (`plans/no-progress-lift-a.md`).
|
||||||
|
- **About 04:14.** After the fresh session reproduced the confabulation,
|
||||||
|
excluded `z-ai/glm-5.3-flash` from routing. It had no picks after 04:13:57,
|
||||||
|
and traffic moved to deepseek and mimo. At the operator's direction, Claude
|
||||||
|
aborted the confabulating session.
|
||||||
|
- **~15:18.** Watchdog timer enabled (`systemctl --user enable --now
|
||||||
|
llm-router-watchdog.timer`).
|
||||||
|
- **15:19.** Plugin-loader root cause found in `opencode.log`; plugin fix
|
||||||
|
installed and opencode restarted.
|
||||||
|
- **15:19:57.** Router restarted.
|
||||||
|
- **After 15:19.** Conversation identity confirmed arriving: 17 of 17 requests
|
||||||
|
after the restart carry `c:` keys and an agent.
|
||||||
|
|
||||||
|
### Status
|
||||||
|
|
||||||
|
**Partly mitigated; detection not built; leading theory (3) unconfirmed.**
|
||||||
|
|
||||||
|
**Lift A shipped.** PR #102 (a24de1b) landed the surfacing half of North Star #4
|
||||||
|
(conversation identity and watchdog), live ~15:18 EDT. **Surfacing half of North
|
||||||
|
Star #4 is only partly shipped** — the identity half is confirmed (17 of 17
|
||||||
|
requests after the restart carry `c:` keys), but the detection half still waits
|
||||||
|
on Lift B (`plans/no-progress-detection.md`).
|
||||||
|
|
||||||
|
**Confounded by design, and say so.** Three mitigations landed within about
|
||||||
|
25 minutes:
|
||||||
|
- pinch off at 03:50
|
||||||
|
- `auto` capped at 200k at 04:02
|
||||||
|
- `z-ai/glm-5.3-flash` excluded at about 04:14
|
||||||
|
|
||||||
|
So "sessions look healthier since" cannot say which one mattered, or whether
|
||||||
|
it is several interacting: model degradation near a full window, oversized
|
||||||
|
sessions, pruning, and stale orchestration state. The one clean data point is
|
||||||
|
the 04:04 reproduction (glm, pinch off, small context, still confabulated),
|
||||||
|
which is why theory 3 leads.
|
||||||
|
|
||||||
|
To attribute properly, change one variable at a time and watch the watchdog's
|
||||||
|
stall rate. For example, re-admit glm with everything else held, or re-enable
|
||||||
|
pinch with a budget above 200k. Until then, treat every theory here as
|
||||||
|
unproven.
|
||||||
|
Also missing: a **behavior breaker**. The availability breaker trips on 5xx and
|
||||||
|
PR #76 trips on malformed output; nothing trips on well-formed output that
|
||||||
|
confabulates or loops. Planned for Lift B (`plans/no-progress-detection.md`,
|
||||||
|
section 8).
|
||||||
|
- In place:
|
||||||
|
- pinch off, runtime and persisted
|
||||||
|
- `auto` limit 200,000, active for opencode sessions started after a restart
|
||||||
|
- #99 and #100 deployed at 03:15
|
||||||
|
- Pending, operator:
|
||||||
|
- install `router-link.js` in place of `router-outcome.js`
|
||||||
|
- restart opencode
|
||||||
|
- Pending, lift: `plans/no-progress-detection.md`, the watchdog first.
|
||||||
|
- If pinch is re-enabled, its budget must sit above the 200k cap, and it must
|
||||||
|
never cut recent results.
|
||||||
|
|
||||||
|
**Rule:** steady spend is not the failure; **spend with no concrete change
|
||||||
|
landing is.** Until the detector ships:
|
||||||
|
- Check an unattended agent run by whether commits are landing, not by the
|
||||||
|
quota page.
|
||||||
|
- Verify each worker commit's `git show --stat` against its message; a worker's
|
||||||
|
"done" is a claim.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## The pattern
|
## The pattern
|
||||||
|
|
||||||
The first five are defensible local decisions — free a port, stop a process,
|
The first five are defensible local decisions — free a port, stop a process,
|
||||||
@@ -373,6 +635,14 @@ file by mistake; #7 was a *correct*, deliberate, reviewed code change that
|
|||||||
was still incomplete, because "every consumer" silently excluded a consumer
|
was still incomplete, because "every consumer" silently excluded a consumer
|
||||||
that git cannot see.
|
that git cannot see.
|
||||||
|
|
||||||
|
#8 breaks the pattern from the other side: the router did what it was
|
||||||
|
configured to do, routing each request to a model that fit. It was the one
|
||||||
|
component that saw all the spend while being blind to whether any of it produced
|
||||||
|
anything, so it measured cost when the thing that mattered was progress. Whether
|
||||||
|
its own `pinch` pruning drove the loops is still a theory (see #8, Theory 1). If
|
||||||
|
confirmed, #8 joins the family of defensible settings, a 50k budget chosen for
|
||||||
|
~130k sessions, that turned harmful when the input changed shape.
|
||||||
|
|
||||||
The generalisable fix is the same each time: **compute the thing that is
|
The generalisable fix is the same each time: **compute the thing that is
|
||||||
actually true, and surface it where the operator is already looking.** That is
|
actually true, and surface it where the operator is already looking.** That is
|
||||||
what the `/metrics` warnings and the inline admin warning are for, and it is the
|
what the `/metrics` warnings and the inline admin warning are for, and it is the
|
||||||
|
|||||||
408
docs/watchdog.md
Normal file
408
docs/watchdog.md
Normal file
@@ -0,0 +1,408 @@
|
|||||||
|
> Deep dive into the watchdog subsystem — opencode session loop detection. Back to
|
||||||
|
> [README](../README.md).
|
||||||
|
|
||||||
|
Watchdog is a periodic scanner that probes running opencode sessions for looping
|
||||||
|
patterns — repeated tool calls without progress. It runs as a systemd oneshot
|
||||||
|
every 5 minutes (`deploy/llm-router-watchdog.timer`) and, when a session flags,
|
||||||
|
writes into a local SQLite table and fires desktop notifications through the
|
||||||
|
configured channel pipeline.
|
||||||
|
|
||||||
|
The detector itself is a pure module (`src/progress_detect.py`) that returns a
|
||||||
|
boolean verdict plus a reason dict. The orchestrator (`src/watchdog.py`) reads
|
||||||
|
opencode session data over HTTP, runs the detector, and manages the alert
|
||||||
|
state machine. The notifier (`src/notifier.py`) routes events through
|
||||||
|
configured channels to desktop alerts (`notify-send`).
|
||||||
|
|
||||||
|
## What it catches and doesn't
|
||||||
|
|
||||||
|
Watchdog looks for **looping** — sessions that make repeated tool calls against
|
||||||
|
the same targets without landing meaningful changes. It detects four signal
|
||||||
|
types:
|
||||||
|
|
||||||
|
- **Dup signal**: a call appears more than 25% of the time in the sliding
|
||||||
|
window. The call matcher first normalises `bash` commands (strips env var
|
||||||
|
prefixes and `cd` prefixes, joins lines, collapses comments), then checks
|
||||||
|
tool name + canonicalised JSON args (sorted keys) for equality.
|
||||||
|
|
||||||
|
- **Top signal**: a single target is called 12+ times in the window (or 8+
|
||||||
|
for read-only agents like explore/librarian/oracle). Targets merge calls on
|
||||||
|
the same `(tool, basename)` for file reads, `(bash, first-two-words)` for
|
||||||
|
commands, and `(tool, pattern[:60])` for grep/glob.
|
||||||
|
|
||||||
|
- **Slow signal**: a single tool call accounts for 15+ of the session's total
|
||||||
|
calls, AND no call in the entire session history has landed (tree write or
|
||||||
|
git commit). This catches sessions stuck on a single read/analysis.
|
||||||
|
|
||||||
|
- **Coverage signal**: the session rereads one file at 4x its line count
|
||||||
|
within the window. The coverage score divides total bytes read by file length,
|
||||||
|
capped at the file's actual lines. A file read 4x its length means the model
|
||||||
|
is rereading without making progress.
|
||||||
|
|
||||||
|
**NOT steady spend.** A session that steadily calls different files on a
|
||||||
|
difficult refactor is not flagged — each call lands a unique `(tool, args)`
|
||||||
|
and the top target never crosses the threshold. The detector only fires when
|
||||||
|
repetition outpaces the call budget, not when a session is quietly burning
|
||||||
|
tokens across many distinct tools.
|
||||||
|
|
||||||
|
### Landed calls
|
||||||
|
|
||||||
|
A call counts as **landed** when it:
|
||||||
|
- is an `edit`, `write`, or `patch` with a non-empty `metadata.diff` field, or
|
||||||
|
- is a `bash` command containing `git commit` with `metadata.exit == 0`
|
||||||
|
|
||||||
|
Landed calls break both the slow and coverage signals. A session that reads a
|
||||||
|
file 4x, then edits it once, clears all signals immediately — the landed check
|
||||||
|
exits early and the verdict returns `False`.
|
||||||
|
|
||||||
|
### Ancestral trees
|
||||||
|
|
||||||
|
Sessions can be parented: the opencode API returns `parentID` on child sessions.
|
||||||
|
Watchdog resolves root sessions and evaluates EACH session in the tree on its
|
||||||
|
own calls, NOT merged into the root. A child's flagged verdict bubbles up to the
|
||||||
|
root level, which is what the alert dedup key uses (`opencode-loop:{root_session_id}`).
|
||||||
|
|
||||||
|
For the landed-time check, a session sees NOT only its own landed calls but also
|
||||||
|
its transitive descendants' landed calls. This prevents a parent session from
|
||||||
|
being flagged just because its child landed — even though the parent may still
|
||||||
|
be looping independently.
|
||||||
|
|
||||||
|
## Signals and thresholds
|
||||||
|
|
||||||
|
Thresholds live in `DetectConfig` (`src/progress_detect.py:38-53`) and are
|
||||||
|
configurable under `watchdog.detector.*` in `config/config.yaml`.
|
||||||
|
|
||||||
|
| Key | Default | Meaning |
|
||||||
|
|---|---|---|
|
||||||
|
| `watchdog.detector.dup_min` | 0.25 | Fraction: calls must repeat at this rate to flag |
|
||||||
|
| `watchdog.detector.top_min` | 12 | Minimum calls to a single target (normal agents) |
|
||||||
|
| `watchdog.detector.top_min_ro` | 8 | Minimum calls to a single target (read-only agents) |
|
||||||
|
| `watchdog.detector.cum_min` | 15 | Same call must account for this many total calls |
|
||||||
|
| `watchdog.detector.cover_min` | 4.0 | File must be reread this many times its length |
|
||||||
|
| `watchdog.detector.window` | 60 | Sliding window in number of tool calls (not seconds) |
|
||||||
|
| `watchdog.detector.min_calls` | 40 | Minimum total calls before evaluation triggers |
|
||||||
|
| `watchdog.detector.read_only_agent_keywords` | `["explore", "librarian", "oracle"]` | Title keywords that make an agent read-only (lowered `top_min` from 12 to 8) |
|
||||||
|
|
||||||
|
Read-only agents get a lowered top threshold because agents whose job is
|
||||||
|
exploration or library work are expected to call the same targets repeatedly —
|
||||||
|
the floor is 8 instead of 12.
|
||||||
|
|
||||||
|
## Alert lifecycle
|
||||||
|
|
||||||
|
The alert state machine is per-root-session and tracked in the
|
||||||
|
`watchdog_alerts` table.
|
||||||
|
|
||||||
|
### Trigger (initial detection)
|
||||||
|
|
||||||
|
When the detector flags a root session's tree for the first time:
|
||||||
|
|
||||||
|
1. The watchdog writes a `watchdog_alerts` row with `state='open'`.
|
||||||
|
2. The `local_llm_enabled` path asks a local Ollama model (defaulting to
|
||||||
|
`verification.model` or `qwen2.5-coder-router:14b`) for a second opinion.
|
||||||
|
It receives the agent name and the reason dict as context, and the model
|
||||||
|
must respond with only "yes" or "no" — no explanation.
|
||||||
|
3. If the local LLM answers "yes", severity is set to `critical`. If it
|
||||||
|
answers "no" or times out (returns `None`), severity is set to `warning`.
|
||||||
|
4. The notifier dispatches an `AlertEvent` through every matching channel.
|
||||||
|
5. The `flagged_ticks` counter starts at 1.
|
||||||
|
|
||||||
|
A cap of 3 simultaneous LLM second-opinion calls (`_MAX_LLM = 3`) prevents a
|
||||||
|
burst of flagged sessions from flooding the local model.
|
||||||
|
|
||||||
|
### Escalate (persistent looping)
|
||||||
|
|
||||||
|
On every subsequent tick where the session is still flagged:
|
||||||
|
|
||||||
|
1. `flagged_ticks` is incremented.
|
||||||
|
2. If `flagged_ticks >= 3` (15 minutes at 5-minute ticks) OR the LLM answers
|
||||||
|
"yes" again, the alert escalates to `severity='critical'` and the escalation
|
||||||
|
event is dispatched.
|
||||||
|
3. If the alert is already critical, ticking continues but no duplicate
|
||||||
|
escalation fires.
|
||||||
|
|
||||||
|
### Resolve (session is gone or fixed)
|
||||||
|
|
||||||
|
A root session alert resolves in two conditions:
|
||||||
|
|
||||||
|
- **Session disappeared**: the root session ID is no longer returned by the
|
||||||
|
opencode `/session` API. The watchdog writes `state='resolved'`,
|
||||||
|
`resolved_at=<now>`, and `severity='info'`.
|
||||||
|
- **Session fixed**: the root session's tree no longer has any flagged sessions
|
||||||
|
in the detector's verdict. The alert is resolved the same way and a resolve
|
||||||
|
event is dispatched.
|
||||||
|
|
||||||
|
Resolve events always bypass the rate limit — they must clear the previous
|
||||||
|
notification so the operator knows the alert is over.
|
||||||
|
|
||||||
|
### No-opencode
|
||||||
|
|
||||||
|
When no `rc-servers.json` or no opencode server answers, the tick writes an
|
||||||
|
`outcome='no_opencode'` record and returns immediately with zero alerts. This
|
||||||
|
uses the filesystem lock (`router.db/.watchdog.lock`) so two watchdog instances
|
||||||
|
cannot run simultaneously.
|
||||||
|
|
||||||
|
## Admin surfaces
|
||||||
|
|
||||||
|
Watchdog data surfaces through three admin pages, each serving a different
|
||||||
|
operational question.
|
||||||
|
|
||||||
|
### Loops panel — index.html#loops
|
||||||
|
|
||||||
|
On the main dashboard (`admin/frontend/index.html`), the **Loops** card
|
||||||
|
(id="loops") shows all currently open alerts:
|
||||||
|
|
||||||
|
- **Severity badge**: the alert's current severity (`critical` in red,
|
||||||
|
`warning` in orange-yellow, `info` in grey)
|
||||||
|
- **Dedup key**: the raw `opencode-loop:{root_session_id}` string, so the
|
||||||
|
operator can cross-reference with `journalctl` or `router.db`
|
||||||
|
- **Opened at**: `opened_at` ISO timestamp from the first trigger
|
||||||
|
- State: `open` or `resolved`; the panel filters to `resolved_at IS NULL`
|
||||||
|
|
||||||
|
The panel is client-side only — it polls `/api/watchdog/loops` which queries
|
||||||
|
`watchdog_alerts` for all unresolved rows.
|
||||||
|
|
||||||
|
### Watchdog card — controls.html
|
||||||
|
|
||||||
|
The Controls page (`admin/frontend/controls.html`) includes a dedicated
|
||||||
|
**Watchdog** card with operational status:
|
||||||
|
|
||||||
|
- **Last tick**: timestamp of the most recent `watchdog_ticks` row
|
||||||
|
- **Sessions seen**: how many opencode sessions were enumerated on that tick
|
||||||
|
- **Flagged verdicts**: count of flagged + total from the latest tick's
|
||||||
|
`watchdog_verdicts` rows
|
||||||
|
- **Open alerts**: count from `watchdog_alerts WHERE resolved_at IS NULL`
|
||||||
|
- **Notification channels**: list of configured channels with toggle controls
|
||||||
|
(enabled/min_severity), editable via `POST /api/watchdog/channels`
|
||||||
|
- **Refresh** button: reloads the watchdog status from the API
|
||||||
|
- **Send test alert** button: fires a test `AlertEvent` through enabled channels
|
||||||
|
|
||||||
|
The status endpoints are:
|
||||||
|
- `GET /api/watchdog/status` — last tick, verdict counts, open alert count
|
||||||
|
- `GET /api/watchdog/channels` — per-channel settings from `watchdog_channel_settings`
|
||||||
|
- `POST /api/watchdog/channels` — upsert channel settings (enabled, min_severity)
|
||||||
|
- `POST /api/watchdog/test-alert` — deliver a test event (severity + optional
|
||||||
|
channel_name filter) to a live Notifier instance
|
||||||
|
|
||||||
|
### Per-model stall rollup — models.html
|
||||||
|
|
||||||
|
The `watchdog_verdicts` table has `model_id`/`provider` columns and indexes on
|
||||||
|
them, but no per-model stall rollup, no Block button beside evidence, and no
|
||||||
|
Blocked list are built yet. `blocked` is only a value in the Models override
|
||||||
|
dropdown (see docs/admin-portal.md).
|
||||||
|
|
||||||
|
## Database schema
|
||||||
|
|
||||||
|
Four tables live in `router.db`, created idempotently by both
|
||||||
|
`config/schema.sql` and `watchdog_store.py`:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
watchdog_ticks -- one row per watchdog tick
|
||||||
|
watchdog_verdicts -- per-session verdict at each tick (joined to ticks)
|
||||||
|
watchdog_alerts -- alert lifecycle state machine (dedup_key PK)
|
||||||
|
watchdog_channel_settings -- per-channel toggle + severity gate
|
||||||
|
```
|
||||||
|
|
||||||
|
Indexes on `watchdog_verdicts` cover `model_id`, `created_at`, `flagged`, and
|
||||||
|
`session_root` — the columns the admin endpoint queries most frequently.
|
||||||
|
|
||||||
|
The ticks table carries `ticked_at`, `sessions_seen`, and `outcome`
|
||||||
|
(`'ok'`, `'flagged'`, or `'no_opencode'`). The verdicts table joins to ticks
|
||||||
|
via `tick_id` and carries the full reason dict fields: `dup`, `top`,
|
||||||
|
`top_what`, `landed`, `slow`, `coverage`, plus `calls_since_landed`,
|
||||||
|
`cost_since_landed_usd`, and the LLM second-opinion answer.
|
||||||
|
|
||||||
|
Alerts are keyed by `dedup_key` (deduplicated per root session) and carry
|
||||||
|
`opened_at`, `last_fired_at`, and `resolved_at`. The transition logic lives in
|
||||||
|
the `_fire_alert()` helper: trigger inserts only if no open row exists,
|
||||||
|
escalate increments `flagged_ticks` and bumps severity, and resolve writes
|
||||||
|
`resolved_at` and downgrades to `severity='info'`.
|
||||||
|
|
||||||
|
## Knobs
|
||||||
|
|
||||||
|
All watchdog configuration lives under `watchdog:` in `config/config.yaml`:
|
||||||
|
|
||||||
|
| Key | Default | Required |
|
||||||
|
|---|---|---|
|
||||||
|
| `watchdog.enabled` | `true` | Gates the entire subsystem |
|
||||||
|
| `watchdog.local_llm_enabled` | `true` | Whether to call Ollama for second opinions |
|
||||||
|
| `watchdog.model` | `null` (uses `verification.model`) | Local model name |
|
||||||
|
| `watchdog.detector.window` | 60 | Sliding window size in calls |
|
||||||
|
| `watchdog.detector.dup_min` | 0.25 | Dup threshold (fraction) |
|
||||||
|
| `watchdog.detector.top_min` | 12 | Top target threshold (normal agents) |
|
||||||
|
| `watchdog.detector.top_min_ro` | 8 | Top target threshold (read-only agents) |
|
||||||
|
| `watchdog.detector.cum_min` | 15 | Slow-signal threshold |
|
||||||
|
| `watchdog.detector.cover_min` | 4.0 | Coverage-signal multiplier |
|
||||||
|
| `watchdog.detector.min_calls` | 40 | Minimum eval calls |
|
||||||
|
| `watchdog.read_only_agents` | `["explore", "librarian", "oracle"]` | Agent title keywords for reduced threshold |
|
||||||
|
| `watchdog.dashboard_base_url` | `"http://127.0.0.1:8080/admin"` | Base URL for alert links |
|
||||||
|
|
||||||
|
All seven `watchdog.detector.*` keys plus `watchdog.enabled` and
|
||||||
|
`watchdog.local_llm_enabled` (nine allowlisted keys total) are allowlisted for
|
||||||
|
live editing through the
|
||||||
|
admin portal (`POST /admin/api/config/{key}`), which writes to
|
||||||
|
`config.local.yaml` (the gitignored machine-local overlay). The changes are
|
||||||
|
validated by `RouterConfig` before reaching disk, and take effect at the next
|
||||||
|
tick — no service restart required since `detect_config_from_pydantic()` reads
|
||||||
|
cfg fresh each invocation.
|
||||||
|
|
||||||
|
## Install
|
||||||
|
|
||||||
|
The watchdog ships as two systemd user units:
|
||||||
|
|
||||||
|
- `llm-router-watchdog.service` — a `Type=oneshot` that runs
|
||||||
|
`PYTHONPATH=%h/llm-router/src .venv/bin/python -m watchdog --once`
|
||||||
|
- `llm-router-watchdog.timer` — fires 5 min after boot, then every 5 min
|
||||||
|
|
||||||
|
Installation follows the same pattern as the other deploy units. The shipped
|
||||||
|
unit files use `%h/llm-router` placeholders that need rewriting to your actual
|
||||||
|
repo path:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
REPO=$(pwd)
|
||||||
|
for u in deploy/llm-router-watchdog.{service,timer}; do
|
||||||
|
sed "s|%h/llm-router|${REPO}|g" "$u" \
|
||||||
|
> ~/.config/systemd/user/"$(basename "$u")"
|
||||||
|
done
|
||||||
|
systemctl --user daemon-reload
|
||||||
|
systemctl --user enable --now llm-router-watchdog.timer
|
||||||
|
```
|
||||||
|
|
||||||
|
The service unit has **no `EnvironmentFile`** — watchdog needs no provider API
|
||||||
|
keys because it only queries the local opencode session APIs and the local
|
||||||
|
Ollama instance.
|
||||||
|
|
||||||
|
Check installation:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
systemctl --user status llm-router-watchdog.timer
|
||||||
|
journalctl --user -u llm-router-watchdog.service --no-pager
|
||||||
|
```
|
||||||
|
|
||||||
|
## Run --once
|
||||||
|
|
||||||
|
For ad-hoc debugging or pre-deploy smoke test:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
PYTHONPATH=src .venv/bin/python -m watchdog --once
|
||||||
|
```
|
||||||
|
|
||||||
|
The `--once` flag runs a single tick and exits 0 once the tick has run,
|
||||||
|
however many alerts it fired (the count is logged as `watchdog tick done: N
|
||||||
|
alert(s) fired`); systemd would otherwise mark the oneshot unit failed on every
|
||||||
|
real alert. It connects to the same `router.db` that the timed
|
||||||
|
service uses, acquires the filesystem lock, reads `rc-servers.json`, probes
|
||||||
|
opencode servers, and follows the full evaluate-alert-resolve pipeline.
|
||||||
|
|
||||||
|
Add `--config /path/to/config.yaml` to point at a non-default config file.
|
||||||
|
|
||||||
|
### Debug output
|
||||||
|
|
||||||
|
The `--once` run emits structured `watchdog=` log lines for every state
|
||||||
|
transition:
|
||||||
|
|
||||||
|
```
|
||||||
|
2026-01-15 14:30:01 INFO watchdog: watchdog=2026-01-15T14:30:01+00:00 dedup='opencode-loop:ses_xxxxx' state=trigger severity=warning title='opencode loop: explore'
|
||||||
|
```
|
||||||
|
|
||||||
|
The `title` field carries the agent name derived from the session title. The
|
||||||
|
`dedup` field is the root session's dedup key, matching `journalctl` traces
|
||||||
|
and admin panel rows.
|
||||||
|
|
||||||
|
## Backtest
|
||||||
|
|
||||||
|
`scripts/progress_backtest.py` replays labelled sessions through the detector
|
||||||
|
to validate threshold calibration. It reads from a fixture JSON file and runs
|
||||||
|
a sliding-window evaluation with step 5.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
PYTHONPATH=src .venv/bin/python -m scripts.progress_backtest --fixture \
|
||||||
|
tests/fixtures/progress/fixture.json
|
||||||
|
```
|
||||||
|
|
||||||
|
The fixture ships in `tests/fixtures/progress/fixture.json` and contains
|
||||||
|
15 labelled sessions: 8 marked `must_flag` and 7 marked `must_not_flag`.
|
||||||
|
Expected output after a successful run:
|
||||||
|
|
||||||
|
```
|
||||||
|
# backtest complete: 8/15 flagged
|
||||||
|
```
|
||||||
|
|
||||||
|
At least 8 sessions should trigger a flag. The fixture is generated from live
|
||||||
|
opencode sessions (see [export fixture](#re-export-fixture) below) and scrubbed
|
||||||
|
to protect file contents — content is SHA-1 hashed except for the few keys the
|
||||||
|
detector actually inspects (`filePath`, `command`, `pattern`, etc.).
|
||||||
|
|
||||||
|
## Re-export fixture
|
||||||
|
|
||||||
|
The fixture generator pulls sessions from the opencode SQLite database
|
||||||
|
(`~/.local/share/opencode/opencode.db`) and writes the scrubbed test fixture:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/export_progress_fixture.py
|
||||||
|
```
|
||||||
|
|
||||||
|
Output goes to `tests/fixtures/progress/fixture.json`. The script takes a
|
||||||
|
session ID to label map (`LABELS` dict), extracts call history, resolves file
|
||||||
|
line counts (via `git show` for deleted worktree paths), and scrubs all strings
|
||||||
|
to SHA-1 prefixes, preserving only the detector's target keys.
|
||||||
|
|
||||||
|
The fixture generation script hardcodes paths to the operator's home directory
|
||||||
|
and a specific worktree commit (`WORKTREE_COMMIT = "827c408"`). When
|
||||||
|
regenerating, update those paths to match the current environment.
|
||||||
|
|
||||||
|
## Notifier pipeline
|
||||||
|
|
||||||
|
Alert events flow through `src/notifier.py` before reaching the operator. Each
|
||||||
|
alert is an `AlertEvent` dataclass mirroring PagerDuty Events v2 shape:
|
||||||
|
`dedup_key`, `severity` (info/warning/critical), `state` (trigger/escalate/
|
||||||
|
resolve), `title`, `summary`, `details`, and `source` (hardcoded to
|
||||||
|
`"6krrt-watchdog"`).
|
||||||
|
|
||||||
|
Each configured channel has its own `min_severity` gate — a critical event
|
||||||
|
passes through a channel with `min_severity=warning`, but a warning event does
|
||||||
|
not pass a channel with `min_severity=critical`. Each channel also has a rate
|
||||||
|
limit window of 300 seconds (5 minutes). Resolve events always bypass the rate
|
||||||
|
limit to ensure clear-the-alert notifications are delivered.
|
||||||
|
|
||||||
|
Currently only `type=desktop` channels are implemented; they invoke
|
||||||
|
`notify-send -u {urgency} {title} {summary}`. Unknown channel types are
|
||||||
|
logged as warnings and skipped. Missing `notify-send` (no `DISPLAY` or
|
||||||
|
`libnotify` installed) is logged at warning level — the notifier never raises.
|
||||||
|
|
||||||
|
Channel settings (`watchdog_channel_settings`) are editable at runtime through
|
||||||
|
the admin portal and stored per-channel in the database, separate from the
|
||||||
|
base config.
|
||||||
|
|
||||||
|
## Known limits
|
||||||
|
|
||||||
|
- **Under 40 calls: invisible.** `min_calls=40` means the detector does not
|
||||||
|
evaluate sessions with fewer tool calls. Short troubleshooting sessions that
|
||||||
|
loop on 10 calls pass through undetected.
|
||||||
|
|
||||||
|
- **Atlas at cover_min 4.0 exactly.** A session whose last 60 calls read a
|
||||||
|
single file exactly 4x its line count hits the coverage threshold at the
|
||||||
|
boundary. There is no margin: `>=` comparison means exactly 4.0x flags.
|
||||||
|
This matters for sessions that genuinely need to re-read a reference file
|
||||||
|
repeatedly — the coverage signal cannot be tuned per-file.
|
||||||
|
|
||||||
|
- **Stale localhost:4096 probes.** Watchdog reads `rc-servers.json` from
|
||||||
|
`~/.local/share/opencode/rc-servers.json`, which may contain stale entries
|
||||||
|
for opencode sessions that have since ended. These produce empty call lists
|
||||||
|
and are silently skipped — but they add network round-trip latency to each
|
||||||
|
tick. Only the first answering server's sessions are evaluated (the watchdog
|
||||||
|
picks `answering[0]` from the list of servers that successfully return a
|
||||||
|
session list).
|
||||||
|
|
||||||
|
- **No streaming path.** `POST /outcome` is the only ground truth signal for
|
||||||
|
streaming traffic; the watchdog has no equivalent endpoint to report streaming
|
||||||
|
session outcomes. It can only see tool-call repetition, not whether the answer
|
||||||
|
was useful.
|
||||||
|
|
||||||
|
- **Local LLM gate.** Second-opinion calls are gated by both
|
||||||
|
`local_compute.enabled` (global) and `watchdog.local_llm_enabled`
|
||||||
|
(per-subsystem). If either is false, no LLM calls are made and severity stays
|
||||||
|
as `warning` on initial trigger regardless of model quality. The cap of 3
|
||||||
|
simultaneous LLM calls means that if 5 sessions flag, 2 wait without opinion.
|
||||||
|
|
||||||
|
- **No provider calls.** Watchdog does not touch the dispatch path. It only
|
||||||
|
reads opencode session data and runs local heuristics. No provider calls, no
|
||||||
|
quota consume, no cost incurred.
|
||||||
@@ -13,7 +13,7 @@
|
|||||||
"auto": {
|
"auto": {
|
||||||
"name": "auto (router picks, interactive)",
|
"name": "auto (router picks, interactive)",
|
||||||
"limit": {
|
"limit": {
|
||||||
"context": 782324,
|
"context": 200000,
|
||||||
"output": 16384
|
"output": 16384
|
||||||
},
|
},
|
||||||
"modalities": {
|
"modalities": {
|
||||||
|
|||||||
397
plans/cockpit-brainstorm.md
Normal file
397
plans/cockpit-brainstorm.md
Normal file
@@ -0,0 +1,397 @@
|
|||||||
|
# Cockpit brainstorm: the router as an operator's instrument panel
|
||||||
|
|
||||||
|
Status: reference -- brainstorm, not a queue item; its six quick wins shipped in PR #101
|
||||||
|
|
||||||
|
**Date:** 2026-09-23
|
||||||
|
**Read against:** the repo tarball as of this date (code, config, plans,
|
||||||
|
screenshots). **Not** read against the live `router.db`, so every "gap" below
|
||||||
|
is a gap in what the code surfaces, not a claim about what the data shows.
|
||||||
|
Section 5 is the exception: it was read against the live portal.
|
||||||
|
|
||||||
|
## The framing
|
||||||
|
|
||||||
|
6krrt as a local LLM router with a cockpit. The operator:
|
||||||
|
|
||||||
|
1. adjusts dials and knobs in flight,
|
||||||
|
2. sees where tokens are wasted,
|
||||||
|
3. watches proficiency, and
|
||||||
|
4. sees poor results from a carrier or model surface, so they can decide
|
||||||
|
whether to preclude it (with the circuit breaker as the automatic version
|
||||||
|
of that decision for outages).
|
||||||
|
|
||||||
|
The router makes the per-request decision. The cockpit is where the operator
|
||||||
|
makes the slower decisions: which models and carriers are trusted, for what,
|
||||||
|
and at what settings. Under this framing, the question for every feature is
|
||||||
|
whether it helps the operator notice something and act on it.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## What the cockpit already has
|
||||||
|
|
||||||
|
| job | what exists | where |
|
||||||
|
|---|---|---|
|
||||||
|
| in-flight dials | ~11 runtime knobs (bool/float/int tables in `admin.py`), persisted config writes via `/api/config/{key}`, classifier card, profiles CRUD, gaming mode | Controls, Profiles |
|
||||||
|
| knob coverage | North Star rule #1 + `test_admin_knob_coverage.py` | tests |
|
||||||
|
| token waste | pinch savings card; `cache_rate_series` + warnings; incumbent cache pricing dial; `cost_estimate_calibration` in `/metrics` | Dashboard, Controls, `/metrics` |
|
||||||
|
| proficiency | model × category matrix with evidence status (measured / inherited / thin, n per cell); client outcome log with attributable + applied flags | Proficiency |
|
||||||
|
| poor results | `content_fault_warnings` (malformed output rate); `rejection_warnings`; availability override (active / deprecated / stale); allowlist editor | warnings bell, Models, Providers |
|
||||||
|
| breaker | passive per-(model, provider) breaker on 5xx; on/off knob | Controls (toggle only) |
|
||||||
|
|
||||||
|
The foundation is good. What's mostly missing is the link between a signal
|
||||||
|
and the action it should prompt, plus any record of which actions were taken.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Cross-cutting ideas (these help all four jobs)
|
||||||
|
|
||||||
|
### X1. Flight recorder: a change log
|
||||||
|
|
||||||
|
**Gap.** `admin_schema.sql` has only `admin_model_overrides`. Nothing records
|
||||||
|
when a knob, override, allowlist entry, profile or feedback fold changed, what
|
||||||
|
it changed from, or why. Runtime knobs are "NEVER persisted", so a restart
|
||||||
|
silently reverts them, and nothing notes that the revert happened.
|
||||||
|
|
||||||
|
**Idea.**
|
||||||
|
- An `admin_changes` table: `(at, surface, key, old, new, reason, source)`,
|
||||||
|
where `source` is `runtime | persisted | override | allowlist | profile |
|
||||||
|
feedback_fold | restart`. Every admin write path appends a row.
|
||||||
|
- An `operator_epoch` id stamped on each `route_decisions` row, incremented
|
||||||
|
on each **operator action** (a row in `admin_changes`). Every metric can then
|
||||||
|
be split before and after a change without joining on timestamps. Scope it
|
||||||
|
to operator actions on purpose: "any change that affects routing" would also
|
||||||
|
cover catalog price polls and every feedback fold, so the epoch would tick
|
||||||
|
constantly and split nothing.
|
||||||
|
- Change markers drawn on the dashboard history charts.
|
||||||
|
- A drift badge when a runtime value differs from its persisted twin
|
||||||
|
("this will revert on restart").
|
||||||
|
- Restart reverts are loggable without persisting runtime knobs: the change
|
||||||
|
log already holds each knob's last runtime value, so at startup, compare it
|
||||||
|
with the loaded config and write one `restart` row per value discarded.
|
||||||
|
|
||||||
|
**Why first.** Adjusting knobs in flight produces no learning if you can't
|
||||||
|
later tell what a change did. The other ideas here lean on this one.
|
||||||
|
|
||||||
|
### X2. Structured warnings with actions attached
|
||||||
|
|
||||||
|
**Gap.** `/metrics` warnings are plain strings. The dashboard decides which
|
||||||
|
page fixes a warning, and how severe it is, by regex on the text
|
||||||
|
(`warningTarget` and `warningSeverity` in `index.html`). Rewording a warning
|
||||||
|
silently breaks its link and its severity.
|
||||||
|
|
||||||
|
**Idea.** Emit `{class, severity, subject: {model, provider, category?},
|
||||||
|
evidence_url, actions: [...]}`. The warning registry in
|
||||||
|
`test_tui_warnings.py` already lists every class, so it can serve as the
|
||||||
|
schema. The bell can then offer an action in place: "exclude from
|
||||||
|
`tool_use_agentic`", "quarantine 24h", "open decisions filtered to this
|
||||||
|
model".
|
||||||
|
|
||||||
|
### X3. Preview before apply (replay)
|
||||||
|
|
||||||
|
**Gap.** The reviews replayed thousands of real decisions to test a setting,
|
||||||
|
but only as one-off scripts pasted into plan documents.
|
||||||
|
|
||||||
|
**Idea.** For any change that affects ranking (`quality_tolerance`, the
|
||||||
|
incumbent dial, profile edits, availability overrides, graded exclusions
|
||||||
|
from P2 below), the portal replays the last N decisions through
|
||||||
|
`select_candidates` and `rank_candidates` with the proposed value, then shows:
|
||||||
|
- a winner-shift matrix: old winner → new winner, with counts,
|
||||||
|
- estimated cost delta,
|
||||||
|
- decisions that would become unroutable (the 2026-09-01 and 2026-09-04
|
||||||
|
incidents were both deprecations that emptied a candidate set).
|
||||||
|
|
||||||
|
This is cheap because the ranking modules are pure. Its output is an estimate
|
||||||
|
of routing change, not quality change, and should be labeled that way.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Dials in flight
|
||||||
|
|
||||||
|
- **D1. Timed changes.** "Apply for 2h, then revert." Turns a knob change into
|
||||||
|
a bounded experiment and lowers the risk of forgetting a test setting.
|
||||||
|
Needs X1 to record both the apply and the revert.
|
||||||
|
- **D2. Blast radius on hover.** How many of the last 24h of decisions this
|
||||||
|
knob would have touched. It's the X3 replay reduced to one number.
|
||||||
|
- **D3. Session-scoped A/B.** Assign new sessions to setting A or B and
|
||||||
|
compare cost and outcome rate per arm. It has to be per session, not per
|
||||||
|
turn, because a mid-session switch costs cache (token-waste Wave 2
|
||||||
|
measured 0.919 → 0.348). Larger build; only worth it once X1 exists and
|
||||||
|
outcome volume can support a comparison.
|
||||||
|
- **D4. Grouping by effect.** Group controls by what they move (cost,
|
||||||
|
quality, latency, safety, measurement-only) rather than by config section,
|
||||||
|
and mark the few that actually change routing. Knob coverage guarantees
|
||||||
|
every control exists; this makes them usable.
|
||||||
|
|
||||||
|
## 2. Token waste
|
||||||
|
|
||||||
|
**Gap.** Waste signals are spread across pinch, cache rate, calibration and
|
||||||
|
the verifications table, and none is in dollars in one place.
|
||||||
|
`cost_estimate_calibration` is in `/metrics` but not in the portal.
|
||||||
|
|
||||||
|
- **W1. Waste ledger.** One panel with dollars per waste class over a window,
|
||||||
|
each with a trend:
|
||||||
|
| class | how it's computed |
|
||||||
|
|---|---|
|
||||||
|
| switch cache loss | billed on switch turns − same prompt at the session's same-model cache rate |
|
||||||
|
| retries | re-billed prompt tokens from `iteration.py` attempts |
|
||||||
|
| paid for failure | billed cost of responses later reported `ok:false` via `/outcome` |
|
||||||
|
| empty-200 / malformed | billed cost of responses with a structural `malformed` verdict |
|
||||||
|
| truncation | responses stopped by `finish_reason: length` with no client cap |
|
||||||
|
| pinch (negative) | dollars saved, from `pinch_summary` |
|
||||||
|
|
||||||
|
"Paid for failure" counts only decisions that received an outcome, so it
|
||||||
|
shows n and outcome coverage (the share of billed decisions with a report)
|
||||||
|
next to the dollar figure.
|
||||||
|
- **W2. Why did it switch?** Record a switch reason on each decision: category
|
||||||
|
changed, breaker open, exploration, override removed the incumbent, profile
|
||||||
|
change. Then rank sessions by switch cost and drill into the turns. Without
|
||||||
|
a reason, a switch is only a cost; with one, it points to a knob.
|
||||||
|
- **W3. Calibration panel.** Surface `cost_estimate_calibration` per model,
|
||||||
|
headlining the **spread** (as its docstring argues), and flag pairs where
|
||||||
|
the estimator and the bill disagree on order. Where that happens, the cost
|
||||||
|
tiebreak picks the wrong model.
|
||||||
|
- **W4. Cost per successful outcome.** Per (model, provider, category):
|
||||||
|
billed dollars ÷ `ok:true` outcomes. A cheap model that fails often isn't
|
||||||
|
cheap. This is the one waste figure that includes quality.
|
||||||
|
|
||||||
|
Show n and outcome coverage next to it. Models that carry more traffic
|
||||||
|
collect more outcomes (the same exposure bias the feedback fold had to
|
||||||
|
correct), and if coverage stays hidden, a thinly reported cheap model will
|
||||||
|
look better than it is.
|
||||||
|
|
||||||
|
## 3. Proficiency monitoring
|
||||||
|
|
||||||
|
**Gap.** The matrix shows current state well. It doesn't show movement, and
|
||||||
|
it doesn't show which thin cells actually matter.
|
||||||
|
|
||||||
|
- **P1. Trend and drift per cell.** A sparkline of the outcome rate over time.
|
||||||
|
Compare a recent window against the long-run rate with an interval, and
|
||||||
|
flag drops that fall outside it. Carriers change serving setups (quant,
|
||||||
|
engine, hardware) without notice, and a drift flag is how that shows up.
|
||||||
|
- **P2. Evidence priority.** Rank cells by `traffic share × uncertainty`. A
|
||||||
|
thin cell that routing never consults doesn't matter; a thin cell deciding
|
||||||
|
30% of traffic does. The ranking tells you where to spend `eval_proficiency`
|
||||||
|
runs or exploration budget.
|
||||||
|
- **P3. The tolerance band per category.** Show which models sit within
|
||||||
|
`quality_tolerance` of the leader, so it's clear where cost is deciding and
|
||||||
|
where quality is.
|
||||||
|
- **P4. Label provenance per cell.** The share of each cell's outcomes whose
|
||||||
|
category came from a fresh classification versus a cached or borrowed one.
|
||||||
|
North Star rule #2 exists because borrowed labels trained the matrix;
|
||||||
|
this makes it visible cell by cell.
|
||||||
|
|
||||||
|
## 4. Poor results → operator preclusion (and the breaker)
|
||||||
|
|
||||||
|
**Gap.** The operator's only preclusion tools are binary: an availability
|
||||||
|
override (active / deprecated / stale) or removing a model from the
|
||||||
|
allowlist. The breaker is in-memory, trips only on 5xx, forgets everything
|
||||||
|
on restart, and shows nothing but an on/off toggle.
|
||||||
|
|
||||||
|
- **Q1. Carrier/model scorecard.** One sortable table per (model, provider)
|
||||||
|
over a window: outcome fail rate (with n), malformed / empty-200 rate,
|
||||||
|
5xx and timeout count, breaker trips, p95 latency, estimate-vs-bill error,
|
||||||
|
cache rate, cost per success (W4). This is the "who's misbehaving" view the
|
||||||
|
preclusion decision needs.
|
||||||
|
- **Q2. Same model, different carriers.** Group scorecard rows by base model,
|
||||||
|
e.g. `qwen3.6-35b` on NeuralWatt vs OpenRouter. This separates "the model is
|
||||||
|
bad at this" from "this carrier serves it badly", which calls for a
|
||||||
|
different fix: drop the carrier's row, or restrict the model.
|
||||||
|
- **Q3. Graded actions** in place of deprecate-or-not. Each action records a
|
||||||
|
reason and links its evidence in X1:
|
||||||
|
1. exclude from category X. This **cannot** reuse `eligible_categories`:
|
||||||
|
that field is restrict-only (`admin.py`, the probe docstring), and NULL
|
||||||
|
means "every category", so a deny needs its own column,
|
||||||
|
2. exclude when the request carries tools (the `deepseek-v4-flash` case).
|
||||||
|
This may be the way out of the frozen `tool_use_agentic` data:
|
||||||
|
`min_tool_proficiency` can only read scores that stopped moving on
|
||||||
|
2026-09-15, while a per-model manual rule lets the operator decide from
|
||||||
|
Q1 scorecard evidence instead,
|
||||||
|
3. restrict to batch / `-flex` only,
|
||||||
|
4. quarantine: exploration-only, so it keeps collecting evidence without
|
||||||
|
carrying traffic (see the open question below; deferred until Q1 shows
|
||||||
|
it is needed),
|
||||||
|
5. timed ban, e.g. 24h, auto-expiring and logged,
|
||||||
|
6. deprecate (what exists today).
|
||||||
|
|
||||||
|
Each action goes through at least the minimal X3 check (would this empty a
|
||||||
|
candidate set for any recent decision?) before committing. That check alone
|
||||||
|
would have caught both the 2026-09-01 and 2026-09-04 incidents; the full
|
||||||
|
winner-shift replay can come later.
|
||||||
|
- **Q4. Breaker visibility.** Show open circuits with `down_until`, current
|
||||||
|
cooldown and trip count; keep a trip history (persist it, since `_store`
|
||||||
|
is process-lifetime); add manual force-open (drain a model on purpose) and
|
||||||
|
force-close (reset after a known fix).
|
||||||
|
- **Q5. A quality breaker, separate from availability.** Trip on content
|
||||||
|
failures (empty-200, mangled output, a burst of `ok:false`) using the
|
||||||
|
novelty-or-rate rule already in `rejection_warnings`. This is roughly what
|
||||||
|
`plans/mangled-output-detection.md` specs (status: planned). Keep it a
|
||||||
|
separate breaker and panel so a quality trip doesn't read as an outage.
|
||||||
|
**One doc owns it:** fold Q5 into `mangled-output-detection.md` rather than
|
||||||
|
specifying it twice, and let this file point there.
|
||||||
|
- **Q6. Carrier-level breaker.** When several models on one provider trip
|
||||||
|
inside a short window, open the provider rather than walking its models one
|
||||||
|
cooldown at a time. Account exhaustion already has its own path; this is for
|
||||||
|
partial carrier outages.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Quality of life: ergonomics and readouts
|
||||||
|
|
||||||
|
Unlike the sections above, this one **was** read against the live portal
|
||||||
|
(8080, view-only, 2026-09-23), plus the frontend source. Each item names the
|
||||||
|
thing observed.
|
||||||
|
|
||||||
|
### Readouts that mislead today
|
||||||
|
|
||||||
|
- **R1. Home tiles use different windows.** Quota is "this period",
|
||||||
|
Decisions 7d, Models and Pinch 30d. Busiest models shows `kimi-k2.7-code`
|
||||||
|
at 10,028 while the Decisions tile says 2,887 for everything, and both are
|
||||||
|
correct. Label each tile's window on its face, or drive all of them from
|
||||||
|
the Activity card's 24h / 7d / 30d selector.
|
||||||
|
- **R2. The Decisions tile counts verifications, and mixes diagnostics with
|
||||||
|
ground truth.** It reads "2,887 verified in 7 days: 132 ok, 93 failed", but
|
||||||
|
`index.html` sums `verdict_mix`, which counts `verifications` rows, not
|
||||||
|
decisions. Checked read-only against the live DB on 2026-09-23:
|
||||||
|
| | tile | actually |
|
||||||
|
|---|---|---|
|
||||||
|
| headline | 2,887 | 2,678 route decisions; 2,887 is verification rows, 2,660 of them `unverifiable` |
|
||||||
|
| "ok" | 132 | 96 client `succeeded` + 29 `local_llm` ok + 7 structural ok |
|
||||||
|
| "failed" | 93 | 82 client `failed` + 11 `malformed` (local_llm + structural) |
|
||||||
|
The tile adds structural and `local_llm` verdicts, which CLAUDE.md calls
|
||||||
|
diagnostics only, to `/outcome` reports, the only ground truth. Show route
|
||||||
|
decisions as the headline, then client outcomes on their own: "178 client
|
||||||
|
reports (6.6%): 96 ok, 82 failed (46%)". Diagnostics, if shown at all, go on
|
||||||
|
a separate line.
|
||||||
|
- **R3. The Activity chart smooths across gaps.** The decisions series is a
|
||||||
|
spline through sparse points, so it draws continuous traffic across hours
|
||||||
|
that had none. Break the line at empty buckets or use bars. Separately,
|
||||||
|
billed requests exceed decisions at several peaks; the chart should say what
|
||||||
|
the excess is (cloud classifier calls? retries? unrouted pins?) rather than
|
||||||
|
leave two lines that disagree unexplained.
|
||||||
|
- **R4. Sparklines have no values.** The tile sparklines carry no axis and no
|
||||||
|
hover. Add a hover value, or min and max labels.
|
||||||
|
- **R5. The Classifier card flashes a wrong value.** On first paint the mode
|
||||||
|
select reads `local_llm` and only switches to `local_encoder` once config
|
||||||
|
loads. For a second, the card shows a value that isn't configured. Render
|
||||||
|
a loading state instead of the default.
|
||||||
|
- **R6. Config echo vs live state.** The Classifier card shows what the
|
||||||
|
overlay says (`local_encoder`, device `cpu`), not what the process actually
|
||||||
|
loaded. Show the resolved model, device and recent p50 latency from the
|
||||||
|
running classifier. (It also exposed doc drift: CLAUDE.md says the live
|
||||||
|
deployment runs `device: cuda`; the overlay says `cpu`.)
|
||||||
|
- **R7. Proficiency colour encodes evidence, not score.** Green / amber mean
|
||||||
|
measured / thin, so the best model in a column isn't visible without reading
|
||||||
|
every number. Mark each column's leader, outline the `quality_tolerance`
|
||||||
|
band (P3), make columns sortable, and add a "routable only" toggle so rows
|
||||||
|
that are deprecated or not allowlisted stop padding the matrix.
|
||||||
|
|
||||||
|
### Decisions page ergonomics
|
||||||
|
|
||||||
|
- **E1. Collapse runs.** The top 40 rows are one session: same category,
|
||||||
|
profile, tier, source and model, and only ctx and cost move. Fold
|
||||||
|
consecutive same-session, same-model rows into one expandable row: "38 turns,
|
||||||
|
ctx 75k to 98k, $0.27 total". Switches then stand out as row boundaries,
|
||||||
|
which is what W2 wants to surface.
|
||||||
|
- **E2. Session view.** A per-session rollup: turns, total cost, context
|
||||||
|
growth, switches, outcome count. Context climbing about 1k per turn is
|
||||||
|
visible in the raw table and invisible everywhere else.
|
||||||
|
- **E3. Filters in the URL.** `decisions.html` never reads `URLSearchParams`,
|
||||||
|
so no other page can link to "decisions for this model" or "this session".
|
||||||
|
X2 actions, the Q1 scorecard and the model modal all need that link.
|
||||||
|
- **E4. Row drill-down.** A row click does nothing. Open a panel with the
|
||||||
|
decision's rejected candidates, verification verdict, `/outcome` report,
|
||||||
|
and estimated vs billed cost from its energy observation.
|
||||||
|
- **E5. Server-side filtering.** The page loads 1,000 of 32,540 rows and
|
||||||
|
filters and searches only those, with a footnote saying so. Push filters to
|
||||||
|
the API so that "All" means all.
|
||||||
|
- **E6. Model, provider and source as filters.** Today they're reachable only
|
||||||
|
through free-text search.
|
||||||
|
|
||||||
|
### Controls page
|
||||||
|
|
||||||
|
- **C1. One knob table, not two.** Runtime Knobs and Persisted Config list
|
||||||
|
largely the same knobs in two columns, in different orders, with persisted
|
||||||
|
labels truncated (`objective.incumbent_cache_pri…`). Checking drift means
|
||||||
|
matching rows by eye. Use one table: knob, live value, persisted value,
|
||||||
|
layer (base / overlay), drift badge. That table is also X1's natural home.
|
||||||
|
- **C2. Separate Restart Service.** It sits in the same button row as Refresh
|
||||||
|
Catalog and Seed Energy. Move it apart, and have its confirmation list which
|
||||||
|
runtime values differ from persisted and will revert.
|
||||||
|
- **C3. Prose in cards.** Local Compute carries three paragraphs; the portal
|
||||||
|
style rule is no paragraphs. Keep one line plus a docs link or tooltip.
|
||||||
|
- **C4. Default and last change per knob.** Show the code default beside
|
||||||
|
each value, plus "changed 2h ago from 0.3" once X1 exists.
|
||||||
|
|
||||||
|
### Portal-wide
|
||||||
|
|
||||||
|
- **G1. Nav lives in eight files.** Each page hardcodes the same `<ul>` of
|
||||||
|
nav links. `navbar.js` already notes that adding Quota was an eight-file
|
||||||
|
edit; have it render the links too.
|
||||||
|
- **G2. Width.** Most pages sit in `container-xl` and use well under half of
|
||||||
|
a wide screen, while Decisions, the densest table, is the most cramped.
|
||||||
|
Proficiency already goes full-width. Go fluid on the data-heavy pages.
|
||||||
|
- **G3. Dismissals don't stick on rate warnings.** Dismissal is keyed on the
|
||||||
|
warning's exact text, on purpose (a changed condition resurfaces it). But
|
||||||
|
warnings that embed a live rate ("3.2x pace") change text every poll, so a
|
||||||
|
dismissal never holds. X2's `class` + `subject` is the right key; resurface
|
||||||
|
on a severity change, not a digit change.
|
||||||
|
- **G4. Keyboard.** `/` to focus search, `Esc` to close modals, and `j`/`k`
|
||||||
|
on the Decisions table. Cheap, and this is a page someone lives in.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## One thing the cockpit framing makes more urgent
|
||||||
|
|
||||||
|
The portal is loopback-only with no auth, and it already includes
|
||||||
|
`/api/restart-service`, `/api/apply-feedback` and provider deletes. The more
|
||||||
|
it becomes the place decisions are made, the stronger the pull to reach it
|
||||||
|
from other machines, and moving the router onto a Proxmox box on the LAN is
|
||||||
|
exactly that move. Add auth before binding anything but `127.0.0.1`. Same
|
||||||
|
warning CLAUDE.md already gives for the API, with more at stake.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## A possible order
|
||||||
|
|
||||||
|
| # | item | why here |
|
||||||
|
|---|---|---|
|
||||||
|
| 0 | portal auth | **only if** the router is leaving `127.0.0.1` (the Proxmox move); blocks that move, not this list |
|
||||||
|
| 1 | X1 change log + operator epoch | everything else reads it |
|
||||||
|
| 2 | Q1 scorecard + Q4 breaker visibility | the preclusion decision needs one view first |
|
||||||
|
| 3 | Q3 graded actions (via X1) + minimal X3 (empties-a-candidate-set check) | turns the scorecard into decisions without repeating 09-01 / 09-04 |
|
||||||
|
| 4 | X2 structured warnings | lets warnings trigger Q3 actions directly |
|
||||||
|
| 5 | W1 waste ledger + W2 switch reasons | token waste in dollars, with causes |
|
||||||
|
| 6 | X3 full replay preview (winner shift, cost delta) | makes knob changes safe to try |
|
||||||
|
| 7 | P1 drift, P2 evidence priority | proficiency monitoring over time |
|
||||||
|
| 8 | Q5 quality breaker (owned by `mangled-output-detection.md`), D1 timed changes | automate what the operator has been doing by hand |
|
||||||
|
| 9 | D3 session A/B, Q6 carrier breaker | only after the above prove out |
|
||||||
|
|
||||||
|
Section 5 sits outside this order: most of it is small and independent.
|
||||||
|
Quick wins that also unblock the numbered items: **E3** (URL filters, needed
|
||||||
|
by X2 and Q1), **R2** (outcome coverage, the same honesty W1 and W4 need),
|
||||||
|
**C1** (one knob table, X1's home), **G3** (fixed by X2's class keys). R5 and
|
||||||
|
G1 are fixes of a few lines each.
|
||||||
|
|
||||||
|
**North Star #1 applies to every item here.** Q5 thresholds, D1 durations,
|
||||||
|
timed-ban lengths and any quarantine settings are new config knobs, so each
|
||||||
|
ships with its admin control or `test_admin_knob_coverage.py` fails. Scope the
|
||||||
|
control into the item, not a follow-up.
|
||||||
|
|
||||||
|
### Open questions, with proposed answers
|
||||||
|
|
||||||
|
- **Is the change log append-only history, or also an undo stack?**
|
||||||
|
Proposed: append-only, plus a per-row "revert to old value" that writes one
|
||||||
|
key and appends its own row. That is not the shape CLAUDE.md rejected for
|
||||||
|
gaming mode: that concern was one flag writing five keys and drifting apart.
|
||||||
|
A one-key revert can't drift, and the history stays intact.
|
||||||
|
- **Should graded exclusions live in the DB or in `config.local.yaml`?**
|
||||||
|
Proposed: the DB, next to `admin_model_overrides` (extend it or add a
|
||||||
|
sibling table). Timed bans need an expiry, and an expiry belongs in a DB
|
||||||
|
column. The overlay is per-machine and not in git, so it's the wrong home
|
||||||
|
for judgments about a carrier. And availability overrides already live in
|
||||||
|
the DB. The category exclusion still needs a new deny column (see Q3.1).
|
||||||
|
- **Does quarantine conflict with session-scoped exploration?** Yes, and the
|
||||||
|
weight goes up: a quarantined model would carry whole sessions, not single
|
||||||
|
turns. There's a second problem too. `exploration.py` only chooses among
|
||||||
|
models that already passed the hard filters and were ranked, so quarantine
|
||||||
|
needs a new state, "eligible to explore, barred from winning", rather than
|
||||||
|
reusing anything existing. Defer it until Q1 shows it's needed.
|
||||||
217
plans/cockpit-quick-wins.md
Normal file
217
plans/cockpit-quick-wins.md
Normal file
@@ -0,0 +1,217 @@
|
|||||||
|
# Cockpit quick wins: six small admin-portal fixes
|
||||||
|
|
||||||
|
Status: done -- shipped in PR #101
|
||||||
|
Date: 2026-09-23
|
||||||
|
Source: `plans/cockpit-brainstorm.md` section 5 (items R2, R5, E3, C1, G1, G3).
|
||||||
|
Each item below was observed on the live portal or checked read-only against
|
||||||
|
the live `router.db` on 2026-09-23.
|
||||||
|
|
||||||
|
## Ground rules (all items)
|
||||||
|
|
||||||
|
- **Worktree, not the main checkout.** Create a worktree from `main` HEAD
|
||||||
|
(`c87e675` or later) on branch `feat/cockpit-quick-wins`. The main checkout
|
||||||
|
at `/home/alee/Sources/6krrt` has UNCOMMITTED edits from other work,
|
||||||
|
including `admin/frontend/controls.html` (a `confidence_threshold` ->
|
||||||
|
`confidence_min` rename around lines 1026 and 1150) and
|
||||||
|
`tests/test_admin_frontend.py`. Do not touch, stash, or commit them. C1 edits
|
||||||
|
a different region of `controls.html`; keep it that way so the later merge
|
||||||
|
is clean.
|
||||||
|
- Commit by explicit path. Never `git add -A` or `git add .`.
|
||||||
|
- **Never touch port 8080** (production). For visual checks, run a throwaway
|
||||||
|
instance on **8081** from the worktree, with a temp copy of the DB, never the
|
||||||
|
live `router.db` for writes.
|
||||||
|
- Never edit `config/config.yaml` or `config/config.local.yaml`.
|
||||||
|
- ASCII only in new code, comments, and UI strings. No middle-dot separators
|
||||||
|
(U+00B7); relate facts with layout, or with `:` or a rephrase.
|
||||||
|
- No paragraphs of prose in the UI. One line per hint at most.
|
||||||
|
- One commit per item, in the order listed. `pytest` (offline) and
|
||||||
|
`ruff check` stay green after each commit.
|
||||||
|
- For every UI change, take a screenshot on 8081 and check the changed area
|
||||||
|
zoomed in, not just the item's checklist.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. R5: Classifier card shows a wrong value on first paint
|
||||||
|
|
||||||
|
**Problem.** `admin/frontend/controls.html:264-268`: the
|
||||||
|
`#classifier-mode-select` `<select>` renders with `local_llm` selected by
|
||||||
|
default (first `<option>`). Until `loadClassifierConfig` (sets `.value` at
|
||||||
|
~line 1116) returns, the card claims `local_llm` when the live mode is
|
||||||
|
`local_encoder`.
|
||||||
|
|
||||||
|
**Fix.** Render the select disabled, with a leading
|
||||||
|
`<option value="" selected>loading...</option>`, and the Save button
|
||||||
|
disabled. On load: remove the placeholder option, set the value, enable both.
|
||||||
|
On fetch failure: keep them disabled and show an inline one-line error.
|
||||||
|
|
||||||
|
**Test.** In `tests/test_admin_frontend.py`: assert that the static HTML of the
|
||||||
|
select has a selected placeholder and the `disabled` attribute, and that no
|
||||||
|
real mode option is `selected` in the static markup.
|
||||||
|
|
||||||
|
## 2. G1: Nav links hardcoded in eight files
|
||||||
|
|
||||||
|
**Problem.** Every page (`index.html`, `models.html`, `profiles.html`,
|
||||||
|
`proficiency.html`, `decisions.html`, `quota.html`, `controls.html`,
|
||||||
|
`providers.html`) hardcodes the same `<ul>` of seven `nav-item` links, and
|
||||||
|
sets `active` by hand. `navbar.js` (the header comment) already records that
|
||||||
|
adding Quota was an eight-file edit.
|
||||||
|
|
||||||
|
**Fix.** Define the link list once in `navbar.js`
|
||||||
|
(`[{href, label}]`, same order as today: Models, Profiles, Proficiency,
|
||||||
|
Decisions, Quota, Controls, Providers). Render it into a placeholder
|
||||||
|
(e.g. `<ul class="navbar-nav" id="nav-links"></ul>`) on each page. Mark
|
||||||
|
`active` by matching `location.pathname`, with `/admin` and `/admin/` mapping
|
||||||
|
to Home (no link active). Remove the hardcoded `<li>`s from all eight pages.
|
||||||
|
|
||||||
|
Keep the rendered markup identical: same classes (`nav-item`, `nav-link`,
|
||||||
|
`active`), same hrefs, same order. Pages must still render their nav when
|
||||||
|
`navbar.js` is loaded, and it is already loaded on every page.
|
||||||
|
|
||||||
|
**Test.** One test that parses all eight HTML files and asserts none contains
|
||||||
|
a hardcoded `class="nav-link" href="/admin/...` list. One test asserting that
|
||||||
|
the `navbar.js` link list has exactly the seven hrefs above.
|
||||||
|
|
||||||
|
## 3. E3: Decisions filters are not in the URL
|
||||||
|
|
||||||
|
**Problem.** `admin/frontend/decisions.html` never reads or writes
|
||||||
|
`URLSearchParams`. No other page can link to "decisions for this model" or
|
||||||
|
"this session", and a filtered view cannot be shared or reloaded.
|
||||||
|
|
||||||
|
**Fix.**
|
||||||
|
- On load, after the filter `<select>`s are populated (~lines 405-420), read
|
||||||
|
`kind`, `category`, `profile`, `tier`, `q` (search) and `size` from
|
||||||
|
`location.search` and apply them. If a value is not among a select's
|
||||||
|
options, add it as an option rather than dropping it (a link may name a
|
||||||
|
category that has no rows in the loaded window).
|
||||||
|
- On any filter change (`resetPageAndRender`, the Clear button, page size),
|
||||||
|
write the current filters with `history.replaceState`. Omit empty values.
|
||||||
|
Don't push history entries.
|
||||||
|
- `q` also accepts a model id or a session id, since search already matches
|
||||||
|
both (`search` at ~line 433).
|
||||||
|
|
||||||
|
**Test.** A static test that `decisions.html` reads `URLSearchParams` and calls
|
||||||
|
`history.replaceState`. On 8081, open
|
||||||
|
`/admin/decisions?category=diff_checking&q=deepseek`, confirm the filters and
|
||||||
|
the rows match, and screenshot it.
|
||||||
|
|
||||||
|
## 4. R2: The Decisions tile counts verifications and mixes diagnostics with ground truth
|
||||||
|
|
||||||
|
**Problem.** `admin/frontend/index.html:597-610`: the Home "Decisions" tile
|
||||||
|
sums `_snap.verdict_mix`, which is `metrics.verdict_mix()` and counts
|
||||||
|
`verifications` rows by verdict. Live DB, last 7 days:
|
||||||
|
|
||||||
|
| verdict | kind | n |
|
||||||
|
|---|---|---|
|
||||||
|
| unverifiable | structural | 2,660 |
|
||||||
|
| succeeded | client_outcome | 96 |
|
||||||
|
| failed | client_outcome | 82 |
|
||||||
|
| ok | local_llm | 29 |
|
||||||
|
| ok | structural | 7 |
|
||||||
|
| malformed | local_llm / structural | 4 / 7 |
|
||||||
|
| truncated | structural | 2 |
|
||||||
|
|
||||||
|
The tile shows "2,887 verified in 7 days: 132 ok, 93 failed". Actual route
|
||||||
|
decisions in 7 days: **2,678**. "ok" adds client `succeeded` to structural and
|
||||||
|
`local_llm` ok. "failed" adds client `failed` to `malformed`. Per CLAUDE.md,
|
||||||
|
structural and `local_llm` verdicts are diagnostics only; `/outcome`
|
||||||
|
(`kind = 'client_outcome'`) is the only ground truth.
|
||||||
|
|
||||||
|
**Fix.**
|
||||||
|
- Backend: in `src/metrics.py`, add `decision_outcome_summary(conn,
|
||||||
|
since_days=7)` returning
|
||||||
|
`{decisions, client_reports, client_ok, client_failed}`:
|
||||||
|
`decisions` = `route_decisions` rows with `observed_at` inside the window;
|
||||||
|
`client_*` = `verifications` rows with `kind = 'client_outcome'` inside the
|
||||||
|
window (`succeeded` = ok, `failed` = failed). Include it in the admin
|
||||||
|
snapshot (`src/admin.py` near line 2131) as `decision_outcomes`. Leave
|
||||||
|
`verdict_mix` unchanged; the TUI uses it.
|
||||||
|
- Frontend: the tile headline becomes `decisions`. Qualifier lines:
|
||||||
|
- `routed in 7 days`
|
||||||
|
- `178 client reports (6.6%): 96 ok, 82 failed` (coverage =
|
||||||
|
`client_reports / decisions`; the failed count uses the warn colour when
|
||||||
|
the fail rate among reports is above 10%, matching the existing Signal
|
||||||
|
quality card threshold).
|
||||||
|
- With zero reports: `no client reports in 7 days`.
|
||||||
|
- Drop the word "verified".
|
||||||
|
- Use `metrics.py`'s existing window idiom:
|
||||||
|
`julianday(observed_at) > julianday('now', '-' || ? || ' days')`.
|
||||||
|
|
||||||
|
**Test.** `tests/`: a metrics unit test with a seeded in-memory DB that
|
||||||
|
includes structural, `local_llm` and `client_outcome` rows, asserting only
|
||||||
|
`client_outcome` rows are counted and that `decisions` comes from
|
||||||
|
`route_decisions`. A frontend static test that the tile no longer reads
|
||||||
|
`verdict_mix`.
|
||||||
|
|
||||||
|
## 5. G3: Warning dismissals don't stick on rate warnings
|
||||||
|
|
||||||
|
**Problem.** `admin/frontend/index.html:1010-1040`: a dismissal is keyed on
|
||||||
|
the warning's exact text. That's deliberate ("when the underlying numbers move
|
||||||
|
the text moves with them... the warning comes back"). But warnings that embed
|
||||||
|
a live rate or count ("3.2x pace", "n=4") change text on nearly every poll, so
|
||||||
|
a dismissal never holds.
|
||||||
|
|
||||||
|
**Fix (interim, until structured warnings exist).** Key each dismissal on
|
||||||
|
`digit-normalized text + severity`: replace every run of digits (and decimal
|
||||||
|
points inside a number) with `#`, then append `|` + `warningSeverity(text)`.
|
||||||
|
The warning comes back when its wording or its severity changes, not when a
|
||||||
|
digit moves. This is the same normalization `rejection_warnings` already uses
|
||||||
|
for grouping (CLAUDE.md, the reactive detector section).
|
||||||
|
- Migrate the existing stored keys on load: normalize each and dedupe, same
|
||||||
|
pattern as the existing one-time carry-over from `dismissedWarnings`.
|
||||||
|
- Update the block comment to state the new rule and why it changed.
|
||||||
|
|
||||||
|
**Test.** A static or JS-extracted unit test: two texts differing only in
|
||||||
|
digits produce the same key; the same text at different severity produces a
|
||||||
|
different key.
|
||||||
|
|
||||||
|
## 6. C1: One knob table instead of two columns
|
||||||
|
|
||||||
|
**Problem.** `admin/frontend/controls.html`: "Runtime Knobs" (`#runtime-list`,
|
||||||
|
`renderRuntime` ~line 687) and "Persisted Config" (`#config-list`,
|
||||||
|
`renderConfig` ~line 821) list largely the same knobs in two side-by-side
|
||||||
|
cards, in different orders, and persisted labels truncate
|
||||||
|
(`objective.incumbent_cache_pri...`). Checking whether a runtime value has
|
||||||
|
drifted from its persisted value means matching rows by eye.
|
||||||
|
|
||||||
|
**Fix.**
|
||||||
|
- One card, **Knobs**, full width, as a table:
|
||||||
|
`knob | live | persisted | layer | ` with the columns meaning:
|
||||||
|
- `knob`: the dotted config key, never truncated (wrap if needed).
|
||||||
|
- `live`: the runtime control, if the knob has one, else empty.
|
||||||
|
- `persisted`: the persisted-config control, if the key is in the persisted
|
||||||
|
allowlist, else empty.
|
||||||
|
- `layer`: the existing `base` / `overlay` badge.
|
||||||
|
- A **drift** badge in the row when live differs from persisted
|
||||||
|
("reverts on restart").
|
||||||
|
- Pair rows through the runtime-to-config mapping that already exists in
|
||||||
|
`src/admin.py` (`_BOOL_KNOBS` at ~line 348, plus the non-bool runtime knob
|
||||||
|
tables beside it). Expose that mapping to the frontend: add the dotted
|
||||||
|
config key to each item in the `GET /admin/api/runtime` response, rather
|
||||||
|
than duplicating the table in JS.
|
||||||
|
- Keep the existing save semantics exactly: runtime controls POST to
|
||||||
|
`api/runtime/{knob}` immediately, as today; persisted controls stay dirty
|
||||||
|
until the Save button, as today. Keep the existing hint badges (for example
|
||||||
|
"enables the reuse window").
|
||||||
|
- Keep the help lines ("Live on the next request...", "Applied on
|
||||||
|
restart...") as one line each, placed as column-header tooltips or a single
|
||||||
|
line under the table.
|
||||||
|
- Don't touch the classifier-card region (lines ~1000-1160); another branch
|
||||||
|
edits it.
|
||||||
|
|
||||||
|
**Test.** `test_admin_knob_coverage.py` must still pass unchanged. Add a
|
||||||
|
test that every runtime knob in the API response carries a config key, and
|
||||||
|
that the key resolves to a real `RouterConfig` path. Visual check on 8081:
|
||||||
|
set one runtime knob so it differs from persisted (on the 8081 instance
|
||||||
|
only), then confirm the drift badge shows. Screenshot the whole table and a
|
||||||
|
zoomed crop of one drifted row.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Done means
|
||||||
|
|
||||||
|
- Six commits on `feat/cockpit-quick-wins`, one per item, in the order above.
|
||||||
|
- `pytest` and `ruff check` green on the worktree.
|
||||||
|
- Screenshots from 8081 for items 1, 3, 4 and 6.
|
||||||
|
- The 8081 instance stopped, by the PID you started.
|
||||||
|
- A short report listing each commit, the tests added, and anything
|
||||||
|
deferred, with the reason.
|
||||||
167
plans/deploy-separation.md
Normal file
167
plans/deploy-separation.md
Normal file
@@ -0,0 +1,167 @@
|
|||||||
|
# Deploy separation: production runs from its own clone, not the dev checkout
|
||||||
|
|
||||||
|
Status: planned -- spec and runbook ready, dry run done (see "Dry run"); the cutover is an operator step not yet run
|
||||||
|
Date: 2026-09-26
|
||||||
|
Why: incident #8, and #4, #5 and #7 before it. Production has been one `git pull`,
|
||||||
|
one uncommitted edit or one agent session away from the dev checkout.
|
||||||
|
|
||||||
|
## Current state (measured 2026-09-26)
|
||||||
|
|
||||||
|
- **Every installed user unit runs from `~/Sources/6krrt`**, the dev checkout:
|
||||||
|
- `llm-router.service`
|
||||||
|
- `llm-router-poller`, `-seed`, `-backup`, `-offsite`
|
||||||
|
- `llm-router-baseline-report`
|
||||||
|
- **What they take from it:** `WorkingDirectory`, `PYTHONPATH`,
|
||||||
|
`EnvironmentFile=.env`, `.venv/bin/*`, and `HF_HOME=.hf-cache`.
|
||||||
|
`ReadWritePaths` covers the whole source tree, so the live service can write
|
||||||
|
anywhere in the repo.
|
||||||
|
- **Everything else follows the working directory,** because the code resolves
|
||||||
|
paths against it: `config/config.yaml`, the `config/config.local.yaml`
|
||||||
|
overlay, and `database.path: "router.db"`. Moving the working directory moves
|
||||||
|
all of it, with no code change.
|
||||||
|
- **The repo's own unit templates** (`deploy/*.service`) already target a
|
||||||
|
separate `%h/llm-router` directory, except `-backup` and `-offsite`. The split
|
||||||
|
was the original design. The installed units drifted onto the dev checkout,
|
||||||
|
and `~/llm-router` does not exist.
|
||||||
|
- **Sizes:** `router.db` 48 MB, `.hf-cache` 2.8 GB, `.venv` 7.2 GB. `.env`
|
||||||
|
holds `NEURALWATT_API_KEY` and `OPENROUTER_API_KEY`.
|
||||||
|
|
||||||
|
## Target layout
|
||||||
|
|
||||||
|
```
|
||||||
|
~/.local/share/6krrt/app production clone of the Gitea repo, detached at a
|
||||||
|
merged origin/main commit. Nobody edits here; no
|
||||||
|
agent works here; nothing here ever runs `git clean`.
|
||||||
|
router.db live DB (gitignored; checkout never touches it)
|
||||||
|
config/config.local.yaml live overlay (gitignored)
|
||||||
|
.env live keys, mode 600 (gitignored)
|
||||||
|
.venv/ production venv, built from the clone's requirements
|
||||||
|
~/.cache/6krrt-hf shared HF model cache (moved once from
|
||||||
|
~/Sources/6krrt/.hf-cache; production and the dev
|
||||||
|
sandbox both point HF_HOME here)
|
||||||
|
~/Sources/6krrt dev checkout: source only. No live DB, no
|
||||||
|
production .env. The 8081 sandbox uses DB copies.
|
||||||
|
```
|
||||||
|
|
||||||
|
Units: every `llm-router*` unit's paths change from `%h/Sources/6krrt` to
|
||||||
|
`%h/.local/share/6krrt/app`, and `ReadWritePaths` narrows to that directory.
|
||||||
|
The backup and offsite scripts already take `REPO` from the environment; set it
|
||||||
|
to the app directory.
|
||||||
|
|
||||||
|
## `scripts/deploy.sh` (to build; not needed for the first cutover)
|
||||||
|
|
||||||
|
The only way code reaches production after the cutover:
|
||||||
|
1. **Show the change.** `git -C $APP fetch origin`, then
|
||||||
|
`git log --oneline HEAD..origin/main`.
|
||||||
|
2. **Preflight, without touching the running service:**
|
||||||
|
- load `origin/main`'s config against the real overlay (catches a #7-style
|
||||||
|
rename)
|
||||||
|
- check the requirements are satisfied in `$APP/.venv`
|
||||||
|
- `import dispatcher` from a temp export of `origin/main`
|
||||||
|
3. **Cut over.** Record the current SHA, `git checkout --detach origin/main`,
|
||||||
|
`uv pip install` only if the requirements changed, restart, and poll
|
||||||
|
`/health` for 30 s.
|
||||||
|
4. **Roll back on failure.** Check out the recorded SHA, restart, and alert.
|
||||||
|
5. **Record the deployed SHA** in `$APP/.deployed`, and expose it on `/health`
|
||||||
|
if that endpoint gains a field, so "what is running" has one answer.
|
||||||
|
|
||||||
|
## Cutover runbook
|
||||||
|
|
||||||
|
Downtime is about 1 minute. The venv is built BEFORE stopping anything. Run
|
||||||
|
the commands from a shell, not from an agent session.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
APP=~/.local/share/6krrt/app
|
||||||
|
DEV=~/Sources/6krrt
|
||||||
|
|
||||||
|
# 0. Build the clone and venv while production keeps running.
|
||||||
|
mkdir -p ~/.local/share/6krrt
|
||||||
|
git clone ssh://git@git.adlee.work:2222/alee/6krrt.git "$APP"
|
||||||
|
git -C "$APP" checkout --detach origin/main
|
||||||
|
(cd "$APP" && uv venv .venv && uv pip install --python .venv/bin/python \
|
||||||
|
-r requirements.txt -r requirements-encoder.txt)
|
||||||
|
|
||||||
|
# 1. Share the HF cache (one move; the dev sandbox follows in step 6).
|
||||||
|
mkdir -p ~/.cache && mv "$DEV/.hf-cache" ~/.cache/6krrt-hf
|
||||||
|
|
||||||
|
# 2. Stop production and every timer that touches the DB.
|
||||||
|
systemctl --user stop llm-router-poller.timer llm-router-seed.timer \
|
||||||
|
llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer
|
||||||
|
systemctl --user stop llm-router.service
|
||||||
|
|
||||||
|
# 3. Copy state; originals stay put until step 7.
|
||||||
|
sqlite3 "$DEV/router.db" ".backup '$APP/router.db'"
|
||||||
|
command cp -p "$DEV/config/config.local.yaml" "$APP/config/config.local.yaml"
|
||||||
|
command cp -p "$DEV/.env" "$APP/.env" && chmod 600 "$APP/.env"
|
||||||
|
sqlite3 "$APP/router.db" 'PRAGMA integrity_check;' # must print: ok
|
||||||
|
|
||||||
|
# 4. Repoint the units (backups first), then reload.
|
||||||
|
cd ~/.config/systemd/user
|
||||||
|
mkdir -p ~/.local/share/6krrt/unit-backup && command cp -p llm-router* ~/.local/share/6krrt/unit-backup/ -r
|
||||||
|
sed -i "s#%h/Sources/6krrt#%h/.local/share/6krrt/app#g; s#/home/alee/Sources/6krrt#/home/alee/.local/share/6krrt/app#g" llm-router*.service
|
||||||
|
sed -i "s#^Environment=HF_HOME=.*#Environment=HF_HOME=%h/.cache/6krrt-hf#" llm-router.service
|
||||||
|
systemctl --user daemon-reload
|
||||||
|
|
||||||
|
# 5. Start and verify.
|
||||||
|
systemctl --user start llm-router.service
|
||||||
|
sleep 3; curl -s localhost:8080/health
|
||||||
|
systemctl --user show llm-router.service -p WorkingDirectory
|
||||||
|
curl -s localhost:8080/admin/api/runtime | head -c 200 # pinch persisted false = overlay loaded
|
||||||
|
systemctl --user start llm-router-poller.timer llm-router-seed.timer \
|
||||||
|
llm-router-backup.timer llm-router-offsite.timer llm-router-baseline-report.timer
|
||||||
|
|
||||||
|
# 6. Point the dev sandbox at the shared HF cache and the new live DB.
|
||||||
|
# (edit scripts/sandbox.sh: HF_HOME and the live-DB path; see "Follow-ups")
|
||||||
|
|
||||||
|
# 7. Only after a day of healthy running: retire the dev copies.
|
||||||
|
mv "$DEV/router.db" "$DEV/router.db.pre-cutover-$(date +%Y%m%d)"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Rollback, before step 7:** restore the unit backups from
|
||||||
|
`~/.local/share/6krrt/unit-backup/`, run `daemon-reload`, start the service.
|
||||||
|
The dev copies of the DB, overlay and `.env` were never moved, so production is
|
||||||
|
back where it started. Any decisions logged after the cutover stay in the
|
||||||
|
app-directory DB.
|
||||||
|
|
||||||
|
## Follow-ups that must land with or right after the cutover
|
||||||
|
|
||||||
|
- **`CLAUDE.md` North Star #3** and the "Measure the live `router.db`" memory
|
||||||
|
name `/home/alee/Sources/6krrt/router.db` as the live DB. Update both to
|
||||||
|
`~/.local/share/6krrt/app/router.db`, or agents will measure a stale file.
|
||||||
|
- **`scripts/sandbox.sh`:** `HF_HOME` goes to `~/.cache/6krrt-hf`. Its DB copy
|
||||||
|
instructions change to the app directory's DB.
|
||||||
|
- **`deploy/*.service` templates:** make the repo's templates say
|
||||||
|
`%h/.local/share/6krrt/app` so reinstalling from the repo reproduces this
|
||||||
|
layout, and fix `-backup` and `-offsite` to match.
|
||||||
|
- **`deploy/README.md`:** replace the install section with this layout and
|
||||||
|
`deploy.sh`.
|
||||||
|
- **The admin portal's "Restart Service" trigger** keeps working (it calls
|
||||||
|
`systemctl`), but any trigger that runs a script by repo-relative path should
|
||||||
|
be checked against the new working directory.
|
||||||
|
|
||||||
|
## Dry run
|
||||||
|
|
||||||
|
See the result appended below.
|
||||||
|
|
||||||
|
**Result, 2026-09-26 04:3x, against `origin/main` `827c408`:**
|
||||||
|
- **Clone from Gitea and detached checkout:** OK.
|
||||||
|
- **Fresh venv, `uv venv` plus both requirements files:** 2 s (all from `uv`'s
|
||||||
|
cache); `torch 2.9.1+cu128` and `transformers` import. Building the venv
|
||||||
|
does not affect downtime, and step 0 can run hours ahead.
|
||||||
|
- **State:** the DB copied by sqlite backup with the source opened `mode=ro`,
|
||||||
|
and `integrity_check` = ok; overlay copied; `.env` replaced with dummy keys
|
||||||
|
for the dry run.
|
||||||
|
- **Served on 8081** from the clone. `/proc/<pid>/cwd` confirmed the clone
|
||||||
|
directory, with `HF_HOME` on the dev cache read-only:
|
||||||
|
- `/health` 200; `/admin/controls` 200
|
||||||
|
- `/admin/api/runtime` showed `pinch_enabled` `persisted: False`, so the
|
||||||
|
copied overlay loaded
|
||||||
|
- the snapshot read the copied DB: 3,994 decisions, 200 client reports over 7
|
||||||
|
days
|
||||||
|
- `/route` considered 41 candidates and picked `deepseek/deepseek-v4-flash`
|
||||||
|
- **Stopped** by the recorded PID; 8081 free; the dummy `.env` deleted.
|
||||||
|
|
||||||
|
**Not exercised by the dry run:** the systemd unit edits (step 4); the timers
|
||||||
|
under the new `ReadWritePaths`; the HF cache move (step 1); and real provider
|
||||||
|
keys. The unit edit is the step to watch during the cutover. `systemctl --user
|
||||||
|
show llm-router.service -p WorkingDirectory` in step 5 confirms it took.
|
||||||
105
plans/docs-lift-a-pass.md
Normal file
105
plans/docs-lift-a-pass.md
Normal file
@@ -0,0 +1,105 @@
|
|||||||
|
# Docs pass after Lift A (PR #102)
|
||||||
|
|
||||||
|
Status: done -- executed on branch docs/lift-a-docs-pass
|
||||||
|
Date: 2026-09-26
|
||||||
|
This brief is the contract. Keep the plan small: one todo per doc, each with
|
||||||
|
exact section anchors. Do NOT read `plans/no-progress-detection.md` (reference
|
||||||
|
only).
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Bring the docs in line with what PR #102 shipped and what is running now. The
|
||||||
|
docs describe the code as it is on this branch; where a doc and the code
|
||||||
|
disagree, the code wins, and the todo says which line it checked.
|
||||||
|
|
||||||
|
## Where the work happens
|
||||||
|
|
||||||
|
- Worktree: `/home/alee/Sources/6krrt-worktrees/docs-lift-a-pass`, branch
|
||||||
|
`docs/lift-a-docs-pass`, based on `origin/main` `a24de1b` plus three commits
|
||||||
|
that already carry North Star #4 in `CLAUDE.md`, the incident #8 rewrite in
|
||||||
|
`docs/incidents.md`, and the plans. Build on them; do not rewrite them.
|
||||||
|
- **Docs only.** No code, config, test or unit-file changes.
|
||||||
|
- **Do not touch** `README.md`, `AGENTS.md`, `deploy/README.md`: the owner has
|
||||||
|
uncommitted edits to them in the main checkout. Anything that belongs in
|
||||||
|
`deploy/README.md` goes in `docs/watchdog.md` instead, and the done report
|
||||||
|
lists it for the owner to fold in.
|
||||||
|
|
||||||
|
## What shipped (verify each against the code before writing it)
|
||||||
|
|
||||||
|
- `src/progress_detect.py`: pure detector. Signals dup, top, slow, coverage;
|
||||||
|
landed gating; read-only agents; canonical (sorted-key) args; bash command
|
||||||
|
normalization (comment lines, `cd dir &&`, `VAR=x` prefixes dropped).
|
||||||
|
- `src/watchdog.py` + `deploy/llm-router-watchdog.{service,timer}`: 5-minute
|
||||||
|
oneshot. Judges only sessions with a tool call since the last tick, on their
|
||||||
|
full history, per session; landed counts the session's own descendants only.
|
||||||
|
Alert lifecycle in `watchdog_alerts`: trigger once, escalate once (LLM yes or
|
||||||
|
3 ticks), resolve once. Quiet `no_opencode` tick when opencode is closed.
|
||||||
|
Optional local-model second opinion (does not veto).
|
||||||
|
- `src/notifier.py`: `desktop` channel (`notify-send`), PagerDuty-Events-v2
|
||||||
|
shaped `AlertEvent`, per-channel `min_severity`, rate limit backed by the DB.
|
||||||
|
- `src/watchdog_store.py`: tables `watchdog_ticks`, `watchdog_verdicts`,
|
||||||
|
`watchdog_alerts`, `watchdog_channel_settings`.
|
||||||
|
- Admin: Loops panel (`index.html#loops`), Watchdog card on Controls (last
|
||||||
|
tick, verdicts, channels, Send test alert), per-model stall rollup and
|
||||||
|
Block/Unblock on Models. New availability value `blocked`, excluded from
|
||||||
|
routing and from pinned requests via `_admin_excluded_models`.
|
||||||
|
- Endpoints: `GET /admin/api/watchdog/status`, `POST /admin/api/watchdog/channels`,
|
||||||
|
`POST /admin/api/watchdog/test-alert`; availability endpoint accepts `blocked`.
|
||||||
|
- Config: `watchdog.*` (incl. `detector.*`, `dashboard_base_url`),
|
||||||
|
`notifications.channels`. Watchdog knobs are persisted-only (separate process).
|
||||||
|
- `deploy/opencode-plugin/router-link.js`: the `parentCache` export made
|
||||||
|
opencode 1.18 reject the whole plugin ("Plugin export is not a function" in
|
||||||
|
`~/.local/share/opencode/log/opencode.log`), so identity headers never
|
||||||
|
reached the router from #99 until this fix. A loader-contract test now asserts
|
||||||
|
every export is a function.
|
||||||
|
- `scripts/progress_backtest.py` (`--fixture` replays the 15 labelled sessions;
|
||||||
|
expect 8/15 flagged) and `scripts/export_progress_fixture.py`.
|
||||||
|
- Operator state 2026-09-26: timer enabled ~15:18 EDT with unit paths repointed
|
||||||
|
from `%h/llm-router` to `%h/Sources/6krrt` (production still runs from the
|
||||||
|
dev checkout; see `plans/deploy-separation.md`); plugin fix installed and
|
||||||
|
opencode restarted 15:19; router restarted 15:19:57.
|
||||||
|
|
||||||
|
## Todos (one per doc)
|
||||||
|
|
||||||
|
1. **New `docs/watchdog.md`**: what it catches and does not (steady spend is not
|
||||||
|
waste), signals and thresholds, alert lifecycle, the admin surfaces, knobs,
|
||||||
|
install (the `sed` repoint + `enable --now`), how to run `--once`, backtest,
|
||||||
|
re-exporting the fixture, and known limits: sessions under 40 calls are
|
||||||
|
invisible; the Atlas must-flag case sits exactly at `cover_min` 4.0; the
|
||||||
|
stale `localhost:4096` entry in `rc-servers.json` logs a harmless probe
|
||||||
|
warning.
|
||||||
|
2. **`CLAUDE.md`**: "What's built and working" (anchor `## What's built and
|
||||||
|
working`): add the modules above in the existing one-bullet style, linking
|
||||||
|
`docs/watchdog.md`; update the `tests/` bullet count from the real
|
||||||
|
`pytest --collect-only -q` total. "Run as a service" (anchor `## Run as a
|
||||||
|
service`): one short paragraph on the watchdog timer. Do not edit the North
|
||||||
|
Star section.
|
||||||
|
3. **`docs/incidents.md` #8**: under Status, add that Lift A shipped (PR #102,
|
||||||
|
`a24de1b`) and the timer is live; add the plugin-loader finding as Evidence
|
||||||
|
(with the log line) and to the Human response timeline; add a symptom-table
|
||||||
|
row "opencode plugin does nothing -> grep opencode.log for `failed to load
|
||||||
|
plugin`". Keep the Evidence / labelled Theory convention in the file header.
|
||||||
|
4. **`docs/admin-portal.md`**: sections for the Loops panel, the Watchdog card,
|
||||||
|
and the Models rollup with Block/Unblock (where `blocked` differs from
|
||||||
|
`deprecated`: an operator's reasoned stop, one click to undo).
|
||||||
|
5. **`docs/api.md`**: the three watchdog endpoints and `blocked` on the
|
||||||
|
availability endpoint (request/response shapes read from `src/admin.py`).
|
||||||
|
6. **`docs/data-model.md`**: the four watchdog tables (from `config/schema.sql`),
|
||||||
|
and `blocked` in the `availability` row at the `models` section (line ~34)
|
||||||
|
and wherever `admin_model_overrides` is described.
|
||||||
|
7. **`docs/clients.md`** (anchor `## Session-directory attribution (opencode
|
||||||
|
plugin)`) and `docs/api.md` near the `X-Router-*` table (line ~151): the
|
||||||
|
plugin loader contract (every export must be a function) and how to confirm
|
||||||
|
the plugin loaded.
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- ASCII only; no middle dots. Match each file's existing style and heading
|
||||||
|
depth. Admin-portal docs describe the UI, they do not add prose to it.
|
||||||
|
- Numbers and names come from the code on this branch, never from this brief.
|
||||||
|
- Commit per todo, by explicit path. Never `git add -A`.
|
||||||
|
- Do not push and do not open a PR. The done report lists commits and the
|
||||||
|
items for `deploy/README.md`.
|
||||||
|
- If a worker's context passes ~150k tokens, stop it and start a fresh one.
|
||||||
|
|
||||||
|
Write `.omo/plans/docs-lift-a-pass.md`, then stop and report.
|
||||||
241
plans/no-progress-detection-prototype.py
Normal file
241
plans/no-progress-detection-prototype.py
Normal file
@@ -0,0 +1,241 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Interim opencode loop watcher (stand-in until the no-progress lift ships).
|
||||||
|
|
||||||
|
Signals per session, over a sliding window of its last WINDOW tool calls:
|
||||||
|
dup share of calls that repeat an earlier (tool, args) in the window
|
||||||
|
top count of the most repeated (tool, args) in the window
|
||||||
|
landed any edit/write with a non-empty diff, or a `git commit` with exit 0,
|
||||||
|
by the session OR any descendant within the window's time span
|
||||||
|
|
||||||
|
Flag when the window is full enough and nothing landed and
|
||||||
|
dup >= DUP_MIN or top >= TOP_MIN.
|
||||||
|
Read-only agents (explore/librarian/oracle) never land, so for them only
|
||||||
|
top >= TOP_MIN_RO counts.
|
||||||
|
|
||||||
|
Modes:
|
||||||
|
backtest : replay every session's history, report worst window per session
|
||||||
|
watch : poll live sessions; on a new flag print it, notify-send, exit 0
|
||||||
|
"""
|
||||||
|
import json, sys, time, subprocess, urllib.request, collections, os
|
||||||
|
|
||||||
|
U = os.environ.get("OC_URL", "http://127.0.0.1:4097")
|
||||||
|
WINDOW, MIN_CALLS = 60, 40
|
||||||
|
DUP_MIN, TOP_MIN, TOP_MIN_RO = 0.25, 12, 8
|
||||||
|
CUM_MIN = 15 # one exact call repeated this often over the whole session: a slow loop
|
||||||
|
COVER_MIN = 4.0 # lines read from one file in the window / its length: re-reading, whatever the offsets
|
||||||
|
READ_ONLY = ("@explore", "@librarian", "@oracle", "explore subagent")
|
||||||
|
|
||||||
|
|
||||||
|
def get(path):
|
||||||
|
return json.load(urllib.request.urlopen(U + path, timeout=30))
|
||||||
|
|
||||||
|
|
||||||
|
def calls_of(sid):
|
||||||
|
out = []
|
||||||
|
for m in get(f"/session/{sid}/message"):
|
||||||
|
for p in m["parts"]:
|
||||||
|
if p.get("type") != "tool":
|
||||||
|
continue
|
||||||
|
st = p.get("state") or {}
|
||||||
|
inp = st.get("input") or {}
|
||||||
|
md = st.get("metadata") or {}
|
||||||
|
t = (st.get("time") or {}).get("start") or m["info"].get("time", {}).get("created", 0)
|
||||||
|
tool = p.get("tool")
|
||||||
|
landed = (tool in ("edit", "write", "patch") and bool(md.get("diff"))) or (
|
||||||
|
tool == "bash" and "git commit" in (inp.get("command") or "") and md.get("exit") == 0)
|
||||||
|
out.append((t, tool, json.dumps(inp, sort_keys=True), landed))
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
import re
|
||||||
|
_PATHISH = re.compile(r"[\w./-]+\.(?:json|md|py|js|html|yaml|yml|toml|txt|sql|sh|mjs|css)\b")
|
||||||
|
|
||||||
|
|
||||||
|
_CD_PREFIX = re.compile(r"^\s*cd\s+\S+\s*(?:&&|;)\s*")
|
||||||
|
_ENV_PREFIX = re.compile(r"^\s*(?:[A-Z_][A-Z0-9_]*=\S+\s+)+")
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_bash(cmd):
|
||||||
|
"""Drop comment lines and leading `cd dir &&` / `VAR=x` prefixes, so
|
||||||
|
differently-worded commands under one prefix do not collapse to one key."""
|
||||||
|
lines = [ln for ln in cmd.splitlines() if not ln.lstrip().startswith("#")]
|
||||||
|
c = " ".join(lines).strip()
|
||||||
|
for _ in range(3):
|
||||||
|
c2 = _ENV_PREFIX.sub("", _CD_PREFIX.sub("", c))
|
||||||
|
if c2 == c:
|
||||||
|
break
|
||||||
|
c = c2
|
||||||
|
return c
|
||||||
|
|
||||||
|
|
||||||
|
def target_of(tool, args_json):
|
||||||
|
"""Coarse target: the file a call is about, ignoring offsets and the exact
|
||||||
|
command wording, so 28 differently-worded `cat boulder.json` count as one."""
|
||||||
|
try:
|
||||||
|
a = json.loads(args_json)
|
||||||
|
except Exception:
|
||||||
|
return (tool, args_json[:80])
|
||||||
|
if a.get("filePath"):
|
||||||
|
# A read keys on its slice: walking a big file in pieces is progress,
|
||||||
|
# re-reading the same slice is not.
|
||||||
|
return ("file", os.path.basename(a["filePath"]), a.get("offset"), a.get("limit")) if tool == "read" \
|
||||||
|
else ("file", os.path.basename(a["filePath"]))
|
||||||
|
if tool == "bash":
|
||||||
|
cmd = _normalize_bash(a.get("command") or "")
|
||||||
|
files = sorted({os.path.basename(f) for f in _PATHISH.findall(cmd)})
|
||||||
|
if files:
|
||||||
|
# Numbers in the command are the slice (sed -n 100,160p, head -40):
|
||||||
|
# walking a file is progress; the same slice again is not.
|
||||||
|
nums = tuple(sorted(set(re.findall(r"\b\d+\b", cmd))))
|
||||||
|
return ("bash-files", ",".join(files[:3]), nums)
|
||||||
|
return ("bash", " ".join(cmd.split()[:2]))
|
||||||
|
if tool in ("grep", "glob"):
|
||||||
|
return (tool, str(a.get("pattern"))[:60])
|
||||||
|
return (tool, args_json[:80])
|
||||||
|
|
||||||
|
|
||||||
|
def window_stats(win):
|
||||||
|
fp = collections.Counter((c[1], c[2]) for c in win)
|
||||||
|
dup = sum(v - 1 for v in fp.values() if v > 1) / max(1, len(win))
|
||||||
|
tg = collections.Counter(target_of(c[1], c[2]) for c in win)
|
||||||
|
(top_k, top_n) = tg.most_common(1)[0] if tg else (("", ""), 0)
|
||||||
|
return dup, top_n, top_k
|
||||||
|
|
||||||
|
|
||||||
|
_LEN = {}
|
||||||
|
|
||||||
|
|
||||||
|
def file_lines(path):
|
||||||
|
if path not in _LEN:
|
||||||
|
try:
|
||||||
|
with open(path, "rb") as fh:
|
||||||
|
_LEN[path] = max(1, fh.read().count(b"\n"))
|
||||||
|
except OSError:
|
||||||
|
_LEN[path] = None
|
||||||
|
return _LEN[path]
|
||||||
|
|
||||||
|
|
||||||
|
def coverage(win):
|
||||||
|
"""Max over files of (lines requested in the window / file length)."""
|
||||||
|
req = collections.Counter()
|
||||||
|
for c in win:
|
||||||
|
if c[1] != "read":
|
||||||
|
continue
|
||||||
|
a = json.loads(c[2])
|
||||||
|
fp = a.get("filePath")
|
||||||
|
n = file_lines(fp) if fp else None
|
||||||
|
if not n:
|
||||||
|
continue
|
||||||
|
req[fp] += min(a.get("limit") or n, n)
|
||||||
|
best = max(((v / file_lines(k), k) for k, v in req.items()), default=(0.0, ""))
|
||||||
|
return round(best[0], 1), os.path.basename(best[1])
|
||||||
|
|
||||||
|
|
||||||
|
def is_ro(title):
|
||||||
|
t = (title or "").lower()
|
||||||
|
return any(k in t for k in READ_ONLY)
|
||||||
|
|
||||||
|
|
||||||
|
def evaluate(sess, calls, landed_times_tree, end_idx=None):
|
||||||
|
calls = calls if end_idx is None else calls[:end_idx]
|
||||||
|
win = calls[-WINDOW:]
|
||||||
|
if len(win) < MIN_CALLS:
|
||||||
|
return None
|
||||||
|
t0, t1 = win[0][0], win[-1][0]
|
||||||
|
landed = any(t0 <= lt <= t1 for lt in landed_times_tree)
|
||||||
|
dup, top_n, top_k = window_stats(win)
|
||||||
|
ro = is_ro(sess.get("title"))
|
||||||
|
cum_k, cum_n = collections.Counter((c[1], c[2]) for c in calls).most_common(1)[0]
|
||||||
|
slow = (not landed) and cum_n >= CUM_MIN
|
||||||
|
cov, cov_file = coverage(win)
|
||||||
|
reread = (not landed) and cov >= COVER_MIN
|
||||||
|
if reread and not (top_n >= TOP_MIN):
|
||||||
|
top_n, top_k = int(cov), ("coverage", f"{cov_file} read {cov}x its length")
|
||||||
|
if ro:
|
||||||
|
flag = top_n >= TOP_MIN_RO or slow or reread
|
||||||
|
else:
|
||||||
|
flag = ((not landed) and (dup >= DUP_MIN or top_n >= TOP_MIN)) or slow or reread
|
||||||
|
if slow and not (top_n >= TOP_MIN):
|
||||||
|
top_n, top_k = cum_n, ("exact-total", cum_k[0], cum_k[1][:80])
|
||||||
|
return dict(flag=flag, dup=round(dup, 2), top=top_n, top_what=" ".join(str(x) for x in top_k)[:120],
|
||||||
|
landed=landed, ro=ro, n=len(calls), t1=t1)
|
||||||
|
|
||||||
|
|
||||||
|
def tree_landed(sessions, calls_by):
|
||||||
|
kids = collections.defaultdict(list)
|
||||||
|
for s in sessions:
|
||||||
|
if s.get("parentID"):
|
||||||
|
kids[s["parentID"]].append(s["id"])
|
||||||
|
memo = {}
|
||||||
|
def lt(sid):
|
||||||
|
if sid in memo:
|
||||||
|
return memo[sid]
|
||||||
|
own = [c[0] for c in calls_by.get(sid, []) if c[3]]
|
||||||
|
for k in kids.get(sid, []):
|
||||||
|
own += lt(k)
|
||||||
|
memo[sid] = own
|
||||||
|
return own
|
||||||
|
return lt
|
||||||
|
|
||||||
|
|
||||||
|
def backtest(since_ms):
|
||||||
|
sessions = [s for s in get("/session") if s.get("time", {}).get("updated", 0) >= since_ms]
|
||||||
|
calls_by = {s["id"]: calls_of(s["id"]) for s in sessions}
|
||||||
|
lt = tree_landed(sessions, calls_by)
|
||||||
|
for s in sessions:
|
||||||
|
cs = calls_by[s["id"]]
|
||||||
|
worst, first_flag = None, None
|
||||||
|
for i in range(MIN_CALLS, len(cs) + 1, 5):
|
||||||
|
r = evaluate(s, cs, lt(s["id"]), i)
|
||||||
|
if r and (worst is None or (r["flag"], r["dup"], r["top"]) > (worst["flag"], worst["dup"], worst["top"])):
|
||||||
|
worst = r
|
||||||
|
if r and r["flag"] and first_flag is None:
|
||||||
|
first_flag = i
|
||||||
|
if worst:
|
||||||
|
print(f"{'FLAG' if worst['flag'] else 'ok '} calls={len(cs):4d} first_flag_at={first_flag} "
|
||||||
|
f"dup={worst['dup']} top={worst['top']} landed={worst['landed']} ro={worst['ro']} | {s.get('title','')[:50]}"
|
||||||
|
+ (f"\n top: {worst['top_what'][:110]}" if worst['flag'] else ""))
|
||||||
|
|
||||||
|
|
||||||
|
def watch(minutes, state_path):
|
||||||
|
seen = set()
|
||||||
|
if os.path.exists(state_path):
|
||||||
|
seen = set(json.load(open(state_path)))
|
||||||
|
deadline = time.time() + minutes * 60
|
||||||
|
while time.time() < deadline:
|
||||||
|
try:
|
||||||
|
now = time.time() * 1000
|
||||||
|
sessions = get("/session")
|
||||||
|
status = get("/session/status")
|
||||||
|
recent = [s for s in sessions if s["id"] in status or now - s.get("time", {}).get("updated", 0) < 5 * 60e3]
|
||||||
|
calls_by = {s["id"]: calls_of(s["id"]) for s in recent}
|
||||||
|
# Judge only sessions doing work now: busy, or a tool call in the
|
||||||
|
# last 5 min. A session's "updated" stamp moves on abort, restart
|
||||||
|
# and title edits, so it cannot stand in for activity.
|
||||||
|
active = [s for s in recent if s["id"] in status
|
||||||
|
or (calls_by[s["id"]] and now - calls_by[s["id"]][-1][0] < 5 * 60e3)]
|
||||||
|
lt = tree_landed(sessions, calls_by)
|
||||||
|
for s in active:
|
||||||
|
r = evaluate(s, calls_by[s["id"]], lt(s["id"]))
|
||||||
|
key = f"{s['id']}:{len(calls_by[s['id']]) // 60}"
|
||||||
|
if r and r["flag"] and key not in seen:
|
||||||
|
seen.add(key)
|
||||||
|
json.dump(sorted(seen), open(state_path, "w"))
|
||||||
|
msg = (f"{s.get('title','')[:60]} ({s['id']}): {r['n']} calls, dup {r['dup']}, "
|
||||||
|
f"top x{r['top']}: {r['top_what']}, landed={r['landed']}, busy={s['id'] in status}")
|
||||||
|
print("LOOP SUSPECT:", msg)
|
||||||
|
subprocess.run(["notify-send", "-u", "critical", "opencode loop suspected", msg[:250]],
|
||||||
|
check=False, timeout=5)
|
||||||
|
return 0
|
||||||
|
except Exception as e: # opencode down or restarting: keep watching
|
||||||
|
print(f"[{time.strftime('%H:%M:%S')}] poll error: {e}", file=sys.stderr)
|
||||||
|
time.sleep(120)
|
||||||
|
print("no loops seen in", minutes, "min")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
if sys.argv[1] == "backtest":
|
||||||
|
backtest(int(sys.argv[2]))
|
||||||
|
else:
|
||||||
|
sys.exit(watch(int(sys.argv[2]), sys.argv[3]))
|
||||||
572
plans/no-progress-detection.md
Normal file
572
plans/no-progress-detection.md
Normal file
@@ -0,0 +1,572 @@
|
|||||||
|
# No-progress detection: catch an agent that is spinning, not one that is working
|
||||||
|
|
||||||
|
Status: reference -- parent spec; Lift A shipped in PR #102, Lift B (router-side judge, modes, behavior breaker) not started
|
||||||
|
Date: 2026-09-26
|
||||||
|
Trigger: the `cockpit-quick-wins` Atlas run, 2026-09-25 19:00 to 2026-09-26 02:00.
|
||||||
|
Incident write-up: `docs/incidents.md` #8.
|
||||||
|
|
||||||
|
## What happened, measured
|
||||||
|
|
||||||
|
Read-only against the live `router.db` and opencode's own session store
|
||||||
|
(`http://127.0.0.1:4097`), 2026-09-26.
|
||||||
|
|
||||||
|
Spend, 19:00 to 02:00 local:
|
||||||
|
- 323M prompt tokens (95% cached) and about $7.81, over roughly 8 hours.
|
||||||
|
- 79% of the money went to OpenRouter `z-ai/glm-5.3-flash`: 1,368 calls
|
||||||
|
averaging 203k prompt tokens.
|
||||||
|
- 921 calls carried 120k+ prompt tokens and cost $6.57. NeuralWatt `glm-5.3`
|
||||||
|
took 47 calls averaging 673k context, max 682,679, because nothing else fit.
|
||||||
|
- Every decision was `classification_source = classifier` on the default
|
||||||
|
profile. The router did what it was told; nothing was pinned.
|
||||||
|
|
||||||
|
**Spend rate is the wrong signal.** Tonight's worst hour ($2.42) and worst
|
||||||
|
3-hour window ($4.13) sit inside the prior 30 days' normal range (hourly p90
|
||||||
|
$1.46, max $4.83; worst 3h $9.65). A rate alarm quiet enough for a healthy
|
||||||
|
all-day agent run would have missed this entirely, and one tight enough to catch
|
||||||
|
it would fire on good work. The operator's framing is the right one: steady
|
||||||
|
usage is fine. What is not fine is spend with **no concrete change landing**.
|
||||||
|
|
||||||
|
**The sessions that wasted money are visibly different in their tool calls:**
|
||||||
|
|
||||||
|
| session (opencode) | tool calls | exact-duplicate calls | worst repeat | landed a change? |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Item 2 worker, 2nd attempt | 423 | 30% | `read navbar.js` x61 | no (worktree unchanged) |
|
||||||
|
| Item 3 worker, 1st attempt | 132 | 23% | `read test_admin_js_units.py` x16 | no |
|
||||||
|
| Item 4 worker | 134 | 19% | `read .omo/plans/...` x8 | partly (code, no tests) |
|
||||||
|
| Item 2 worker, 1st attempt | 307 | 17% | `read navbar.js` x21 | no |
|
||||||
|
| Atlas orchestrator | 576 | 11% | playwright console check x10 | yes (via workers) |
|
||||||
|
| healthy workers (items 1, 5, 6, 3-retry) | 30-96 | 0-6% | 1-3 | yes |
|
||||||
|
|
||||||
|
**A second loop shape, 2026-09-26 02:21 to about 02:56:** the Atlas orchestrator
|
||||||
|
itself, after its plan was complete:
|
||||||
|
- `.omo/boulder.json` carried a stale `status: completed` and `pr_url` from the
|
||||||
|
previous plan.
|
||||||
|
- oh-my-openagent's continuation hook injected "continue" turns; all 3 user
|
||||||
|
turns in the window were injected.
|
||||||
|
- Atlas, working from a trimmed context, re-derived its state every time:
|
||||||
|
54 turns, about 19M cached-read tokens, `boulder.json` read 28 times, the
|
||||||
|
plan's checkboxes grepped 16 times, nothing landed.
|
||||||
|
|
||||||
|
The operator caught it only by watching. The judge must treat this as
|
||||||
|
`stalling`, and Phase 0 must include it as a calibration case. The plugin can
|
||||||
|
also see injected continuation turns (`chat.message`), which is a useful extra
|
||||||
|
signal: "continuation turns keep arriving and nothing lands" is this shape
|
||||||
|
exactly.
|
||||||
|
|
||||||
|
**A third shape, 2026-09-26 around 03:05 to 03:17: a read-only agent.**
|
||||||
|
Prometheus's `explore` subagent reached 322 tool calls, 19% exact duplicates,
|
||||||
|
and read `deploy/opencode-plugin/router-outcome.js` 20 times. The file had been
|
||||||
|
deleted from under it: the production checkout was synced to `origin/main`,
|
||||||
|
where #99 retired it. The subagent was aborted.
|
||||||
|
|
||||||
|
Read-only agents (`explore`, `librarian`, `oracle`, and the like) never land a
|
||||||
|
change by design, so "no landed change" cannot be their signal. For them the
|
||||||
|
judge uses:
|
||||||
|
- target novelty: new files or queries per N calls, and
|
||||||
|
- a failure streak on the same target, such as repeated reads of a path that
|
||||||
|
no longer exists.
|
||||||
|
|
||||||
|
Identify them by the `X-Router-Agent` value, via a configurable list of
|
||||||
|
read-only agent names.
|
||||||
|
|
||||||
|
"Exact duplicate" means the same tool with byte-identical arguments. Duplicate
|
||||||
|
share alone does not separate the orchestrator (11%, productive) from a stuck
|
||||||
|
worker (17%). Duplicates **plus** no landed change does.
|
||||||
|
|
||||||
|
## Why nothing caught it
|
||||||
|
|
||||||
|
- `circuit_breaker.py` is availability-only: it trips on provider 5xx.
|
||||||
|
- The only spend alarm is `runway_low_warning`: balance / burn < 6 h, with burn
|
||||||
|
averaged over a 24 h balance window. OpenRouter read $24.73 at $0.25/h, so 97 h
|
||||||
|
of runway. It answers "will I run out", not "is this being wasted".
|
||||||
|
- The router cannot see progress at all. It sees prompts, not whether an edit
|
||||||
|
landed or a command keeps failing. The opencode plugin sees both.
|
||||||
|
- The DEPLOYED router cannot tell conversations apart either. `session_key`
|
||||||
|
merges concurrent conversations under one multi-agent client, and `/outcome`
|
||||||
|
attribution falls back to a directory match that 409s when several sessions
|
||||||
|
share a cwd, which is exactly the Atlas-plus-workers shape. **The fix already
|
||||||
|
exists on `origin/main` (PR #99, conversation-identity) but is not deployed:**
|
||||||
|
production runs from a local checkout 29 commits behind `origin/main`, the
|
||||||
|
live `route_decisions` has no `agent` / `parent_key` columns, and the
|
||||||
|
installed plugin is still `router-outcome.js` rather than #99's
|
||||||
|
`router-link.js`.
|
||||||
|
|
||||||
|
## Design
|
||||||
|
|
||||||
|
Split the job where the information lives. **The plugin is the sensor; the
|
||||||
|
router is the judge.** No raw task text leaves opencode, and the router never
|
||||||
|
stores any; the "never store raw task text" rule is unchanged.
|
||||||
|
|
||||||
|
### 1. Exact session identity: ALREADY BUILT upstream (#99), build on it
|
||||||
|
|
||||||
|
PR #99 (`conversation-identity`, merged to `origin/main` as `fe8c813`) did this.
|
||||||
|
Reuse it; do not rebuild it:
|
||||||
|
- **Plugin.** `deploy/opencode-plugin/router-link.js` replaces
|
||||||
|
`router-outcome.js`. Its `chat.headers` hook sends `X-Router-Conversation`
|
||||||
|
(the opencode session id), `X-Router-Agent` (slugged to the router's charset)
|
||||||
|
and `X-Router-Parent`, with a bounded parent-lookup cache.
|
||||||
|
- **Router.** `src/conversation_identity.py` resolves identity from those
|
||||||
|
headers. `session_key` becomes `'c:' + conversation id`. `route_decisions`
|
||||||
|
gained `agent` and `parent_key` columns. `/outcome` attributes to the exact
|
||||||
|
conversation.
|
||||||
|
|
||||||
|
What remains for this lift:
|
||||||
|
- **Deploy #99.** It is merged but not running. Syncing the production checkout
|
||||||
|
to `origin/main` and installing `router-link.js` are operator steps: the main
|
||||||
|
checkout has uncommitted work, and part of it, the `confidence_min` rename,
|
||||||
|
also landed upstream via #100.
|
||||||
|
- **Fix the exit code in `router-link.js`.** It kept `router-outcome.js`'s
|
||||||
|
`output?.exitCode ?? output?.exit_code` (line 83 on `origin/main`); see "Fixes
|
||||||
|
that ride along".
|
||||||
|
- **Build the progress events (section 2) into `router-link.js`,** keyed by the
|
||||||
|
same conversation and parent ids, so the judge rolls children up to parents
|
||||||
|
through `parent_key`.
|
||||||
|
|
||||||
|
Everywhere below, "session" means #99's conversation. Where this spec says
|
||||||
|
`X-Opencode-Session`, read `X-Router-Conversation`.
|
||||||
|
|
||||||
|
### 2. Progress events (plugin `tool.execute.after` -> router `POST /progress`)
|
||||||
|
|
||||||
|
The plugin sends one small event per tool call, fire-and-forget with a short
|
||||||
|
timeout, never blocking the session:
|
||||||
|
|
||||||
|
```
|
||||||
|
{session_id, parent_session_id, agent, tool, call_id,
|
||||||
|
fingerprint, # sha256 of tool + normalized args; never the args themselves
|
||||||
|
target, # coarse, normalized: the file path for read/edit/write, the
|
||||||
|
# command's first two tokens for bash; no contents
|
||||||
|
exit, # bash: output.metadata.exit
|
||||||
|
landed} # bool, see below
|
||||||
|
```
|
||||||
|
|
||||||
|
`landed` is the load-bearing definition, and it is deliberately narrow:
|
||||||
|
- `edit` / `write` / `patch` with a non-empty `metadata.diff`
|
||||||
|
- `bash` whose command is `git commit` (or `git commit --amend`) with exit 0
|
||||||
|
- A test or lint command (the plugin's existing `TEST_COMMAND` list) that PASSES
|
||||||
|
after the session's previous run of the same fingerprint FAILED, i.e. a fix
|
||||||
|
landed.
|
||||||
|
|
||||||
|
Reads, greps, todo writes, sub-agent dispatches and passing-again tests are not
|
||||||
|
progress. An orchestrator's progress is its children's: roll child `landed`
|
||||||
|
events up to the parent via `parent_session_id`.
|
||||||
|
|
||||||
|
Storage: a new `progress_events` table, pruned to a rolling window (knob). It
|
||||||
|
holds fingerprints, targets and booleans, no content.
|
||||||
|
|
||||||
|
### 3. The judge (router, per client session)
|
||||||
|
|
||||||
|
Signals per session, over a rolling window of its last N tool events:
|
||||||
|
- `turns_since_landed`: LLM requests since the last landed event (self plus
|
||||||
|
children for a parent).
|
||||||
|
- `usd_since_landed`: billed spend since then, joined through `request_id`.
|
||||||
|
- `dup_share`: exact-duplicate fingerprint share in the window.
|
||||||
|
- `fail_streak`: consecutive nonzero exits on the same fingerprint (the
|
||||||
|
"sandbox is broken and the agent keeps retrying" case).
|
||||||
|
- `context_tokens`: the latest request's size, to report alongside, since a
|
||||||
|
stall at 600k context costs far more per turn than one at 30k.
|
||||||
|
|
||||||
|
Verdicts:
|
||||||
|
- `ok`: something landed recently.
|
||||||
|
- `stalling` (warn): no landed change for `stall_turns` turns or
|
||||||
|
`stall_usd` dollars, AND (`dup_share >= dup_warn` OR `fail_streak >= fail_warn`).
|
||||||
|
- `looping` (act, if enabled): the stall condition at the higher `loop_*`
|
||||||
|
thresholds.
|
||||||
|
|
||||||
|
Both conditions have to hold. Long-running healthy work with no duplicates (a
|
||||||
|
big read-heavy exploration) stays `ok` until it hits the plain turns/spend
|
||||||
|
ceiling, which is a separate, looser knob.
|
||||||
|
|
||||||
|
**Calibrate before trusting it.** Phase 0 replays the judge offline over
|
||||||
|
opencode's session store (the 11 sessions above plus older history), and reports
|
||||||
|
the verdict each session would have received and when. Thresholds are picked
|
||||||
|
from that, not guessed. The Atlas orchestrator (11% duplicates, productive
|
||||||
|
through children) is the key false-positive case, and it must stay `ok`.
|
||||||
|
|
||||||
|
### 4. Optional: a local model as a second opinion
|
||||||
|
|
||||||
|
The deterministic signals should carry this. Where they are ambiguous, for
|
||||||
|
example near-duplicates (the same file read at shifting offsets, or the same
|
||||||
|
failing test with a different `-k`), there are two cheaper steps before any
|
||||||
|
model:
|
||||||
|
- normalize `read` fingerprints to the path, and bash to command plus target
|
||||||
|
- the local encoder already loaded for `classifier.mode: local_encoder` can
|
||||||
|
embed the `target` strings to cluster "similar-ish" events
|
||||||
|
|
||||||
|
A local LLM judge ("does this digest look like a loop?") runs only on a session
|
||||||
|
already flagged `stalling`, on a digest of tool names, targets and exit codes,
|
||||||
|
never content. It is gated by `local_compute.enabled`, since the operator pays
|
||||||
|
the local power bill. Treat this as a later phase, justified by Phase 0 showing
|
||||||
|
ambiguous cases the deterministic signals miss.
|
||||||
|
|
||||||
|
### 5. What happens on a detection: modes
|
||||||
|
|
||||||
|
One knob, `progress.mode`, picks the response. **Every mode warns:**
|
||||||
|
- A `/metrics` warning (admin bell and TUI). It names the session, agent, turns
|
||||||
|
and $ since the last landed change, and the top repeated target.
|
||||||
|
- An opencode toast via `POST /tui/show-toast` in the window the operator is
|
||||||
|
watching.
|
||||||
|
- An **alert** through the router's notifier (section 6), which reaches the
|
||||||
|
operator when nobody is looking at either screen.
|
||||||
|
|
||||||
|
| mode | on detection |
|
||||||
|
|---|---|
|
||||||
|
| `warn` | Warnings only. The neutral default: today's behaviour plus a warning. |
|
||||||
|
| `auto_compact` | Compact the stuck session with the incident report injected, then continue. No refusal, no orchestrator stop. |
|
||||||
|
| `auto_limit_context` | Cap the stuck session's context. No stop. |
|
||||||
|
| `auto_recover` | Two stages, below. |
|
||||||
|
|
||||||
|
**`auto_recover`, stage 1 (first detection in a session tree):**
|
||||||
|
1. The router refuses the stuck session's chat requests (429, scoped by
|
||||||
|
`X-Router-Conversation`), so nothing more is spent while the plugin acts.
|
||||||
|
2. The plugin walks `parentID` to the tree's root, usually Atlas, and calls
|
||||||
|
`POST /session/{id}/abort` on the stuck session and on the root.
|
||||||
|
3. The plugin compacts the root: `POST /session/{root}/summarize`. Its
|
||||||
|
`experimental.session.compacting` hook appends the **incident report** to
|
||||||
|
the compaction prompt, so the report survives into the compacted context.
|
||||||
|
4. `experimental.compaction.autocontinue` stays enabled, so the root restarts
|
||||||
|
from the compacted context with the report in view. The router lifts the
|
||||||
|
refusal when the compaction completes.
|
||||||
|
5. Toast: "recovered <agent>: <one-line reason>".
|
||||||
|
|
||||||
|
**`auto_recover`, stage 2 (detected again in the same tree within
|
||||||
|
`recover_window`):** apply the `auto_limit_context` cap to the tree, and toast
|
||||||
|
again.
|
||||||
|
|
||||||
|
**Stage 3, a hard stop.** Recovery can itself loop: recover, loop, recover. After
|
||||||
|
`max_recoveries` in the window, the router refuses the tree and does not
|
||||||
|
restart it. It toasts and raises a high-severity warning, and the tree waits for
|
||||||
|
an admin Resume. Without this cap, auto-recovery is an unattended retry loop,
|
||||||
|
which is the failure it exists to stop.
|
||||||
|
|
||||||
|
**The incident report** is built deterministically from `progress_events`
|
||||||
|
(fingerprints, targets and exit codes, never content), so it is free and cannot
|
||||||
|
hallucinate. Example:
|
||||||
|
|
||||||
|
```
|
||||||
|
[router] A sub-agent stalled and was stopped.
|
||||||
|
session: Item 2 G1 nav refactor (Sisyphus-Junior), child of this session
|
||||||
|
423 tool calls, 0 landed changes (no edit with a diff, no commit)
|
||||||
|
most repeated: read admin/frontend/navbar.js x61, read tests/... x16
|
||||||
|
spend since last landed change: $1.84, 70M prompt tokens
|
||||||
|
Avoid: re-reading whole files, since reads repeat without edits. Give the
|
||||||
|
worker the exact edit to make, and require an on-disk diff or commit before it
|
||||||
|
reports done.
|
||||||
|
```
|
||||||
|
|
||||||
|
The "Avoid:" line maps from the dominant signal: duplicates, a failure streak,
|
||||||
|
or context size. A local model may rephrase it later (section 4); it is not
|
||||||
|
needed to produce it.
|
||||||
|
|
||||||
|
**"Limit context"** has two candidate mechanisms. Phase 0 must verify which works
|
||||||
|
before stage 2 is built:
|
||||||
|
- **(a) Router-side, per-session tighter `pinch` budget.** The machinery
|
||||||
|
exists (`context_prune.py`), but pruning rewrites the prefix, and
|
||||||
|
`plans/token-waste-waves.md` Wave 3 measured rewrites costing the cache. So
|
||||||
|
this trades cache hits for size. Measure it.
|
||||||
|
- **(b) The router answers an over-cap request with a context-overflow-shaped
|
||||||
|
error**, so opencode runs its own overflow compaction. The autocontinue hook
|
||||||
|
receives `overflow: boolean`, so opencode does have such a path. Unverified:
|
||||||
|
whether it recognizes a router error as an overflow. Test this on 8081 before
|
||||||
|
relying on it.
|
||||||
|
|
||||||
|
Also, the Controls or Home page gets a live "Sessions" readout: session, agent,
|
||||||
|
parent, verdict, mode stage, turns and $ since landed, and the top repeated
|
||||||
|
target. This fits `plans/cockpit-brainstorm.md`. Each tree with a refusal gets
|
||||||
|
a Resume button.
|
||||||
|
|
||||||
|
Every threshold and the mode are knobs with admin controls (North Star #1).
|
||||||
|
Per the "build for re-tuning" rule, the neutral default is `warn`, with Phase
|
||||||
|
0-calibrated thresholds.
|
||||||
|
|
||||||
|
Without the plugin (another client, or a stale install), only the router half
|
||||||
|
works: warnings, alerts and the scoped 429. Abort, compact and restart need the
|
||||||
|
plugin. The Sessions readout says which sessions have a live plugin (from the
|
||||||
|
headers in section 1), so a mode that silently cannot act is visible.
|
||||||
|
|
||||||
|
### 6. Alerting: a notifier with pluggable channels
|
||||||
|
|
||||||
|
Toasts and the admin bell only help someone who is looking. Incident #8 ran
|
||||||
|
overnight. Alerts come from the **router**, not the plugin: the router is the
|
||||||
|
judge and is always running, while the plugin dies with the opencode process
|
||||||
|
it lives in.
|
||||||
|
|
||||||
|
**Ship now:** one channel, `desktop` (`notify-send`). **Design for later:** SMS
|
||||||
|
(Twilio or similar), RingCentral, PagerDuty, a generic webhook (ntfy, Slack).
|
||||||
|
Each of those should be a small adapter added without touching the detector.
|
||||||
|
|
||||||
|
Shape:
|
||||||
|
- **An alert is an event with a lifecycle**, not a message. Fields:
|
||||||
|
- `dedup_key`: for this detector, the tree root session id
|
||||||
|
- `severity`: `info`, `warning` or `critical`
|
||||||
|
- `state`: `trigger`, `escalate` or `resolve`
|
||||||
|
- `title` and a one-line `summary`
|
||||||
|
- `details`, the incident report from section 5
|
||||||
|
- `source`: `no-progress`, so other detectors can reuse the notifier later
|
||||||
|
|
||||||
|
This maps one-to-one onto PagerDuty Events v2 (`dedup_key`,
|
||||||
|
`event_action: trigger/resolve`), so that adapter is thin. Channels without
|
||||||
|
a lifecycle (SMS, desktop) render `trigger` and `escalate` as a message and
|
||||||
|
`resolve` as a short "recovered" message, or skip it (per-channel knob).
|
||||||
|
- **Transitions, not repeats.** An alert fires when a verdict *changes*, e.g.
|
||||||
|
`ok -> stalling`, `stalling -> looping`, a recovery, the hard stop, or
|
||||||
|
`-> ok`. It never fires per request. Per-channel rate limit and a quiet
|
||||||
|
window stop a flapping session from paging anyone at 3am every minute.
|
||||||
|
- **Severity routing per channel:** each channel has `min_severity`. The
|
||||||
|
intended setup: desktop gets `warning` and up; an SMS or PagerDuty channel
|
||||||
|
added later gets only `critical`, i.e. the hard stop, or `looping` in `warn`
|
||||||
|
mode.
|
||||||
|
- **Delivery never blocks routing.** Alerts go through a small queue with a
|
||||||
|
timeout and one retry. A failed delivery is logged and shown in the admin
|
||||||
|
UI; it never slows or fails a chat request.
|
||||||
|
- **Secrets stay out of config files.** Channel credentials (API keys, routing
|
||||||
|
keys, phone numbers) are read from environment variables named in config,
|
||||||
|
the same `api_key_env` pattern `dispatch_providers` already uses, and loaded
|
||||||
|
from `.env` by the systemd unit. `config.yaml` and `config.local.yaml` never
|
||||||
|
hold a secret.
|
||||||
|
- **Config:** a `notifications.channels` list, each entry
|
||||||
|
`{name, type, enabled, min_severity, resolve_messages, ...type-specific}`.
|
||||||
|
Only `type: desktop` exists in this lift. An unknown `type` fails config load
|
||||||
|
(StrictModel), so a typo cannot silently disable alerting.
|
||||||
|
- **Admin (North Star #1):** per channel, an enabled toggle, a `min_severity`
|
||||||
|
select, and a **Send test alert** button. The test button is the only way to
|
||||||
|
know a channel works before the night it matters. Recent deliveries and
|
||||||
|
failures show under the Sessions readout.
|
||||||
|
|
||||||
|
**`desktop` specifics, to verify in Phase 0.** The router runs as a systemd
|
||||||
|
*user* service, so `notify-send` needs the session D-Bus
|
||||||
|
(`DBUS_SESSION_BUS_ADDRESS`, normally `unix:path=/run/user/<uid>/bus`). Check
|
||||||
|
that the unit's environment has it and that its sandboxing (`ProtectHome` and
|
||||||
|
friends) allows the socket. If not, say so in the plan and fix it in
|
||||||
|
`deploy/`, not by weakening the sandbox wholesale. Use urgency `critical` for
|
||||||
|
`critical` alerts so they persist on screen.
|
||||||
|
|
||||||
|
### 7. The watchdog timer: detection that needs nobody watching
|
||||||
|
|
||||||
|
**Why.** All three loops in incident #8 were caught by a human or by a Claude
|
||||||
|
Code session reading opencode's session store by hand. The operator does not
|
||||||
|
want to depend on either. The watchdog is an independent, local, scheduled
|
||||||
|
sensor. It runs whether or not anyone is at the screen, whether or not a Claude
|
||||||
|
session is open, and whether or not the plugin loaded. **It calls no cloud
|
||||||
|
model.**
|
||||||
|
|
||||||
|
**What.** A systemd user timer, `deploy/llm-router-watchdog.{service,timer}`,
|
||||||
|
shaped like the existing `llm-router-*` timers, every `watchdog.interval`
|
||||||
|
(proposed 5 min). Each tick:
|
||||||
|
1. **Find opencode.** Read `~/.local/share/opencode/rc-servers.json`, written
|
||||||
|
by the `session-registry` plugin. Probe each recorded `serverUrl`. If none
|
||||||
|
answers, exit quietly; no opencode means nothing to watch.
|
||||||
|
2. **List sessions active since the last tick** via `GET /session`, and read
|
||||||
|
each one's recent messages via `GET /session/{id}/message`.
|
||||||
|
3. **Deterministic check, free, on every active session.** The same signals as
|
||||||
|
the judge (section 3): exact-duplicate share, the most-repeated target, the
|
||||||
|
failure streak, landed changes (for read-only agents, target novelty
|
||||||
|
instead), and injected continuation turns. This is Phase 0's replay code
|
||||||
|
run live. One implementation, imported by both, not a copy.
|
||||||
|
4. **Local LLM second opinion, only on sessions step 3 flags.** Skipped when
|
||||||
|
`local_compute.enabled` is off, which keeps gaming mode honest. The model is
|
||||||
|
whatever is configured: a `watchdog.model` knob that defaults to
|
||||||
|
`verification.model`, on that section's base URL. Today that is
|
||||||
|
`qwen2.5-coder-router:14b` on the local Ollama (`localhost:11434`), already
|
||||||
|
kept warm by local verification, so a call pays no model-load delay. It
|
||||||
|
runs on a digest of tool names, targets and exit codes, never content.
|
||||||
|
Phase 0 measures this model's verdicts on the three incident #8 cases
|
||||||
|
(must flag: the stuck workers and the post-completion Atlas loop; must not
|
||||||
|
flag: productive Atlas) before its answer is trusted. It asks "is this agent looping without progress? yes or no,
|
||||||
|
with a one-line reason". A healthy tick costs zero inference.
|
||||||
|
5. **Report, do not act.** POST the verdict to the router (the `/progress`
|
||||||
|
judge or a sibling endpoint), tagged `source: watchdog`. The **router**
|
||||||
|
applies the configured mode, alerts through the notifier, and enforces
|
||||||
|
`max_recoveries` and Resume. The watchdog never aborts or compacts anything
|
||||||
|
itself, so there is exactly one recovery path and one place that decides.
|
||||||
|
|
||||||
|
**If the router is down,** the watchdog still sends a `critical` desktop
|
||||||
|
alert directly with `notify-send`, and nothing else. A dead router is exactly
|
||||||
|
when an unattended run needs a human.
|
||||||
|
|
||||||
|
**Failure modes to design for:**
|
||||||
|
- A tick that overruns the interval: use a lock file and skip the tick, do not
|
||||||
|
stack them.
|
||||||
|
- An opencode API shape change. The v2 port breaks this the same way it breaks
|
||||||
|
`oc_work_start.py`. Detect the unexpected shape and raise one
|
||||||
|
`warning`-level alert saying the watchdog is blind, rather than silently
|
||||||
|
reporting "all healthy".
|
||||||
|
- The watchdog's own health. Record `last_tick_at` and the result (via the
|
||||||
|
router, or a state file), and show it in the admin Sessions readout. A
|
||||||
|
watchdog that stopped running must be visible.
|
||||||
|
|
||||||
|
**Knobs (North Star #1):**
|
||||||
|
- `watchdog.enabled`
|
||||||
|
- `watchdog.interval`
|
||||||
|
- `watchdog.local_llm_enabled` (default on, still gated by `local_compute`)
|
||||||
|
- `watchdog.model` (default: `verification.model`)
|
||||||
|
- `watchdog.read_only_agents`, the list from the read-only case above
|
||||||
|
|
||||||
|
The timer ships enabled, since detection plus a warning is the neutral,
|
||||||
|
low-risk default; the actions follow `progress.mode`.
|
||||||
|
|
||||||
|
### 8. The behavior breaker (Lift B): trip a model that stalls sessions
|
||||||
|
|
||||||
|
**Gap.** The breaker family covers availability (`circuit_breaker.py`, 5xx) and
|
||||||
|
malformed output (PR #76: mojibake, empty 200s). Nothing trips on
|
||||||
|
**well-formed output that does not do the job**. In incident #8,
|
||||||
|
`z-ai/glm-5.3-flash` produced fluent text that confabulated truncation and
|
||||||
|
re-read in loops, even with pinch off and a small context. It was only
|
||||||
|
excluded by hand.
|
||||||
|
|
||||||
|
**Trigger.** Verdicts from this detector (the watchdog, and the router judge
|
||||||
|
once built), attributed to the model(s) that served each flagged session in its
|
||||||
|
window. Trip when:
|
||||||
|
- stall verdicts on one model span at least `min_sessions` distinct sessions
|
||||||
|
(proposed 3) within `window` (proposed 60 min), AND
|
||||||
|
- that model's stall rate is well above the other models' on the same window's
|
||||||
|
traffic, so one hard task cannot blame a model.
|
||||||
|
|
||||||
|
**Response.** The same shape as the existing breakers: a passive routing skip
|
||||||
|
with cooldown and backoff, plus a critical alert through the notifier and an
|
||||||
|
admin Resume. Key it on the model, not the provider, per
|
||||||
|
`plans/mangled-output-detection.md`'s scope reasoning: here one model misbehaved
|
||||||
|
while its siblings on the same provider did not. Record every trip with its
|
||||||
|
evidence (sessions, verdicts, rates) so a human can review it.
|
||||||
|
|
||||||
|
**Blocked on identity.** Joining "this session stalled" to "these models served
|
||||||
|
it" needs #99's conversation headers on each request. As of 2026-09-26 they do
|
||||||
|
not arrive: `router-link.js` is installed and passes its 29 tests, and the router
|
||||||
|
reads the headers, but production decisions still carry fingerprint
|
||||||
|
`session_key`s with no `agent`. Lift B starts by finding out why.
|
||||||
|
|
||||||
|
## Fixes that ride along
|
||||||
|
|
||||||
|
- **The plugin never reads a real exit code.** Both `router-outcome.js`
|
||||||
|
(deployed) and its replacement `router-link.js` (#99, line 83) have this. In
|
||||||
|
1.18,
|
||||||
|
`tool.execute.after` gives `output = {title, output, metadata}`, and bash's
|
||||||
|
exit is `output.metadata.exit` (verified on a live session). `looksFailed`
|
||||||
|
checks `output.exitCode` / `output.exit_code`, which do not exist, so every
|
||||||
|
verdict has come from the failure-text regex alone. Read `metadata.exit`
|
||||||
|
first.
|
||||||
|
- **The installed plugin differs from the repo copy.** The main checkout has an
|
||||||
|
uncommitted comment block. It is harmless, but the lift should add a check
|
||||||
|
(install script or a `/health` field reporting the plugin's version string
|
||||||
|
from a header) so a stale install is visible.
|
||||||
|
- **`opencode.json` advertised `auto` with `limit.context: 782324`**, the
|
||||||
|
largest window in the catalog, so opencode only compacted near 782k. Tonight's
|
||||||
|
contexts reached 682k and were resent every turn. Lowered to 200,000 on
|
||||||
|
2026-09-26 by owner decision (see the end). It caps what `auto` is asked to
|
||||||
|
carry, so Phase 0 should also measure how often a real agent turn now
|
||||||
|
compacts, and whether any task needed more.
|
||||||
|
|
||||||
|
## Ground rules for the executor
|
||||||
|
|
||||||
|
- **Base on `origin/main`, NOT the local `main`.** The local `main` checkout
|
||||||
|
(`c87e675`) is 29 commits behind `origin/main` (`f554fc1` as of 2026-09-26)
|
||||||
|
and 2 ahead with local-only commits. It lacks #99 (conversation identity,
|
||||||
|
`router-link.js`) and #100 (local-encoder rebuild). `git fetch origin` first,
|
||||||
|
then branch `feat/no-progress-detection` from `origin/main` in its own
|
||||||
|
worktree. Read code from that worktree or `git show origin/main:<path>`, never
|
||||||
|
from the main checkout's files. Do not touch, stash or commit anything in the
|
||||||
|
main checkout; it has uncommitted work from other efforts.
|
||||||
|
- The plugin to extend is `deploy/opencode-plugin/router-link.js` (with its
|
||||||
|
`router-link.test.mjs`). `router-outcome.js` was replaced by #99.
|
||||||
|
- **Port 8080 is production; never touch it.** Integration checks run on 8081
|
||||||
|
via `scripts/sandbox.sh`, against a copy of the live DB made with the sqlite
|
||||||
|
backup API (source opened `mode=ro`).
|
||||||
|
- Never edit `config/config.yaml` or `config/config.local.yaml`. New knobs get
|
||||||
|
defaults in `config.yaml` only as part of this lift's own tracked changes.
|
||||||
|
- **The installed plugin at `~/.config/opencode/plugins/` is live for every
|
||||||
|
opencode session on this machine,** including the one running the lift.
|
||||||
|
Develop and test against the repo copy. Installing a new version is an
|
||||||
|
operator step, stated in the done report, not something the executor does.
|
||||||
|
- **Commit by explicit path. Never `git add -A` or `git add .`** Incident #8's
|
||||||
|
item 4 swept an `.omc/` state file into a commit that way.
|
||||||
|
- **Every worker claim is verified by the orchestrator** against
|
||||||
|
`git show --stat` and the actual test files before acceptance. Incident #8
|
||||||
|
had a commit message listing tests that did not exist.
|
||||||
|
- **Read large files in targeted slices** (grep for the anchor, then read
|
||||||
|
around it). Incident #8's stuck workers re-read whole files dozens of times.
|
||||||
|
- ASCII only in new code, comments and UI strings; no middle-dot separators.
|
||||||
|
No prose paragraphs in the admin UI.
|
||||||
|
- Lint gate: pinned `uvx ruff@0.16.9`, no NEW findings in touched files
|
||||||
|
compared with the base commit (same recipe as
|
||||||
|
`.omo/plans/cockpit-quick-wins.md` runbook step 5). Never `ruff --fix`
|
||||||
|
outside the lines the item changes.
|
||||||
|
- `pytest` offline stays green after every commit. Record the baseline on
|
||||||
|
`origin/main` first. On `c87e675`, one test,
|
||||||
|
`tests/test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity`,
|
||||||
|
failed before any change; check whether it still does on `origin/main` rather
|
||||||
|
than assuming.
|
||||||
|
- New `route_decisions` columns are registered in
|
||||||
|
`tests/test_tui_schema_drift.py`; every new knob gets an admin control or a
|
||||||
|
recorded `DELIBERATELY_NOT_IN_ADMIN` reason (`tests/test_admin_knob_coverage.py`).
|
||||||
|
|
||||||
|
## Phases
|
||||||
|
|
||||||
|
0. **Calibrate offline.** Replay over opencode session history; produce the
|
||||||
|
verdict timeline per session and a proposed threshold table. No router or
|
||||||
|
plugin changes. **Start from `plans/no-progress-detection-prototype.py`,**
|
||||||
|
an interim watcher tuned on 2026-09-26 against the 15 sessions of incident
|
||||||
|
#8. On that set it flags every known loop (both item-2 workers, the item-3
|
||||||
|
first attempt, item 4, the post-completion Atlas stretch, the explore
|
||||||
|
helper stuck on a deleted file, a helper that printed the spec 12 times)
|
||||||
|
and none of the healthy sessions. Four findings from tuning it:
|
||||||
|
- Exact `(tool, args)` fingerprints alone miss reworded commands: 28
|
||||||
|
differently-worded `cat boulder.json` never matched.
|
||||||
|
- Collapsing to the bare file over-flags ordinary sliced reading, which is
|
||||||
|
what the ground rules tell workers to do. Key a `read` on file plus
|
||||||
|
offset/limit, and a `bash` target on its referenced files plus the
|
||||||
|
numbers in the command (the slice).
|
||||||
|
- Slow loops spread across 300+ calls escape a 60-call window. Add a
|
||||||
|
whole-session rule: one exact call repeated 15+ times with nothing landed
|
||||||
|
recently.
|
||||||
|
- Read-only agents need the separate rule (repeats of one target), since
|
||||||
|
"nothing landed" is always true for them.
|
||||||
|
Its thresholds (`WINDOW=60`, `DUP_MIN=0.25`, `TOP_MIN=12`, `TOP_MIN_RO=8`,
|
||||||
|
`CUM_MIN=15`) are a starting point fitted on one night, not a result.
|
||||||
|
1. **Identity: build on #99, do not rebuild it.** The `metadata.exit` fix in
|
||||||
|
`router-link.js` (and its `.test.mjs`), plus whatever #99 lacks for
|
||||||
|
progress roll-up, if anything. Deploying #99 itself (syncing the
|
||||||
|
production checkout, installing `router-link.js`) is an operator step. It
|
||||||
|
goes in the done report as a prerequisite for Phase 2 to see real traffic,
|
||||||
|
not an executor task.
|
||||||
|
2. **Progress events, the judge and the watchdog, `warn` mode.** `/progress`,
|
||||||
|
`progress_events`, verdicts, the `/metrics` warning, toast, Sessions readout,
|
||||||
|
and the knobs. Also the notifier (section 6) with the `desktop` channel and
|
||||||
|
the admin Send test alert button, and the watchdog timer (section 7) with
|
||||||
|
its local-LLM second opinion. **Order within the phase: build the watchdog
|
||||||
|
and the notifier first.** Together with Phase 0's detection code, they
|
||||||
|
protect unattended runs before the plugin half is even deployed.
|
||||||
|
3. **`auto_compact` and `auto_recover` stage 1.** Abort, compact with the
|
||||||
|
injected report, restart; the scoped 429; the `max_recoveries` hard stop and
|
||||||
|
admin Resume. The incident report builder lands here.
|
||||||
|
4. **`auto_limit_context` and `auto_recover` stage 2**, using whichever
|
||||||
|
mechanism Phase 0 verified.
|
||||||
|
5. **Local-model second opinion inside the router's live judge** (section 4),
|
||||||
|
only if earlier phases show ambiguous cases. The watchdog already has one
|
||||||
|
(section 7); this would add it to the per-request path.
|
||||||
|
|
||||||
|
The v2 port (`project_opencode_v2_pinned`) changes hook registration. Keep the
|
||||||
|
plugin's logic in plain functions so the v1 and v2 shells stay thin.
|
||||||
|
|
||||||
|
## Owner decisions
|
||||||
|
|
||||||
|
Settled 2026-09-26: the response is a mode (`warn`, `auto_compact`,
|
||||||
|
`auto_limit_context`, `auto_recover` with two stages), and every mode warns.
|
||||||
|
`auto_recover` stops and compacts the tree's root (Atlas), not only the stuck
|
||||||
|
worker.
|
||||||
|
|
||||||
|
Also settled 2026-09-26:
|
||||||
|
- **The default mode is `warn`.**
|
||||||
|
- **`max_recoveries` is 2** per tree per window, then the hard stop.
|
||||||
|
- **The `auto` context limit is 200,000.** Applied the same day, ahead of this
|
||||||
|
lift, in the repo `opencode.json` and the global
|
||||||
|
`~/.config/opencode/opencode.json` (backup at
|
||||||
|
`opencode.json.bak-20260926-auto-context`). The global file pins every agent
|
||||||
|
(atlas, sisyphus-junior, prometheus, ...) to `llm-router/auto`. opencode
|
||||||
|
reads config at startup, so it takes effect for sessions started after a
|
||||||
|
restart. `auto:batch` stays at 782,324 on purpose: batch is where long
|
||||||
|
contexts are expected.
|
||||||
|
|
||||||
|
- **Alerting ships with `desktop` (`notify-send`) only**, behind a notifier
|
||||||
|
built for more channels. SMS, RingCentral, PagerDuty and webhooks are later
|
||||||
|
adapters (section 6), not part of this lift.
|
||||||
|
|
||||||
|
No owner decisions remain open. Thresholds come from Phase 0.
|
||||||
227
plans/no-progress-lift-a.md
Normal file
227
plans/no-progress-lift-a.md
Normal file
@@ -0,0 +1,227 @@
|
|||||||
|
# Lift A: loop watchdog, desktop alerts, and the shared detector
|
||||||
|
|
||||||
|
Status: done -- shipped in PR #102; the admin surfacing (component 5) shipped only partly
|
||||||
|
Date: 2026-09-26
|
||||||
|
Parent spec: `plans/no-progress-detection.md` (reference only; do NOT read it
|
||||||
|
end to end). This brief is the contract for Lift A.
|
||||||
|
Incident: `docs/incidents.md` #8.
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Catch an opencode agent session that is spinning without progress, and alert
|
||||||
|
the operator on the desktop, with nobody watching and no cloud model involved.
|
||||||
|
|
||||||
|
## Why this slice first, and why it is small
|
||||||
|
|
||||||
|
Incident #8's loops were caught only by a human or a Claude session reading
|
||||||
|
opencode's session store by hand. Lift A replaces that with a local timer.
|
||||||
|
|
||||||
|
It is deliberately narrow because the last planning attempt failed on size:
|
||||||
|
- a 540-line spec
|
||||||
|
- a planning context that grew to about 600k tokens
|
||||||
|
- the model reporting "elided" reads that were stored intact, and re-reading
|
||||||
|
the spec 84 times (68x its length)
|
||||||
|
|
||||||
|
Keep this lift's inputs, todos and contexts small.
|
||||||
|
|
||||||
|
**Out of scope (Lift B):**
|
||||||
|
- router-side progress events, the live judge, and `/progress`
|
||||||
|
- recovery modes (`auto_compact`, `auto_recover`, `auto_limit_context`)
|
||||||
|
- scoped 429s
|
||||||
|
- the `router-link.js` changes
|
||||||
|
- SMS, PagerDuty and other alert channels
|
||||||
|
|
||||||
|
## Components
|
||||||
|
|
||||||
|
### 1. Shared detector: `src/progress_detect.py` (pure, stdlib only)
|
||||||
|
|
||||||
|
Port the rules from `plans/no-progress-detection-prototype.py`. It is tuned on
|
||||||
|
the incident #8 sessions; keep its behaviour and make it a clean, tested
|
||||||
|
module. Input: a session's tool calls as `(time, tool, args_json, landed)`.
|
||||||
|
Output: a verdict with the reason. Signals, over a window of the last
|
||||||
|
`window` calls:
|
||||||
|
- `dup`: exact `(tool, args)` duplicate share
|
||||||
|
- `top`: the most-repeated target. A `read` target is the file plus
|
||||||
|
offset/limit. A `bash` target is the files it names plus the numbers in the
|
||||||
|
command (the slice).
|
||||||
|
- `slow`: one exact call repeated `cum_min`+ times over the whole session
|
||||||
|
- `coverage`: lines requested from one file in the window divided by its
|
||||||
|
length. This catches re-reading a file through shifting offsets, which the
|
||||||
|
other signals miss.
|
||||||
|
- `landed`: an edit or write with a non-empty diff, or a `git commit` with
|
||||||
|
exit 0, by the session or any descendant (`parentID`) within the window.
|
||||||
|
|
||||||
|
Verdict:
|
||||||
|
- Normal agents are flagged when nothing landed and any of dup, top, slow or
|
||||||
|
coverage crosses its threshold.
|
||||||
|
- Read-only agents (explore, librarian, oracle; a configurable list) never
|
||||||
|
land, so for them top, slow or coverage alone decides.
|
||||||
|
- Thresholds start at the prototype's values: `window` 60, `dup_min` 0.25,
|
||||||
|
`top_min` 12, `top_min_ro` 8, `cum_min` 15, `cover_min` 4.0.
|
||||||
|
|
||||||
|
### 2. Calibration backtest: `scripts/progress_backtest.py`
|
||||||
|
|
||||||
|
It replays sessions from opencode's HTTP API through `progress_detect` and
|
||||||
|
prints a per-session verdict with the first flag point. Commit a small
|
||||||
|
**fixture** of the labelled sessions below, exported as tool-call fingerprints
|
||||||
|
(tool name, args, landed; never message text), under `tests/fixtures/`. A test
|
||||||
|
asserts every label. Live opencode is not needed in tests.
|
||||||
|
|
||||||
|
MUST FLAG:
|
||||||
|
|
||||||
|
| session | what it was |
|
||||||
|
|---|---|
|
||||||
|
| `ses_f24cc0258ffepbJPoOzoiLZi9s` | item 2 worker, read navbar.js x61 |
|
||||||
|
| `ses_f24fcf651ffeOm9O4E14QSZeJ4` | item 2 worker, 1st attempt |
|
||||||
|
| `ses_f24890640ffewKnIPqjjAfFvOk` | item 3 worker, 1st attempt |
|
||||||
|
| `ses_f245607eaffeGxPH8F1tvvdvkG` | item 4 worker |
|
||||||
|
| `ses_f237e7e31ffeOzz4kUKu4CWzMu` | explore helper re-reading a deleted file |
|
||||||
|
| `ses_f235d755bffeo0uWtD1g2HwQUg` | helper that printed the spec 12x |
|
||||||
|
| `ses_f2506ac70ffebLUxBsq0aauOn9` | Atlas; must flag in its post-02:07 stretch |
|
||||||
|
| `ses_f2380dd4cffeVwcZn3Kv3nQN7v` | planner that re-read the spec 68x its length |
|
||||||
|
|
||||||
|
MUST NOT FLAG:
|
||||||
|
|
||||||
|
| session | what it was |
|
||||||
|
|---|---|
|
||||||
|
| `ses_f246e699fffeCHUPTLr5ajBhWC` | item 3 retry |
|
||||||
|
| `ses_f245f78f8ffeGMn3PgArpOwhFy` | item 3 compact |
|
||||||
|
| `ses_f243ea08effemMwXdMRE1gpnhK` | item 5 |
|
||||||
|
| `ses_f242e5e28ffehF97ZxxdbTeylQ` | item 6 |
|
||||||
|
| `ses_f236bbdaaffezwnvH10X2rgrxc` | explore, sliced reads of dispatcher.py |
|
||||||
|
| `ses_f236bc7baffeG5LQsr2K1c46xO` | explore |
|
||||||
|
| `ses_f2e81929fffeaVbWui0X9QjmRQ` | cockpit planner |
|
||||||
|
|
||||||
|
The Atlas session also must NOT flag before its 02:07 stretch.
|
||||||
|
|
||||||
|
Export the fixture FIRST, in the first todo. opencode sessions can be deleted,
|
||||||
|
and the fixture is the ground truth.
|
||||||
|
|
||||||
|
### 3. Notifier: `src/notifier.py`, `desktop` channel only
|
||||||
|
|
||||||
|
- **An alert is an event:**
|
||||||
|
- `dedup_key` (the session tree root)
|
||||||
|
- `severity` (`info`, `warning` or `critical`)
|
||||||
|
- `state` (`trigger`, `escalate` or `resolve`)
|
||||||
|
- `title`, `summary`, `details`, `source`
|
||||||
|
|
||||||
|
The shape maps one-to-one onto PagerDuty Events v2, so later adapters are
|
||||||
|
thin. Only the `desktop` channel ships.
|
||||||
|
- **`desktop` runs `notify-send`,** with urgency `critical` for critical
|
||||||
|
alerts.
|
||||||
|
- **Fire on transitions only,** with a per-channel rate limit and a quiet
|
||||||
|
window.
|
||||||
|
- **A delivery failure is logged, never raised.**
|
||||||
|
- **Config:** a `notifications.channels` list of
|
||||||
|
`{name, type, enabled, min_severity}`. An unknown `type` fails config load
|
||||||
|
(StrictModel). Secrets would come from env vars named in config; desktop has
|
||||||
|
none.
|
||||||
|
|
||||||
|
### 4. Watchdog: `src/watchdog.py` + `deploy/llm-router-watchdog.{service,timer}`
|
||||||
|
|
||||||
|
A systemd user timer, shaped like the existing `llm-router-*` timers, every 5
|
||||||
|
min. Each tick:
|
||||||
|
1. **Find opencode.** Read `~/.local/share/opencode/rc-servers.json` and probe
|
||||||
|
each `serverUrl`. If none answers, exit 0 quietly.
|
||||||
|
2. **Read activity.** List sessions active since the last tick and read their
|
||||||
|
tool calls.
|
||||||
|
3. **Detect.** Run `progress_detect` on each active session.
|
||||||
|
4. **Second opinion, flagged sessions only.** Ask the configured local model
|
||||||
|
(`watchdog.model`, default `verification.model`, today
|
||||||
|
`qwen2.5-coder-router:14b` on the local Ollama) "looping without progress?
|
||||||
|
yes/no + one line", on a digest of tool names, targets and exit codes, never
|
||||||
|
content. Skip this when `local_compute.enabled` is false or
|
||||||
|
`watchdog.local_llm_enabled` is false. Report the model's answer alongside
|
||||||
|
the deterministic verdict; it does not veto it in Lift A.
|
||||||
|
5. **Alert.** Alert through the notifier directly, and write rows to the
|
||||||
|
router DB: a `watchdog_ticks` table (time, sessions seen, outcome) and a
|
||||||
|
`watchdog_verdicts` table. The router need not be up; SQLite WAL handles the
|
||||||
|
second writer.
|
||||||
|
|
||||||
|
Also:
|
||||||
|
- A lock file prevents overlapping ticks.
|
||||||
|
- An unexpected opencode API shape raises one `warning` alert ("watchdog is
|
||||||
|
blind"), never a silent all-clear.
|
||||||
|
- Admin (North Star #1): a small Watchdog card showing the last tick, recent
|
||||||
|
verdicts, the channel list with enabled and `min_severity` controls, and a
|
||||||
|
**Send test alert** button. Every new knob gets a control or a recorded
|
||||||
|
reason; `tests/test_admin_knob_coverage.py` enforces it.
|
||||||
|
- Knobs:
|
||||||
|
- `watchdog.enabled`, `watchdog.local_llm_enabled`, `watchdog.model`,
|
||||||
|
`watchdog.read_only_agents`
|
||||||
|
- the detector thresholds from section 1
|
||||||
|
- `notifications.channels`
|
||||||
|
|
||||||
|
The tick interval lives in the timer unit; document it next to the knobs.
|
||||||
|
|
||||||
|
## Ground rules
|
||||||
|
|
||||||
|
- **Base on `origin/main`.** Branch `feat/no-progress-lift-a` in its own
|
||||||
|
worktree. Do not touch, stash or commit anything in the main checkout.
|
||||||
|
- **Port 8080 is production.** Integration checks run on 8081
|
||||||
|
(`scripts/sandbox.sh`), against a sqlite-backup copy of the live DB.
|
||||||
|
- **Never edit** `config/config.yaml`'s existing values or
|
||||||
|
`config/config.local.yaml`. New keys get defaults in `config.yaml`.
|
||||||
|
- **Do not install or enable** the timer on this machine; that is an operator
|
||||||
|
step in the done report.
|
||||||
|
- **Commit by explicit path; never `git add -A`.** The orchestrator checks each
|
||||||
|
worker's `git show --stat` against its claims and the plan's paths before
|
||||||
|
accepting it.
|
||||||
|
- **Read in targeted slices.** Grep for the anchor, read around it. Do NOT read
|
||||||
|
`plans/no-progress-detection.md` whole; this brief is the contract.
|
||||||
|
- **If a worker's context passes about 150k tokens, stop it** and start a fresh
|
||||||
|
worker with a smaller prompt, rather than letting it grind.
|
||||||
|
- **ASCII only; no middle dots.** No prose paragraphs in the admin UI.
|
||||||
|
- **Lint:** pinned `uvx ruff@0.16.9`, no NEW findings in touched files against
|
||||||
|
the base, never `ruff --fix` outside changed lines.
|
||||||
|
- **pytest:** record the `origin/main` baseline first; it stays green apart
|
||||||
|
from the baseline.
|
||||||
|
|
||||||
|
## Done means
|
||||||
|
|
||||||
|
- Every must-flag and must-not-flag label passes in tests, and the live
|
||||||
|
backtest prints the same verdicts.
|
||||||
|
- One real `notify-send` from the admin test button on 8081.
|
||||||
|
- The watchdog runs once by hand (`python -m watchdog --once`) against live
|
||||||
|
opencode and prints its verdicts.
|
||||||
|
- The done report lists each commit, the tests, and the operator steps
|
||||||
|
(install and enable the timer).
|
||||||
|
|
||||||
|
## Addendum, 2026-09-26: surfacing and one-click block (North Star #4)
|
||||||
|
|
||||||
|
The owner set a new priority: surface waste in the admin portal, and make
|
||||||
|
stopping it one click (`CLAUDE.md` North Star #4). This lift therefore also
|
||||||
|
ships:
|
||||||
|
|
||||||
|
### 5. Loops panel and one-click model block (admin portal)
|
||||||
|
|
||||||
|
- **A Loops panel** on the Home page (or the Controls page, whichever the
|
||||||
|
cockpit layout fits), reading `watchdog_verdicts`: each flagged session with
|
||||||
|
agent, verdict, turns and $ since the last landed change, the top repeated
|
||||||
|
target, and when it was flagged.
|
||||||
|
- **A per-model rollup:** stall verdicts per model over the last hour, next to
|
||||||
|
each model's share of traffic. It needs session-to-model attribution; see
|
||||||
|
item 0.
|
||||||
|
- **A Block model button** beside each model in the rollup. It writes
|
||||||
|
`admin_model_overrides` with `availability = 'blocked'` and a reason (default
|
||||||
|
"stalled N sessions", editable), through the existing
|
||||||
|
`POST /admin/api/models/{model_id}/{provider}/availability` path. Verify that
|
||||||
|
routing excludes `blocked` exactly like `deprecated` (`load_candidates`), and
|
||||||
|
add a test for it.
|
||||||
|
- **A Blocked models list** with reason, time and a one-click Unblock
|
||||||
|
(`DELETE .../availability`).
|
||||||
|
- **Desktop alerts link to the panel** (`http://127.0.0.1:8080/admin/` plus an
|
||||||
|
anchor).
|
||||||
|
|
||||||
|
### 0. First todo: why do the identity headers not arrive?
|
||||||
|
|
||||||
|
`router-link.js` (#99) is installed in `~/.config/opencode/plugins/`, is
|
||||||
|
identical to the repo copy, and passes its 29 tests. The router reads
|
||||||
|
`X-Router-Conversation` (`src/dispatcher.py`, `resolve_identity(request.headers,
|
||||||
|
...)`). Yet production `route_decisions` after the 04:02 opencode restart still
|
||||||
|
show fingerprint `session_key`s and `agent = NULL`. Find out why, and fix it if
|
||||||
|
it is a small change.
|
||||||
|
|
||||||
|
The per-model rollup depends on it. If it cannot be fixed in this lift, the
|
||||||
|
rollup ships disabled with a visible "needs conversation identity" note, and
|
||||||
|
everything else ships.
|
||||||
@@ -790,7 +790,8 @@ def _get_db_path(cfg: Any) -> str:
|
|||||||
return "router.db"
|
return "router.db"
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
def main(argv: list[str] | None = None) -> int:
|
||||||
|
"""Run one watchdog tick; return the process exit code (0 when it ran)."""
|
||||||
ap = argparse.ArgumentParser(
|
ap = argparse.ArgumentParser(
|
||||||
description="Watchdog orchestrator for opencode loop detection",
|
description="Watchdog orchestrator for opencode loop detection",
|
||||||
)
|
)
|
||||||
@@ -800,7 +801,7 @@ if __name__ == "__main__":
|
|||||||
ap.add_argument(
|
ap.add_argument(
|
||||||
"--config", default="config/config.yaml", help="config.yaml path",
|
"--config", default="config/config.yaml", help="config.yaml path",
|
||||||
)
|
)
|
||||||
args = ap.parse_args()
|
args = ap.parse_args(argv)
|
||||||
|
|
||||||
logging.basicConfig(
|
logging.basicConfig(
|
||||||
level=logging.INFO,
|
level=logging.INFO,
|
||||||
@@ -841,6 +842,14 @@ if __name__ == "__main__":
|
|||||||
last_fired_cb = _make_last_fired(conn)
|
last_fired_cb = _make_last_fired(conn)
|
||||||
notifier = Notifier(notifier_cfg, last_fired=last_fired_cb)
|
notifier = Notifier(notifier_cfg, last_fired=last_fired_cb)
|
||||||
|
|
||||||
rc = tick(cfg, conn, notifier=notifier)
|
alerts_fired = tick(cfg, conn, notifier=notifier)
|
||||||
conn.close()
|
conn.close()
|
||||||
sys.exit(rc or 0)
|
logger.info("watchdog tick done: %d alert(s) fired", alerts_fired)
|
||||||
|
# A completed tick exits 0 however many alerts it fired. Returning the
|
||||||
|
# alert count made systemd mark the oneshot unit FAILED on every real
|
||||||
|
# alert (and wrap past 255), so a working alert looked like a crash.
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
|
|||||||
@@ -1180,3 +1180,21 @@ class TestDescendantLandedScope:
|
|||||||
)
|
)
|
||||||
finally:
|
finally:
|
||||||
_close_conn(conn, path)
|
_close_conn(conn, path)
|
||||||
|
|
||||||
|
|
||||||
|
class TestMainExitCode:
|
||||||
|
"""`python -m watchdog --once` runs as a systemd oneshot: a tick that ran
|
||||||
|
and fired alerts is a success, so the exit code must not carry the alert
|
||||||
|
count (systemd would mark the unit FAILED on every real alert)."""
|
||||||
|
|
||||||
|
def test_main_returns_zero_when_alerts_fire(self, tmp_path):
|
||||||
|
import watchdog
|
||||||
|
|
||||||
|
cfg = MagicMock()
|
||||||
|
cfg.database.path = str(tmp_path / "router.db")
|
||||||
|
cfg.notifications.channels = []
|
||||||
|
with patch("watchdog.load_config", return_value=cfg), \
|
||||||
|
patch("watchdog.tick", return_value=3) as fake_tick:
|
||||||
|
rc = watchdog.main(["--once", "--config", "unused.yaml"])
|
||||||
|
assert fake_tick.called
|
||||||
|
assert rc == 0
|
||||||
|
|||||||
Reference in New Issue
Block a user