Files
6krrt/docs/incidents.md

34 KiB

Incidents: how this router has broken, and how to tell which one it is

Eight times now, something around the router has silently degraded either the agent depending on it or the operator trying to see it clearly. They share a shape worth naming: none of them announce themselves as router problems. Five of the eight presented as an opaque client-side error — a connection refused, an "Unprocessable Content", an "internal server error" — and diagnosing each meant knowing which log or table to look in. The other three didn't error toward the client at all: #5 destroyed data outright, #6 just went quiet in one corner of the admin UI while /health stayed green the whole time, and #8 spent money for hours while every check stayed green.

#8 is the only one so far that cost money rather than capability, and the router behaved as configured throughout. Read it before leaving an agent run unattended.

#5 is the only one so far that destroyed data, and the only one where recovery depended on luck rather than design. Read it before running any cleanup command in this repo.

#7 is the only one so far to take the whole service down from a clean config-rename PR, not from an agent's own mistake — read it before renaming or removing any config key.

This page exists so the next one takes minutes rather than hours. Start with the symptom table, then read only the relevant section.

How to write an entry (from #8 on).

  • Separate Evidence (measured, with where it was measured) from Theory (inferred). Anything not established by a measurement or a reproduction goes under Theory. It states what supports it and what would confirm or refute it, and stays there until that test is done.
  • Include a Human response timeline: what the operator noticed, decided and did. Mark actions an assistant took, and at whose direction.
  • Entries #1 to #7 predate this convention.

opencode.json points opencode's own model traffic at http://127.0.0.1:8080/v1, so on this machine the router is the coding agent's inference supply. Breaking it breaks the thing you would use to fix it. That is why these keep happening, and why the diagnostics below are worth having to hand.

Symptom → first check

What you see Likely One-line check
ConnectionError: Connection refused, retrying #1 killed by name/port systemctl --user status llm-router — recently restarted, or inactive
active (running) but nothing answers #2 shutdown hang curl -s localhost:8080/health fails while systemd says active
Agent dies mid-task on an opaque 4xx #3 candidate set empty sqlite3 router.db "select task_tier, required_context_tokens, rejected_reason from route_decisions where selected_model is null order by id desc limit 5;"
"internal server error", dashboard blank #4 config/db path journalctl --user -u llm-router --since '10 min ago' | grep -c '" 500' then git diff config/config.yaml
Prices/windows look wrong, nothing errors catalog frozen sqlite3 router.db "select max(last_updated) from models;"
500s + no such table, venv/.env gone #5 git clean -fdx ls -la router.db .env .venv — a 0-byte db and a missing .venv is conclusive
A newly-shipped admin knob just isn't there, /health is fine #6 backend running stale code systemctl --user status llm-router uptime vs. git log -1 --format=%cd on the commit that added the knob
activating (auto-restart), pydantic_core...ValidationError: ...extra_forbidden in the journal #7 renamed key vs. un-migrated overlay journalctl --user -u llm-router --since '5 min ago' | grep extra_forbidden then grep -n <old key name> config/config.local.yaml
Spend is high after an agent run, nothing errored, no warning fired #8 agent looping with no progress sqlite3 router.db "select session_key, strftime('%H',observed_at,'localtime') h, round(sum(prompt_tokens)/1e6) mtok, round(sum(cost_usd),2) usd from energy_observations where observed_at > datetime('now','-12 hours') group by 1,2 order by usd desc limit 8;" then check whether that run's commits actually landed
opencode plugin does nothing #8 plugin loader rejects non-function export grep 'failed to load plugin' ~/.local/share/opencode/log/opencode.log

That last row is not an incident yet — it is the silent-staleness failure described in the "Run as a service" section of CLAUDE.md. An unpolled catalog fails open: it keeps routing on data that may be weeks old and every row still reads active.


#1 — Killed by name or port (2026-08-29)

Symptom. The dispatcher goes unreachable; the client retries against a refused connection.

Cause. An agent ran pkill -f "uvicorn dispatcher:app" (also .*dispatcher, also plain "8080") to free the port for its own throwaway instance. When the port came back 5s later — Restart=always resurrecting the supervised service — it read that as "the kill didn't work" and escalated to pkill -9.

Root cause was documentation, not code. AGENTS.md offered a bare python -m uvicorn dispatcher:app --reload as an equally-valid way to bring the router up, with no warning that 8080 is normally already held. An agent following that instruction and finding the port busy has no way to know the right move is systemctl --user restart.

Fixed in AGENTS.md: ad hoc runs bind --port 8081 explicitly, and the section says outright never to send a kill signal to anything matched by name or port. Full audit trail in plans/router-unreachable-signal-investigation.md — the syscall-level extension to pidfd_send_signal is what finally caught it, after kill/tgkill both came back clean.

Convention that came out of it: 8080 is production, always. Throwaway instances bind 8081. Never signal a process matched by name or port rather than by a PID you started yourself.

#2 — Bare SIGTERM hangs the process forever (2026-08-29)

Symptom. systemctl --user status reports active (running); the socket is closed and nothing answers. The process logged Waiting for connections to close and never got past it.

Cause. SSE clients holding /events/decisions open (the TUI, the admin dashboard) never disconnect, so uvicorn's graceful shutdown has nothing to wait out. That alone is only slow. What made it permanent: systemd enforces TimeoutStopUSec only when it is running the stop job, so a signal delivered outside that path leaves the unit active forever — the main PID never exits, so Restart= never fires either.

Fixed in deploy/llm-router.service: --timeout-graceful-shutdown 5 caps the drain regardless of who sends the signal. This also fixed ordinary restarts, which were silently taking the full 10s-then-SIGKILL path for the same reason.

Follow-up the fix exposed: once the process could exit cleanly on SIGTERM, Restart=on-failure excluded it from auto-restart — systemd assumes SIGTERM means someone deliberately asked it to stop. Changed to Restart=always. Deliberate systemctl stop/restart are still honoured; systemd tracks those separately from the Restart= decision.

Recovery: systemctl --user restart llm-router.service. Since a hung process was never in a tracked stop job, this issues a fresh cycle that systemd does enforce the timeout on.

#3 — Admin overrides collapsed the tier-3 context ceiling (2026-09-01)

Symptom. An agent failed mid-task with "Unprocessable Content". Surfaced ~19 hours after the cause.

Cause. Seven expensive models were deprecated through /admin — a reasonable cost decision in isolation. The models table still read active (the poller refreshes it every 2h), but admin_model_overrides overlays deprecated and _admin_deprecated_models feeds that into routing's hard filters. Effect on the maximum servable required_context_tokens:

tier before during eligible models
1 782,324 782,324 6
2 782,324 782,324 5
3 782,324 94,196 1

Tier-3 traffic has an observed max of 268,168 tokens, so every tier-3 request above 94k returned 422 No model satisfies the hard filters. Nothing warned; the portal reported the change as a plain success.

The obvious check for this is wrong. Warning when a higher tier's ceiling sits below a lower tier's is a theorem, not a fault: ceiling(T) is the max effective_context_window over models with tier >= T, and tier is a capability floor, so the eligible set shrinks monotonically and ceiling(1) >= ceiling(2) >= ceiling(3) holds for every catalog. Such a warning fires always and means nothing. The real detector compares the ceiling against observed demand — silent on all three tiers today, fires on the outage state.

Fixed by the /metrics ceiling warnings and an inline warning at the admin availability toggle, so the cost of a deprecation is visible while looking at the switch. See plans/catalog-staleness-and-poller-failure-modes.md §4.4.

Recovery: re-activate enough large-window rows to make the ladder continuous. Today: 94,196 → 192,500 → 782,324.

#4 — Tracked config pointed at a nonexistent database (2026-09-01)

Symptom. "Internal server error" in the client; the admin dashboard blank.

Cause. A QA step edited the tracked config/config.yaml, changing database.path from router.db to /tmp/router-qa/router.db, created the directory but never a database in it, and left the edit in the working tree. config/config.yaml is the file the live systemd service reads.

1753  200
 201  500 Internal Server Error
sqlite3.OperationalError: unable to open database file

115 GET /metrics · 65 GET /health · 14 /admin/api/snapshot
 14 /admin/api/history · 7 POST /v1/chat/completions

Those 7 chat completions are real inference failing, not dashboard noise.

Fixed by reverting the one line and restarting. It was never committed.

Rule: never point config/config.yaml at test fixtures. To run against a throwaway database, pass a different config file, monkeypatch cfg.database.path in-process, or use a temp copy. If you must touch a file the live service reads, restore it in the same step and verify curl -s localhost:8080/health before moving on. A QA step that leaves production broken has not passed.

#5 — git clean -fdx destroyed the database, key and venv (2026-09-04)

Symptom. Router returned 500 on every request. Journal showed sqlite3.OperationalError: no such table: energy_observations while systemd reported the service active.

Cause. An agent ran git clean -fdx in the repo. The -x flag removes ignored files as well as untracked ones, and everything this deployment needs to run is ignored by design:

lost what it was
router.db truncated to 0 bytes — 22,776 energy observations, 17,321 route decisions, 148 proficiency rows
.env the NeuralWatt API key
.venv the virtualenv the systemd unit's ExecStart runs from
config/config.local.yaml the operator's electricity tariff
node_modules

This is worse than it looks from the command. git clean -fd is a reasonable thing for an agent to run to get a clean tree. Adding -x turns it from "discard my scratch files" into "delete the deployment", and nothing in the repo warns you.

Recovery was luck, not design. There is no backup of router.db by policy. What saved it was that the plan running at the time had made a QA copy at /tmp/qa-config-local-overlay-r2/ 25 seconds before the wipe, and that copy happened to include .env. Integrity check passed and the restore was effectively lossless. Had that plan been a different one, the entire measurement history of the project would be gone.

Recovery steps, in order:

# 1. find a surviving copy — QA/scratch dirs are the likely place
find /home/alee /tmp -name "router.db" -size +0
sqlite3 <candidate> "select count(*) from energy_observations;"
sqlite3 <candidate> "select integrity_check from pragma_integrity_check limit 1;"

# 2. restore data, key, overlay
cp -f <candidate> router.db
cp -f <candidate-dir>/.env .env && chmod 600 .env
#    config/config.local.yaml from your own copy

# 3. rebuild the venv (pinned, so this is deterministic)
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt

# 4. restart and verify
systemctl --user restart llm-router.service
curl -s localhost:8080/health

What actually limited the damage was two unrelated decisions made minutes earlier: the in-flight plan's 11 files had just been committed rather than left uncommitted, and the operator's tariff had been parked outside the repo instead of restored in place. Both were reactions to the same file being clobbered repeatedly that day — the seventh time is what prompted moving it out of reach.

Known gap this leaves open. router.db has no backup policy. It holds every energy/cost observation, every routing decision, and all proficiency scores — none of which can be rebuilt without re-running evals that cost real money, and the historical observations cannot be rebuilt at all. A periodic snapshot is cheap insurance and does not exist.

Rules that came out of it:

  • Never git clean -x in this repo. Use targeted paths. If you need a clean tree, git stash preserves; clean destroys.
  • Never clean untracked files you did not create. An untracked file in this tree is as likely to be operator data as build residue.
  • Before any destructive git command, ask what is ignored, not just what is untracked. Here that list is the database, the API key and the runtime.

#6 — Admin knobs shipped to disk, invisible until the process restarted (2026-09-17)

Symptom. Runtime knobs recently added to the admin portal (the incumbent- gate dial from PR #89, the session-cache-window controls from PR #90) were not displaying, despite both being merged and present in the working tree. /health and every other page reported normal the entire time.

Cause. admin/frontend/controls.html and its sibling pages are served fresh off disk on every load — confirmed by the GET /admin/api/frontend-version digest added for incident-adjacent work on 2026-09-13, which exists precisely because frontend files are read live, not baked into the process. src/admin.py and src/config.py are not: they are imported once at process start and stay exactly as they were until the process restarts, no matter what lands on disk afterward. Both PR #89 and PR #90 shipped their frontend half and backend half together, correctly, per this project's own north-star rule that a knob needs both — but the two halves have different deploy timing. The frontend half takes effect on the very next page load. The backend half takes effect only on the next restart. In the window between "merged and pulled" and "service restarted," a fresh page load runs new frontend code against an old backend process: the new knob's controls call admin API fields and endpoints that do not exist yet in the running process's memory, so they render empty rather than erroring, while everything unrelated keeps working — the mismatch is scoped to exactly the new surface, which is what makes it easy to miss.

Fixed by systemctl --user restart llm-router.service — mechanically the same recovery as #1 and #4, for a different reason: the process was not dead or misconfigured, it was correct for a version of the code that no longer matched what was on disk.

The reverse case was already caught; this direction was not. /admin/api/frontend-version detects exactly the opposite mismatch — an open browser tab holding stale frontend JS against a newer backend — and prompts a reload. Nothing detects a fresh frontend load against a stale backend process, because nothing exposes what code the running process actually has loaded: no version stamp, no commit hash, no process-start time surfaced anywhere in the admin UI.

Known gap this leaves open. Every future PR that ships an admin knob (frontend + backend together, exactly as required) reopens this same silent window between merge and restart. A GET /admin/api/backend-version returning the process's start time or the commit it was launched against — mirroring the existing frontend-version check, compared against git rev-parse HEAD on disk — would close it: the admin UI could show a "restart to pick up code changes" banner the same way it already handles the reverse direction. Not built.

Rule: after merging any PR that touches src/admin.py, src/config.py, src/dispatcher.py, or anything else the systemd unit imports at start, restart llm-router.service before trusting what the admin portal shows. Merging code and deploying it are different steps here, and today only one of them happens automatically.

#7 — A clean config rename crash-looped production via the un-migrated overlay (2026-09-18)

Symptom. Minutes after merging and fast-forwarding a routine refactor PR, systemctl --user status llm-router showed activating (auto-restart) with Result: exit-code, cycling every ~6 seconds. Every attempt logged the same traceback, admin.py/dispatcher.py never got past load_config.

Cause. The merged PR renamed session_cache.staleness_minutes to session_cache.staleness_seconds — a straight rename, no deprecated alias, matching this project's own stated convention (see the quota-shape removal: "every consumer was updated in the same change, so there are no deprecated aliases"). Every tracked consumer was updated in that same change: config/ config.yaml, src/config.py, src/admin.py, src/dispatcher.py, tests, docs. One consumer isn't tracked and can't be: config/config.local.yaml, the gitignored, machine-local overlay, which had staleness_minutes: 2 sitting in it from an earlier tuning session. StrictModel's extra="forbid" rejected that now-unknown key on every single startup:

pydantic_core._pydantic_core.ValidationError: 1 validation error for RouterConfig
session_cache.staleness_minutes
  Extra inputs are not permitted [type=extra_forbidden, input_value=2, input_type=int]

Unlike #6, this wasn't a stale-process-vs-fresh-disk mismatch — the process could not start at all, because dispatcher.py's module-level cfg = load_config(...) runs at import time, before the app can serve anything. git diff/git log show nothing wrong, because nothing tracked is wrong; the break lives entirely in a file no diff will ever show.

Fixed by migrating the overlay by hand: backed up config/config.local.yaml, diffed before/after to confirm only the one line changed, converted the value (staleness_minutes: 2 → staleness_seconds: 120, same real duration), then restarted. Recovered on the first clean attempt — about 7 crash-restart cycles, under a minute of total downtime.

This was foreseen and still happened. The PR that did the rename was explicitly told not to touch config/config.local.yaml and that its migration would happen "separately, directly, by the operator... after this PR merges" — correct instructions, followed correctly, and the gap between "PR merges" and "operator migrates the overlay" was still a live window a production restart landed inside. Knowing the risk existed did not close it; only the migration itself did.

Rule: a rename or removal of any config key is only complete when config/config.local.yaml has been checked, not just the tracked files — grep -n <old key name> config/config.local.yaml before merging, or at minimum before the next restart. This is a corollary of the existing "config.local.yaml is irreplaceable" guardrail, extended to include a config key's name as part of what merging code can silently invalidate, not only its content. A rename PR that ships without an accompanying "does anything override this?" check on the overlay is incomplete, even when every tracked file is correct.


#8 — Agent sessions looped for hours with no progress, and nothing noticed (2026-09-25)

Symptom. Nothing errored. The router stayed healthy, /metrics showed zero warnings, and the quota alarm read kind: none. The operator noticed early on 2026-09-26 that an opencode run had "burned a bunch of tokens over the hours".

Evidence

Measured read-only against the live router.db and opencode's own session store. The window is 2026-09-25 19:00 to 2026-09-26 02:00 local unless noted.

Spend.

prompt tokens 323M, 95% cached
billed about $7.81, roughly 3x the busiest of the previous ten days ($2.66 on 09-20)
largest share OpenRouter z-ai/glm-5.3-flash: 1,368 calls averaging 203k prompt tokens, $6.16
largest contexts NeuralWatt glm-5.3: 47 calls averaging 673k, max 682,679, the only model whose window still fit
120k+ prompt calls 921 of them, $6.57 of the total
routing every decision classification_source = classifier, default profile; nothing pinned

What the agents did. An Atlas run (.omo/plans/cockpit-quick-wins.md) delegated each item to a sub-agent worker:

worker session tool calls exact-duplicate calls worst repeat landed
item 2, 2nd attempt 423 30% read navbar.js x61 nothing
item 3, 1st attempt 132 23% read test_admin_js_units.py x16 nothing
item 4 134 19% read of the plan x8 code only
item 2, 1st attempt 307 17% read navbar.js x21 nothing
healthy workers 30 to 96 0 to 6% 1 to 3 yes
  • Item 4's commit message claimed four test files it never wrote, and it committed with git add -A.
  • Atlas later said its own turns were "just re-reading the plan and state over and over instead of doing work".
  • A planning session described its reads as "heavily elided". The stored tool outputs in opencode were intact.

Pinch was rewriting agent context.

  • pinch.budget_tokens: 50000 pruned 1,915 of 1,992 decisions over nine hours, keeping 48% of the prompt on average.
  • The planning session's 527k-token prompts went out at about 125k.
  • src/context_prune.py replaces older tool results with [read: result omitted], or keeps a 1,500-character head and tail around a [N chars trimmed...] marker. Results inside the protected window get the same cut above protected_max_chars (20,000).
  • Pinch's config had not changed since 2026-09-06. It had pruned 90 to 100% of requests daily for two weeks without incident. The input changed:
09-10 to 09-23 09-25 09-26
avg prompt into pinch 100k to 150k 236k 545k
p90 prompt 170k to 260k 528k 1.31M
  • opencode.json advertised auto with limit.context: 782324, so opencode did not compact until near that size.

Routing drifted with size.

  • deepseek/deepseek-v4-flash has 384k effective context; z-ai/glm-5.3-flash has 655k.
  • Their proficiency was tied (diff_checking 0.933 vs 0.935) and unchanged since 09-10.
  • Deepseek went from 290 picks on 09-23 to 0 on 09-25, while appearing 851 times as runner-up.

A clean reproduction, 2026-09-26 04:04 to 04:14. A fresh planning session, on the Lift A brief (188 lines):

  • Conditions: context about 88k tokens, pinch off, 0 of 29 tool parts marked compacted by opencode, every stored read intact.
  • What it did: it still reported its reads "elided with [...]", wrote fake "Earlier tool responses received ... (counts, not proof of task completion)" blocks into its own replies (assistant text parts, not tool output), invented session IDs, and called the brief "393 lines".
  • Where the phrase is not: in opencode, oh-my-openagent (JS bundle and native binary), or this repo.
  • Model: z-ai/glm-5.3-flash served 27 of 34 chat decisions in the 15 minutes before it was stopped.

Found along the way.

  • deploy/opencode-plugin/router-outcome.js has never read a real exit code. On opencode 1.18, bash's exit is output.metadata.exit; the plugin reads output.exitCode.
  • PR #99 (conversation identity) was merged to origin/main but not running. Production ran from a local checkout 29 commits behind.
  • Plugin loader rejected router-link.js. The plugin exported parentCache as a module-level const. Opencode 1.18's plugin loader discards any plugin with a non-function export — the entire file is silently dropped. The fix was to make parentCache a property on the factory function instead (L224, router-link.js in PR #102, a24de1b). Confirmed by grepping ~/.local/share/opencode/log/opencode.log for 'failed to load plugin'.
  • How much of the $7.81 was waste cannot be computed. The run did land six reviewed commits, and the deployed router cannot tell conversations apart.

Theory

These are inferences, not measurements. Each lists what would confirm or refute it.

  1. Pinch cutting oversized sessions drove, or worsened, the re-read loops. Weakened by the clean reproduction, which looped with pinch off. As sessions grew past about 250k, the fixed 50k budget removed most of each agent's working memory, recent reads included. The agent saw its reads cut, re-read them, grew the context, and got cut harder.
    • Supported by: the agents' own descriptions ("elided", "drowning in tool-output truncation"), intact stored outputs, the prune rates, and the timing of the size jump.
    • Not tested: no run has been observed with pinch off.
    • Confirm: with pinch off and auto at 200k, the duplicate-read signature should not recur on comparable runs. The interim watcher (plans/no-progress-detection-prototype.py) and Phase 0 data can show it.
  2. Sessions got that large because long unattended orchestration ran without compaction. Atlas made 668 tool calls, one worker 423, under a 782k limit, and the re-read loop in theory 1 fed the growth. The limit is a fact; that it explains the growth is theory.
  3. Leading theory: z-ai/glm-5.3-flash, or its OpenRouter path, confabulates truncation and then loops. The clean reproduction above removed pinch, context size and input corruption, and the behaviour stayed. Theories 1 and 2 remain plausible aggravators, not the cause.
    • Confirm: with glm excluded, comparable agent runs show no elision claims and no duplicate-read signature. A direct A/B on one prompt would settle model versus provider path.

Why nothing caught it

Each existing guard answers a different question:

  • circuit_breaker.py trips on provider 5xx. There were none.
  • runway_low_warning fires when balance / burn drops under 6 hours. Burn is averaged over a 24 h balance window: OpenRouter read $24.73 at $0.25/h, which is 97 hours of runway. It asks "will I run out", not "is this being wasted".
  • Spend rate was not abnormal. The worst hour ($2.42) and worst 3-hour window ($4.13) sit inside the prior 30 days' range (hourly p90 $1.46, max $4.83; worst 3 h $9.65).
  • The router cannot see progress, and the deployed build could not tell conversations apart.

Recurrences

  • 02:21 to about 02:56: Atlas looped after its plan was complete. Every checkbox was ticked, but .omo/boulder.json still carried status: completed and pr_url: .../pulls/77 from the previous plan (evidence: the file). The continuation hook injected "continue" turns, 3 of 3 user turns in the window (evidence: the session). Atlas re-derived its state each time: 54 turns, about 19M cached-read tokens, boulder.json read 28 times, checkboxes grepped 16 times. Cause: stale state, not pruning. The branch was never pushed, so PR 77 was not touched.
  • About 03:05 to 03:17: a read-only explore subagent re-read router-outcome.js 20 times (19% duplicate calls across 322). The production sync had just deleted that file. Read-only agents never land changes, so the detector needs a separate signal for them.

Human response

Times are local, 2026-09-26.

  • About 02:10. Noticed the spend and asked for a mechanism in 6krrt to catch it. Rejected a spend-rate alarm, since steady agent usage is legitimate, and framed the target as "are concrete changes landing".
  • About 02:20 to 02:40. Decided the design for plans/no-progress-detection.md:
    • response modes (warn, auto_compact, auto_limit_context, a two-stage auto_recover), with every mode warning
    • warn as the default
    • at most 2 recoveries per session tree
    • notify-send now, with pluggable SMS, RingCentral and PagerDuty channels later
    • a local-model watchdog timer, "so I don't have to rely on Claude"
    • auto's context limit lowered to 200,000 in the repo and global opencode.json
  • About 02:50. Caught the post-completion Atlas loop by watching. At the operator's direction, Claude aborted the session at 02:56 and cleared the stale pr_url in .omo/boulder.json.
  • About 03:10. Merged PR #101. A first git pull aborted on local changes, and the restart on that line reloaded the old code. Then ran the prepared sync (back up, drop the changes already upstream, reset --keep origin/main, re-apply the rest) and restarted onto 827c408 at 03:15.
  • About 03:30. Asked Claude to watch opencode for loops until the detector ships (the interim watcher).
  • About 03:45. Reported the planning session's "truncation hell". At the operator's direction ("hot fix that"), Claude set pinch.enabled false at runtime and persisted it to config/config.local.yaml (backup config.local.yaml.bak-20260926-pinch). The operator restarted the service at 03:52 and chose not to open a PR for the switch flip.
  • About 04:02. Installed router-link.js in place of router-outcome.js and restarted opencode (200k auto limit now active). Kicked off a fresh, small Lift A planning session (plans/no-progress-lift-a.md).
  • About 04:14. After the fresh session reproduced the confabulation, excluded z-ai/glm-5.3-flash from routing. It had no picks after 04:13:57, and traffic moved to deepseek and mimo. At the operator's direction, Claude aborted the confabulating session.
  • ~15:18. Watchdog timer enabled (systemctl --user enable --now llm-router-watchdog.timer).
  • 15:19. Plugin-loader root cause found in opencode.log; plugin fix installed and opencode restarted.
  • 15:19:57. Router restarted.
  • After 15:19. Conversation identity confirmed arriving: 17 of 17 requests after the restart carry c: keys and an agent.

Status

Partly mitigated; detection not built; leading theory (3) unconfirmed.

Lift A shipped. PR #102 (a24de1b) landed the surfacing half of North Star #4 (conversation identity and watchdog), live ~15:18 EDT. Surfacing half of North Star #4 is only partly shipped — the identity half is confirmed (17 of 17 requests after the restart carry c: keys), but the detection half still waits on Lift B (plans/no-progress-detection.md).

Confounded by design, and say so. Three mitigations landed within about 25 minutes:

  • pinch off at 03:50
  • auto capped at 200k at 04:02
  • z-ai/glm-5.3-flash excluded at about 04:14

So "sessions look healthier since" cannot say which one mattered, or whether it is several interacting: model degradation near a full window, oversized sessions, pruning, and stale orchestration state. The one clean data point is the 04:04 reproduction (glm, pinch off, small context, still confabulated), which is why theory 3 leads.

To attribute properly, change one variable at a time and watch the watchdog's stall rate. For example, re-admit glm with everything else held, or re-enable pinch with a budget above 200k. Until then, treat every theory here as unproven. Also missing: a behavior breaker. The availability breaker trips on 5xx and PR #76 trips on malformed output; nothing trips on well-formed output that confabulates or loops. Planned for Lift B (plans/no-progress-detection.md, section 8).

  • In place:
    • pinch off, runtime and persisted
    • auto limit 200,000, active for opencode sessions started after a restart
    • #99 and #100 deployed at 03:15
  • Pending, operator:
    • install router-link.js in place of router-outcome.js
    • restart opencode
  • Pending, lift: plans/no-progress-detection.md, the watchdog first.
  • If pinch is re-enabled, its budget must sit above the 200k cap, and it must never cut recent results.

Rule: steady spend is not the failure; spend with no concrete change landing is. Until the detector ships:

  • Check an unattended agent run by whether commits are landing, not by the quota page.
  • Verify each worker commit's git show --stat against its message; a worker's "done" is a claim.

The pattern

The first five are defensible local decisions — free a port, stop a process, deprecate an expensive model, point at a test database, clean the working tree — that silently removed capability while the router kept reporting itself healthy. Three of the five were caused by an agent working on the router, using the router.

#5 breaks the pattern in one way worth noting: it did not degrade quietly, it failed loudly and immediately. What it removed was not capability but state — and state, unlike capability, does not come back when you fix the code.

#6 breaks it a different way again: nobody made a decision at all. Nothing was disabled, deprecated, or misconfigured — the code and the running process simply disagreed about what existed, for exactly as long as it took someone to notice and restart. It is ordinary deploy lag wearing this project's usual costume: a real change, present on disk, invisible where the operator was looking.

#7 is the sharpest version yet of the same family as #4: a file the live service reads was wrong, and nothing in the repo's own diff could show it, because the file itself isn't in the repo. #4 was a QA edit left in a tracked file by mistake; #7 was a correct, deliberate, reviewed code change that was still incomplete, because "every consumer" silently excluded a consumer that git cannot see.

#8 breaks the pattern from the other side: the router did what it was configured to do, routing each request to a model that fit. It was the one component that saw all the spend while being blind to whether any of it produced anything, so it measured cost when the thing that mattered was progress. Whether its own pinch pruning drove the loops is still a theory (see #8, Theory 1). If confirmed, #8 joins the family of defensible settings, a 50k budget chosen for ~130k sessions, that turned harmful when the input changed shape.

The generalisable fix is the same each time: compute the thing that is actually true, and surface it where the operator is already looking. That is what the /metrics warnings and the inline admin warning are for, and it is the reason the catalog-age and ceiling checks exist at all. #6 and #7 are the two cases here where that thing has not been computed yet — see their "known gap" and "rule" respectively.