First incident in this project with real data loss, and the first where recovery depended on luck rather than design. An agent ran `git clean -fdx`. The -x flag removes IGNORED files as well as untracked ones, and everything this deployment needs to run is ignored by design: router.db (truncated to 0 bytes -- 22,776 energy observations, 17,321 route decisions, 148 proficiency rows), .env (the NeuralWatt API key), .venv (the virtualenv systemd's ExecStart runs from), and config/config.local.yaml (the operator's tariff). Router returned 500 with "no such table: energy_observations" while systemd reported active. `git clean -fd` is a reasonable thing for an agent to run. Adding -x turns it from "discard my scratch files" into "delete the deployment", and nothing in the repo warned about that. Recovery was luck: the plan running at the time had made a QA copy in /tmp 25 seconds before the wipe, and that copy happened to include .env. Restore was effectively lossless. A different plan and the entire measurement history would be gone. Records the recovery runbook, and the gap it exposes: router.db has no backup policy despite holding data that either costs money to rebuild (proficiency, via eval runs) or cannot be rebuilt at all (historical observations). Also notes what limited the damage -- committing the in-flight plan's 11 files minutes earlier, and parking the tariff outside the repo -- both of which were reactions to that same file being clobbered seven times the same day. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
12 KiB
Incidents: how this router has broken, and how to tell which one it is
Five times now, a change that looked local to the router has silently degraded the agent depending on it. They share a shape worth naming: none of them announce themselves as router problems. Four of the five presented as an opaque client-side error — a connection refused, an "Unprocessable Content", an "internal server error" — and diagnosing each meant knowing which log or table to look in.
#5 is the only one so far that destroyed data, and the only one where recovery depended on luck rather than design. Read it before running any cleanup command in this repo.
This page exists so the next one takes minutes rather than hours. Start with the symptom table, then read only the relevant section.
opencode.json points opencode's own model traffic at http://127.0.0.1:8080/v1,
so on this machine the router is the coding agent's inference supply. Breaking
it breaks the thing you would use to fix it. That is why these keep happening,
and why the diagnostics below are worth having to hand.
Symptom → first check
| What you see | Likely | One-line check |
|---|---|---|
ConnectionError: Connection refused, retrying |
#1 killed by name/port | systemctl --user status llm-router — recently restarted, or inactive |
active (running) but nothing answers |
#2 shutdown hang | curl -s localhost:8080/health fails while systemd says active |
| Agent dies mid-task on an opaque 4xx | #3 candidate set empty | sqlite3 router.db "select task_tier, required_context_tokens, rejected_reason from route_decisions where selected_model is null order by id desc limit 5;" |
| "internal server error", dashboard blank | #4 config/db path | journalctl --user -u llm-router --since '10 min ago' | grep -c '" 500' then git diff config/config.yaml |
| Prices/windows look wrong, nothing errors | catalog frozen | sqlite3 router.db "select max(last_updated) from models;" |
500s + no such table, venv/.env gone |
#5 git clean -fdx |
ls -la router.db .env .venv — a 0-byte db and a missing .venv is conclusive |
That last row is not an incident yet — it is the silent-staleness failure
described in the "Run as a service" section of CLAUDE.md. An unpolled catalog
fails open: it keeps routing on data that may be weeks old and every row
still reads active.
#1 — Killed by name or port (2026-08-29)
Symptom. The dispatcher goes unreachable; the client retries against a refused connection.
Cause. An agent ran pkill -f "uvicorn dispatcher:app" (also .*dispatcher,
also plain "8080") to free the port for its own throwaway instance. When the
port came back 5s later — Restart=always resurrecting the supervised service —
it read that as "the kill didn't work" and escalated to pkill -9.
Root cause was documentation, not code. AGENTS.md offered a bare
python -m uvicorn dispatcher:app --reload as an equally-valid way to bring the
router up, with no warning that 8080 is normally already held. An agent
following that instruction and finding the port busy has no way to know the
right move is systemctl --user restart.
Fixed in AGENTS.md: ad hoc runs bind --port 8081 explicitly, and the
section says outright never to send a kill signal to anything matched by name or
port. Full audit trail in
plans/router-unreachable-signal-investigation.md — the syscall-level extension
to pidfd_send_signal is what finally caught it, after kill/tgkill both came
back clean.
Convention that came out of it: 8080 is production, always. Throwaway instances bind 8081. Never signal a process matched by name or port rather than by a PID you started yourself.
#2 — Bare SIGTERM hangs the process forever (2026-08-29)
Symptom. systemctl --user status reports active (running); the socket is
closed and nothing answers. The process logged Waiting for connections to close and never got past it.
Cause. SSE clients holding /events/decisions open (the TUI, the admin
dashboard) never disconnect, so uvicorn's graceful shutdown has nothing to wait
out. That alone is only slow. What made it permanent: systemd enforces
TimeoutStopUSec only when it is running the stop job, so a signal delivered
outside that path leaves the unit active forever — the main PID never exits,
so Restart= never fires either.
Fixed in deploy/llm-router.service: --timeout-graceful-shutdown 5 caps
the drain regardless of who sends the signal. This also fixed ordinary restarts,
which were silently taking the full 10s-then-SIGKILL path for the same reason.
Follow-up the fix exposed: once the process could exit cleanly on SIGTERM,
Restart=on-failure excluded it from auto-restart — systemd assumes SIGTERM
means someone deliberately asked it to stop. Changed to Restart=always.
Deliberate systemctl stop/restart are still honoured; systemd tracks those
separately from the Restart= decision.
Recovery: systemctl --user restart llm-router.service. Since a hung
process was never in a tracked stop job, this issues a fresh cycle that systemd
does enforce the timeout on.
#3 — Admin overrides collapsed the tier-3 context ceiling (2026-09-01)
Symptom. An agent failed mid-task with "Unprocessable Content". Surfaced ~19 hours after the cause.
Cause. Seven expensive models were deprecated through /admin — a
reasonable cost decision in isolation. The models table still read active
(the poller refreshes it every 2h), but admin_model_overrides overlays
deprecated and _admin_deprecated_models feeds that into routing's hard
filters. Effect on the maximum servable required_context_tokens:
| tier | before | during | eligible models |
|---|---|---|---|
| 1 | 782,324 | 782,324 | 6 |
| 2 | 782,324 | 782,324 | 5 |
| 3 | 782,324 | 94,196 | 1 |
Tier-3 traffic has an observed max of 268,168 tokens, so every tier-3 request
above 94k returned 422 No model satisfies the hard filters. Nothing warned;
the portal reported the change as a plain success.
The obvious check for this is wrong. Warning when a higher tier's ceiling
sits below a lower tier's is a theorem, not a fault: ceiling(T) is the max
effective_context_window over models with tier >= T, and tier is a capability
floor, so the eligible set shrinks monotonically and
ceiling(1) >= ceiling(2) >= ceiling(3) holds for every catalog. Such a warning
fires always and means nothing. The real detector compares the ceiling against
observed demand — silent on all three tiers today, fires on the outage state.
Fixed by the /metrics ceiling warnings and an inline warning at the admin
availability toggle, so the cost of a deprecation is visible while looking at the
switch. See plans/catalog-staleness-and-poller-failure-modes.md §4.4.
Recovery: re-activate enough large-window rows to make the ladder continuous. Today: 94,196 → 192,500 → 782,324.
#4 — Tracked config pointed at a nonexistent database (2026-09-01)
Symptom. "Internal server error" in the client; the admin dashboard blank.
Cause. A QA step edited the tracked config/config.yaml, changing
database.path from router.db to /tmp/router-qa/router.db, created the
directory but never a database in it, and left the edit in the working tree.
config/config.yaml is the file the live systemd service reads.
1753 200
201 500 Internal Server Error
sqlite3.OperationalError: unable to open database file
115 GET /metrics · 65 GET /health · 14 /admin/api/snapshot
14 /admin/api/history · 7 POST /v1/chat/completions
Those 7 chat completions are real inference failing, not dashboard noise.
Fixed by reverting the one line and restarting. It was never committed.
Rule: never point config/config.yaml at test fixtures. To run against a
throwaway database, pass a different config file, monkeypatch
cfg.database.path in-process, or use a temp copy. If you must touch a file the
live service reads, restore it in the same step and verify
curl -s localhost:8080/health before moving on. A QA step that leaves
production broken has not passed.
#5 — git clean -fdx destroyed the database, key and venv (2026-09-04)
Symptom. Router returned 500 on every request. Journal showed
sqlite3.OperationalError: no such table: energy_observations while systemd
reported the service active.
Cause. An agent ran git clean -fdx in the repo. The -x flag removes
ignored files as well as untracked ones, and everything this deployment
needs to run is ignored by design:
| lost | what it was |
|---|---|
router.db |
truncated to 0 bytes — 22,776 energy observations, 17,321 route decisions, 148 proficiency rows |
.env |
the NeuralWatt API key |
.venv |
the virtualenv the systemd unit's ExecStart runs from |
config/config.local.yaml |
the operator's electricity tariff |
node_modules |
This is worse than it looks from the command. git clean -fd is a reasonable
thing for an agent to run to get a clean tree. Adding -x turns it from
"discard my scratch files" into "delete the deployment", and nothing in the
repo warns you.
Recovery was luck, not design. There is no backup of router.db by policy.
What saved it was that the plan running at the time had made a QA copy at
/tmp/qa-config-local-overlay-r2/ 25 seconds before the wipe, and that
copy happened to include .env. Integrity check passed and the restore was
effectively lossless. Had that plan been a different one, the entire
measurement history of the project would be gone.
Recovery steps, in order:
# 1. find a surviving copy — QA/scratch dirs are the likely place
find /home/alee /tmp -name "router.db" -size +0
sqlite3 <candidate> "select count(*) from energy_observations;"
sqlite3 <candidate> "select integrity_check from pragma_integrity_check limit 1;"
# 2. restore data, key, overlay
cp -f <candidate> router.db
cp -f <candidate-dir>/.env .env && chmod 600 .env
# config/config.local.yaml from your own copy
# 3. rebuild the venv (pinned, so this is deterministic)
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
# 4. restart and verify
systemctl --user restart llm-router.service
curl -s localhost:8080/health
What actually limited the damage was two unrelated decisions made minutes earlier: the in-flight plan's 11 files had just been committed rather than left uncommitted, and the operator's tariff had been parked outside the repo instead of restored in place. Both were reactions to the same file being clobbered repeatedly that day — the seventh time is what prompted moving it out of reach.
Known gap this leaves open. router.db has no backup policy. It holds
every energy/cost observation, every routing decision, and all proficiency
scores — none of which can be rebuilt without re-running evals that cost real
money, and the historical observations cannot be rebuilt at all. A periodic
snapshot is cheap insurance and does not exist.
Rules that came out of it:
- Never
git clean -xin this repo. Use targeted paths. If you need a clean tree,git stashpreserves;cleandestroys. - Never clean untracked files you did not create. An untracked file in this tree is as likely to be operator data as build residue.
- Before any destructive git command, ask what is ignored, not just what is untracked. Here that list is the database, the API key and the runtime.
The pattern
All five are defensible local decisions — free a port, stop a process, deprecate an expensive model, point at a test database, clean the working tree — that silently removed capability while the router kept reporting itself healthy. Three of the five were caused by an agent working on the router, using the router.
#5 breaks the pattern in one way worth noting: it did not degrade quietly, it failed loudly and immediately. What it removed was not capability but state — and state, unlike capability, does not come back when you fix the code.
The generalisable fix is the same each time: compute the thing that is
actually true, and surface it where the operator is already looking. That is
what the /metrics warnings and the inline admin warning are for, and it is the
reason the catalog-age and ceiling checks exist at all.