Two incidents from the last day's deploy cycle, same family as #4 -- a file the live service depends on was wrong in a way no tracked diff can show. #6: PR #89/#90 shipped admin knobs (frontend + backend together, correctly), but the running process only picks up the backend half on restart while the frontend half is served fresh every request -- so a fresh page load against an unrestarted process silently omitted the new controls, with /health staying green throughout. #7: the session_cache.staleness_minutes -> staleness_seconds rename updated every tracked consumer in the same change, per this project's no-alias convention, but config/config.local.yaml (gitignored, machine-local) still held the old key. StrictModel's extra_forbidden rejected it on every startup attempt, crash-looping production for ~7 restart cycles until the overlay was migrated by hand. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0161UAZyxtvmKaVzasJsnqJc
20 KiB
Incidents: how this router has broken, and how to tell which one it is
Seven times now, a change that looked local to the router has silently
degraded either the agent depending on it or the operator trying to see it
clearly. They share a shape worth naming: none of them announce themselves
as router problems. Five of the seven presented as an opaque client-side
error — a connection refused, an "Unprocessable Content", an "internal server
error" — and diagnosing each meant knowing which log or table to look in. The
other two didn't error toward the client at all: #5 destroyed data outright,
and #6 just went quiet in one corner of the admin UI while /health stayed
green the whole time.
#5 is the only one so far that destroyed data, and the only one where recovery depended on luck rather than design. Read it before running any cleanup command in this repo.
#7 is the only one so far to take the whole service down from a clean config-rename PR, not from an agent's own mistake — read it before renaming or removing any config key.
This page exists so the next one takes minutes rather than hours. Start with the symptom table, then read only the relevant section.
opencode.json points opencode's own model traffic at http://127.0.0.1:8080/v1,
so on this machine the router is the coding agent's inference supply. Breaking
it breaks the thing you would use to fix it. That is why these keep happening,
and why the diagnostics below are worth having to hand.
Symptom → first check
| What you see | Likely | One-line check |
|---|---|---|
ConnectionError: Connection refused, retrying |
#1 killed by name/port | systemctl --user status llm-router — recently restarted, or inactive |
active (running) but nothing answers |
#2 shutdown hang | curl -s localhost:8080/health fails while systemd says active |
| Agent dies mid-task on an opaque 4xx | #3 candidate set empty | sqlite3 router.db "select task_tier, required_context_tokens, rejected_reason from route_decisions where selected_model is null order by id desc limit 5;" |
| "internal server error", dashboard blank | #4 config/db path | journalctl --user -u llm-router --since '10 min ago' | grep -c '" 500' then git diff config/config.yaml |
| Prices/windows look wrong, nothing errors | catalog frozen | sqlite3 router.db "select max(last_updated) from models;" |
500s + no such table, venv/.env gone |
#5 git clean -fdx |
ls -la router.db .env .venv — a 0-byte db and a missing .venv is conclusive |
A newly-shipped admin knob just isn't there, /health is fine |
#6 backend running stale code | systemctl --user status llm-router uptime vs. git log -1 --format=%cd on the commit that added the knob |
activating (auto-restart), pydantic_core...ValidationError: ...extra_forbidden in the journal |
#7 renamed key vs. un-migrated overlay | journalctl --user -u llm-router --since '5 min ago' | grep extra_forbidden then grep -n <old key name> config/config.local.yaml |
That last row is not an incident yet — it is the silent-staleness failure
described in the "Run as a service" section of CLAUDE.md. An unpolled catalog
fails open: it keeps routing on data that may be weeks old and every row
still reads active.
#1 — Killed by name or port (2026-08-29)
Symptom. The dispatcher goes unreachable; the client retries against a refused connection.
Cause. An agent ran pkill -f "uvicorn dispatcher:app" (also .*dispatcher,
also plain "8080") to free the port for its own throwaway instance. When the
port came back 5s later — Restart=always resurrecting the supervised service —
it read that as "the kill didn't work" and escalated to pkill -9.
Root cause was documentation, not code. AGENTS.md offered a bare
python -m uvicorn dispatcher:app --reload as an equally-valid way to bring the
router up, with no warning that 8080 is normally already held. An agent
following that instruction and finding the port busy has no way to know the
right move is systemctl --user restart.
Fixed in AGENTS.md: ad hoc runs bind --port 8081 explicitly, and the
section says outright never to send a kill signal to anything matched by name or
port. Full audit trail in
plans/router-unreachable-signal-investigation.md — the syscall-level extension
to pidfd_send_signal is what finally caught it, after kill/tgkill both came
back clean.
Convention that came out of it: 8080 is production, always. Throwaway instances bind 8081. Never signal a process matched by name or port rather than by a PID you started yourself.
#2 — Bare SIGTERM hangs the process forever (2026-08-29)
Symptom. systemctl --user status reports active (running); the socket is
closed and nothing answers. The process logged Waiting for connections to close and never got past it.
Cause. SSE clients holding /events/decisions open (the TUI, the admin
dashboard) never disconnect, so uvicorn's graceful shutdown has nothing to wait
out. That alone is only slow. What made it permanent: systemd enforces
TimeoutStopUSec only when it is running the stop job, so a signal delivered
outside that path leaves the unit active forever — the main PID never exits,
so Restart= never fires either.
Fixed in deploy/llm-router.service: --timeout-graceful-shutdown 5 caps
the drain regardless of who sends the signal. This also fixed ordinary restarts,
which were silently taking the full 10s-then-SIGKILL path for the same reason.
Follow-up the fix exposed: once the process could exit cleanly on SIGTERM,
Restart=on-failure excluded it from auto-restart — systemd assumes SIGTERM
means someone deliberately asked it to stop. Changed to Restart=always.
Deliberate systemctl stop/restart are still honoured; systemd tracks those
separately from the Restart= decision.
Recovery: systemctl --user restart llm-router.service. Since a hung
process was never in a tracked stop job, this issues a fresh cycle that systemd
does enforce the timeout on.
#3 — Admin overrides collapsed the tier-3 context ceiling (2026-09-01)
Symptom. An agent failed mid-task with "Unprocessable Content". Surfaced ~19 hours after the cause.
Cause. Seven expensive models were deprecated through /admin — a
reasonable cost decision in isolation. The models table still read active
(the poller refreshes it every 2h), but admin_model_overrides overlays
deprecated and _admin_deprecated_models feeds that into routing's hard
filters. Effect on the maximum servable required_context_tokens:
| tier | before | during | eligible models |
|---|---|---|---|
| 1 | 782,324 | 782,324 | 6 |
| 2 | 782,324 | 782,324 | 5 |
| 3 | 782,324 | 94,196 | 1 |
Tier-3 traffic has an observed max of 268,168 tokens, so every tier-3 request
above 94k returned 422 No model satisfies the hard filters. Nothing warned;
the portal reported the change as a plain success.
The obvious check for this is wrong. Warning when a higher tier's ceiling
sits below a lower tier's is a theorem, not a fault: ceiling(T) is the max
effective_context_window over models with tier >= T, and tier is a capability
floor, so the eligible set shrinks monotonically and
ceiling(1) >= ceiling(2) >= ceiling(3) holds for every catalog. Such a warning
fires always and means nothing. The real detector compares the ceiling against
observed demand — silent on all three tiers today, fires on the outage state.
Fixed by the /metrics ceiling warnings and an inline warning at the admin
availability toggle, so the cost of a deprecation is visible while looking at the
switch. See plans/catalog-staleness-and-poller-failure-modes.md §4.4.
Recovery: re-activate enough large-window rows to make the ladder continuous. Today: 94,196 → 192,500 → 782,324.
#4 — Tracked config pointed at a nonexistent database (2026-09-01)
Symptom. "Internal server error" in the client; the admin dashboard blank.
Cause. A QA step edited the tracked config/config.yaml, changing
database.path from router.db to /tmp/router-qa/router.db, created the
directory but never a database in it, and left the edit in the working tree.
config/config.yaml is the file the live systemd service reads.
1753 200
201 500 Internal Server Error
sqlite3.OperationalError: unable to open database file
115 GET /metrics · 65 GET /health · 14 /admin/api/snapshot
14 /admin/api/history · 7 POST /v1/chat/completions
Those 7 chat completions are real inference failing, not dashboard noise.
Fixed by reverting the one line and restarting. It was never committed.
Rule: never point config/config.yaml at test fixtures. To run against a
throwaway database, pass a different config file, monkeypatch
cfg.database.path in-process, or use a temp copy. If you must touch a file the
live service reads, restore it in the same step and verify
curl -s localhost:8080/health before moving on. A QA step that leaves
production broken has not passed.
#5 — git clean -fdx destroyed the database, key and venv (2026-09-04)
Symptom. Router returned 500 on every request. Journal showed
sqlite3.OperationalError: no such table: energy_observations while systemd
reported the service active.
Cause. An agent ran git clean -fdx in the repo. The -x flag removes
ignored files as well as untracked ones, and everything this deployment
needs to run is ignored by design:
| lost | what it was |
|---|---|
router.db |
truncated to 0 bytes — 22,776 energy observations, 17,321 route decisions, 148 proficiency rows |
.env |
the NeuralWatt API key |
.venv |
the virtualenv the systemd unit's ExecStart runs from |
config/config.local.yaml |
the operator's electricity tariff |
node_modules |
This is worse than it looks from the command. git clean -fd is a reasonable
thing for an agent to run to get a clean tree. Adding -x turns it from
"discard my scratch files" into "delete the deployment", and nothing in the
repo warns you.
Recovery was luck, not design. There is no backup of router.db by policy.
What saved it was that the plan running at the time had made a QA copy at
/tmp/qa-config-local-overlay-r2/ 25 seconds before the wipe, and that
copy happened to include .env. Integrity check passed and the restore was
effectively lossless. Had that plan been a different one, the entire
measurement history of the project would be gone.
Recovery steps, in order:
# 1. find a surviving copy — QA/scratch dirs are the likely place
find /home/alee /tmp -name "router.db" -size +0
sqlite3 <candidate> "select count(*) from energy_observations;"
sqlite3 <candidate> "select integrity_check from pragma_integrity_check limit 1;"
# 2. restore data, key, overlay
cp -f <candidate> router.db
cp -f <candidate-dir>/.env .env && chmod 600 .env
# config/config.local.yaml from your own copy
# 3. rebuild the venv (pinned, so this is deterministic)
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
# 4. restart and verify
systemctl --user restart llm-router.service
curl -s localhost:8080/health
What actually limited the damage was two unrelated decisions made minutes earlier: the in-flight plan's 11 files had just been committed rather than left uncommitted, and the operator's tariff had been parked outside the repo instead of restored in place. Both were reactions to the same file being clobbered repeatedly that day — the seventh time is what prompted moving it out of reach.
Known gap this leaves open. router.db has no backup policy. It holds
every energy/cost observation, every routing decision, and all proficiency
scores — none of which can be rebuilt without re-running evals that cost real
money, and the historical observations cannot be rebuilt at all. A periodic
snapshot is cheap insurance and does not exist.
Rules that came out of it:
- Never
git clean -xin this repo. Use targeted paths. If you need a clean tree,git stashpreserves;cleandestroys. - Never clean untracked files you did not create. An untracked file in this tree is as likely to be operator data as build residue.
- Before any destructive git command, ask what is ignored, not just what is untracked. Here that list is the database, the API key and the runtime.
#6 — Admin knobs shipped to disk, invisible until the process restarted (2026-09-17)
Symptom. Runtime knobs recently added to the admin portal (the incumbent-
gate dial from PR #89, the session-cache-window controls from PR #90) were not
displaying, despite both being merged and present in the working tree.
/health and every other page reported normal the entire time.
Cause. admin/frontend/controls.html and its sibling pages are served
fresh off disk on every load — confirmed by the GET /admin/api/frontend-version digest added for incident-adjacent work on
2026-09-13, which exists precisely because frontend files are read live, not
baked into the process. src/admin.py and src/config.py are not: they are
imported once at process start and stay exactly as they were until the
process restarts, no matter what lands on disk afterward. Both PR #89 and
PR #90 shipped their frontend half and backend half together, correctly, per
this project's own north-star rule that a knob needs both — but the two
halves have different deploy timing. The frontend half takes effect on the
very next page load. The backend half takes effect only on the next restart.
In the window between "merged and pulled" and "service restarted," a fresh
page load runs new frontend code against an old backend process: the new
knob's controls call admin API fields and endpoints that do not exist yet in
the running process's memory, so they render empty rather than erroring,
while everything unrelated keeps working — the mismatch is scoped to exactly
the new surface, which is what makes it easy to miss.
Fixed by systemctl --user restart llm-router.service — mechanically the
same recovery as #1 and #4, for a different reason: the process was not dead
or misconfigured, it was correct for a version of the code that no longer
matched what was on disk.
The reverse case was already caught; this direction was not.
/admin/api/frontend-version detects exactly the opposite mismatch — an
open browser tab holding stale frontend JS against a newer backend — and
prompts a reload. Nothing detects a fresh frontend load against a stale
backend process, because nothing exposes what code the running process
actually has loaded: no version stamp, no commit hash, no process-start time
surfaced anywhere in the admin UI.
Known gap this leaves open. Every future PR that ships an admin knob
(frontend + backend together, exactly as required) reopens this same silent
window between merge and restart. A GET /admin/api/backend-version
returning the process's start time or the commit it was launched against —
mirroring the existing frontend-version check, compared against git rev-parse HEAD on disk — would close it: the admin UI could show a "restart
to pick up code changes" banner the same way it already handles the reverse
direction. Not built.
Rule: after merging any PR that touches src/admin.py, src/config.py,
src/dispatcher.py, or anything else the systemd unit imports at start,
restart llm-router.service before trusting what the admin portal shows.
Merging code and deploying it are different steps here, and today only one of
them happens automatically.
#7 — A clean config rename crash-looped production via the un-migrated overlay (2026-09-18)
Symptom. Minutes after merging and fast-forwarding a routine refactor PR,
systemctl --user status llm-router showed activating (auto-restart) with
Result: exit-code, cycling every ~6 seconds. Every attempt logged the same
traceback, admin.py/dispatcher.py never got past load_config.
Cause. The merged PR renamed session_cache.staleness_minutes to
session_cache.staleness_seconds — a straight rename, no deprecated alias,
matching this project's own stated convention (see the quota-shape removal:
"every consumer was updated in the same change, so there are no deprecated
aliases"). Every tracked consumer was updated in that same change: config/ config.yaml, src/config.py, src/admin.py, src/dispatcher.py, tests,
docs. One consumer isn't tracked and can't be: config/config.local.yaml, the
gitignored, machine-local overlay, which had staleness_minutes: 2 sitting in
it from an earlier tuning session. StrictModel's extra="forbid" rejected
that now-unknown key on every single startup:
pydantic_core._pydantic_core.ValidationError: 1 validation error for RouterConfig
session_cache.staleness_minutes
Extra inputs are not permitted [type=extra_forbidden, input_value=2, input_type=int]
Unlike #6, this wasn't a stale-process-vs-fresh-disk mismatch — the process
could not start at all, because dispatcher.py's module-level cfg = load_config(...) runs at import time, before the app can serve anything.
git diff/git log show nothing wrong, because nothing tracked is wrong;
the break lives entirely in a file no diff will ever show.
Fixed by migrating the overlay by hand: backed up
config/config.local.yaml, diffed before/after to confirm only the one line
changed, converted the value (staleness_minutes: 2 → staleness_seconds: 120, same real duration), then restarted. Recovered on the first clean
attempt — about 7 crash-restart cycles, under a minute of total downtime.
This was foreseen and still happened. The PR that did the rename was
explicitly told not to touch config/config.local.yaml and that its
migration would happen "separately, directly, by the operator... after this
PR merges" — correct instructions, followed correctly, and the gap between
"PR merges" and "operator migrates the overlay" was still a live window a
production restart landed inside. Knowing the risk existed did not close it;
only the migration itself did.
Rule: a rename or removal of any config key is only complete when
config/config.local.yaml has been checked, not just the tracked files —
grep -n <old key name> config/config.local.yaml before merging, or at
minimum before the next restart. This is a corollary of the existing
"config.local.yaml is irreplaceable" guardrail, extended to include a config
key's name as part of what merging code can silently invalidate, not only
its content. A rename PR that ships without an accompanying "does anything
override this?" check on the overlay is incomplete, even when every tracked
file is correct.
The pattern
The first five are defensible local decisions — free a port, stop a process, deprecate an expensive model, point at a test database, clean the working tree — that silently removed capability while the router kept reporting itself healthy. Three of the five were caused by an agent working on the router, using the router.
#5 breaks the pattern in one way worth noting: it did not degrade quietly, it failed loudly and immediately. What it removed was not capability but state — and state, unlike capability, does not come back when you fix the code.
#6 breaks it a different way again: nobody made a decision at all. Nothing was disabled, deprecated, or misconfigured — the code and the running process simply disagreed about what existed, for exactly as long as it took someone to notice and restart. It is ordinary deploy lag wearing this project's usual costume: a real change, present on disk, invisible where the operator was looking.
#7 is the sharpest version yet of the same family as #4: a file the live service reads was wrong, and nothing in the repo's own diff could show it, because the file itself isn't in the repo. #4 was a QA edit left in a tracked file by mistake; #7 was a correct, deliberate, reviewed code change that was still incomplete, because "every consumer" silently excluded a consumer that git cannot see.
The generalisable fix is the same each time: compute the thing that is
actually true, and surface it where the operator is already looking. That is
what the /metrics warnings and the inline admin warning are for, and it is the
reason the catalog-age and ceiling checks exist at all. #6 and #7 are the two
cases here where that thing has not been computed yet — see their "known gap"
and "rule" respectively.