# Incidents: how this router has broken, and how to tell which one it is Seven times now, a change that looked local to the router has silently degraded either the agent depending on it or the operator trying to see it clearly. They share a shape worth naming: **none of them announce themselves as router problems.** Five of the seven presented as an opaque client-side error — a connection refused, an "Unprocessable Content", an "internal server error" — and diagnosing each meant knowing which log or table to look in. The other two didn't error toward the client at all: #5 destroyed data outright, and #6 just went quiet in one corner of the admin UI while `/health` stayed green the whole time. **#5 is the only one so far that destroyed data**, and the only one where recovery depended on luck rather than design. Read it before running any cleanup command in this repo. **#7 is the only one so far to take the whole service down from a clean config-rename PR**, not from an agent's own mistake — read it before renaming or removing any config key. This page exists so the next one takes minutes rather than hours. Start with the symptom table, then read only the relevant section. `opencode.json` points opencode's own model traffic at `http://127.0.0.1:8080/v1`, so on this machine the router *is* the coding agent's inference supply. Breaking it breaks the thing you would use to fix it. That is why these keep happening, and why the diagnostics below are worth having to hand. ## Symptom → first check | What you see | Likely | One-line check | |---|---|---| | `ConnectionError: Connection refused`, retrying | #1 killed by name/port | `systemctl --user status llm-router` — recently restarted, or inactive | | `active (running)` but nothing answers | #2 shutdown hang | `curl -s localhost:8080/health` fails while systemd says active | | Agent dies mid-task on an opaque 4xx | #3 candidate set empty | `sqlite3 router.db "select task_tier, required_context_tokens, rejected_reason from route_decisions where selected_model is null order by id desc limit 5;"` | | "internal server error", dashboard blank | #4 config/db path | `journalctl --user -u llm-router --since '10 min ago' \| grep -c '" 500'` then `git diff config/config.yaml` | | Prices/windows look wrong, nothing errors | catalog frozen | `sqlite3 router.db "select max(last_updated) from models;"` | | 500s + `no such table`, venv/.env gone | #5 `git clean -fdx` | `ls -la router.db .env .venv` — a 0-byte db and a missing `.venv` is conclusive | | A newly-shipped admin knob just isn't there, `/health` is fine | #6 backend running stale code | `systemctl --user status llm-router` uptime vs. `git log -1 --format=%cd` on the commit that added the knob | | `activating (auto-restart)`, `pydantic_core...ValidationError: ...extra_forbidden` in the journal | #7 renamed key vs. un-migrated overlay | `journalctl --user -u llm-router --since '5 min ago' \| grep extra_forbidden` then `grep -n config/config.local.yaml` | That last row is not an incident yet — it is the silent-staleness failure described in the "Run as a service" section of `CLAUDE.md`. An unpolled catalog fails **open**: it keeps routing on data that may be weeks old and every row still reads `active`. --- ## #1 — Killed by name or port (2026-08-29) **Symptom.** The dispatcher goes unreachable; the client retries against a refused connection. **Cause.** An agent ran `pkill -f "uvicorn dispatcher:app"` (also `.*dispatcher`, also plain `"8080"`) to free the port for its own throwaway instance. When the port came back 5s later — `Restart=always` resurrecting the supervised service — it read that as "the kill didn't work" and escalated to `pkill -9`. **Root cause was documentation, not code.** `AGENTS.md` offered a bare `python -m uvicorn dispatcher:app --reload` as an equally-valid way to bring the router up, with no warning that 8080 is normally already held. An agent following that instruction and finding the port busy has no way to know the right move is `systemctl --user restart`. **Fixed** in `AGENTS.md`: ad hoc runs bind `--port 8081` explicitly, and the section says outright never to send a kill signal to anything matched by name or port. Full audit trail in `plans/router-unreachable-signal-investigation.md` — the syscall-level extension to `pidfd_send_signal` is what finally caught it, after `kill`/`tgkill` both came back clean. **Convention that came out of it:** 8080 is production, always. Throwaway instances bind 8081. Never signal a process matched by name or port rather than by a PID you started yourself. ## #2 — Bare SIGTERM hangs the process forever (2026-08-29) **Symptom.** `systemctl --user status` reports `active (running)`; the socket is closed and nothing answers. The process logged `Waiting for connections to close` and never got past it. **Cause.** SSE clients holding `/events/decisions` open (the TUI, the admin dashboard) never disconnect, so uvicorn's graceful shutdown has nothing to wait out. That alone is only slow. What made it permanent: systemd enforces `TimeoutStopUSec` only when *it* is running the stop job, so a signal delivered outside that path leaves the unit `active` forever — the main PID never exits, so `Restart=` never fires either. **Fixed** in `deploy/llm-router.service`: `--timeout-graceful-shutdown 5` caps the drain regardless of who sends the signal. This also fixed ordinary restarts, which were silently taking the full 10s-then-SIGKILL path for the same reason. **Follow-up the fix exposed:** once the process could exit cleanly on SIGTERM, `Restart=on-failure` excluded it from auto-restart — systemd assumes SIGTERM means someone deliberately asked it to stop. Changed to `Restart=always`. Deliberate `systemctl stop`/`restart` are still honoured; systemd tracks those separately from the `Restart=` decision. **Recovery:** `systemctl --user restart llm-router.service`. Since a hung process was never in a tracked stop job, this issues a fresh cycle that systemd *does* enforce the timeout on. ## #3 — Admin overrides collapsed the tier-3 context ceiling (2026-09-01) **Symptom.** An agent failed mid-task with "Unprocessable Content". Surfaced ~19 hours after the cause. **Cause.** Seven expensive models were deprecated through `/admin` — a reasonable cost decision in isolation. The `models` table still read `active` (the poller refreshes it every 2h), but `admin_model_overrides` overlays `deprecated` and `_admin_deprecated_models` feeds that into routing's hard filters. Effect on the maximum servable `required_context_tokens`: | tier | before | during | eligible models | |---|---|---|---| | 1 | 782,324 | 782,324 | 6 | | 2 | 782,324 | 782,324 | 5 | | 3 | 782,324 | **94,196** | **1** | Tier-3 traffic has an observed max of 268,168 tokens, so every tier-3 request above 94k returned `422 No model satisfies the hard filters`. Nothing warned; the portal reported the change as a plain success. **The obvious check for this is wrong.** Warning when a higher tier's ceiling sits below a lower tier's is a *theorem*, not a fault: `ceiling(T)` is the max `effective_context_window` over models with `tier >= T`, and tier is a capability floor, so the eligible set shrinks monotonically and `ceiling(1) >= ceiling(2) >= ceiling(3)` holds for every catalog. Such a warning fires always and means nothing. The real detector compares the ceiling against **observed demand** — silent on all three tiers today, fires on the outage state. **Fixed** by the `/metrics` ceiling warnings and an inline warning at the admin availability toggle, so the cost of a deprecation is visible while looking at the switch. See `plans/catalog-staleness-and-poller-failure-modes.md` §4.4. **Recovery:** re-activate enough large-window rows to make the ladder continuous. Today: 94,196 → 192,500 → 782,324. ## #4 — Tracked config pointed at a nonexistent database (2026-09-01) **Symptom.** "Internal server error" in the client; the admin dashboard blank. **Cause.** A QA step edited the **tracked** `config/config.yaml`, changing `database.path` from `router.db` to `/tmp/router-qa/router.db`, created the directory but never a database in it, and left the edit in the working tree. `config/config.yaml` is the file the live systemd service reads. ``` 1753 200 201 500 Internal Server Error sqlite3.OperationalError: unable to open database file 115 GET /metrics · 65 GET /health · 14 /admin/api/snapshot 14 /admin/api/history · 7 POST /v1/chat/completions ``` Those 7 chat completions are real inference failing, not dashboard noise. **Fixed** by reverting the one line and restarting. It was never committed. **Rule:** never point `config/config.yaml` at test fixtures. To run against a throwaway database, pass a different config file, monkeypatch `cfg.database.path` in-process, or use a temp copy. If you must touch a file the live service reads, restore it in the same step and verify `curl -s localhost:8080/health` before moving on. A QA step that leaves production broken has not passed. ## #5 — `git clean -fdx` destroyed the database, key and venv (2026-09-04) **Symptom.** Router returned 500 on every request. Journal showed `sqlite3.OperationalError: no such table: energy_observations` while systemd reported the service `active`. **Cause.** An agent ran `git clean -fdx` in the repo. The `-x` flag removes **ignored** files as well as untracked ones, and everything this deployment needs to run is ignored by design: | lost | what it was | |---|---| | `router.db` | truncated to 0 bytes — 22,776 energy observations, 17,321 route decisions, 148 proficiency rows | | `.env` | the NeuralWatt API key | | `.venv` | the virtualenv the systemd unit's `ExecStart` runs from | | `config/config.local.yaml` | the operator's electricity tariff | | `node_modules` | | This is worse than it looks from the command. `git clean -fd` is a reasonable thing for an agent to run to get a clean tree. Adding `-x` turns it from "discard my scratch files" into "delete the deployment", and nothing in the repo warns you. **Recovery was luck, not design.** There is no backup of `router.db` by policy. What saved it was that the plan running at the time had made a QA copy at `/tmp/qa-config-local-overlay-r2/` **25 seconds before the wipe**, and that copy happened to include `.env`. Integrity check passed and the restore was effectively lossless. Had that plan been a different one, the entire measurement history of the project would be gone. **Recovery steps, in order:** ```bash # 1. find a surviving copy — QA/scratch dirs are the likely place find /home/alee /tmp -name "router.db" -size +0 sqlite3 "select count(*) from energy_observations;" sqlite3 "select integrity_check from pragma_integrity_check limit 1;" # 2. restore data, key, overlay cp -f router.db cp -f /.env .env && chmod 600 .env # config/config.local.yaml from your own copy # 3. rebuild the venv (pinned, so this is deterministic) python3 -m venv .venv && .venv/bin/pip install -r requirements.txt # 4. restart and verify systemctl --user restart llm-router.service curl -s localhost:8080/health ``` **What actually limited the damage** was two unrelated decisions made minutes earlier: the in-flight plan's 11 files had just been committed rather than left uncommitted, and the operator's tariff had been parked *outside* the repo instead of restored in place. Both were reactions to the same file being clobbered repeatedly that day — the seventh time is what prompted moving it out of reach. **Known gap this leaves open.** `router.db` has no backup policy. It holds every energy/cost observation, every routing decision, and all proficiency scores — none of which can be rebuilt without re-running evals that cost real money, and the historical observations cannot be rebuilt at all. A periodic snapshot is cheap insurance and does not exist. **Rules that came out of it:** - **Never `git clean -x` in this repo.** Use targeted paths. If you need a clean tree, `git stash` preserves; `clean` destroys. - **Never clean untracked files you did not create.** An untracked file in this tree is as likely to be operator data as build residue. - Before any destructive git command, ask what is *ignored*, not just what is untracked. Here that list is the database, the API key and the runtime. ## #6 — Admin knobs shipped to disk, invisible until the process restarted (2026-09-17) **Symptom.** Runtime knobs recently added to the admin portal (the incumbent- gate dial from PR #89, the session-cache-window controls from PR #90) were not displaying, despite both being merged and present in the working tree. `/health` and every other page reported normal the entire time. **Cause.** `admin/frontend/controls.html` and its sibling pages are served fresh off disk on every load — confirmed by the `GET /admin/api/frontend-version` digest added for incident-adjacent work on 2026-09-13, which exists precisely because frontend files are read live, not baked into the process. `src/admin.py` and `src/config.py` are not: they are imported once at process start and stay exactly as they were until the process restarts, no matter what lands on disk afterward. Both PR #89 and PR #90 shipped their frontend half and backend half together, correctly, per this project's own north-star rule that a knob needs both — but the two halves have different deploy timing. The frontend half takes effect on the very next page load. The backend half takes effect only on the next restart. In the window between "merged and pulled" and "service restarted," a fresh page load runs new frontend code against an old backend process: the new knob's controls call admin API fields and endpoints that do not exist yet in the running process's memory, so they render empty rather than erroring, while everything unrelated keeps working — the mismatch is scoped to exactly the new surface, which is what makes it easy to miss. **Fixed** by `systemctl --user restart llm-router.service` — mechanically the same recovery as #1 and #4, for a different reason: the process was not dead or misconfigured, it was correct for a version of the code that no longer matched what was on disk. **The reverse case was already caught; this direction was not.** `/admin/api/frontend-version` detects exactly the opposite mismatch — an *open browser tab* holding stale frontend JS against a *newer* backend — and prompts a reload. Nothing detects a *fresh* frontend load against a *stale* backend process, because nothing exposes what code the running process actually has loaded: no version stamp, no commit hash, no process-start time surfaced anywhere in the admin UI. **Known gap this leaves open.** Every future PR that ships an admin knob (frontend + backend together, exactly as required) reopens this same silent window between merge and restart. A `GET /admin/api/backend-version` returning the process's start time or the commit it was launched against — mirroring the existing frontend-version check, compared against `git rev-parse HEAD` on disk — would close it: the admin UI could show a "restart to pick up code changes" banner the same way it already handles the reverse direction. Not built. **Rule:** after merging any PR that touches `src/admin.py`, `src/config.py`, `src/dispatcher.py`, or anything else the systemd unit imports at start, restart `llm-router.service` before trusting what the admin portal shows. Merging code and deploying it are different steps here, and today only one of them happens automatically. ## #7 — A clean config rename crash-looped production via the un-migrated overlay (2026-09-18) **Symptom.** Minutes after merging and fast-forwarding a routine refactor PR, `systemctl --user status llm-router` showed `activating (auto-restart)` with `Result: exit-code`, cycling every ~6 seconds. Every attempt logged the same traceback, `admin.py`/`dispatcher.py` never got past `load_config`. **Cause.** The merged PR renamed `session_cache.staleness_minutes` to `session_cache.staleness_seconds` — a straight rename, no deprecated alias, matching this project's own stated convention (see the quota-shape removal: "every consumer was updated in the same change, so there are no deprecated aliases"). Every *tracked* consumer was updated in that same change: `config/ config.yaml`, `src/config.py`, `src/admin.py`, `src/dispatcher.py`, tests, docs. One consumer isn't tracked and can't be: `config/config.local.yaml`, the gitignored, machine-local overlay, which had `staleness_minutes: 2` sitting in it from an earlier tuning session. `StrictModel`'s `extra="forbid"` rejected that now-unknown key on every single startup: ``` pydantic_core._pydantic_core.ValidationError: 1 validation error for RouterConfig session_cache.staleness_minutes Extra inputs are not permitted [type=extra_forbidden, input_value=2, input_type=int] ``` Unlike #6, this wasn't a stale-process-vs-fresh-disk mismatch — the process could not start at all, because `dispatcher.py`'s module-level `cfg = load_config(...)` runs at import time, before the app can serve anything. `git diff`/`git log` show nothing wrong, because nothing tracked *is* wrong; the break lives entirely in a file no diff will ever show. **Fixed** by migrating the overlay by hand: backed up `config/config.local.yaml`, diffed before/after to confirm only the one line changed, converted the value (`staleness_minutes: 2` → `staleness_seconds: 120`, same real duration), then restarted. Recovered on the first clean attempt — about 7 crash-restart cycles, under a minute of total downtime. **This was foreseen and still happened.** The PR that did the rename was explicitly told not to touch `config/config.local.yaml` and that its migration would happen "separately, directly, by the operator... after this PR merges" — correct instructions, followed correctly, and the gap between "PR merges" and "operator migrates the overlay" was still a live window a production restart landed inside. Knowing the risk existed did not close it; only the migration itself did. **Rule:** a rename or removal of any config key is only complete when `config/config.local.yaml` has been checked, not just the tracked files — `grep -n config/config.local.yaml` before merging, or at minimum before the next restart. This is a corollary of the existing "`config.local.yaml` is irreplaceable" guardrail, extended to include a config key's *name* as part of what merging code can silently invalidate, not only its content. A rename PR that ships without an accompanying "does anything override this?" check on the overlay is incomplete, even when every tracked file is correct. --- ## The pattern The first five are defensible local decisions — free a port, stop a process, deprecate an expensive model, point at a test database, clean the working tree — that silently removed capability while the router kept reporting itself healthy. Three of the five were caused by an agent working on the router, using the router. #5 breaks the pattern in one way worth noting: it did not degrade quietly, it failed loudly and immediately. What it removed was not capability but *state* — and state, unlike capability, does not come back when you fix the code. #6 breaks it a different way again: nobody made a decision at all. Nothing was disabled, deprecated, or misconfigured — the code and the running process simply disagreed about what existed, for exactly as long as it took someone to notice and restart. It is ordinary deploy lag wearing this project's usual costume: a real change, present on disk, invisible where the operator was looking. #7 is the sharpest version yet of the same family as #4: a file the live service reads was wrong, and nothing in the repo's own diff could show it, because the file itself isn't in the repo. #4 was a QA edit left in a tracked file by mistake; #7 was a *correct*, deliberate, reviewed code change that was still incomplete, because "every consumer" silently excluded a consumer that git cannot see. The generalisable fix is the same each time: **compute the thing that is actually true, and surface it where the operator is already looking.** That is what the `/metrics` warnings and the inline admin warning are for, and it is the reason the catalog-age and ceiling checks exist at all. #6 and #7 are the two cases here where that thing has not been computed yet — see their "known gap" and "rule" respectively.