plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
15 KiB
Catalog staleness: the fuse is real, the timer is not what arms it
Status: done -- rejection_warnings and the demand-ceiling detector in metrics.py
Date: 2026-09-01
Status: FINAL — ready to implement. Small; one commit.
Prompted by: "is stale_after_days: 3 too tight?"
Answer: no — and it is also not the knob that governs the failure the docs
warn about. Both CLAUDE.md and README.md described this backwards and have
been corrected in the same pass.
1. What was actually checked
| poll cadence | every 2 hours, llm-router-poller.timer, healthy |
| poller errors, last 7 days | 0 |
| last run | upserted 19/19 models, 22 min before this was written |
| rows currently stale | 1 — deepseek-ai/DeepSeek-V4-Flash, genuinely gone from the catalog since 2026-08-24. Correct. |
tests covering fetch_neuralwatt or mark_stale |
none |
Against a 2-hourly poll, stale_after_days: 3 is 36 consecutive successful
polls of margin. Nothing has expired that should not have.
Recovery is automatic as well: upsert writes
availability = excluded.availability (poller.py:251), so a single good poll
flips every stale row back to active. The blast radius of a false stale-out is
one 2-hour cycle, not manual intervention.
2. The documented failure mode is not reachable
Both docs said, in substance: an unpolled catalog marks every row stale and the router returns zero candidates.
It cannot. mark_stale (poller.py:283) runs only inside poller.main(),
and main() returns early on requests.RequestException — before upsert
and before mark_stale. So:
| event | what actually happens |
|---|---|
| timer stopped / disabled | nothing runs, so nothing is marked. Catalog freezes at last-known-good; every row still reads active. |
| provider API down, DNS fails, network out | early return at the fetch. Same: catalog frozen, nothing marked. |
The real behaviour is silent and fail-open — the router keeps routing on prices, context windows and capability flags that may be weeks out of date, and nothing anywhere says so. That is arguably the better default (a soft failure beats a hard one), but it is the opposite of what the docs told you to watch for, so attention was pointed at the wrong risk.
3. The path that can empty the candidate set
fetch_neuralwatt (poller.py:161):
payload = resp.json()
rows = []
for m in payload.get("data", []):
No floor on row count. A 200 response carrying an empty or truncated data
array — a partial provider outage, a schema change, an auth path degrading to
an empty list — clears raise_for_status(), returns [], upserts nothing, and
then mark_stale runs anyway because main() never saw an exception.
Three days of that and every row is stale; exclude_stale: true then leaves
zero candidates for everything.
So the fuse exists — but stale_after_days only sets its length. It neither
arms nor disarms it. Lengthening it buys a longer fuse on something that
should not be armed at all.
Note this also breaks mark_stale's own docstring, which claims it flags "rows
that weren't touched by this poll run". After a zero-row fetch that is
vacuously every row, which is not what the sentence means.
4. The fix
4.1 Sanity floor in fetch_neuralwatt
Refuse an implausible catalog by raising, so main() takes its existing
early-return path — which is already the correct behaviour for a failed fetch.
- Zero rows → raise. Unconditional. An empty catalog is never a real answer from a provider that has 19 models.
- Below ~50% of the current row count → log a warning and proceed. A legitimate mass deprecation should not be blocked outright, but it should not pass unremarked either.
Raise a RequestException subclass (or let main() catch a new
CatalogTooSmall alongside it) so the existing handler covers it and the exit
code stays 1 — systemd then reports the unit as failed, which is the loud signal
this path currently lacks.
The row count to compare against is SELECT COUNT(*) FROM models WHERE provider = 'neuralwatt', so this needs the connection. Either pass it in, or move the
check into main() between fetch_neuralwatt and upsert — the latter keeps
fetch_neuralwatt free of DB access, which matches the module's existing shape.
Prefer the check in main().
4.2 Then leave stale_after_days at 3 — or shorten it
Once the floor exists, availability = 'stale' means exactly one thing: this
row was absent from a catalog we successfully fetched. That is a direct
observation, not something needing a three-day timer to establish. A genuinely
removed public model currently stays routable for three days after it vanishes;
with the floor in place, 1 day (or two consecutive polls) is defensible and
tighter.
This is a judgement call with no measurement behind it either way, so the recommendation is: land the floor, leave the value at 3, revisit only if a removed model actually causes a dispatch failure. Changing both at once would make the next incident ambiguous.
4.3 Surface catalog age, since the real failure is silence
metrics.py already builds a warnings list (metrics.py:131) that the admin
dashboard's warnings bell and the TUI both render. Add one:
catalog last polled 4.2 days ago — prices, context windows and capability
flags may be out of date. Check: systemctl --user status llm-router-poller.timer
Threshold at max(1 day, 2 x stale_after_days / 3) or simply 1 day — the poll
is 2-hourly, so anything past a day means the timer is not running. This is the
change that actually addresses the fail-open behaviour in §2: it converts a
silent freeze into something visible without changing the (correct) decision to
keep serving.
4.4 Surface context-ceiling collapse — the other path that empties the candidate set
Added 2026-09-01, after this exact failure took the router down mid-task.
§3 covers one way to reach zero candidates. Admin availability overrides are another, and unlike the poller path it fires instantly, silently, and by a deliberate operator action that gives no indication of what it just cost.
What happened. Seven models were deprecated through /admin at 03:11-03:12
on 2026-09-01 — the expensive ones, a reasonable cost decision in isolation. The
models table still read active (the poller refreshes it every 2h), but
admin_model_overrides overlays deprecated and _admin_deprecated_models
(dispatcher.py:1977) feeds that into routing's hard filters. Measured effect on
the maximum servable required_context_tokens, interactive:
| tier | ceiling before | ceiling during | eligible models |
|---|---|---|---|
| 1 | 782,324 | 782,324 | 6 |
| 2 | 782,324 | 782,324 | 5 |
| 3 | 782,324 | 94,196 | 1 |
Tier 3's ceiling fell 8.3x while tiers 1 and 2 were untouched. Observed
tier-3 traffic has a max of 268,168 tokens and a p95 of 179,356, so every
tier-3 request above 94,196 returned 422 No model satisfies the hard filters.
Nothing warned, nothing logged above INFO, and the admin portal reported the
change as a plain success. It surfaced ~19 hours later as an agent failing
mid-task with an opaque "Unprocessable Content", and took a route_decisions
query to diagnose.
CORRECTION, 2026-09-01 — an earlier draft of this section proposed an
"inversion" check. It was wrong, and the error is worth keeping because it is
easy to make again. That draft said: warn when tier T's ceiling is below
tier T-1's. But ceiling(T) = max(effective_context_window) over models with
tier >= T, and tier is a capability FLOOR, so as T rises the eligible set
shrinks monotonically. Therefore ceiling(1) >= ceiling(2) >= ceiling(3)
holds for every possible catalog. A "higher tier has a lower ceiling" warning
is not a fault detector — it is a theorem, and it would fire on every healthy
catalog while never indicating anything. Verified live: 782,324 / 782,324 /
782,324 today, and 782,324 / 782,324 / 94,196 during the outage — descending in
both cases, exactly as the monotonicity predicts.
The real harm has nothing to do with ordering. It is that the ceiling at some tier fell below what that tier is actually asked to serve, and the only way to know that is to compare against traffic.
Add two warnings. Both compute the ceiling with the same filters routing
uses (tier floor, access_level, availability, admin overrides, latency_class
vs latency_tolerance), so they cannot drift from real routing behaviour:
- Ceiling below observed demand — the primary check. Tier T's ceiling is
below the max (or p95)
required_context_tokensamongroute_decisionsrows withkind = 'chat'andtask_tier = Tover the last N days. Self-calibrating: it measures the router against what it is actually asked for rather than an invented constant. Verified against live data — silent on all three tiers today, and it fires on the outage state (tier 3 ceiling 94,196 vs observed max 268,168). - Zero eligible models for any
(tier, latency_tolerance). Cheap, and it catches total collapse — but note it would not have fired on 2026-09-01, because tier 3 still had one eligible model. It is a backstop, not the detector.
Optionally, an escalation-hazard warning. escalation.enabled is true with
max_tier: 3, so a request can be bumped a tier at runtime. When
ceiling(T+1) is below the p95 demand at tier T, escalating a servable
request can make it unservable. That is the real hazard the word "inversion"
was reaching for, stated in a form that is actually checkable.
Computing the ceiling correctly matters too. Call select_candidates with
required_context_tokens=0 — i.e. apply every hard filter except context —
then take ceiling = max(effective_context_window) and
count = len(candidates). Filtering by the max window first (an easy mistake)
yields the right ceiling by luck but leaves count holding only the models tied
at the top, which silently breaks the zero-eligible check.
Both belong in the existing metrics.py warnings list (metrics.py:131),
next to §4.3's catalog-age warning — same rationale, same render path (admin
warnings bell, TUI).
And warn at the moment of causation, which matters more than the dashboard.
admin.py's POST /api/models/{model_id}/{provider}/availability
(admin.py:502) is where the damage is done. After the write, recompute the
ceilings and, if the change introduced an inversion or dropped a ceiling below
observed demand, (a) log at WARNING and (b) return the warning in the response
body so the portal can render it inline against the toggle that caused it. The
operator should learn that deprecating kimi-k2.7-code just cost tier 3 its
context headroom while looking at the switch, not nineteen hours later.
Do not block the change — deprecating a model is a legitimate operator action and there are good reasons to accept a lower ceiling temporarily. Warn loudly, refuse nothing.
Recovery, for the record: re-activating kimi-k2.7-code (192,500 eff)
restored the 94k-192k band and glm-5.3 (782,324 eff) restored everything
above it, giving a continuous ladder of 94,196 → 192,500 → 782,324. The
kimi-k3 family stays deprecated deliberately: glm-5.3 covers the identical
782k band at half the prompt price and a third of the completion price.
5. File-by-file
src/poller.pymain()— afterfetch_neuralwatt()and beforeupsert(), comparelen(rows)against the current row count; raise/exit(1) on zero, warn below half. Keep the existingRequestExceptionhandler as the single exit path.mark_staledocstring — say that it assumes the caller has already established the response was plausible, and point at the check.
src/metrics.py— the catalog-age warning (§4.3) and the two context-ceiling warnings (§4.4) in the existingwarningslist. Factor the ceiling computation into one helper —context_ceilings(conn, cfg) -> {(tier, latency_tolerance): (ceiling, eligible_count)}— reusingrouting.select_candidatesrather than re-implementing the filters, so it cannot drift from real routing.src/admin.py—POST /api/models/{model_id}/{provider}/availability(:502) recomputes ceilings after the write; logs at WARNING and returns the warning in the response body when the change introduces an inversion, drops a ceiling below observed demand, or empties a tier. Never blocks the change.admin/frontend/models.html— render that inline warning against the toggle that caused it. This is the highest-value half of §4.4; a dashboard warning the operator has to go looking for is what failed here.tests/test_poller_parsing.py(or a newtests/test_poller_freshness.py) — there is currently no coverage of either function. Pin:- a zero-row payload exits non-zero and marks nothing stale;
- a payload at 40% of current rows proceeds but warns;
- a full payload marks a row absent from it stale, and only that row;
- a previously-stale row returns to
activeafter one good poll (the auto-recovery in §1, currently untested); main()onRequestExceptionmarks nothing.
CLAUDE.md"Run as a service" andREADME.md"As a systemd service" — already corrected in this pass; no further edit needed.
Second entry point to keep in mind
admin.py:734 POST /api/refresh-catalog shells out to python -m poller, so
it shares main() and inherits the floor for free. Worth a manual click after
the change to confirm the admin trigger still reports success correctly when the
poller now exits 1 on a degraded fetch.
6. Summary
Is stale_after_days: 3 too tight? |
No. 36 successful polls of margin at the current 2-hourly cadence, and a false stale-out self-heals in one cycle. |
Does the plan in proficiency-exposure-bias-and-exploration.md cover it? |
No — different subsystem entirely (that one touches proficiency and candidate ranking). |
| What is actually wrong | fetch_neuralwatt has no floor on row count, so a 200-with-no-data arms a catalog-wide stale-out on a 3-day fuse. And an unpolled catalog fails silently open, which is the opposite of what the docs claimed. |
| Second path to zero candidates (§4.4) | Admin availability overrides. Deprecating 7 models on 2026-09-01 collapsed tier 3's context ceiling from 782,324 to 94,196 — an 8.3x drop, below the 268,168 max that tier actually receives — with no warning anywhere. Took the router down mid-task and surfaced ~19h later as an opaque agent failure. |
| Fix | A row-count floor in main(); a catalog-age warning and two context-ceiling warnings in /metrics; an inline warning at the admin toggle that causes a ceiling collapse; and the first tests this module has ever had. |
The through-line across all of §3, §4.3 and §4.4: every failure here is the router quietly losing the ability to serve requests while continuing to report itself healthy. None of them are wrong decisions — a frozen catalog, a deprecated expensive model, and a tight budget are all defensible. What is wrong is that each one is invisible until an agent fails with an error that names none of them. The fix in every case is the same shape: compute the thing that is actually true, and say so where the operator is already looking.