Files
6krrt/plans/catalog-staleness-and-poller-failure-modes.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

15 KiB

Catalog staleness: the fuse is real, the timer is not what arms it

Status: done -- rejection_warnings and the demand-ceiling detector in metrics.py

Date: 2026-09-01 Status: FINAL — ready to implement. Small; one commit. Prompted by: "is stale_after_days: 3 too tight?" Answer: no — and it is also not the knob that governs the failure the docs warn about. Both CLAUDE.md and README.md described this backwards and have been corrected in the same pass.


1. What was actually checked

poll cadence every 2 hours, llm-router-poller.timer, healthy
poller errors, last 7 days 0
last run upserted 19/19 models, 22 min before this was written
rows currently stale 1 — deepseek-ai/DeepSeek-V4-Flash, genuinely gone from the catalog since 2026-08-24. Correct.
tests covering fetch_neuralwatt or mark_stale none

Against a 2-hourly poll, stale_after_days: 3 is 36 consecutive successful polls of margin. Nothing has expired that should not have.

Recovery is automatic as well: upsert writes availability = excluded.availability (poller.py:251), so a single good poll flips every stale row back to active. The blast radius of a false stale-out is one 2-hour cycle, not manual intervention.


2. The documented failure mode is not reachable

Both docs said, in substance: an unpolled catalog marks every row stale and the router returns zero candidates.

It cannot. mark_stale (poller.py:283) runs only inside poller.main(), and main() returns early on requests.RequestException — before upsert and before mark_stale. So:

event what actually happens
timer stopped / disabled nothing runs, so nothing is marked. Catalog freezes at last-known-good; every row still reads active.
provider API down, DNS fails, network out early return at the fetch. Same: catalog frozen, nothing marked.

The real behaviour is silent and fail-open — the router keeps routing on prices, context windows and capability flags that may be weeks out of date, and nothing anywhere says so. That is arguably the better default (a soft failure beats a hard one), but it is the opposite of what the docs told you to watch for, so attention was pointed at the wrong risk.


3. The path that can empty the candidate set

fetch_neuralwatt (poller.py:161):

payload = resp.json()
rows = []
for m in payload.get("data", []):

No floor on row count. A 200 response carrying an empty or truncated data array — a partial provider outage, a schema change, an auth path degrading to an empty list — clears raise_for_status(), returns [], upserts nothing, and then mark_stale runs anyway because main() never saw an exception.

Three days of that and every row is stale; exclude_stale: true then leaves zero candidates for everything.

So the fuse exists — but stale_after_days only sets its length. It neither arms nor disarms it. Lengthening it buys a longer fuse on something that should not be armed at all.

Note this also breaks mark_stale's own docstring, which claims it flags "rows that weren't touched by this poll run". After a zero-row fetch that is vacuously every row, which is not what the sentence means.


4. The fix

4.1 Sanity floor in fetch_neuralwatt

Refuse an implausible catalog by raising, so main() takes its existing early-return path — which is already the correct behaviour for a failed fetch.

  • Zero rows → raise. Unconditional. An empty catalog is never a real answer from a provider that has 19 models.
  • Below ~50% of the current row count → log a warning and proceed. A legitimate mass deprecation should not be blocked outright, but it should not pass unremarked either.

Raise a RequestException subclass (or let main() catch a new CatalogTooSmall alongside it) so the existing handler covers it and the exit code stays 1 — systemd then reports the unit as failed, which is the loud signal this path currently lacks.

The row count to compare against is SELECT COUNT(*) FROM models WHERE provider = 'neuralwatt', so this needs the connection. Either pass it in, or move the check into main() between fetch_neuralwatt and upsert — the latter keeps fetch_neuralwatt free of DB access, which matches the module's existing shape. Prefer the check in main().

4.2 Then leave stale_after_days at 3 — or shorten it

Once the floor exists, availability = 'stale' means exactly one thing: this row was absent from a catalog we successfully fetched. That is a direct observation, not something needing a three-day timer to establish. A genuinely removed public model currently stays routable for three days after it vanishes; with the floor in place, 1 day (or two consecutive polls) is defensible and tighter.

This is a judgement call with no measurement behind it either way, so the recommendation is: land the floor, leave the value at 3, revisit only if a removed model actually causes a dispatch failure. Changing both at once would make the next incident ambiguous.

4.3 Surface catalog age, since the real failure is silence

metrics.py already builds a warnings list (metrics.py:131) that the admin dashboard's warnings bell and the TUI both render. Add one:

catalog last polled 4.2 days ago — prices, context windows and capability
flags may be out of date. Check: systemctl --user status llm-router-poller.timer

Threshold at max(1 day, 2 x stale_after_days / 3) or simply 1 day — the poll is 2-hourly, so anything past a day means the timer is not running. This is the change that actually addresses the fail-open behaviour in §2: it converts a silent freeze into something visible without changing the (correct) decision to keep serving.

4.4 Surface context-ceiling collapse — the other path that empties the candidate set

Added 2026-09-01, after this exact failure took the router down mid-task.

§3 covers one way to reach zero candidates. Admin availability overrides are another, and unlike the poller path it fires instantly, silently, and by a deliberate operator action that gives no indication of what it just cost.

What happened. Seven models were deprecated through /admin at 03:11-03:12 on 2026-09-01 — the expensive ones, a reasonable cost decision in isolation. The models table still read active (the poller refreshes it every 2h), but admin_model_overrides overlays deprecated and _admin_deprecated_models (dispatcher.py:1977) feeds that into routing's hard filters. Measured effect on the maximum servable required_context_tokens, interactive:

tier ceiling before ceiling during eligible models
1 782,324 782,324 6
2 782,324 782,324 5
3 782,324 94,196 1

Tier 3's ceiling fell 8.3x while tiers 1 and 2 were untouched. Observed tier-3 traffic has a max of 268,168 tokens and a p95 of 179,356, so every tier-3 request above 94,196 returned 422 No model satisfies the hard filters. Nothing warned, nothing logged above INFO, and the admin portal reported the change as a plain success. It surfaced ~19 hours later as an agent failing mid-task with an opaque "Unprocessable Content", and took a route_decisions query to diagnose.

CORRECTION, 2026-09-01 — an earlier draft of this section proposed an "inversion" check. It was wrong, and the error is worth keeping because it is easy to make again. That draft said: warn when tier T's ceiling is below tier T-1's. But ceiling(T) = max(effective_context_window) over models with tier >= T, and tier is a capability FLOOR, so as T rises the eligible set shrinks monotonically. Therefore ceiling(1) >= ceiling(2) >= ceiling(3) holds for every possible catalog. A "higher tier has a lower ceiling" warning is not a fault detector — it is a theorem, and it would fire on every healthy catalog while never indicating anything. Verified live: 782,324 / 782,324 / 782,324 today, and 782,324 / 782,324 / 94,196 during the outage — descending in both cases, exactly as the monotonicity predicts.

The real harm has nothing to do with ordering. It is that the ceiling at some tier fell below what that tier is actually asked to serve, and the only way to know that is to compare against traffic.

Add two warnings. Both compute the ceiling with the same filters routing uses (tier floor, access_level, availability, admin overrides, latency_class vs latency_tolerance), so they cannot drift from real routing behaviour:

  1. Ceiling below observed demand — the primary check. Tier T's ceiling is below the max (or p95) required_context_tokens among route_decisions rows with kind = 'chat' and task_tier = T over the last N days. Self-calibrating: it measures the router against what it is actually asked for rather than an invented constant. Verified against live data — silent on all three tiers today, and it fires on the outage state (tier 3 ceiling 94,196 vs observed max 268,168).
  2. Zero eligible models for any (tier, latency_tolerance). Cheap, and it catches total collapse — but note it would not have fired on 2026-09-01, because tier 3 still had one eligible model. It is a backstop, not the detector.

Optionally, an escalation-hazard warning. escalation.enabled is true with max_tier: 3, so a request can be bumped a tier at runtime. When ceiling(T+1) is below the p95 demand at tier T, escalating a servable request can make it unservable. That is the real hazard the word "inversion" was reaching for, stated in a form that is actually checkable.

Computing the ceiling correctly matters too. Call select_candidates with required_context_tokens=0 — i.e. apply every hard filter except context — then take ceiling = max(effective_context_window) and count = len(candidates). Filtering by the max window first (an easy mistake) yields the right ceiling by luck but leaves count holding only the models tied at the top, which silently breaks the zero-eligible check.

Both belong in the existing metrics.py warnings list (metrics.py:131), next to §4.3's catalog-age warning — same rationale, same render path (admin warnings bell, TUI).

And warn at the moment of causation, which matters more than the dashboard. admin.py's POST /api/models/{model_id}/{provider}/availability (admin.py:502) is where the damage is done. After the write, recompute the ceilings and, if the change introduced an inversion or dropped a ceiling below observed demand, (a) log at WARNING and (b) return the warning in the response body so the portal can render it inline against the toggle that caused it. The operator should learn that deprecating kimi-k2.7-code just cost tier 3 its context headroom while looking at the switch, not nineteen hours later.

Do not block the change — deprecating a model is a legitimate operator action and there are good reasons to accept a lower ceiling temporarily. Warn loudly, refuse nothing.

Recovery, for the record: re-activating kimi-k2.7-code (192,500 eff) restored the 94k-192k band and glm-5.3 (782,324 eff) restored everything above it, giving a continuous ladder of 94,196 → 192,500 → 782,324. The kimi-k3 family stays deprecated deliberately: glm-5.3 covers the identical 782k band at half the prompt price and a third of the completion price.


5. File-by-file

  • src/poller.py
    • main() — after fetch_neuralwatt() and before upsert(), compare len(rows) against the current row count; raise/exit(1) on zero, warn below half. Keep the existing RequestException handler as the single exit path.
    • mark_stale docstring — say that it assumes the caller has already established the response was plausible, and point at the check.
  • src/metrics.py — the catalog-age warning (§4.3) and the two context-ceiling warnings (§4.4) in the existing warnings list. Factor the ceiling computation into one helper — context_ceilings(conn, cfg) -> {(tier, latency_tolerance): (ceiling, eligible_count)} — reusing routing.select_candidates rather than re-implementing the filters, so it cannot drift from real routing.
  • src/admin.py — POST /api/models/{model_id}/{provider}/availability (:502) recomputes ceilings after the write; logs at WARNING and returns the warning in the response body when the change introduces an inversion, drops a ceiling below observed demand, or empties a tier. Never blocks the change.
  • admin/frontend/models.html — render that inline warning against the toggle that caused it. This is the highest-value half of §4.4; a dashboard warning the operator has to go looking for is what failed here.
  • tests/test_poller_parsing.py (or a new tests/test_poller_freshness.py) — there is currently no coverage of either function. Pin:
    • a zero-row payload exits non-zero and marks nothing stale;
    • a payload at 40% of current rows proceeds but warns;
    • a full payload marks a row absent from it stale, and only that row;
    • a previously-stale row returns to active after one good poll (the auto-recovery in §1, currently untested);
    • main() on RequestException marks nothing.
  • CLAUDE.md "Run as a service" and README.md "As a systemd service" — already corrected in this pass; no further edit needed.

Second entry point to keep in mind

admin.py:734 POST /api/refresh-catalog shells out to python -m poller, so it shares main() and inherits the floor for free. Worth a manual click after the change to confirm the admin trigger still reports success correctly when the poller now exits 1 on a degraded fetch.


6. Summary

Is stale_after_days: 3 too tight? No. 36 successful polls of margin at the current 2-hourly cadence, and a false stale-out self-heals in one cycle.
Does the plan in proficiency-exposure-bias-and-exploration.md cover it? No — different subsystem entirely (that one touches proficiency and candidate ranking).
What is actually wrong fetch_neuralwatt has no floor on row count, so a 200-with-no-data arms a catalog-wide stale-out on a 3-day fuse. And an unpolled catalog fails silently open, which is the opposite of what the docs claimed.
Second path to zero candidates (§4.4) Admin availability overrides. Deprecating 7 models on 2026-09-01 collapsed tier 3's context ceiling from 782,324 to 94,196 — an 8.3x drop, below the 268,168 max that tier actually receives — with no warning anywhere. Took the router down mid-task and surfaced ~19h later as an opaque agent failure.
Fix A row-count floor in main(); a catalog-age warning and two context-ceiling warnings in /metrics; an inline warning at the admin toggle that causes a ceiling collapse; and the first tests this module has ever had.

The through-line across all of §3, §4.3 and §4.4: every failure here is the router quietly losing the ability to serve requests while continuing to report itself healthy. None of them are wrong decisions — a frozen catalog, a deprecated expensive model, and a tight budget are all defensible. What is wrong is that each one is invisible until an agent fails with an error that names none of them. The fix in every case is the same shape: compute the thing that is actually true, and say so where the operator is already looking.