Files
6krrt/plans/catalog-staleness-and-poller-failure-modes.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

295 lines
15 KiB
Markdown

# Catalog staleness: the fuse is real, the timer is not what arms it
Status: done -- rejection_warnings and the demand-ceiling detector in metrics.py
**Date:** 2026-09-01
**Status:** **FINAL — ready to implement.** Small; one commit.
**Prompted by:** "is `stale_after_days: 3` too tight?"
**Answer:** no — and it is also not the knob that governs the failure the docs
warn about. Both `CLAUDE.md` and `README.md` described this backwards and have
been corrected in the same pass.
---
## 1. What was actually checked
| | |
|---|---|
| poll cadence | **every 2 hours**, `llm-router-poller.timer`, healthy |
| poller errors, last 7 days | **0** |
| last run | upserted 19/19 models, 22 min before this was written |
| rows currently stale | **1** — `deepseek-ai/DeepSeek-V4-Flash`, genuinely gone from the catalog since 2026-08-24. Correct. |
| tests covering `fetch_neuralwatt` or `mark_stale` | **none** |
Against a 2-hourly poll, `stale_after_days: 3` is **36 consecutive successful
polls** of margin. Nothing has expired that should not have.
Recovery is automatic as well: `upsert` writes
`availability = excluded.availability` (`poller.py:251`), so a single good poll
flips every stale row back to `active`. The blast radius of a false stale-out is
one 2-hour cycle, not manual intervention.
---
## 2. The documented failure mode is not reachable
Both docs said, in substance: *an unpolled catalog marks every row stale and the
router returns zero candidates.*
It cannot. `mark_stale` (`poller.py:283`) runs only **inside** `poller.main()`,
and `main()` returns early on `requests.RequestException` — **before** `upsert`
and **before** `mark_stale`. So:
| event | what actually happens |
|---|---|
| timer stopped / disabled | nothing runs, so nothing is marked. Catalog freezes at last-known-good; every row still reads `active`. |
| provider API down, DNS fails, network out | early return at the fetch. Same: catalog frozen, nothing marked. |
The real behaviour is **silent and fail-open** — the router keeps routing on
prices, context windows and capability flags that may be weeks out of date, and
nothing anywhere says so. That is arguably the better default (a soft failure
beats a hard one), but it is the *opposite* of what the docs told you to watch
for, so attention was pointed at the wrong risk.
---
## 3. The path that can empty the candidate set
`fetch_neuralwatt` (`poller.py:161`):
```python
payload = resp.json()
rows = []
for m in payload.get("data", []):
```
No floor on row count. A **200 response carrying an empty or truncated `data`
array** — a partial provider outage, a schema change, an auth path degrading to
an empty list — clears `raise_for_status()`, returns `[]`, upserts nothing, and
then `mark_stale` runs anyway because `main()` never saw an exception.
Three days of that and every row is stale; `exclude_stale: true` then leaves
zero candidates for everything.
So the fuse exists — but `stale_after_days` only sets its **length**. It neither
arms nor disarms it. **Lengthening it buys a longer fuse on something that
should not be armed at all.**
Note this also breaks `mark_stale`'s own docstring, which claims it flags "rows
that weren't touched by this poll run". After a zero-row fetch that is
vacuously every row, which is not what the sentence means.
---
## 4. The fix
### 4.1 Sanity floor in `fetch_neuralwatt`
Refuse an implausible catalog by **raising**, so `main()` takes its existing
early-return path — which is already the correct behaviour for a failed fetch.
- **Zero rows → raise.** Unconditional. An empty catalog is never a real answer
from a provider that has 19 models.
- **Below ~50% of the current row count → log a warning and proceed.** A
legitimate mass deprecation should not be blocked outright, but it should not
pass unremarked either.
Raise a `RequestException` subclass (or let `main()` catch a new
`CatalogTooSmall` alongside it) so the existing handler covers it and the exit
code stays 1 — systemd then reports the unit as failed, which is the loud signal
this path currently lacks.
The row count to compare against is `SELECT COUNT(*) FROM models WHERE provider
= 'neuralwatt'`, so this needs the connection. Either pass it in, or move the
check into `main()` between `fetch_neuralwatt` and `upsert` — the latter keeps
`fetch_neuralwatt` free of DB access, which matches the module's existing shape.
**Prefer the check in `main()`.**
### 4.2 Then leave `stale_after_days` at 3 — or shorten it
Once the floor exists, `availability = 'stale'` means exactly one thing: *this
row was absent from a catalog we successfully fetched.* That is a direct
observation, not something needing a three-day timer to establish. A genuinely
removed public model currently stays routable for three days after it vanishes;
with the floor in place, 1 day (or two consecutive polls) is defensible and
tighter.
This is a judgement call with no measurement behind it either way, so the
recommendation is: **land the floor, leave the value at 3, revisit only if a
removed model actually causes a dispatch failure.** Changing both at once would
make the next incident ambiguous.
### 4.3 Surface catalog age, since the real failure is silence
`metrics.py` already builds a `warnings` list (`metrics.py:131`) that the admin
dashboard's warnings bell and the TUI both render. Add one:
```
catalog last polled 4.2 days ago — prices, context windows and capability
flags may be out of date. Check: systemctl --user status llm-router-poller.timer
```
Threshold at `max(1 day, 2 x stale_after_days / 3)` or simply 1 day — the poll
is 2-hourly, so anything past a day means the timer is not running. This is the
change that actually addresses the fail-open behaviour in §2: it converts a
silent freeze into something visible without changing the (correct) decision to
keep serving.
### 4.4 Surface context-ceiling collapse — the other path that empties the candidate set
**Added 2026-09-01, after this exact failure took the router down mid-task.**
§3 covers one way to reach zero candidates. Admin availability overrides are
another, and unlike the poller path it fires instantly, silently, and by a
deliberate operator action that gives no indication of what it just cost.
**What happened.** Seven models were deprecated through `/admin` at 03:11-03:12
on 2026-09-01 — the expensive ones, a reasonable cost decision in isolation. The
`models` table still read `active` (the poller refreshes it every 2h), but
`admin_model_overrides` overlays `deprecated` and `_admin_deprecated_models`
(`dispatcher.py:1977`) feeds that into routing's hard filters. Measured effect on
the maximum servable `required_context_tokens`, interactive:
| tier | ceiling before | ceiling during | eligible models |
|---|---|---|---|
| 1 | 782,324 | 782,324 | 6 |
| 2 | 782,324 | 782,324 | 5 |
| 3 | 782,324 | **94,196** | **1** |
Tier 3's ceiling fell **8.3x** while tiers 1 and 2 were untouched. Observed
tier-3 traffic has a max of 268,168 tokens and a p95 of 179,356, so every
tier-3 request above 94,196 returned `422 No model satisfies the hard filters`.
Nothing warned, nothing logged above INFO, and the admin portal reported the
change as a plain success. It surfaced ~19 hours later as an agent failing
mid-task with an opaque "Unprocessable Content", and took a `route_decisions`
query to diagnose.
**CORRECTION, 2026-09-01 — an earlier draft of this section proposed an
"inversion" check. It was wrong, and the error is worth keeping because it is
easy to make again.** That draft said: warn when tier *T*'s ceiling is below
tier *T-1*'s. But `ceiling(T) = max(effective_context_window)` over models with
`tier >= T`, and tier is a capability FLOOR, so as *T* rises the eligible set
**shrinks monotonically**. Therefore `ceiling(1) >= ceiling(2) >= ceiling(3)`
holds for *every possible catalog*. A "higher tier has a lower ceiling" warning
is not a fault detector — it is a theorem, and it would fire on every healthy
catalog while never indicating anything. Verified live: 782,324 / 782,324 /
782,324 today, and 782,324 / 782,324 / 94,196 during the outage — descending in
both cases, exactly as the monotonicity predicts.
The real harm has nothing to do with ordering. It is that **the ceiling at some
tier fell below what that tier is actually asked to serve**, and the only way to
know that is to compare against traffic.
**Add two warnings.** Both compute the ceiling with the same filters routing
uses (tier floor, `access_level`, `availability`, admin overrides, `latency_class`
vs `latency_tolerance`), so they cannot drift from real routing behaviour:
1. **Ceiling below observed demand** — the primary check. Tier *T*'s ceiling is
below the max (or p95) `required_context_tokens` among `route_decisions` rows
with `kind = 'chat'` and `task_tier = T` over the last N days.
Self-calibrating: it measures the router against what it is actually asked
for rather than an invented constant. Verified against live data — silent on
all three tiers today, and it fires on the outage state (tier 3 ceiling
94,196 vs observed max 268,168).
2. **Zero eligible models** for any `(tier, latency_tolerance)`. Cheap, and it
catches total collapse — but note it would *not* have fired on 2026-09-01,
because tier 3 still had one eligible model. It is a backstop, not the
detector.
**Optionally, an escalation-hazard warning.** `escalation.enabled` is true with
`max_tier: 3`, so a request can be bumped a tier at runtime. When
`ceiling(T+1)` is below the p95 demand at tier *T*, escalating a servable
request can make it unservable. That is the real hazard the word "inversion"
was reaching for, stated in a form that is actually checkable.
**Computing the ceiling correctly matters too.** Call `select_candidates` with
`required_context_tokens=0` — i.e. apply every hard filter *except* context —
then take `ceiling = max(effective_context_window)` and
`count = len(candidates)`. Filtering by the max window first (an easy mistake)
yields the right ceiling by luck but leaves `count` holding only the models tied
at the top, which silently breaks the zero-eligible check.
Both belong in the existing `metrics.py` `warnings` list (`metrics.py:131`),
next to §4.3's catalog-age warning — same rationale, same render path (admin
warnings bell, TUI).
**And warn at the moment of causation, which matters more than the dashboard.**
`admin.py`'s `POST /api/models/{model_id}/{provider}/availability`
(`admin.py:502`) is where the damage is done. After the write, recompute the
ceilings and, if the change introduced an inversion or dropped a ceiling below
observed demand, (a) log at WARNING and (b) return the warning in the response
body so the portal can render it inline against the toggle that caused it. The
operator should learn that deprecating `kimi-k2.7-code` just cost tier 3 its
context headroom **while looking at the switch**, not nineteen hours later.
Do **not** block the change — deprecating a model is a legitimate operator
action and there are good reasons to accept a lower ceiling temporarily. Warn
loudly, refuse nothing.
**Recovery, for the record:** re-activating `kimi-k2.7-code` (192,500 eff)
restored the 94k-192k band and `glm-5.3` (782,324 eff) restored everything
above it, giving a continuous ladder of 94,196 → 192,500 → 782,324. The
`kimi-k3` family stays deprecated deliberately: `glm-5.3` covers the identical
782k band at half the prompt price and a third of the completion price.
---
## 5. File-by-file
- `src/poller.py`
- `main()` — after `fetch_neuralwatt()` and before `upsert()`, compare
`len(rows)` against the current row count; raise/exit(1) on zero, warn below
half. Keep the existing `RequestException` handler as the single exit path.
- `mark_stale` docstring — say that it assumes the caller has already
established the response was plausible, and point at the check.
- `src/metrics.py` — the catalog-age warning (§4.3) and the two
context-ceiling warnings (§4.4) in the existing `warnings` list. Factor the
ceiling computation into one helper — `context_ceilings(conn, cfg) ->
{(tier, latency_tolerance): (ceiling, eligible_count)}` — reusing
`routing.select_candidates` rather than re-implementing the filters, so it
cannot drift from real routing.
- `src/admin.py` — `POST /api/models/{model_id}/{provider}/availability`
(`:502`) recomputes ceilings after the write; logs at WARNING and returns the
warning in the response body when the change introduces an inversion, drops a
ceiling below observed demand, or empties a tier. Never blocks the change.
- `admin/frontend/models.html` — render that inline warning against the toggle
that caused it. This is the highest-value half of §4.4; a dashboard warning
the operator has to go looking for is what failed here.
- `tests/test_poller_parsing.py` (or a new `tests/test_poller_freshness.py`) —
there is currently **no** coverage of either function. Pin:
- a zero-row payload exits non-zero and marks **nothing** stale;
- a payload at 40% of current rows proceeds but warns;
- a full payload marks a row absent from it stale, and only that row;
- a previously-stale row returns to `active` after one good poll
(the auto-recovery in §1, currently untested);
- `main()` on `RequestException` marks nothing.
- `CLAUDE.md` "Run as a service" and `README.md` "As a systemd service" —
**already corrected in this pass**; no further edit needed.
### Second entry point to keep in mind
`admin.py:734` `POST /api/refresh-catalog` shells out to `python -m poller`, so
it shares `main()` and inherits the floor for free. Worth a manual click after
the change to confirm the admin trigger still reports success correctly when the
poller now exits 1 on a degraded fetch.
---
## 6. Summary
| | |
|---|---|
| Is `stale_after_days: 3` too tight? | **No.** 36 successful polls of margin at the current 2-hourly cadence, and a false stale-out self-heals in one cycle. |
| Does the plan in `proficiency-exposure-bias-and-exploration.md` cover it? | **No** — different subsystem entirely (that one touches `proficiency` and candidate ranking). |
| What is actually wrong | `fetch_neuralwatt` has no floor on row count, so a 200-with-no-data arms a catalog-wide stale-out on a 3-day fuse. And an unpolled catalog fails **silently open**, which is the opposite of what the docs claimed. |
| Second path to zero candidates (§4.4) | Admin availability overrides. Deprecating 7 models on 2026-09-01 collapsed tier 3's context ceiling from 782,324 to **94,196** — an 8.3x drop, below the 268,168 max that tier actually receives — with no warning anywhere. Took the router down mid-task and surfaced ~19h later as an opaque agent failure. |
| Fix | A row-count floor in `main()`; a catalog-age warning and two context-ceiling warnings in `/metrics`; an inline warning at the admin toggle that causes a ceiling collapse; and the first tests this module has ever had. |
**The through-line across all of §3, §4.3 and §4.4:** every failure here is the
router quietly losing the ability to serve requests while continuing to report
itself healthy. None of them are wrong decisions — a frozen catalog, a
deprecated expensive model, and a tight budget are all defensible. What is wrong
is that each one is invisible until an agent fails with an error that names none
of them. The fix in every case is the same shape: compute the thing that is
actually true, and say so where the operator is already looking.