plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
295 lines
15 KiB
Markdown
295 lines
15 KiB
Markdown
# Catalog staleness: the fuse is real, the timer is not what arms it
|
|
|
|
Status: done -- rejection_warnings and the demand-ceiling detector in metrics.py
|
|
|
|
**Date:** 2026-09-01
|
|
**Status:** **FINAL — ready to implement.** Small; one commit.
|
|
**Prompted by:** "is `stale_after_days: 3` too tight?"
|
|
**Answer:** no — and it is also not the knob that governs the failure the docs
|
|
warn about. Both `CLAUDE.md` and `README.md` described this backwards and have
|
|
been corrected in the same pass.
|
|
|
|
---
|
|
|
|
## 1. What was actually checked
|
|
|
|
| | |
|
|
|---|---|
|
|
| poll cadence | **every 2 hours**, `llm-router-poller.timer`, healthy |
|
|
| poller errors, last 7 days | **0** |
|
|
| last run | upserted 19/19 models, 22 min before this was written |
|
|
| rows currently stale | **1** — `deepseek-ai/DeepSeek-V4-Flash`, genuinely gone from the catalog since 2026-08-24. Correct. |
|
|
| tests covering `fetch_neuralwatt` or `mark_stale` | **none** |
|
|
|
|
Against a 2-hourly poll, `stale_after_days: 3` is **36 consecutive successful
|
|
polls** of margin. Nothing has expired that should not have.
|
|
|
|
Recovery is automatic as well: `upsert` writes
|
|
`availability = excluded.availability` (`poller.py:251`), so a single good poll
|
|
flips every stale row back to `active`. The blast radius of a false stale-out is
|
|
one 2-hour cycle, not manual intervention.
|
|
|
|
---
|
|
|
|
## 2. The documented failure mode is not reachable
|
|
|
|
Both docs said, in substance: *an unpolled catalog marks every row stale and the
|
|
router returns zero candidates.*
|
|
|
|
It cannot. `mark_stale` (`poller.py:283`) runs only **inside** `poller.main()`,
|
|
and `main()` returns early on `requests.RequestException` — **before** `upsert`
|
|
and **before** `mark_stale`. So:
|
|
|
|
| event | what actually happens |
|
|
|---|---|
|
|
| timer stopped / disabled | nothing runs, so nothing is marked. Catalog freezes at last-known-good; every row still reads `active`. |
|
|
| provider API down, DNS fails, network out | early return at the fetch. Same: catalog frozen, nothing marked. |
|
|
|
|
The real behaviour is **silent and fail-open** — the router keeps routing on
|
|
prices, context windows and capability flags that may be weeks out of date, and
|
|
nothing anywhere says so. That is arguably the better default (a soft failure
|
|
beats a hard one), but it is the *opposite* of what the docs told you to watch
|
|
for, so attention was pointed at the wrong risk.
|
|
|
|
---
|
|
|
|
## 3. The path that can empty the candidate set
|
|
|
|
`fetch_neuralwatt` (`poller.py:161`):
|
|
|
|
```python
|
|
payload = resp.json()
|
|
rows = []
|
|
for m in payload.get("data", []):
|
|
```
|
|
|
|
No floor on row count. A **200 response carrying an empty or truncated `data`
|
|
array** — a partial provider outage, a schema change, an auth path degrading to
|
|
an empty list — clears `raise_for_status()`, returns `[]`, upserts nothing, and
|
|
then `mark_stale` runs anyway because `main()` never saw an exception.
|
|
|
|
Three days of that and every row is stale; `exclude_stale: true` then leaves
|
|
zero candidates for everything.
|
|
|
|
So the fuse exists — but `stale_after_days` only sets its **length**. It neither
|
|
arms nor disarms it. **Lengthening it buys a longer fuse on something that
|
|
should not be armed at all.**
|
|
|
|
Note this also breaks `mark_stale`'s own docstring, which claims it flags "rows
|
|
that weren't touched by this poll run". After a zero-row fetch that is
|
|
vacuously every row, which is not what the sentence means.
|
|
|
|
---
|
|
|
|
## 4. The fix
|
|
|
|
### 4.1 Sanity floor in `fetch_neuralwatt`
|
|
|
|
Refuse an implausible catalog by **raising**, so `main()` takes its existing
|
|
early-return path — which is already the correct behaviour for a failed fetch.
|
|
|
|
- **Zero rows → raise.** Unconditional. An empty catalog is never a real answer
|
|
from a provider that has 19 models.
|
|
- **Below ~50% of the current row count → log a warning and proceed.** A
|
|
legitimate mass deprecation should not be blocked outright, but it should not
|
|
pass unremarked either.
|
|
|
|
Raise a `RequestException` subclass (or let `main()` catch a new
|
|
`CatalogTooSmall` alongside it) so the existing handler covers it and the exit
|
|
code stays 1 — systemd then reports the unit as failed, which is the loud signal
|
|
this path currently lacks.
|
|
|
|
The row count to compare against is `SELECT COUNT(*) FROM models WHERE provider
|
|
= 'neuralwatt'`, so this needs the connection. Either pass it in, or move the
|
|
check into `main()` between `fetch_neuralwatt` and `upsert` — the latter keeps
|
|
`fetch_neuralwatt` free of DB access, which matches the module's existing shape.
|
|
**Prefer the check in `main()`.**
|
|
|
|
### 4.2 Then leave `stale_after_days` at 3 — or shorten it
|
|
|
|
Once the floor exists, `availability = 'stale'` means exactly one thing: *this
|
|
row was absent from a catalog we successfully fetched.* That is a direct
|
|
observation, not something needing a three-day timer to establish. A genuinely
|
|
removed public model currently stays routable for three days after it vanishes;
|
|
with the floor in place, 1 day (or two consecutive polls) is defensible and
|
|
tighter.
|
|
|
|
This is a judgement call with no measurement behind it either way, so the
|
|
recommendation is: **land the floor, leave the value at 3, revisit only if a
|
|
removed model actually causes a dispatch failure.** Changing both at once would
|
|
make the next incident ambiguous.
|
|
|
|
### 4.3 Surface catalog age, since the real failure is silence
|
|
|
|
`metrics.py` already builds a `warnings` list (`metrics.py:131`) that the admin
|
|
dashboard's warnings bell and the TUI both render. Add one:
|
|
|
|
```
|
|
catalog last polled 4.2 days ago — prices, context windows and capability
|
|
flags may be out of date. Check: systemctl --user status llm-router-poller.timer
|
|
```
|
|
|
|
Threshold at `max(1 day, 2 x stale_after_days / 3)` or simply 1 day — the poll
|
|
is 2-hourly, so anything past a day means the timer is not running. This is the
|
|
change that actually addresses the fail-open behaviour in §2: it converts a
|
|
silent freeze into something visible without changing the (correct) decision to
|
|
keep serving.
|
|
|
|
### 4.4 Surface context-ceiling collapse — the other path that empties the candidate set
|
|
|
|
**Added 2026-09-01, after this exact failure took the router down mid-task.**
|
|
|
|
§3 covers one way to reach zero candidates. Admin availability overrides are
|
|
another, and unlike the poller path it fires instantly, silently, and by a
|
|
deliberate operator action that gives no indication of what it just cost.
|
|
|
|
**What happened.** Seven models were deprecated through `/admin` at 03:11-03:12
|
|
on 2026-09-01 — the expensive ones, a reasonable cost decision in isolation. The
|
|
`models` table still read `active` (the poller refreshes it every 2h), but
|
|
`admin_model_overrides` overlays `deprecated` and `_admin_deprecated_models`
|
|
(`dispatcher.py:1977`) feeds that into routing's hard filters. Measured effect on
|
|
the maximum servable `required_context_tokens`, interactive:
|
|
|
|
| tier | ceiling before | ceiling during | eligible models |
|
|
|---|---|---|---|
|
|
| 1 | 782,324 | 782,324 | 6 |
|
|
| 2 | 782,324 | 782,324 | 5 |
|
|
| 3 | 782,324 | **94,196** | **1** |
|
|
|
|
Tier 3's ceiling fell **8.3x** while tiers 1 and 2 were untouched. Observed
|
|
tier-3 traffic has a max of 268,168 tokens and a p95 of 179,356, so every
|
|
tier-3 request above 94,196 returned `422 No model satisfies the hard filters`.
|
|
Nothing warned, nothing logged above INFO, and the admin portal reported the
|
|
change as a plain success. It surfaced ~19 hours later as an agent failing
|
|
mid-task with an opaque "Unprocessable Content", and took a `route_decisions`
|
|
query to diagnose.
|
|
|
|
**CORRECTION, 2026-09-01 — an earlier draft of this section proposed an
|
|
"inversion" check. It was wrong, and the error is worth keeping because it is
|
|
easy to make again.** That draft said: warn when tier *T*'s ceiling is below
|
|
tier *T-1*'s. But `ceiling(T) = max(effective_context_window)` over models with
|
|
`tier >= T`, and tier is a capability FLOOR, so as *T* rises the eligible set
|
|
**shrinks monotonically**. Therefore `ceiling(1) >= ceiling(2) >= ceiling(3)`
|
|
holds for *every possible catalog*. A "higher tier has a lower ceiling" warning
|
|
is not a fault detector — it is a theorem, and it would fire on every healthy
|
|
catalog while never indicating anything. Verified live: 782,324 / 782,324 /
|
|
782,324 today, and 782,324 / 782,324 / 94,196 during the outage — descending in
|
|
both cases, exactly as the monotonicity predicts.
|
|
|
|
The real harm has nothing to do with ordering. It is that **the ceiling at some
|
|
tier fell below what that tier is actually asked to serve**, and the only way to
|
|
know that is to compare against traffic.
|
|
|
|
**Add two warnings.** Both compute the ceiling with the same filters routing
|
|
uses (tier floor, `access_level`, `availability`, admin overrides, `latency_class`
|
|
vs `latency_tolerance`), so they cannot drift from real routing behaviour:
|
|
|
|
1. **Ceiling below observed demand** — the primary check. Tier *T*'s ceiling is
|
|
below the max (or p95) `required_context_tokens` among `route_decisions` rows
|
|
with `kind = 'chat'` and `task_tier = T` over the last N days.
|
|
Self-calibrating: it measures the router against what it is actually asked
|
|
for rather than an invented constant. Verified against live data — silent on
|
|
all three tiers today, and it fires on the outage state (tier 3 ceiling
|
|
94,196 vs observed max 268,168).
|
|
2. **Zero eligible models** for any `(tier, latency_tolerance)`. Cheap, and it
|
|
catches total collapse — but note it would *not* have fired on 2026-09-01,
|
|
because tier 3 still had one eligible model. It is a backstop, not the
|
|
detector.
|
|
|
|
**Optionally, an escalation-hazard warning.** `escalation.enabled` is true with
|
|
`max_tier: 3`, so a request can be bumped a tier at runtime. When
|
|
`ceiling(T+1)` is below the p95 demand at tier *T*, escalating a servable
|
|
request can make it unservable. That is the real hazard the word "inversion"
|
|
was reaching for, stated in a form that is actually checkable.
|
|
|
|
**Computing the ceiling correctly matters too.** Call `select_candidates` with
|
|
`required_context_tokens=0` — i.e. apply every hard filter *except* context —
|
|
then take `ceiling = max(effective_context_window)` and
|
|
`count = len(candidates)`. Filtering by the max window first (an easy mistake)
|
|
yields the right ceiling by luck but leaves `count` holding only the models tied
|
|
at the top, which silently breaks the zero-eligible check.
|
|
|
|
Both belong in the existing `metrics.py` `warnings` list (`metrics.py:131`),
|
|
next to §4.3's catalog-age warning — same rationale, same render path (admin
|
|
warnings bell, TUI).
|
|
|
|
**And warn at the moment of causation, which matters more than the dashboard.**
|
|
`admin.py`'s `POST /api/models/{model_id}/{provider}/availability`
|
|
(`admin.py:502`) is where the damage is done. After the write, recompute the
|
|
ceilings and, if the change introduced an inversion or dropped a ceiling below
|
|
observed demand, (a) log at WARNING and (b) return the warning in the response
|
|
body so the portal can render it inline against the toggle that caused it. The
|
|
operator should learn that deprecating `kimi-k2.7-code` just cost tier 3 its
|
|
context headroom **while looking at the switch**, not nineteen hours later.
|
|
|
|
Do **not** block the change — deprecating a model is a legitimate operator
|
|
action and there are good reasons to accept a lower ceiling temporarily. Warn
|
|
loudly, refuse nothing.
|
|
|
|
**Recovery, for the record:** re-activating `kimi-k2.7-code` (192,500 eff)
|
|
restored the 94k-192k band and `glm-5.3` (782,324 eff) restored everything
|
|
above it, giving a continuous ladder of 94,196 → 192,500 → 782,324. The
|
|
`kimi-k3` family stays deprecated deliberately: `glm-5.3` covers the identical
|
|
782k band at half the prompt price and a third of the completion price.
|
|
|
|
---
|
|
|
|
## 5. File-by-file
|
|
|
|
- `src/poller.py`
|
|
- `main()` — after `fetch_neuralwatt()` and before `upsert()`, compare
|
|
`len(rows)` against the current row count; raise/exit(1) on zero, warn below
|
|
half. Keep the existing `RequestException` handler as the single exit path.
|
|
- `mark_stale` docstring — say that it assumes the caller has already
|
|
established the response was plausible, and point at the check.
|
|
- `src/metrics.py` — the catalog-age warning (§4.3) and the two
|
|
context-ceiling warnings (§4.4) in the existing `warnings` list. Factor the
|
|
ceiling computation into one helper — `context_ceilings(conn, cfg) ->
|
|
{(tier, latency_tolerance): (ceiling, eligible_count)}` — reusing
|
|
`routing.select_candidates` rather than re-implementing the filters, so it
|
|
cannot drift from real routing.
|
|
- `src/admin.py` — `POST /api/models/{model_id}/{provider}/availability`
|
|
(`:502`) recomputes ceilings after the write; logs at WARNING and returns the
|
|
warning in the response body when the change introduces an inversion, drops a
|
|
ceiling below observed demand, or empties a tier. Never blocks the change.
|
|
- `admin/frontend/models.html` — render that inline warning against the toggle
|
|
that caused it. This is the highest-value half of §4.4; a dashboard warning
|
|
the operator has to go looking for is what failed here.
|
|
- `tests/test_poller_parsing.py` (or a new `tests/test_poller_freshness.py`) —
|
|
there is currently **no** coverage of either function. Pin:
|
|
- a zero-row payload exits non-zero and marks **nothing** stale;
|
|
- a payload at 40% of current rows proceeds but warns;
|
|
- a full payload marks a row absent from it stale, and only that row;
|
|
- a previously-stale row returns to `active` after one good poll
|
|
(the auto-recovery in §1, currently untested);
|
|
- `main()` on `RequestException` marks nothing.
|
|
- `CLAUDE.md` "Run as a service" and `README.md` "As a systemd service" —
|
|
**already corrected in this pass**; no further edit needed.
|
|
|
|
### Second entry point to keep in mind
|
|
|
|
`admin.py:734` `POST /api/refresh-catalog` shells out to `python -m poller`, so
|
|
it shares `main()` and inherits the floor for free. Worth a manual click after
|
|
the change to confirm the admin trigger still reports success correctly when the
|
|
poller now exits 1 on a degraded fetch.
|
|
|
|
---
|
|
|
|
## 6. Summary
|
|
|
|
| | |
|
|
|---|---|
|
|
| Is `stale_after_days: 3` too tight? | **No.** 36 successful polls of margin at the current 2-hourly cadence, and a false stale-out self-heals in one cycle. |
|
|
| Does the plan in `proficiency-exposure-bias-and-exploration.md` cover it? | **No** — different subsystem entirely (that one touches `proficiency` and candidate ranking). |
|
|
| What is actually wrong | `fetch_neuralwatt` has no floor on row count, so a 200-with-no-data arms a catalog-wide stale-out on a 3-day fuse. And an unpolled catalog fails **silently open**, which is the opposite of what the docs claimed. |
|
|
| Second path to zero candidates (§4.4) | Admin availability overrides. Deprecating 7 models on 2026-09-01 collapsed tier 3's context ceiling from 782,324 to **94,196** — an 8.3x drop, below the 268,168 max that tier actually receives — with no warning anywhere. Took the router down mid-task and surfaced ~19h later as an opaque agent failure. |
|
|
| Fix | A row-count floor in `main()`; a catalog-age warning and two context-ceiling warnings in `/metrics`; an inline warning at the admin toggle that causes a ceiling collapse; and the first tests this module has ever had. |
|
|
|
|
**The through-line across all of §3, §4.3 and §4.4:** every failure here is the
|
|
router quietly losing the ability to serve requests while continuing to report
|
|
itself healthy. None of them are wrong decisions — a frozen catalog, a
|
|
deprecated expensive model, and a tight budget are all defensible. What is wrong
|
|
is that each one is invisible until an agent fails with an error that names none
|
|
of them. The fix in every case is the same shape: compute the thing that is
|
|
actually true, and say so where the operator is already looking.
|