# Capability-aware ceiling warnings, and a reactive rejection detector Status: done -- capability_ceilings in metrics.py **Status: FINAL — decision-complete.** Written 2026-09-04 from a live incident. ## The incident On 2026-09-04 two image requests failed with: ``` 422 tier >= 1; context >= 242486 tokens; interactive; vision-capable model (request carries image(s)) ``` Cause: admin availability overrides had `kimi-k3`, `kimi-k3-fast` and `kimi-k3-flex` deprecated. Those are the ONLY vision-capable rows with enough context (782,324). Every remaining active vision model tops out at 192,500, below the request's 242,486. The intersection of the hard filters was empty. This is [incident #3](../docs/incidents.md) recurring through a dimension the existing detector does not model. Resolved by re-activating `kimi-k3`; verified by recomputing the candidate set through `routing.select_candidates`, which now returns exactly one row. ## What already exists — do not rebuild it Establish this before writing code, because a first pass at this analysis got it wrong by checking for a top-level `warnings` key: - **`/metrics` surfaces warnings under `coverage.warnings`**, not at top level. `/admin/api/snapshot` carries the same block. The plumbing is done. - **`metrics.context_ceilings` already applies admin deprecations** via `exclude_models`, which was the fix after incident #3, and delegates to `routing.select_candidates` so the ceiling matches live routing. - **`metrics.demand_ceiling_warnings` already compares ceiling against observed demand**, which is the *correct* shape. Do not replace it with an absolute tier-ordering check: `ceiling(1) >= ceiling(2) >= ceiling(3)` is a theorem (tier is a capability floor, so the eligible set shrinks monotonically), so such a warning fires always and means nothing. `docs/incidents.md` #3 records this trap; do not re-enter it. - **`route_decisions` already records every rejection** — `selected_model IS NULL` with the full `rejected_reason` string. ## Gap 1 — ceilings are blind to capability-gated subsets `context_ceilings` buckets by `(tier, latency_tolerance)` only. During the incident the tier-1 interactive ceiling across all models was still **782,324** (deepseek and others are large and active), so every existing check was silent while the *vision-capable* ceiling had collapsed to 192,500. Measured during the incident: | subset | ceiling | |---|---| | all active models | 782,324 | | **vision-capable only** | **192,500** | The hard filters that can independently empty the candidate set, from `routing.rejection_reason`, are `require_vision` and `require_json_mode` — both **fail closed on NULL** (an unconfirmed capability is treated as absent), which is what makes them able to shrink the set sharply. **Add capability sub-ceilings.** Extend the bucket key, or compute a small number of additional named buckets: `vision` and `json_mode`. Do NOT add a bucket per capability combination — that is a combinatorial explosion for dimensions that do not interact in practice. Two extra series, compared against the demand actually observed for requests carrying images / requesting JSON (`route_decisions.images`, `.json_mode` are already recorded), is enough. Reuse `demand_ceiling_warnings`' existing demand-relative shape; only the bucketing changes. ## Gap 2 — nothing watches actual rejections This is the more valuable half, and it is simpler. Every hard-filter rejection is already persisted with its reason. Nothing reads them. A detector that says *"N requests were rejected in the last hour"* would have caught this incident within minutes, **without modelling any capability dimension at all** — and would equally catch the next failure through a dimension nobody predicted. Predictive checks (Gap 1) only catch dimensions you thought of. A reactive check catches everything, at the cost of firing after the first failure rather than before. Ship both; the reactive one is the safety net. Add a warning in `metrics.scoring_coverage`'s existing `warnings` list: - Count `route_decisions` rows with `selected_model IS NULL AND rejected_reason IS NOT NULL` in a recent window (1 hour and 24 hours are both useful; pick one, state which). - Group by a **normalized** reason, not the literal string. The raw strings embed the request's token count (`context >= 242486 tokens`), so grouping by the literal yields n=1 per row and hides the pattern. Normalize by replacing digit runs with a placeholder before grouping. - Report the count, the normalized reason, and the most recent timestamp, so an operator can tell a live problem from an old one. - **Zero rejections must produce no warning.** A rejection is not inherently an error — a genuinely impossible request should 422. The signal is a *rate*. ## Non-goals - Do not change `routing.py`. Nothing about the hard filters is wrong; the candidate set was correctly empty. This plan is observability only. - Do not auto-revert admin overrides or "helpfully" re-activate models. The deprecations were deliberate operator cost decisions. Warn; do not act. - Do not add a bucket per capability combination. - Do not replace `demand_ceiling_warnings` with an absolute tier-ordering comparison (see above — it is a theorem). - Do not touch the `local_vision` fallback. It could not have rescued this request anyway: that path receives the raw unpruned message list, and 242,486 tokens against `num_ctx 8192` was never going to fit. Worth a doc note, not a code change. ## Success criteria - A test reproducing the incident state — `kimi-k3*` excluded via `admin_model_overrides`, a vision request above 192,500 tokens — produces a warning naming the vision subset. The same state produces **no** warning from the existing `(tier, latency_tolerance)` checks, proving the gap was real and is now closed. - A test with recent NULL-selection rows produces a rejection-rate warning reporting a normalized reason and a count; a test with zero such rows produces none. - Both warnings appear in `coverage.warnings` on `/metrics` and in `/admin/api/snapshot`, and are visible on the admin dashboard's warnings bell. - `demand_ceiling_warnings` keeps its demand-relative shape; no absolute tier-ordering check is introduced. - Full suite green with `local_energy.enabled` both true and false. - The user's `config/config.local.yaml` is byte-identical after the run (`local_energy.enabled: true`, `tariff_usd_per_kwh: 0.159`). **Corrected 2026-09-05:** this previously named `config/config.yaml`. PR #26 moved deployment values into the gitignored overlay, so `config/config.yaml` is clean and tracked. The file that cannot be recovered is the overlay — it is not in git history at all. ## Superseded — the quota note below was resolved by PR #30 **Corrected 2026-09-05.** An earlier revision of this section said `coverage.warnings` reporting **146% of the 6.25 kWh plan** was "working as designed" and needed an operator decision rather than a code change. That was wrong on both counts and `plans/quota-balance-and-burn-rate.md` (merged as PR #30) replaced it. The warning was not working as designed: it claimed *"a quota is a wall, not a bill — requests fail rather than costing more"* while usage sat past 100% and nothing failed, because overage bills against a credit balance. It also was not the reason credits ran out on 2026-09-02/03 — the balance walking to $0.0071 was, and the router recorded `allowance_remaining_usd` on every response the whole way down while surfacing none of it. `/metrics` now reports `balance_usd`, `burn_rate_usd_per_hour` and `projected_hours_remaining`, and the warning text states that overage is billed against the credit balance and `plan_kwh_per_period` gates nothing. **Why this matters to THIS plan:** it is the same defect class. Both are "the signal was already recorded and nothing surfaced it" — the third and fourth instances. Gap 2 below (watch actual rejections) is the same fix applied to `route_decisions`, and it is the reactive safety net precisely because it does not require anyone to have predicted the dimension.