Files
6krrt/plans/capability-aware-ceiling-warnings.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

162 lines
8.0 KiB
Markdown

# Capability-aware ceiling warnings, and a reactive rejection detector
Status: done -- capability_ceilings in metrics.py
**Status: FINAL — decision-complete.** Written 2026-09-04 from a live incident.
## The incident
On 2026-09-04 two image requests failed with:
```
422 tier >= 1; context >= 242486 tokens; interactive;
vision-capable model (request carries image(s))
```
Cause: admin availability overrides had `kimi-k3`, `kimi-k3-fast` and
`kimi-k3-flex` deprecated. Those are the ONLY vision-capable rows with enough
context (782,324). Every remaining active vision model tops out at 192,500,
below the request's 242,486. The intersection of the hard filters was empty.
This is [incident #3](../docs/incidents.md) recurring through a dimension the
existing detector does not model. Resolved by re-activating `kimi-k3`; verified
by recomputing the candidate set through `routing.select_candidates`, which now
returns exactly one row.
## What already exists — do not rebuild it
Establish this before writing code, because a first pass at this analysis got
it wrong by checking for a top-level `warnings` key:
- **`/metrics` surfaces warnings under `coverage.warnings`**, not at top level.
`/admin/api/snapshot` carries the same block. The plumbing is done.
- **`metrics.context_ceilings` already applies admin deprecations** via
`exclude_models`, which was the fix after incident #3, and delegates to
`routing.select_candidates` so the ceiling matches live routing.
- **`metrics.demand_ceiling_warnings` already compares ceiling against observed
demand**, which is the *correct* shape. Do not replace it with an absolute
tier-ordering check: `ceiling(1) >= ceiling(2) >= ceiling(3)` is a theorem
(tier is a capability floor, so the eligible set shrinks monotonically), so
such a warning fires always and means nothing. `docs/incidents.md` #3 records
this trap; do not re-enter it.
- **`route_decisions` already records every rejection** — `selected_model IS
NULL` with the full `rejected_reason` string.
## Gap 1 — ceilings are blind to capability-gated subsets
`context_ceilings` buckets by `(tier, latency_tolerance)` only. During the
incident the tier-1 interactive ceiling across all models was still **782,324**
(deepseek and others are large and active), so every existing check was silent
while the *vision-capable* ceiling had collapsed to 192,500.
Measured during the incident:
| subset | ceiling |
|---|---|
| all active models | 782,324 |
| **vision-capable only** | **192,500** |
The hard filters that can independently empty the candidate set, from
`routing.rejection_reason`, are `require_vision` and `require_json_mode` —
both **fail closed on NULL** (an unconfirmed capability is treated as absent),
which is what makes them able to shrink the set sharply.
**Add capability sub-ceilings.** Extend the bucket key, or compute a small
number of additional named buckets: `vision` and `json_mode`. Do NOT add a
bucket per capability combination — that is a combinatorial explosion for
dimensions that do not interact in practice. Two extra series, compared against
the demand actually observed for requests carrying images / requesting JSON
(`route_decisions.images`, `.json_mode` are already recorded), is enough.
Reuse `demand_ceiling_warnings`' existing demand-relative shape; only the
bucketing changes.
## Gap 2 — nothing watches actual rejections
This is the more valuable half, and it is simpler.
Every hard-filter rejection is already persisted with its reason. Nothing reads
them. A detector that says *"N requests were rejected in the last hour"* would
have caught this incident within minutes, **without modelling any capability
dimension at all** — and would equally catch the next failure through a
dimension nobody predicted.
Predictive checks (Gap 1) only catch dimensions you thought of. A reactive
check catches everything, at the cost of firing after the first failure rather
than before. Ship both; the reactive one is the safety net.
Add a warning in `metrics.scoring_coverage`'s existing `warnings` list:
- Count `route_decisions` rows with `selected_model IS NULL AND rejected_reason
IS NOT NULL` in a recent window (1 hour and 24 hours are both useful; pick
one, state which).
- Group by a **normalized** reason, not the literal string. The raw strings
embed the request's token count (`context >= 242486 tokens`), so grouping by
the literal yields n=1 per row and hides the pattern. Normalize by replacing
digit runs with a placeholder before grouping.
- Report the count, the normalized reason, and the most recent timestamp, so an
operator can tell a live problem from an old one.
- **Zero rejections must produce no warning.** A rejection is not inherently an
error — a genuinely impossible request should 422. The signal is a *rate*.
## Non-goals
- Do not change `routing.py`. Nothing about the hard filters is wrong; the
candidate set was correctly empty. This plan is observability only.
- Do not auto-revert admin overrides or "helpfully" re-activate models. The
deprecations were deliberate operator cost decisions. Warn; do not act.
- Do not add a bucket per capability combination.
- Do not replace `demand_ceiling_warnings` with an absolute tier-ordering
comparison (see above — it is a theorem).
- Do not touch the `local_vision` fallback. It could not have rescued this
request anyway: that path receives the raw unpruned message list, and 242,486
tokens against `num_ctx 8192` was never going to fit. Worth a doc note, not a
code change.
## Success criteria
- A test reproducing the incident state — `kimi-k3*` excluded via
`admin_model_overrides`, a vision request above 192,500 tokens — produces a
warning naming the vision subset. The same state produces **no** warning from
the existing `(tier, latency_tolerance)` checks, proving the gap was real and
is now closed.
- A test with recent NULL-selection rows produces a rejection-rate warning
reporting a normalized reason and a count; a test with zero such rows produces
none.
- Both warnings appear in `coverage.warnings` on `/metrics` and in
`/admin/api/snapshot`, and are visible on the admin dashboard's warnings bell.
- `demand_ceiling_warnings` keeps its demand-relative shape; no absolute
tier-ordering check is introduced.
- Full suite green with `local_energy.enabled` both true and false.
- The user's `config/config.local.yaml` is byte-identical after the run
(`local_energy.enabled: true`, `tariff_usd_per_kwh: 0.159`). **Corrected
2026-09-05:** this previously named `config/config.yaml`. PR #26 moved
deployment values into the gitignored overlay, so `config/config.yaml` is
clean and tracked. The file that cannot be recovered is the overlay — it is
not in git history at all.
## Superseded — the quota note below was resolved by PR #30
**Corrected 2026-09-05.** An earlier revision of this section said
`coverage.warnings` reporting **146% of the 6.25 kWh plan** was "working as
designed" and needed an operator decision rather than a code change. That was
wrong on both counts and `plans/quota-balance-and-burn-rate.md` (merged as
PR #30) replaced it.
The warning was not working as designed: it claimed *"a quota is a wall, not a
bill — requests fail rather than costing more"* while usage sat past 100% and
nothing failed, because overage bills against a credit balance. It also was
not the reason credits ran out on 2026-09-02/03 — the balance walking to
$0.0071 was, and the router recorded `allowance_remaining_usd` on every
response the whole way down while surfacing none of it.
`/metrics` now reports `balance_usd`, `burn_rate_usd_per_hour` and
`projected_hours_remaining`, and the warning text states that overage is
billed against the credit balance and `plan_kwh_per_period` gates nothing.
**Why this matters to THIS plan:** it is the same defect class. Both are
"the signal was already recorded and nothing surfaced it" — the third and
fourth instances. Gap 2 below (watch actual rejections) is the same fix
applied to `route_decisions`, and it is the reactive safety net precisely
because it does not require anyone to have predicted the dimension.