Files
6krrt/plans/capability-aware-ceiling-warnings.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

8.0 KiB

Capability-aware ceiling warnings, and a reactive rejection detector

Status: done -- capability_ceilings in metrics.py

Status: FINAL — decision-complete. Written 2026-09-04 from a live incident.

The incident

On 2026-09-04 two image requests failed with:

422  tier >= 1; context >= 242486 tokens; interactive;
     vision-capable model (request carries image(s))

Cause: admin availability overrides had kimi-k3, kimi-k3-fast and kimi-k3-flex deprecated. Those are the ONLY vision-capable rows with enough context (782,324). Every remaining active vision model tops out at 192,500, below the request's 242,486. The intersection of the hard filters was empty.

This is incident #3 recurring through a dimension the existing detector does not model. Resolved by re-activating kimi-k3; verified by recomputing the candidate set through routing.select_candidates, which now returns exactly one row.

What already exists — do not rebuild it

Establish this before writing code, because a first pass at this analysis got it wrong by checking for a top-level warnings key:

  • /metrics surfaces warnings under coverage.warnings, not at top level. /admin/api/snapshot carries the same block. The plumbing is done.
  • metrics.context_ceilings already applies admin deprecations via exclude_models, which was the fix after incident #3, and delegates to routing.select_candidates so the ceiling matches live routing.
  • metrics.demand_ceiling_warnings already compares ceiling against observed demand, which is the correct shape. Do not replace it with an absolute tier-ordering check: ceiling(1) >= ceiling(2) >= ceiling(3) is a theorem (tier is a capability floor, so the eligible set shrinks monotonically), so such a warning fires always and means nothing. docs/incidents.md #3 records this trap; do not re-enter it.
  • route_decisions already records every rejection — selected_model IS NULL with the full rejected_reason string.

Gap 1 — ceilings are blind to capability-gated subsets

context_ceilings buckets by (tier, latency_tolerance) only. During the incident the tier-1 interactive ceiling across all models was still 782,324 (deepseek and others are large and active), so every existing check was silent while the vision-capable ceiling had collapsed to 192,500.

Measured during the incident:

subset ceiling
all active models 782,324
vision-capable only 192,500

The hard filters that can independently empty the candidate set, from routing.rejection_reason, are require_vision and require_json_mode — both fail closed on NULL (an unconfirmed capability is treated as absent), which is what makes them able to shrink the set sharply.

Add capability sub-ceilings. Extend the bucket key, or compute a small number of additional named buckets: vision and json_mode. Do NOT add a bucket per capability combination — that is a combinatorial explosion for dimensions that do not interact in practice. Two extra series, compared against the demand actually observed for requests carrying images / requesting JSON (route_decisions.images, .json_mode are already recorded), is enough.

Reuse demand_ceiling_warnings' existing demand-relative shape; only the bucketing changes.

Gap 2 — nothing watches actual rejections

This is the more valuable half, and it is simpler.

Every hard-filter rejection is already persisted with its reason. Nothing reads them. A detector that says "N requests were rejected in the last hour" would have caught this incident within minutes, without modelling any capability dimension at all — and would equally catch the next failure through a dimension nobody predicted.

Predictive checks (Gap 1) only catch dimensions you thought of. A reactive check catches everything, at the cost of firing after the first failure rather than before. Ship both; the reactive one is the safety net.

Add a warning in metrics.scoring_coverage's existing warnings list:

  • Count route_decisions rows with selected_model IS NULL AND rejected_reason IS NOT NULL in a recent window (1 hour and 24 hours are both useful; pick one, state which).
  • Group by a normalized reason, not the literal string. The raw strings embed the request's token count (context >= 242486 tokens), so grouping by the literal yields n=1 per row and hides the pattern. Normalize by replacing digit runs with a placeholder before grouping.
  • Report the count, the normalized reason, and the most recent timestamp, so an operator can tell a live problem from an old one.
  • Zero rejections must produce no warning. A rejection is not inherently an error — a genuinely impossible request should 422. The signal is a rate.

Non-goals

  • Do not change routing.py. Nothing about the hard filters is wrong; the candidate set was correctly empty. This plan is observability only.
  • Do not auto-revert admin overrides or "helpfully" re-activate models. The deprecations were deliberate operator cost decisions. Warn; do not act.
  • Do not add a bucket per capability combination.
  • Do not replace demand_ceiling_warnings with an absolute tier-ordering comparison (see above — it is a theorem).
  • Do not touch the local_vision fallback. It could not have rescued this request anyway: that path receives the raw unpruned message list, and 242,486 tokens against num_ctx 8192 was never going to fit. Worth a doc note, not a code change.

Success criteria

  • A test reproducing the incident state — kimi-k3* excluded via admin_model_overrides, a vision request above 192,500 tokens — produces a warning naming the vision subset. The same state produces no warning from the existing (tier, latency_tolerance) checks, proving the gap was real and is now closed.
  • A test with recent NULL-selection rows produces a rejection-rate warning reporting a normalized reason and a count; a test with zero such rows produces none.
  • Both warnings appear in coverage.warnings on /metrics and in /admin/api/snapshot, and are visible on the admin dashboard's warnings bell.
  • demand_ceiling_warnings keeps its demand-relative shape; no absolute tier-ordering check is introduced.
  • Full suite green with local_energy.enabled both true and false.
  • The user's config/config.local.yaml is byte-identical after the run (local_energy.enabled: true, tariff_usd_per_kwh: 0.159). Corrected 2026-09-05: this previously named config/config.yaml. PR #26 moved deployment values into the gitignored overlay, so config/config.yaml is clean and tracked. The file that cannot be recovered is the overlay — it is not in git history at all.

Superseded — the quota note below was resolved by PR #30

Corrected 2026-09-05. An earlier revision of this section said coverage.warnings reporting 146% of the 6.25 kWh plan was "working as designed" and needed an operator decision rather than a code change. That was wrong on both counts and plans/quota-balance-and-burn-rate.md (merged as PR #30) replaced it.

The warning was not working as designed: it claimed "a quota is a wall, not a bill — requests fail rather than costing more" while usage sat past 100% and nothing failed, because overage bills against a credit balance. It also was not the reason credits ran out on 2026-09-02/03 — the balance walking to $0.0071 was, and the router recorded allowance_remaining_usd on every response the whole way down while surfacing none of it.

/metrics now reports balance_usd, burn_rate_usd_per_hour and projected_hours_remaining, and the warning text states that overage is billed against the credit balance and plan_kwh_per_period gates nothing.

Why this matters to THIS plan: it is the same defect class. Both are "the signal was already recorded and nothing surfaced it" — the third and fourth instances. Gap 2 below (watch actual rejections) is the same fix applied to route_decisions, and it is the reactive safety net precisely because it does not require anyone to have predicted the dimension.