plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
162 lines
8.0 KiB
Markdown
162 lines
8.0 KiB
Markdown
# Capability-aware ceiling warnings, and a reactive rejection detector
|
|
|
|
Status: done -- capability_ceilings in metrics.py
|
|
|
|
**Status: FINAL — decision-complete.** Written 2026-09-04 from a live incident.
|
|
|
|
## The incident
|
|
|
|
On 2026-09-04 two image requests failed with:
|
|
|
|
```
|
|
422 tier >= 1; context >= 242486 tokens; interactive;
|
|
vision-capable model (request carries image(s))
|
|
```
|
|
|
|
Cause: admin availability overrides had `kimi-k3`, `kimi-k3-fast` and
|
|
`kimi-k3-flex` deprecated. Those are the ONLY vision-capable rows with enough
|
|
context (782,324). Every remaining active vision model tops out at 192,500,
|
|
below the request's 242,486. The intersection of the hard filters was empty.
|
|
|
|
This is [incident #3](../docs/incidents.md) recurring through a dimension the
|
|
existing detector does not model. Resolved by re-activating `kimi-k3`; verified
|
|
by recomputing the candidate set through `routing.select_candidates`, which now
|
|
returns exactly one row.
|
|
|
|
## What already exists — do not rebuild it
|
|
|
|
Establish this before writing code, because a first pass at this analysis got
|
|
it wrong by checking for a top-level `warnings` key:
|
|
|
|
- **`/metrics` surfaces warnings under `coverage.warnings`**, not at top level.
|
|
`/admin/api/snapshot` carries the same block. The plumbing is done.
|
|
- **`metrics.context_ceilings` already applies admin deprecations** via
|
|
`exclude_models`, which was the fix after incident #3, and delegates to
|
|
`routing.select_candidates` so the ceiling matches live routing.
|
|
- **`metrics.demand_ceiling_warnings` already compares ceiling against observed
|
|
demand**, which is the *correct* shape. Do not replace it with an absolute
|
|
tier-ordering check: `ceiling(1) >= ceiling(2) >= ceiling(3)` is a theorem
|
|
(tier is a capability floor, so the eligible set shrinks monotonically), so
|
|
such a warning fires always and means nothing. `docs/incidents.md` #3 records
|
|
this trap; do not re-enter it.
|
|
- **`route_decisions` already records every rejection** — `selected_model IS
|
|
NULL` with the full `rejected_reason` string.
|
|
|
|
## Gap 1 — ceilings are blind to capability-gated subsets
|
|
|
|
`context_ceilings` buckets by `(tier, latency_tolerance)` only. During the
|
|
incident the tier-1 interactive ceiling across all models was still **782,324**
|
|
(deepseek and others are large and active), so every existing check was silent
|
|
while the *vision-capable* ceiling had collapsed to 192,500.
|
|
|
|
Measured during the incident:
|
|
|
|
| subset | ceiling |
|
|
|---|---|
|
|
| all active models | 782,324 |
|
|
| **vision-capable only** | **192,500** |
|
|
|
|
The hard filters that can independently empty the candidate set, from
|
|
`routing.rejection_reason`, are `require_vision` and `require_json_mode` —
|
|
both **fail closed on NULL** (an unconfirmed capability is treated as absent),
|
|
which is what makes them able to shrink the set sharply.
|
|
|
|
**Add capability sub-ceilings.** Extend the bucket key, or compute a small
|
|
number of additional named buckets: `vision` and `json_mode`. Do NOT add a
|
|
bucket per capability combination — that is a combinatorial explosion for
|
|
dimensions that do not interact in practice. Two extra series, compared against
|
|
the demand actually observed for requests carrying images / requesting JSON
|
|
(`route_decisions.images`, `.json_mode` are already recorded), is enough.
|
|
|
|
Reuse `demand_ceiling_warnings`' existing demand-relative shape; only the
|
|
bucketing changes.
|
|
|
|
## Gap 2 — nothing watches actual rejections
|
|
|
|
This is the more valuable half, and it is simpler.
|
|
|
|
Every hard-filter rejection is already persisted with its reason. Nothing reads
|
|
them. A detector that says *"N requests were rejected in the last hour"* would
|
|
have caught this incident within minutes, **without modelling any capability
|
|
dimension at all** — and would equally catch the next failure through a
|
|
dimension nobody predicted.
|
|
|
|
Predictive checks (Gap 1) only catch dimensions you thought of. A reactive
|
|
check catches everything, at the cost of firing after the first failure rather
|
|
than before. Ship both; the reactive one is the safety net.
|
|
|
|
Add a warning in `metrics.scoring_coverage`'s existing `warnings` list:
|
|
|
|
- Count `route_decisions` rows with `selected_model IS NULL AND rejected_reason
|
|
IS NOT NULL` in a recent window (1 hour and 24 hours are both useful; pick
|
|
one, state which).
|
|
- Group by a **normalized** reason, not the literal string. The raw strings
|
|
embed the request's token count (`context >= 242486 tokens`), so grouping by
|
|
the literal yields n=1 per row and hides the pattern. Normalize by replacing
|
|
digit runs with a placeholder before grouping.
|
|
- Report the count, the normalized reason, and the most recent timestamp, so an
|
|
operator can tell a live problem from an old one.
|
|
- **Zero rejections must produce no warning.** A rejection is not inherently an
|
|
error — a genuinely impossible request should 422. The signal is a *rate*.
|
|
|
|
## Non-goals
|
|
|
|
- Do not change `routing.py`. Nothing about the hard filters is wrong; the
|
|
candidate set was correctly empty. This plan is observability only.
|
|
- Do not auto-revert admin overrides or "helpfully" re-activate models. The
|
|
deprecations were deliberate operator cost decisions. Warn; do not act.
|
|
- Do not add a bucket per capability combination.
|
|
- Do not replace `demand_ceiling_warnings` with an absolute tier-ordering
|
|
comparison (see above — it is a theorem).
|
|
- Do not touch the `local_vision` fallback. It could not have rescued this
|
|
request anyway: that path receives the raw unpruned message list, and 242,486
|
|
tokens against `num_ctx 8192` was never going to fit. Worth a doc note, not a
|
|
code change.
|
|
|
|
## Success criteria
|
|
|
|
- A test reproducing the incident state — `kimi-k3*` excluded via
|
|
`admin_model_overrides`, a vision request above 192,500 tokens — produces a
|
|
warning naming the vision subset. The same state produces **no** warning from
|
|
the existing `(tier, latency_tolerance)` checks, proving the gap was real and
|
|
is now closed.
|
|
- A test with recent NULL-selection rows produces a rejection-rate warning
|
|
reporting a normalized reason and a count; a test with zero such rows produces
|
|
none.
|
|
- Both warnings appear in `coverage.warnings` on `/metrics` and in
|
|
`/admin/api/snapshot`, and are visible on the admin dashboard's warnings bell.
|
|
- `demand_ceiling_warnings` keeps its demand-relative shape; no absolute
|
|
tier-ordering check is introduced.
|
|
- Full suite green with `local_energy.enabled` both true and false.
|
|
- The user's `config/config.local.yaml` is byte-identical after the run
|
|
(`local_energy.enabled: true`, `tariff_usd_per_kwh: 0.159`). **Corrected
|
|
2026-09-05:** this previously named `config/config.yaml`. PR #26 moved
|
|
deployment values into the gitignored overlay, so `config/config.yaml` is
|
|
clean and tracked. The file that cannot be recovered is the overlay — it is
|
|
not in git history at all.
|
|
|
|
## Superseded — the quota note below was resolved by PR #30
|
|
|
|
**Corrected 2026-09-05.** An earlier revision of this section said
|
|
`coverage.warnings` reporting **146% of the 6.25 kWh plan** was "working as
|
|
designed" and needed an operator decision rather than a code change. That was
|
|
wrong on both counts and `plans/quota-balance-and-burn-rate.md` (merged as
|
|
PR #30) replaced it.
|
|
|
|
The warning was not working as designed: it claimed *"a quota is a wall, not a
|
|
bill — requests fail rather than costing more"* while usage sat past 100% and
|
|
nothing failed, because overage bills against a credit balance. It also was
|
|
not the reason credits ran out on 2026-09-02/03 — the balance walking to
|
|
$0.0071 was, and the router recorded `allowance_remaining_usd` on every
|
|
response the whole way down while surfacing none of it.
|
|
|
|
`/metrics` now reports `balance_usd`, `burn_rate_usd_per_hour` and
|
|
`projected_hours_remaining`, and the warning text states that overage is
|
|
billed against the credit balance and `plan_kwh_per_period` gates nothing.
|
|
|
|
**Why this matters to THIS plan:** it is the same defect class. Both are
|
|
"the signal was already recorded and nothing surfaced it" — the third and
|
|
fourth instances. Gap 2 below (watch actual rejections) is the same fix
|
|
applied to `route_decisions`, and it is the reactive safety net precisely
|
|
because it does not require anyone to have predicted the dimension.
|