From a live incident on 2026-09-04: two image requests 422'd with
"tier >= 1; context >= 242486 tokens; interactive; vision-capable model".
Admin overrides had kimi-k3/-fast/-flex deprecated, and those are the only
vision-capable rows with enough context (782,324). Every remaining active
vision model tops out at 192,500. Incident #3 recurring through a dimension
the existing detector does not model.
The existing machinery is fine and must not be rebuilt: warnings ARE surfaced
(under coverage.warnings, not a top-level key -- a first pass at this analysis
got that wrong), context_ceilings already applies admin deprecations, and
demand_ceiling_warnings already has the correct demand-relative shape.
Gap 1: ceilings bucket by (tier, latency_tolerance) only. During the incident
the tier-1 interactive ceiling was still 782,324 across all models, so every
check stayed silent while the vision-capable ceiling had collapsed to 192,500.
Add vision and json_mode sub-ceilings -- the two hard filters that fail closed
on NULL and can independently empty the set.
Gap 2, and the more valuable half: nothing reads route_decisions' recorded
rejections. A rejection-rate warning would have caught this in minutes without
modelling any capability dimension, and would catch the next failure through a
dimension nobody predicted. Predictive checks only catch what you thought of.
Non-goals are explicit: observability only, no routing changes, and do NOT
auto-revert admin overrides -- those deprecations are deliberate operator cost
decisions. Warn, do not act.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
A profile restricts the candidate set; ranking stays quality-first with cost as
the tiebreak. Filter-only is sufficient for all three named profiles because a
pure-filter 'locality' leaves no cloud row to lose to -- the 0.767-vs-0.95
dormancy only blocks MIXED sets that prefer local without excluding cloud.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
The mechanism is already half-built: auto:batch IS a profile -- a named variant
of auto that changes one hard-filter parameter (latency_tolerance), shipped and
documented for months. This generalises that hard-coded special case rather
than inventing anything.
Three call sites do most of the work already: the wants_routing membership test
(dispatcher.py:2937) becomes a parse; the latency ternary (:2955) becomes one
case of the general mechanism rather than surviving beside it; and
routing.select_candidates already takes exclude_models, task_category,
latency_tolerance and the rest.
Exactly one primitive is missing. select_candidates has a denylist and no
allowlist, while every profile the user named -- locality, onlycheaps,
bigboybritches -- is naturally an allowlist. One `restrict_to` parameter
threaded through rejection_reason -> is_eligible -> select_candidates covers
it, with None meaning unrestricted and its own rejection string so a 422 names
the profile that emptied the set.
One decision is left open for the user and flagged inline: whether a profile is
a candidate filter only, or may also override objective.*. Recommendation is
filter-only -- a pure-filter `locality` already works because restricting to
local rows leaves no cloud row to lose to, which is what the 0.767-vs-0.95
dormancy finding actually blocks.
Also reconciles expand-local-llm-usage's stale checkboxes: todos 1-12 shipped
via PR #22 and todo 13 was executed by hand, verified against six deliverables.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U