feat: capability-aware ceiling warnings + reactive rejection detector #31

Merged
alee merged 2 commits from feat/capability-aware-ceiling-warnings into main 2026-09-05 05:42:54 +00:00
Owner

Capability-aware ceiling warnings + reactive rejection detector

Closes the observability gap behind the 2026-09-04 incident, where the overall tier-1 ceiling looked healthy (782,324, served by non-vision models) while the vision-only sub-ceiling had collapsed to 192,500 — image requests 422'd with zero warnings. Observability only: warn, never act. No routing changes, no auto-revert of admin overrides, no model re-activation.

What's added

1. Capability sub-ceilings + demand warnings (src/metrics.py):

  • capability_ceilings(conn, cfg) — per-dimension {vision, json_mode} × (tier, latency_tolerance) context ceilings, computed through the same routing.select_candidates path live routing uses (admin deprecations included). Returns {dimension: {(tier, lat_tol): {"ceiling", "count"}}}.
  • capability_demand_warnings — demand-relative only: fires vision-capable tier 1 context ceiling (0) is below observed max demand (200000, window: last 30d) for requests carrying images — requests above this may return 422 when observed demand for capability-carrying requests exceeds the sub-ceiling (7d/30d window fallback mirroring demand_ceiling_warnings). No absolute tier-ordering check (that's a theorem), no vision+json_mode combination bucket.

2. Reactive rejection detector (rejection_warnings):

  • Groups rejections (selected_model IS NULL) by (task_tier, normalized_reason) — tier from the structured column, never parsed from the string. Digit normalization (context >= 242486 tokens → context >= N tokens) merges token-count noise; only the column keeps tier >= 1 vs tier >= 3 separable.
  • Signal is disjunctive — novelty OR rate — because a bare threshold provably cannot work: routine traffic on this deployment is ~3 rejections/hr (6 in 24h, 17 all-time), but the incident produced only n=2, below the noise floor. A novel group (zero occurrences in the prior 24h baseline, count ≥ 2 = NOVEL_GROUP_MIN_COUNT, incident citation in the comment) warns as new rejection pattern:; a familiar group warns as rejection rate: only at count ≥ 6 (2× the measured routine peak). n=1 novel is silent (a genuinely impossible request SHOULD 422); zero rejections → zero warnings.
  • Knobs: objective.rejection_warning_window_hours: 1, rejection_warning_baseline_hours: 24, rejection_warning_min_count: 6 (config.yaml comment records the live-DB measurements justifying all three).
  • Known benign false-positive (documented in code, deliberate non-fix): novelty means "zero occurrences in the prior baseline window", so a familiar-but-intermittent pattern can read as novel after a quiet gap and fire one spurious warning — self-clearing; widen rejection_warning_baseline_hours to suppress.

Both warning classes flow through the existing scoring_coverage.warnings list → /metrics, /admin/api/snapshot, the admin bell, and the TUI automatically. 13 new regression tests (incident reproduction with both halves — new check fires AND existing (tier, latency_tolerance) checks stay silent on the same state; tier discrimination; novelty/rate fire+silence matrix; window boundary; custom window; endpoint integration).

Verification evidence (.omo/evidence/capability-aware-ceiling-warnings/)

Fail-on-main first (T3-pytest.log) — every new regression test was run against unmodified main (deeaf80) in a throwaway reference worktree BEFORE passing on the fix:

tests/test_metrics.py  → ImportError: cannot import name 'capability_ceilings' from 'metrics' (collection error, 12 new tests fail)
tests/test_metrics_endpoint.py::test_metrics_coverage_warnings_include_rejection_and_capability → FAILED
  assert any('rejection' in w for w in warnings)  # main's /metrics has no rejection warning on the same seeded state

Suite result (T4-final.txt) — with local_energy.enabled true (worktree overlay copy): 1212 passed; toggled to false (worktree copy only): 1212 passed.

Root overlay integrity — config/config.local.yaml (gitignored, user tariff) bookended by sha256 at every step, final value:

cc3d4c2a12613b069e80cd39765024fa1100e3629f50ac4d40c0c2c744ea5549  (byte-identical throughout)

Real manual QA (F3-manual-qa.txt) — throwaway uvicorn on 127.0.0.1:8081 from isolated run dirs with seeded DBs:

  • Incident state: coverage.warnings = 4 (2 pre-existing + both new types verbatim).
  • Clean state (no overrides, no rejections): only the 2 pre-existing eco/proficiency warnings — neither new type fires.

Review verdicts (Final Verification Wave): F1 plan compliance APPROVE (8/8) · F2 code quality APPROVE (9/9, 2 cosmetic nits) · F3 manual QA APPROVE · F4 scope fidelity APPROVE (exactly 5 files: src/metrics.py, src/config.py, config/config.yaml, tests/test_metrics.py, tests/test_metrics_endpoint.py; routing.py and dispatcher.py byte-identical to main; metrics reads-only — zero INSERT/UPDATE/DELETE).

Notes for reviewers

  • test_rejection_warnings_old_rejections_not_counted seeds two 3-day-old rows instead of the plan's one — a single old row is silent even with broken window filtering (n=1 < novelty floor); two at the floor make the silence provably age-dependent. Test strengthened, assertion unchanged.
  • capability_ceilings delegates to context_ceilings_with_rows, which gained backward-compatible require_vision/require_json_mode params (defaults preserve old behavior) — one copy of the filter logic, identical select_candidates calls to the plan's spec.
  • Deliberately NOT done: no auto-revert of the kimi-k3* deprecations (operator cost decisions), no routing-path change, no auth on the warnings endpoints (existing posture unchanged).
# Capability-aware ceiling warnings + reactive rejection detector Closes the observability gap behind the 2026-09-04 incident, where the overall tier-1 ceiling looked healthy (782,324, served by non-vision models) while the **vision-only** sub-ceiling had collapsed to 192,500 — image requests 422'd with **zero warnings**. Observability only: warn, never act. No routing changes, no auto-revert of admin overrides, no model re-activation. ## What's added **1. Capability sub-ceilings + demand warnings** (`src/metrics.py`): - `capability_ceilings(conn, cfg)` — per-dimension `{vision, json_mode} × (tier, latency_tolerance)` context ceilings, computed through the same `routing.select_candidates` path live routing uses (admin deprecations included). Returns `{dimension: {(tier, lat_tol): {"ceiling", "count"}}}`. - `capability_demand_warnings` — demand-relative only: fires `vision-capable tier 1 context ceiling (0) is below observed max demand (200000, window: last 30d) for requests carrying images — requests above this may return 422` when observed demand for capability-carrying requests exceeds the sub-ceiling (7d/30d window fallback mirroring `demand_ceiling_warnings`). No absolute tier-ordering check (that's a theorem), no vision+json_mode combination bucket. **2. Reactive rejection detector** (`rejection_warnings`): - Groups rejections (`selected_model IS NULL`) by `(task_tier, normalized_reason)` — **tier from the structured column**, never parsed from the string. Digit normalization (`context >= 242486 tokens` → `context >= N tokens`) merges token-count noise; only the column keeps `tier >= 1` vs `tier >= 3` separable. - Signal is **disjunctive — novelty OR rate** — because a bare threshold provably cannot work: routine traffic on this deployment is ~3 rejections/hr (6 in 24h, 17 all-time), but the incident produced only **n=2**, below the noise floor. A **novel** group (zero occurrences in the prior 24h baseline, count ≥ 2 = `NOVEL_GROUP_MIN_COUNT`, incident citation in the comment) warns as `new rejection pattern:`; a **familiar** group warns as `rejection rate:` only at count ≥ 6 (2× the measured routine peak). n=1 novel is silent (a genuinely impossible request SHOULD 422); zero rejections → zero warnings. - Knobs: `objective.rejection_warning_window_hours: 1`, `rejection_warning_baseline_hours: 24`, `rejection_warning_min_count: 6` (config.yaml comment records the live-DB measurements justifying all three). - Known benign false-positive (documented in code, deliberate non-fix): novelty means "zero occurrences in the prior baseline window", so a familiar-but-intermittent pattern can read as novel after a quiet gap and fire one spurious warning — self-clearing; widen `rejection_warning_baseline_hours` to suppress. Both warning classes flow through the existing `scoring_coverage.warnings` list → `/metrics`, `/admin/api/snapshot`, the admin bell, and the TUI automatically. **13 new regression tests** (incident reproduction with both halves — new check fires AND existing `(tier, latency_tolerance)` checks stay silent on the same state; tier discrimination; novelty/rate fire+silence matrix; window boundary; custom window; endpoint integration). ## Verification evidence (`.omo/evidence/capability-aware-ceiling-warnings/`) **Fail-on-main first** (T3-pytest.log) — every new regression test was run against unmodified main (deeaf80) in a throwaway reference worktree BEFORE passing on the fix: ``` tests/test_metrics.py → ImportError: cannot import name 'capability_ceilings' from 'metrics' (collection error, 12 new tests fail) tests/test_metrics_endpoint.py::test_metrics_coverage_warnings_include_rejection_and_capability → FAILED assert any('rejection' in w for w in warnings) # main's /metrics has no rejection warning on the same seeded state ``` **Suite result** (T4-final.txt) — with `local_energy.enabled` **true** (worktree overlay copy): `1212 passed`; toggled to **false** (worktree copy only): `1212 passed`. **Root overlay integrity** — `config/config.local.yaml` (gitignored, user tariff) bookended by sha256 at every step, final value: ``` cc3d4c2a12613b069e80cd39765024fa1100e3629f50ac4d40c0c2c744ea5549 (byte-identical throughout) ``` **Real manual QA** (F3-manual-qa.txt) — throwaway uvicorn on 127.0.0.1:8081 from isolated run dirs with seeded DBs: - Incident state: `coverage.warnings` = 4 (2 pre-existing + both new types verbatim). - Clean state (no overrides, no rejections): only the 2 pre-existing eco/proficiency warnings — neither new type fires. **Review verdicts (Final Verification Wave)**: F1 plan compliance APPROVE (8/8) · F2 code quality APPROVE (9/9, 2 cosmetic nits) · F3 manual QA APPROVE · F4 scope fidelity APPROVE (exactly 5 files: `src/metrics.py`, `src/config.py`, `config/config.yaml`, `tests/test_metrics.py`, `tests/test_metrics_endpoint.py`; `routing.py` and `dispatcher.py` byte-identical to main; metrics reads-only — zero INSERT/UPDATE/DELETE). ## Notes for reviewers - `test_rejection_warnings_old_rejections_not_counted` seeds two 3-day-old rows instead of the plan's one — a single old row is silent even with broken window filtering (n=1 < novelty floor); two at the floor make the silence provably age-dependent. Test strengthened, assertion unchanged. - `capability_ceilings` delegates to `context_ceilings_with_rows`, which gained backward-compatible `require_vision`/`require_json_mode` params (defaults preserve old behavior) — one copy of the filter logic, identical `select_candidates` calls to the plan's spec. - Deliberately NOT done: no auto-revert of the kimi-k3* deprecations (operator cost decisions), no routing-path change, no auth on the warnings endpoints (existing posture unchanged).
alee added 2 commits 2026-09-05 05:34:41 +00:00
alee merged commit 1ae0908bc0 into main 2026-09-05 05:42:54 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#31