Wave 5.3 — premise-expiry checks for quality_tolerance and cost-as-tiebreak #97

Merged
alee merged 1 commits from feat/wave53-premise-expiry into main 2026-09-19 15:57:02 +00:00
Owner

Wave 5.3 — premise-expiry checks for quality_tolerance and cost-as-tiebreak

Two new /metrics detector pairs following the Wave 1.4 cache-rate precedent (series + warning with a minimum-observation floor, wired into scoring_coverage), plus three config.yaml comment rewrites.

Changes

src/metrics.py:

  • proficiency_sample_depth_series / proficiency_sample_depth_warnings: average proficiency sample depth (outcome_samples + self_eval_samples) per category and overall. Warns "quality_tolerance premise expired:" once depth crosses objective.proficiency_depth_warn_min_samples with enough rows on the table — the "2-3 samples per category" justification for the loose tolerance is now enforced by the check, not memory.
  • cumulative_spend_series / cumulative_spend_warnings: all-time SUM(cost_usd) from energy_observations (excluding seed_reference rows, per-provider breakdown). Warns "cost-as-tiebreak premise expired:" once spend passes objective.cumulative_spend_warn_usd — the "$0.07 / fractions of a cent" figure in the blend comment was off by orders of magnitude.
  • Both wired into scoring_coverage() in the same assembly pattern as the cache-rate pair.

src/config.py:

  • Four new Optional knobs on the Objective StrictModel: proficiency_depth_warn_min_samples, proficiency_depth_warn_min_rows, cumulative_spend_warn_usd, cumulative_spend_warn_min_rows, with > 0 validators.

config/config.yaml:

  • Four new config keys under objective: with values and comments.
  • Three comment rewrites: quality_tolerance and the cost-as-tiebreak/blend section now reference the live check function names instead of embedding a stale snapshot number; pinch.relevance.enabled now describes its actual ON-by-default state and the budget-gate rationale from commit 69662f0, correcting the old "off by default... then decide the default" text that was never updated when the decision was made.

tests/test_metrics.py:

  • 15 new tests + 3 seed helpers, covering silent-below-floor, silent-below-threshold, fires-on-expiry, and integration via scoring_coverage.

tests/test_admin_knob_coverage.py:

  • 4 new entries in _CONFIG_ALLOWLIST — the gate requires every scalar leaf of the config to be reachable.

Verification against live production DB

Both checks correctly fire against real data (read-only probe to the shared checkout's router.db):

  • Depth: 3 categories (coding_general avg 42.3, coding_refactor 21.6, tool_use_agentic 21.9) exceed the 20-sample threshold, firing per-category warnings.
  • Spend: cumulative $123.99 across 32,184 priced rows fires both aggregate and per-provider (neuralwatt $119.11) warnings — 1,771× the "$0.07" the old comment cited.

Suite

Full suite: 2163 passed, 0 failed.

## Wave 5.3 — premise-expiry checks for `quality_tolerance` and cost-as-tiebreak Two new `/metrics` detector pairs following the Wave 1.4 cache-rate precedent (series + warning with a minimum-observation floor, wired into `scoring_coverage`), plus three config.yaml comment rewrites. ### Changes **`src/metrics.py`:** - `proficiency_sample_depth_series` / `proficiency_sample_depth_warnings`: average proficiency sample depth (`outcome_samples + self_eval_samples`) per category and overall. Warns `"quality_tolerance premise expired:"` once depth crosses `objective.proficiency_depth_warn_min_samples` with enough rows on the table — the "2-3 samples per category" justification for the loose tolerance is now enforced by the check, not memory. - `cumulative_spend_series` / `cumulative_spend_warnings`: all-time `SUM(cost_usd)` from `energy_observations` (excluding seed_reference rows, per-provider breakdown). Warns `"cost-as-tiebreak premise expired:"` once spend passes `objective.cumulative_spend_warn_usd` — the "$0.07 / fractions of a cent" figure in the blend comment was off by orders of magnitude. - Both wired into `scoring_coverage()` in the same assembly pattern as the cache-rate pair. **`src/config.py`:** - Four new `Optional` knobs on the `Objective` `StrictModel`: `proficiency_depth_warn_min_samples`, `proficiency_depth_warn_min_rows`, `cumulative_spend_warn_usd`, `cumulative_spend_warn_min_rows`, with `> 0` validators. **`config/config.yaml`:** - Four new config keys under `objective:` with values and comments. - Three comment rewrites: `quality_tolerance` and the cost-as-tiebreak/blend section now reference the **live check function names** instead of embedding a stale snapshot number; `pinch.relevance.enabled` now describes its actual ON-by-default state and the budget-gate rationale from commit `69662f0`, correcting the old "off by default... then decide the default" text that was never updated when the decision was made. **`tests/test_metrics.py`:** - 15 new tests + 3 seed helpers, covering silent-below-floor, silent-below-threshold, fires-on-expiry, and integration via `scoring_coverage`. **`tests/test_admin_knob_coverage.py`:** - 4 new entries in `_CONFIG_ALLOWLIST` — the gate requires every scalar leaf of the config to be reachable. ### Verification against live production DB Both checks **correctly fire** against real data (read-only probe to the shared checkout's `router.db`): - **Depth**: 3 categories (`coding_general` avg 42.3, `coding_refactor` 21.6, `tool_use_agentic` 21.9) exceed the 20-sample threshold, firing per-category warnings. - **Spend**: cumulative $123.99 across 32,184 priced rows fires both aggregate and per-provider (`neuralwatt` $119.11) warnings — 1,771× the "$0.07" the old comment cited. ### Suite Full suite: 2163 passed, 0 failed.
alee added 1 commit 2026-09-19 03:49:17 +00:00
Two new /metrics detector pairs following the Wave 1.4 cache-rate precedent
(series + warning with a minimum-observation floor, wired into
scoring_coverage):

- proficiency_sample_depth_series/_warnings: average proficiency sample
  depth (outcome_samples + self_eval_samples) per category and overall.
  Warns 'quality_tolerance premise expired:' once depth crosses
  objective.proficiency_depth_warn_min_samples with enough rows on the
  table — the '2-3 samples per category' justification for the loose
  tolerance is now enforced by the check, not memory.
- cumulative_spend_series/_warnings: all-time SUM(cost_usd) from
  energy_observations excluding seed_reference rows, with per-provider
  breakdown. Warns 'cost-as-tiebreak premise expired:' once spend passes
  objective.cumulative_spend_warn_usd — the '$0.07 / fractions of a cent'
  figure in the blend comment was off by orders of magnitude.

Four objective knobs (Optional, >0 validators, real values + comments in
config.yaml, deliberate-absent entries in the admin-knob gate) and three
config.yaml comment rewrites: the two stale justifications now reference
the live check function names instead of hardcoded numbers, and the
pinch.relevance.enabled comment now describes the actual ON-by-default
state and its budget-gate rationale (commit 69662f0) instead of the
pre-decision text.

Full suite: 2163 passed, 0 failed.
alee merged commit 9f07bfdeb4 into main 2026-09-19 15:57:02 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#97