Files
6krrt/tests/test_tui_warnings.py
adlee-was-taken 6f9f66351d feat(metrics): a cache-rate series, and an expiry check on assumed_cache_rate
Wave 1 item 1.4 of plans/token-waste-waves.md.

`objective.assumed_cache_rate: 0.917` was measured once, on 2026-08-23, over
40.7M tokens, and never re-measured. It is the highest-leverage term in
`routing.estimated_cost` on a 100k-token prompt. Nothing checked whether it
was still true.

`cache_rate_series(conn, cfg)` reports
`sum(cached_prompt_tokens) / sum(prompt_tokens)` per (provider, model) over a
trailing window, token-weighted rather than a mean of per-request rates -- the
bill is denominated in tokens, so one 200k-token turn is not one observation's
worth of evidence against a 1k one. It rides in `coverage` rather than as a
new top-level /metrics key, because /metrics assembles its payload in
dispatcher.py (item 1.3's file) and this is a coverage question anyway.

Two filters, both of which have already been got wrong once in this project:

- Only rows the provider actually reported. `cached_prompt_tokens IS NOT NULL`
  is the operative test, NOT `cached_tokens_source = 'reported'`: by af18009's
  own invariant those are the same set, but the NOT NULL form also keeps rows
  written BEFORE that column existed, whose NULL source means "predates the
  column" and not "the provider said nothing". The live router.db is exactly
  such a database -- it has not restarted since af18009, so the ALTER has not
  run, and a strict `= 'reported'` query returns zero rows there today. The
  source clause is added as a redundant restatement of the invariant so a
  future divergence between the two fails closed instead of quietly widening
  the denominator, and a PRAGMA probe keeps a pre-migration schema from 500ing
  the whole payload (metrics is imported by admin.py too, which can open a
  database that has not been through dispatcher start-up).
- No `seed_reference` rows. Those sweeps never carry a cached count and are
  not user traffic; they were the entire reason NeuralWatt's coverage read
  12.8% against ~98% on real dispatch.

`cache_rate_warnings` compares against the configured constant rather than a
floor. "Is the cache working" is the wrong question -- a deployment at 0.60 is
not broken, it is mispriced -- so the signal is DIVERGENCE, in either
direction, since a real 0.99 underprices every candidate just as surely.
Two classes, the shape `rejection_warnings` uses:

- `cache rate:` -- the aggregate is more than `cache_rate_warn_margin` from
  the assumption. This is the premise-expiry check.
- `cache rate outlier:` -- one (provider, model_id) group is, grouped on the
  structured columns so a provider serving the same base model twice stays two
  groups.

Neither fires on mere presence: both need
`cache_rate_warn_min_observations` reported-cache rows, the way
`classifier_degradation_warning` needs `degraded_warn_min`.

Measured read-only against the live DB, 168h window, 280 observations:

  aggregate                                   0.893
  openrouter  xiaomi/mimo-v2.5      n=136     0.953
  neuralwatt  deepseek-v4-flash     n=39      0.914
  neuralwatt  qwen3.6-35b-fast      n=63      0.870
  openrouter  deepseek/deepseek-v4-flash n=36 0.731
  openrouter  xiaomi/mimo-v2.5-pro  n=4       0.479
  neuralwatt  glm-5.3-flash         n=2       0.436

So the constant still holds in aggregate (0.893 against 0.917, inside the
0.10 margin, nothing fires) while OpenRouter's deepseek route sits 0.186
below it and the outlier class names it. That is the per-group half earning
its keep on the first run: the aggregate was quiet.

Three config keys, all in config.yaml with their reasoning:
`cache_rate_window_hours: 168`, `cache_rate_warn_margin: 0.10`,
`cache_rate_warn_min_observations: 25`.

Two chernobyl seeds and two matchers in tests/test_tui_warnings.py so both
classes render in #warnings-panel. The seeds carry energy_kwh 0.0, a
non-seed_reference category and a NULL cost_usd specifically so they cannot
silence the quota, missing-energy or spend classes the way a careless seed has
before. 15 new tests; no TUI change was needed, the panel already renders the
warnings list generically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-13 11:19:13 -04:00

16 KiB