860 tests: ~26.8s serially, ~7.9s under -n auto. Pinned like every other
dependency -- recreating the venv with unpinned ranges has silently jumped
major versions in this repo before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
score_code's CODE_TIMEOUT_SECONDS is 15, and this test waits the whole
thing out to prove an infinite loop is bounded rather than hung -- 15.08s,
about 39% of the entire suite's runtime, to assert a timeout fires.
Monkeypatching the module attribute to 1 gives the identical
(0.0, 'timeout') result in ~1s. src/eval_proficiency.py is untouched; the
production timeout stays 15s.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
Pinch is the router's only lever on the ~92.6% of cost that is prompt
tokens, and until now there was no way to ask whether it earned its place.
Instrumentation:
- route_decisions gains pinch_original_tokens / pinch_final_tokens, in
config/schema.sql and the code-side ensure_route_decisions migration, so
live databases predating the columns pick them up on next start.
- metrics.pinch_summary aggregates 30-day savings: share pruned, total and
median tokens saved, and estimated dollars saved. The dollar figure is
priced at the SAME blended prompt rate routing.estimated_cost uses
(src/routing.py:363-374), including the fallback when the cached price is
NULL, and reads cfg.objective.assumed_cache_rate so the dashboard and the
router's own cost model cannot disagree.
- Wired into /metrics, /admin/api/snapshot, the TUI model, and an admin
dashboard card. Returns None when pinch is disabled, matching
local_energy_summary's contract; all consumers guard for it.
- The pinch log moves from debug to info, and only when pruning occurred.
Accounting fix:
- prune_context and _relevance_order_for both take extra_fixed_tokens, and
the dispatcher passes the tool-definition overhead. opencode sends ~32k
prompt tokens of tool definitions on a trivial request, and pinch was
excluding all of it from the budget comparison. Threading it through
_relevance_order_for as well matters: without it, relevance ordering would
silently decline to run on exactly the requests that newly need pruning.
Persistence covers the routed success, routed rejection, and passthrough
paths. On passthrough the prune is hoisted above the persist so pinch stats
are recorded rather than NULL -- the persist deliberately stays above
_check_pinned_capabilities, which raises 422, so a rejected pin still writes
the decision row that explains it.
860 tests pass, serially and under pytest-xdist.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
The spec said to price saved tokens at the cached prompt rate. That is wrong
twice: cost_per_1m_prompt_cached is already the discounted per-1M price, so
multiplying it by cache_rate applies the discount a second time, and it drops
the (1 - cache_rate) fraction billed at full price entirely.
Match routing.estimated_cost (src/routing.py:363-374), including its COALESCE
fallback for a NULL cached price. Otherwise the admin dashboard would report a
savings figure computed on a different basis than the cost model the router
actually ranks on.
Caught by Metis during plan gap analysis.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U