Pinch is the router's only lever on the ~92.6% of cost that is prompt tokens, and until now there was no way to ask whether it earned its place. Instrumentation: - route_decisions gains pinch_original_tokens / pinch_final_tokens, in config/schema.sql and the code-side ensure_route_decisions migration, so live databases predating the columns pick them up on next start. - metrics.pinch_summary aggregates 30-day savings: share pruned, total and median tokens saved, and estimated dollars saved. The dollar figure is priced at the SAME blended prompt rate routing.estimated_cost uses (src/routing.py:363-374), including the fallback when the cached price is NULL, and reads cfg.objective.assumed_cache_rate so the dashboard and the router's own cost model cannot disagree. - Wired into /metrics, /admin/api/snapshot, the TUI model, and an admin dashboard card. Returns None when pinch is disabled, matching local_energy_summary's contract; all consumers guard for it. - The pinch log moves from debug to info, and only when pruning occurred. Accounting fix: - prune_context and _relevance_order_for both take extra_fixed_tokens, and the dispatcher passes the tool-definition overhead. opencode sends ~32k prompt tokens of tool definitions on a trivial request, and pinch was excluding all of it from the budget comparison. Threading it through _relevance_order_for as well matters: without it, relevance ordering would silently decline to run on exactly the requests that newly need pruning. Persistence covers the routed success, routed rejection, and passthrough paths. On passthrough the prune is hoisted above the persist so pinch stats are recorded rather than NULL -- the persist deliberately stays above _check_pinned_capabilities, which raises 422, so a rejected pin still writes the decision row that explains it. 860 tests pass, serially and under pytest-xdist. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
7.2 KiB
7.2 KiB