feat(pinch): instrument savings and count tool overhead in the budget #18
Reference in New Issue
Block a user
Delete Branch "feat/pinch-instrumentation-and-token-accounting"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Pinch is the router's only lever on the ~92.6% of request cost that is prompt tokens, and until now there was no way to ask whether it earned its place. This makes it observable and fixes an accounting bug that was hiding most of its work.
Instrumentation
route_decisionsgainspinch_original_tokens/pinch_final_tokens, inconfig/schema.sqland the code-sideensure_route_decisionsmigration, so live databases predating the columns pick them up on next start.metrics.pinch_summaryaggregates 30-day savings: share pruned, total and median tokens saved, estimated dollars saved./metrics,/admin/api/snapshot, the TUI model, and an admin dashboard card. ReturnsNonewhen pinch is disabled, matchinglocal_energy_summary's contract; all three consumers guard for it.debugtoinfo, and only when pruning actually occurred.The dollar figure is priced at the same blended prompt rate
routing.estimated_costuses (src/routing.py:363-374), readingcfg.objective.assumed_cache_rate. An earlier draft of the spec said to price it at the cached rate, which is wrong twice over —cost_per_1m_prompt_cachedis already discounted, so multiplying bycache_rateapplies the discount twice, and it drops the(1 - cache_rate)fraction billed at full price. That would have put a savings number on the dashboard computed on a different basis than the cost model the router ranks on.Accounting fix
prune_contextand_relevance_order_forboth takeextra_fixed_tokens, and the dispatcher passes the tool-definition overhead. opencode sends ~32k prompt tokens of tool definitions on a trivial request, and pinch was excluding all of it from the budget comparison.Threading it through
_relevance_order_foras well is load-bearing: without it, relevance ordering silently declines to run on exactly the requests that newly need pruning — no error, no log, just uniform pruning instead of least-relevant-first.Review finding fixed mid-flight
Persistence covers the routed success, routed rejection, and passthrough paths. On passthrough the persist originally ran 43 lines before the prune, so
kind='passthrough'rows recorded NULL even when pruning occurred. The prune is now hoisted above the persist.The persist deliberately stays above
_check_pinned_capabilities, which raises 422 — so a rejected pin still writes the decision row that explains it. Anyone tidying that order will silently delete rejection records.Guarded by
test_passthrough_records_pinch_columns_when_pruned, which asserts both columns non-NULL andoriginal > final.Verification
pytest-xdist -n auto. Both modes, so xdist is not masking ordering bugs.router.dbbyte-identical before and after the suite (same mtime and size) — no test writes to the live database.config/config.yamluntouched;budget_tokensunchanged.Note on scope
Includes
4cb7a92, a correction to the plan spec inplans/, which is an ancestor of this branch but not yet onorigin/main.Also worth knowing before the next merge:
feat/proficiency-exposure-bias-and-explorationappends columns to the tail of the sameroute_decisionsblock in bothconfig/schema.sqlandensure_route_decisions. Whichever of the two merges second gets a small conflict at theflex_forcedline.🤖 Generated with Claude Code
https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U