fix(pinch): cap outsized tool results even inside the protected window #44

Merged
alee merged 1 commits from fix/pinch-protected-window-size-cap into main 2026-09-06 23:11:09 +00:00

1 Commits

Author SHA1 Message Date
adlee-was-taken
06193b9144 fix(pinch): cap outsized tool results even inside the protected window
keep_last_turns has no size limit inside it: an entire autonomous
tool-call loop with no new user message counts as one protected turn,
so one outsized tool result inside it -- a full verbose test run, a
huge file read -- shipped verbatim regardless of size.

Measured live 2026-09-06: a 324k-token conversation only shrank ~8%
because nearly all of it sat inside the protected window, and even
after pruning was still ~6x pinch.budget_tokens. This matters beyond
raw token cost too -- the pruned size is what feeds
required_context_tokens (dispatcher.py measures it post-prune before
tier/candidate selection), so a poorly-pruned conversation can also
keep a request above a smaller, cheaper model's context ceiling that
a properly-pruned one would have dropped below.

New pinch.protected_max_chars (default 20000, null to disable): any
tool result inside the protected window over this many characters
still gets the same head/tail elision candidates already get.
Deliberately a much higher bar than max_summarize_chars (4000) --
recent results are more likely to still matter -- so this only
catches true outliers, never ordinary recent tool output. The pure
function's own default stays None, so every existing caller/test is
unaffected unless it opts in; PinchConfig supplies the real default
so production gets the fix without a signature change elsewhere.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-06 19:10:24 -04:00