Adds prompt-size awareness to the retry budget (item 4.1 of
plans/token-waste-waves.md). Two coupled changes:
(A) Malformed same-model-first (cache-preserving)
plan_retry() now returns a same-model retry for malformed verdicts
(with same_model=True), keeping the provider prompt cache hot
(~0.92 hit rate vs ~0.35 on switch). Only when retried_same_model
is passed as True (second consecutive malformed) does it escalate
to runners_up[0]. The dispatcher loop tracks retried_same_model
and only drops alternatives[1:] on escalation, not on same-model
retry.
(B) Prompt-size gate on escalation
New parameter max_rebill_prompt_tokens: when prompt_tokens exceeds
this threshold, escalation (cache-destroying cold re-bill) is
suppressed. Same-model retries are always allowed because they
preserve the cache. 0 = unlimited (backward-compatible default).
Config surface:
- iteration.max_rebill_prompt_tokens (int, default 0) in config.yaml
and IterationConfig Pydantic model.
Dispatcher: threads decision.classification.required_context_tokens
as prompt_tokens and cfg.iteration.max_rebill_prompt_tokens as the
threshold into plan_retry(); tracks retried_same_model in the retry
loop.
Tests: 13 existing + 12 new covering:
- same-model malformed retry preferred over escalation
- same-model malformed works with no runners_up
- escalation after same_model exhausted
- escalation blocked/gated on prompt-size threshold (including
boundary and one-over)
- truncation escalation also gated (no change to same-model
truncation)
- zero threshold = unlimited (backward compat)
8.8 KiB
8.8 KiB