Files
6krrt/tests/test_iteration.py
adlee-was-taken c0b3175730 feat(iteration): budget retries in prompt re-bills; prefer cache-preserving same-model retry on malformed
Adds prompt-size awareness to the retry budget (item 4.1 of
plans/token-waste-waves.md). Two coupled changes:

(A) Malformed same-model-first (cache-preserving)
    plan_retry() now returns a same-model retry for malformed verdicts
    (with same_model=True), keeping the provider prompt cache hot
    (~0.92 hit rate vs ~0.35 on switch). Only when retried_same_model
    is passed as True (second consecutive malformed) does it escalate
    to runners_up[0]. The dispatcher loop tracks retried_same_model
    and only drops alternatives[1:] on escalation, not on same-model
    retry.

(B) Prompt-size gate on escalation
    New parameter max_rebill_prompt_tokens: when prompt_tokens exceeds
    this threshold, escalation (cache-destroying cold re-bill) is
    suppressed. Same-model retries are always allowed because they
    preserve the cache. 0 = unlimited (backward-compatible default).

Config surface:
  - iteration.max_rebill_prompt_tokens (int, default 0) in config.yaml
    and IterationConfig Pydantic model.

Dispatcher: threads decision.classification.required_context_tokens
as prompt_tokens and cfg.iteration.max_rebill_prompt_tokens as the
threshold into plan_retry(); tracks retried_same_model in the retry
loop.

Tests: 13 existing + 12 new covering:
  - same-model malformed retry preferred over escalation
  - same-model malformed works with no runners_up
  - escalation after same_model exhausted
  - escalation blocked/gated on prompt-size threshold (including
    boundary and one-over)
  - truncation escalation also gated (no change to same-model
    truncation)
  - zero threshold = unlimited (backward compat)
2026-09-18 00:32:55 -04:00

8.8 KiB