Files
6krrt/docs/pinch.md
adlee-was-taken 2a658e7294 feat(admin): pinch.prefix_probe gets a control, and a warning beside it
The knob was deliberately left off _BOOL_KNOBS when the probe landed in
fb511cd, on the grounds that a toggle whose only effect is to stop collecting
evidence is questionable UX. That is a fair reading and it loses to the
standing rule: a knob reachable only by hand-editing config.local.yaml is
invisible to whoever is actually operating the router, which is the same gap
classifier.cloud_fallback sat in while the portal's own gaming-mode text told
operators they needed it.

So it is added exactly where its two siblings already are -- _BOOL_KNOBS for
the in-memory flip, _CONFIG_ALLOWLIST plus _CONFIG_GET_ORDER for the persisted
write -- and the objection is answered in the UI instead of by omission. One
short line in the meta track of both panels, "off stops the cache-loss
measurement", with the rest of the reasoning in the title attribute and in
docs/pinch.md where it belongs. No new visual pattern: it is the same
.badge.bg-warning the zero-admit profile hint already uses, rendered into the
same track as the "file: <value>" and provenance badges.

The round-trip test is the confidence_threshold lesson applied rather than
relearned. An admin control that writes a value the config loader will later
reject or silently misread is worse than no control -- that one wrote a raw 80
meaning 80% into the overlay and was caught before the restart by luck -- so
the test reads true, writes false, reads back false, and then loads the merged
base + overlay through the real RouterConfig, asserting a genuine False rather
than the string "false" and that nothing else in the pinch block moved. It
runs entirely against a tmp_path copy; the machine-local overlay is never
opened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-13 12:05:39 -04:00

20 KiB

Deep dive into the context-pruning (pinch) feature. Back to README.

Pinch is the router's optional pre-dispatch context-pruning stage. Ported from the MIT-licensed llmrouter "pinch" module, it shrinks the provider-bound conversation once it grows past a token budget, so a long agent session ships fewer prompt tokens upstream — before any paid token is sent. It is on by default (pinch.enabled: true), but is deliberately conservative.

The core is a pure, injectable module (src/context_prune.py): it takes messages and limits and returns pruned messages. dispatcher.py owns reading the config and deciding when to call it.

Where pinch sits

The critical question for a pruning stage is when it runs relative to the decisions that depend on conversation size. Pinch is hoisted before those decisions on both dispatch paths.

Routed path

On the routed path (an auto decision), pinch runs before the measured-context decision — the window/tier/cost selection. The comment in dispatcher.py is explicit: prune ONCE before the measured-context decision, "so the window/tier/cost choice sees the size that will actually ship upstream rather than the raw conversation."

send_messages, pinch_stats = prune_context(
    list(messages),
    budget_tokens=cfg.pinch.budget_tokens,
    keep_last_turns=cfg.pinch.keep_last_turns,
    max_summarize_chars=cfg.pinch.max_summarize_chars,
    relevance_order=_relevance_order_for(
        messages, cfg, budget_tokens=cfg.pinch.budget_tokens,
        extra_fixed_tokens=_tools_overhead,
    ),
    extra_fixed_tokens=_tools_overhead,
)
# ... only then:
measured = estimate_prompt_tokens(send_messages, tools=body.get("tools"))

The pruned list (send_messages) is then reused at dispatch time — it is never pruned twice. When pinch is disabled this whole block is a byte-for-byte no-op and send_messages is the raw messages list, exactly as before.

Pinch trims only tool results, never user/assistant/system messages, so the classification turn above (which used _previous_context from the full untouched messages) is undisturbed — the classifier still sees the whole history.

Passthrough path

On the passthrough path (a pinned model id, dispatched as asked), pinch runs separately with its own prune block. There the ordering matters differently, and is a hard contract (see Passthrough ordering constraint below): prune → persist → _check_pinned_capabilities.

What pinch trims

Pinch's tailoring is deliberate and safe:

  • Only tool results are trimmed or summarized. Older results outside the protected window are the main target, but any outsized result inside the window that is longer than pinch.protected_max_chars is also capped. That threshold is deliberately a much higher bar than max_summarize_chars; recent results are more likely to still matter, so it only catches true outliers. Tool results are where a long agent session's tokens actually live, they are the least likely to still be needed in full by the time a later turn is answered, and replacing one with a short placeholder is reversible at the semantic level — a wrong guess costs context but never breaks the request.
  • Never user / assistant / system messages. Trimming one of those can change what the model is being asked, so they are always kept verbatim.
  • Never the classifier input. The classification turn uses _previous_context from the full untouched messages, and pinch clamps head+tail independently — a separate concern from the classifier's input framing.

Messages are never removed (order and length are preserved): a tool result is either summarized in place (head + tail around a [N chars trimmed...] marker) or replaced with a [name: result omitted] placeholder, so len(pruned) always equals len(messages) and the result still parses as a valid conversation.

Only runs (and only mutates anything) when the estimate actually exceeds the budget; otherwise the original list is returned untouched.

Capping outsized results in the protected window

The protected window, controlled by keep_last_turns, is intentionally generous: anything in the last few user turns stays verbatim because it is most likely to still be needed. But that created a gap. An autonomous tool-call loop (code execution, file reads, test runs) can produce one enormous tool result without a new user message, and that whole result used to count as part of the protected window. Pinch only fired when the total conversation crossed the budget, but once it did, those outsized protected results shipped verbatim regardless of size. In one measured case a 324k-token conversation shrank only ~8% because nearly all of it sat inside the protected window. protected_max_chars closes that gap: tool results inside the window are still preserved unless they cross this outlier threshold, in which case they receive the same head/tail elision applied to older candidates. The default is high enough to avoid touching normal output; it only intervenes when a single result threatens the budget.

Config block

All pinch configuration lives under pinch: in config/config.yaml, validated by PinchConfig in src/config.py. Config is strict (extra="forbid"): unknown keys fail at load.

The full block ships as:

pinch:
  # Relevance-based context pruning. Optional, pre-dispatch stage: when a
  # conversation exceeds budget_tokens, the provider-bound messages are trimmed
  # BEFORE any paid token is sent upstream.
  enabled: true
  budget_tokens: 50000
  keep_last_turns: 4
  max_summarize_chars: 4000
  protected_max_chars: 20000
  relevance:
    enabled: true
    model: "nomic-embed-text"
    base_url: "http://localhost:11434/v1"
    timeout_seconds: 10
    min_candidates: 2
Key Default Meaning
pinch.enabled true Master switch. Off ⇒ the whole prune block is a byte-for-byte no-op.
pinch.budget_tokens 50000 Above this estimated token count the conversation is pruned. Must be > 0.
pinch.keep_last_turns 4 How many recent user turns (plus their assistant replies and tool results) are protected from pruning. Must be > 0.
pinch.max_summarize_chars 4000 Tool results longer than this many chars are summarized in place; shorter ones are collapsed to a placeholder. Must be >= 3000 — below that summarization would grow the message.
pinch.protected_max_chars 20000 keep_last_turns has no size limit inside it — an entire autonomous tool-call loop with no new user message can be one protected turn, so one outsized tool result inside it (a full verbose test run, a huge file read) used to ship verbatim regardless of size. Measured live 2026-09-06: a 324k-token conversation shrank only ~8% because nearly all of it sat inside the protected window. Any tool result inside the protected window over this many chars now gets the same head/tail elision candidates get — deliberately a much higher bar than max_summarize_chars, since recent results are more likely to still matter, so it only catches true outliers. Must be >= 3000, or null to disable.
pinch.prefix_probe true Record where this turn's pruned payload stopped matching the previous turn's (three integers per decision row; hashes live in process memory for one turn and never reach the database). Pure telemetry: off changes no routing and no payload, it only stops collecting the evidence for whether pruning is destroying the provider's prompt cache.
pinch.relevance.enabled true Embedding-model relevance scoring. Requires pinch.enabled too. Any failure reverts to uniform trimming.
pinch.relevance.model nomic-embed-text An embedding model — never classifier.model or verification.model.
pinch.relevance.base_url http://localhost:11434/v1 OpenAI-compatible embeddings endpoint on the same local Ollama.
pinch.relevance.timeout_seconds 10 Embedding call timeout. Must be > 0.
pinch.relevance.min_candidates 2 Below this many trim-eligible candidates, skip the embedding round-trip entirely and fall back to uniform compression. Must be > 0.

PinchRelevanceConfig (src/config.py) defaults mirror the YAML:

Key Default
enabled True
model "nomic-embed-text"
base_url "http://localhost:11434/v1"
timeout_seconds 10
min_candidates 2

PinchRelevanceConfig must point at an embedding model; its docstring is explicit that pointing it at classifier.model or verification.model is wrong, because those are chat models and the embeddings call would fail (gracefully degrading to uniform trimming, but always failing).

Token accounting incl. tool-definition overhead

Token estimation in pinch uses a crude characters-per-token heuristic (CHARS_PER_TOKEN = 3), consistent with the dispatcher:

def estimate_tokens(text):           # src/context_prune.py
    return len(text) // CHARS_PER_TOKEN if text else 0

Tool definitions are a major omitted factor and the reason extra_fixed_tokens exists. Tool definitions ride along in upstream_body on every provider call and count against the same context window, so a trivial opencode request can carry ~32k prompt tokens of tool definitions alone. Before commit d7f051b, pinch excluded all of that from the budget comparison — every prune run undercounted the real conversation, so the budget guard triggered far too late (or not at all) on exactly the tool-heavy requests that newly needed pruning.

Why _relevance_order_for takes extra_fixed_tokens

dispatcher.py computes the tool-definition overhead once per path:

_tools_overhead = len(json.dumps(body.get("tools") or [])) // CHARS_PER_TOKEN

and threads it into both prune_context and _relevance_order_for as extra_fixed_tokens. Inside prune_context:

orig_tokens = sum(estimate_tokens(extract_text(m)) for m in messages) + extra_fixed_tokens
if orig_tokens <= budget_tokens:
    return messages, {"pruned": False, ...}

So the budget guard compares orig_tokens + extra_fixed_tokens against the budget — the same total _relevance_order_for computes for its skip check.

_relevance_order_for takes extra_fixed_tokens for the exact same reason. Its skip check is:

orig_tokens = sum(estimate_tokens(extract_text(m)) for m in messages) + extra_fixed_tokens
if orig_tokens <= budget_tokens:
    return None

Without extra_fixed_tokens, relevance ordering silently declines to run on exactly the requests that newly need pruning: those that exceed the budget only because of tool-definition overhead. The failure mode is silent — no error, no log — it just falls to uniform ranking instead of least-relevant-first. The _relevance_order_for docstring states the contract: "extra_fixed_tokens accounts for overhead that prune_context's budget guard also adds to orig_tokens, so the skip-budget check in here matches the same total."

Where the measured-context estimate fits

estimate_prompt_tokens(send_messages, tools=body.get("tools")) adds len(json.dumps(tools)) to the char count for the routed path's window/tier/cost decision. So the tool definitions are counted twice conceptually — once in pinch's budget guard (_tools_overhead), once in the measured-context estimate (tools=body.get("tools")) — but they represent the same real overhead. Pinch's original_tokens/final_tokens stats deliberately include extra_fixed_tokens so the recorded orig matches what actually ships, and the two numbers stay comparable.

How relevance ordering works

When pinch.relevance.enabled and pinch.enabled are both true, and the conversation is actually over budget, _relevance_order_for makes one batched embeddings call (query + all candidates in a single request) against the Ollama endpoint, then order_by_relevance scores cosine similarity in ascending order — least relevant first. prune_context walks that order, compressing least-relevant candidates first until the token deficit is covered; the more-relevant (protected) candidates stay verbatim. When relevance_order is None (relevance off, too few candidates, over-budget check failed, or the embedding call failed), every candidate is compressed uniformly — byte-for-byte the historical behavior.

Any failure of the embedding round-trip — RequestException, timeout, non-200, unparseable body, missing/malformed vectors — logs relevance_unavailable and returns None, which the pure core already treats as "compress everything". This must never be the thing that breaks a request.

Instrumentation surface

Pinch records into route_decisions and surfaces through /metrics, the admin snapshot, the admin dashboard, and the TUI.

Schema columns

Two INTEGER columns on route_decisions:

Column Notes
pinch_original_tokens Estimated tokens before pruning (includes extra_fixed_tokens). NULL when pinch is disabled or the request predates the columns.
pinch_final_tokens Estimated tokens after pruning. NULL when pinch is disabled or the request predates the columns.

Both columns are created in config/schema.sql and mirrored code-side in ensure_route_decisions for live DBs that predate the feature.

prune_context stats keys

prune_context returns (pruned_messages, stats); on the not-pruned path:

pruned, original_tokens, final_tokens, tokens_saved

and on the pruned path:

pruned, original_tokens, final_tokens, tokens_saved, summarized
Key Meaning
pruned Bool: whether anything was actually trimmed.
original_tokens Budget-guard total: messages + extra_fixed_tokens.
final_tokens Estimated tokens of the returned pruned list.
tokens_saved max(original_tokens - final_tokens, 0) — computed from what the payload actually shrunk by (whole list before vs after), never a per-message estimate.
summarized Count of tool results actually summarized/placeholder-replaced.

metrics.pinch_summary(conn, cfg)

pinch_summary returns None when pinch is disabled (so callers omit the section), otherwise a 30-day aggregate:

Key Meaning
calls_30d Route decisions in the last 30 days.
pruned_calls_30d Decisions where pinch_original_tokens > pinch_final_tokens.
share_pruned pruned_calls_30d / calls_30d, rounded to 4 decimals.
total_tokens_saved Sum of original - final over the window.
median_tokens_saved 50th percentile of per-decision tokens saved (over pruned rows).
dollars_saved_usd_30d Token savings priced at the blended prompt rate, rounded to 6 decimals.
reset_date / next_reset_date The 30-day window start and the upcoming billing-period reset day (only when objective.billing_reset_day is set).

Where it surfaces

  • GET /metrics — the response includes a "pinch" block populated by pinch_summary(conn, cfg).
  • Admin snapshot — admin.py embeds the same "pinch": metrics.pinch_summary(conn, cfg) in /admin/api/snapshot. It also exposes pinch.enabled, pinch.prefix_probe and pinch.relevance.enabled as admin runtime knobs (pinch_enabled, pinch_prefix_probe, pinch_relevance_enabled) and as persisted-config keys. The controls page flags pinch.prefix_probe in both panels, because it is the one knob there whose only effect is on measurement: turned off, routing is unchanged and the cache-loss evidence simply stops accumulating.
  • Admin dashboard — index.html renders a "Pinch savings" card showing share pruned (as a percent), tokens saved, median saved, and 30d dollars saved ($ with 6 decimals). When pinch is disabled the card shows "Pinch is disabled".
  • TUI — the dashboard has pinch rows (share_pruned, total_tokens_saved, median_tokens_saved, dollars_saved_usd_30d) pulled from the snapshot.

Dollars-saved pricing

The dollars_saved_usd_30d figure prices token savings at the blended prompt rate the router itself ranks on, from routing.estimated_cost. It reads cfg.objective.assumed_cache_rate and, per saved row, computes:

cached_price = cost_per_1m_prompt_cached or prompt_price
blended = (1.0 - cache_rate) * prompt_price + cache_rate * cached_price
dollars_saved += saved * blended / 1_000_000

That mirrors estimated_cost in src/routing.py, so the dashboard's dollar figure agrees with the cost model the router uses to rank candidates.

The review-caught error

An earlier version priced savings at the cached rate alone. That was wrong in two compounding ways:

  1. It double-applied the discount — the cached price already reflects the cache discount, so using it as the sole rate applied that discount once for the cached fraction and again across the whole amount.
  2. It dropped the fraction billed at full price — the portion of traffic assumed to miss the cache (the (1.0 - cache_rate) share) was silently billed at the cheaper cached rate too, so savings were overstated.

The fix prices each token at the true blend: the uncached share at full prompt_price, the cached share at cached_price, with assumed_cache_rate splitting the two. Only that blended rate reproduces what the router's estimated_cost computes, so the dashboard figure stays consistent with the cost model the router ranks on.

Passthrough ordering constraint

On the passthrough path the ordering of three steps is a contract — the dispatcher must NOT reorder them:

  1. prune — prune_context runs first.
  2. persist — persist_route_decision records the passthrough with the computed pinch_original_tokens / pinch_final_tokens.
  3. _check_pinned_capabilities — raises 422 when the pinned model cannot satisfy the request.

Why each constraint holds:

  • Prune is hoisted above persist so a pruned passthrough request records real pinch tokens rather than NULL. If persist ran before prune, the pinch_stats (computed only by the prune block) would not exist yet and the row would be written with NULLs, losing the measurement for the very traffic where pruning matters.
  • Persist stays above _check_pinned_capabilities because the check can raise — and a rejected passthrough must not create a partial decision row. _check_pinned_capabilities raises 422 for a pin that cannot possibly work (e.g. an image or JSON-mode request pinned to a model without the capability). If persist ran after the check, a failure would leave nothing behind; if persist ran after and the check raised, the raise would leave the persist unexecuted — but the constraint as written ensures that by the time the capability check can raise, no row has been written for a request that is about to be rejected.
if cfg.pinch.enabled:
    send_messages, pinch_stats = prune_context(...)

persist_route_decision(
    "passthrough", ..., 
    pinch_original_tokens=pinch_stats.get("original_tokens") if pinch_stats is not None else None,
    pinch_final_tokens=pinch_stats.get("final_tokens") if pinch_stats is not None else None,
)

if caps.has_images or caps.require_json_mode:
    _check_pinned_capabilities(requested, caps)

prune → persist → _check_pinned_capabilities is the order the code enforces, and a future reorder would silently corrupt either the measurements or the decision-log integrity.

Operational notes

  • Pinch is on by default but conservative: start with a large budget_tokens and watch route_decisions / pinch stats on real traffic before widening it.
  • Tool results can only be dropped safely because tool outputs are idempotent enough for a placeholder; a wrong guess here loses context, so the default leans small-reduction.
  • The relevance path requires an embedding model pulled on the same Ollama (ollama pull nomic-embed-text) and reachable at pinch.relevance.base_url. Like every other new-and-unproven knob, relevance.enabled ships guarded by pinch.enabled — it has no effect otherwise.