> Deep dive into the context-pruning (pinch) feature. Back to [README](../README.md). Pinch is the router's optional **pre-dispatch context-pruning** stage. Ported from the MIT-licensed llmrouter "pinch" module, it shrinks the provider-bound conversation once it grows past a token budget, so a long agent session ships fewer prompt tokens upstream — before any paid token is sent. It is **on by default** (`pinch.enabled: true`), but is deliberately conservative. The core is a pure, injectable module (`src/context_prune.py`): it takes messages and limits and returns pruned messages. `dispatcher.py` owns reading the config and deciding **when** to call it. ## Where pinch sits The critical question for a pruning stage is *when* it runs relative to the decisions that depend on conversation size. Pinch is hoisted **before** those decisions on both dispatch paths. ### Routed path On the routed path (an `auto` decision), pinch runs **before the measured-context decision** — the window/tier/cost selection. The comment in `dispatcher.py` is explicit: prune ONCE before the measured-context decision, "so the window/tier/cost choice sees the size that will actually ship upstream rather than the raw conversation." ```python send_messages, pinch_stats = prune_context( list(messages), budget_tokens=cfg.pinch.budget_tokens, keep_last_turns=cfg.pinch.keep_last_turns, max_summarize_chars=cfg.pinch.max_summarize_chars, relevance_order=_relevance_order_for( messages, cfg, budget_tokens=cfg.pinch.budget_tokens, extra_fixed_tokens=_tools_overhead, ), extra_fixed_tokens=_tools_overhead, ) # ... only then: measured = estimate_prompt_tokens(send_messages, tools=body.get("tools")) ``` The pruned list (`send_messages`) is then **reused** at dispatch time — it is never pruned twice. When pinch is disabled this whole block is a byte-for-byte no-op and `send_messages` is the raw `messages` list, exactly as before. Pinch trims only **tool results**, never user/assistant/system messages, so the classification turn above (which used `_previous_context` from the **full** untouched messages) is undisturbed — the classifier still sees the whole history. ### Passthrough path On the passthrough path (a pinned model id, dispatched as asked), pinch runs **separately** with its own prune block. There the ordering matters differently, and is a hard contract (see [Passthrough ordering constraint](#passthrough-ordering-constraint) below): prune → persist → `_check_pinned_capabilities`. ## What pinch trims Pinch's tailoring is deliberate and safe: - **Only `tool` results** are trimmed or summarized. Older results outside the protected window are the main target, but any **outsized result inside the window** that is longer than `pinch.protected_max_chars` is also capped. That threshold is deliberately a much higher bar than `max_summarize_chars`; recent results are more likely to still matter, so it only catches true outliers. Tool results are where a long agent session's tokens actually live, they are the least likely to still be needed in full by the time a later turn is answered, and replacing one with a short placeholder is reversible at the semantic level — a wrong guess costs context but never breaks the request. - **Never user / assistant / system messages.** Trimming one of those can change what the model is being asked, so they are always kept verbatim. - **Never the classifier input.** The classification turn uses `_previous_context` from the full untouched messages, and pinch clamps head+tail independently — a separate concern from the classifier's input framing. Messages are **never removed** (order and length are preserved): a tool result is either summarized in place (head + tail around a `[N chars trimmed...]` marker) or replaced with a `[name: result omitted]` placeholder, so `len(pruned)` always equals `len(messages)` and the result still parses as a valid conversation. Only runs (and only mutates anything) when the estimate actually exceeds the budget; otherwise the original list is returned untouched. ### Capping outsized results in the protected window The protected window, controlled by `keep_last_turns`, is intentionally generous: anything in the last few user turns stays verbatim because it is most likely to still be needed. But that created a gap. An autonomous tool-call loop (code execution, file reads, test runs) can produce one enormous tool result without a new user message, and that whole result used to count as part of the protected window. Pinch only fired when the total conversation crossed the budget, but once it did, those outsized protected results shipped verbatim regardless of size. In one measured case a 324k-token conversation shrank only ~8% because nearly all of it sat inside the protected window. `protected_max_chars` closes that gap: tool results inside the window are still preserved unless they cross this outlier threshold, in which case they receive the same head/tail elision applied to older candidates. The default is high enough to avoid touching normal output; it only intervenes when a single result threatens the budget. ## Config block All pinch configuration lives under `pinch:` in `config/config.yaml`, validated by `PinchConfig` in `src/config.py`. Config is strict (`extra="forbid"`): unknown keys fail at load. The full block ships as: ```yaml pinch: # Relevance-based context pruning. Optional, pre-dispatch stage: when a # conversation exceeds budget_tokens, the provider-bound messages are trimmed # BEFORE any paid token is sent upstream. enabled: true budget_tokens: 50000 keep_last_turns: 4 max_summarize_chars: 4000 protected_max_chars: 20000 relevance: enabled: true model: "nomic-embed-text" base_url: "http://localhost:11434/v1" timeout_seconds: 10 min_candidates: 2 ``` | Key | Default | Meaning | |---|---|---| | `pinch.enabled` | `true` | Master switch. Off ⇒ the whole prune block is a byte-for-byte no-op. | | `pinch.budget_tokens` | `50000` | Above this estimated token count the conversation is pruned. Must be `> 0`. | | `pinch.keep_last_turns` | `4` | How many recent user turns (plus their assistant replies and tool results) are protected from pruning. Must be `> 0`. | | `pinch.max_summarize_chars` | `4000` | Tool results longer than this many chars are summarized in place; shorter ones are collapsed to a placeholder. Must be `>= 3000` — below that summarization would grow the message. | | `pinch.protected_max_chars` | `20000` | `keep_last_turns` has no size limit inside it — an entire autonomous tool-call loop with no new user message can be one protected turn, so one outsized tool result inside it (a full verbose test run, a huge file read) used to ship verbatim regardless of size. Measured live 2026-09-06: a 324k-token conversation shrank only ~8% because nearly all of it sat inside the protected window. Any tool result inside the protected window over this many chars now gets the same head/tail elision candidates get — deliberately a much higher bar than `max_summarize_chars`, since recent results are more likely to still matter, so it only catches true outliers. Must be `>= 3000`, or `null` to disable. | | `pinch.prefix_probe` | `true` | Record where this turn's pruned payload stopped matching the previous turn's (three integers per decision row; hashes live in process memory for one turn and never reach the database). Pure telemetry: off changes no routing and no payload, it only stops collecting the evidence for whether pruning is destroying the provider's prompt cache. | | `pinch.relevance.enabled` | `true` | Embedding-model relevance scoring. Requires `pinch.enabled` too. Any failure reverts to uniform trimming. | | `pinch.relevance.model` | `nomic-embed-text` | An **embedding** model — never `classifier.model` or `verification.model`. | | `pinch.relevance.base_url` | `http://localhost:11434/v1` | OpenAI-compatible embeddings endpoint on the same local Ollama. | | `pinch.relevance.timeout_seconds` | `10` | Embedding call timeout. Must be `> 0`. | | `pinch.relevance.min_candidates` | `2` | Below this many trim-eligible candidates, skip the embedding round-trip entirely and fall back to uniform compression. Must be `> 0`. | `PinchRelevanceConfig` (`src/config.py`) defaults mirror the YAML: | Key | Default | |---|---| | `enabled` | `True` | | `model` | `"nomic-embed-text"` | | `base_url` | `"http://localhost:11434/v1"` | | `timeout_seconds` | `10` | | `min_candidates` | `2` | `PinchRelevanceConfig` must point at an **embedding** model; its docstring is explicit that pointing it at `classifier.model` or `verification.model` is wrong, because those are chat models and the embeddings call would fail (gracefully degrading to uniform trimming, but always failing). ## Token accounting incl. tool-definition overhead Token estimation in pinch uses a crude characters-per-token heuristic (`CHARS_PER_TOKEN = 3`), consistent with the dispatcher: ```python def estimate_tokens(text): # src/context_prune.py return len(text) // CHARS_PER_TOKEN if text else 0 ``` Tool definitions are a **major omitted factor** and the reason `extra_fixed_tokens` exists. Tool definitions ride along in `upstream_body` on every provider call and count against the same context window, so a trivial opencode request can carry ~**32k prompt tokens of tool definitions alone**. Before commit `d7f051b`, pinch excluded all of that from the budget comparison — every prune run undercounted the real conversation, so the budget guard triggered far too late (or not at all) on exactly the tool-heavy requests that newly needed pruning. ### Why `_relevance_order_for` takes `extra_fixed_tokens` `dispatcher.py` computes the tool-definition overhead once per path: ```python _tools_overhead = len(json.dumps(body.get("tools") or [])) // CHARS_PER_TOKEN ``` and threads it into both `prune_context` and `_relevance_order_for` as `extra_fixed_tokens`. Inside `prune_context`: ```python orig_tokens = sum(estimate_tokens(extract_text(m)) for m in messages) + extra_fixed_tokens if orig_tokens <= budget_tokens: return messages, {"pruned": False, ...} ``` So the **budget guard** compares `orig_tokens + extra_fixed_tokens` against the budget — the same total `_relevance_order_for` computes for its skip check. `_relevance_order_for` takes `extra_fixed_tokens` for the exact same reason. Its skip check is: ```python orig_tokens = sum(estimate_tokens(extract_text(m)) for m in messages) + extra_fixed_tokens if orig_tokens <= budget_tokens: return None ``` Without `extra_fixed_tokens`, relevance ordering silently declines to run **on exactly the requests that newly need pruning**: those that exceed the budget only because of tool-definition overhead. The failure mode is silent — no error, no log — it just falls to uniform ranking instead of least-relevant-first. The `_relevance_order_for` docstring states the contract: "`extra_fixed_tokens` accounts for overhead that `prune_context`'s budget guard also adds to `orig_tokens`, so the skip-budget check in here matches the same total." ### Where the measured-context estimate fits `estimate_prompt_tokens(send_messages, tools=body.get("tools"))` adds `len(json.dumps(tools))` to the char count for the routed path's window/tier/cost decision. So the tool definitions are counted **twice conceptually** — once in pinch's budget guard (`_tools_overhead`), once in the measured-context estimate (`tools=body.get("tools")`) — but they represent the same real overhead. Pinch's `original_tokens`/`final_tokens` stats deliberately include `extra_fixed_tokens` so the recorded orig matches what actually ships, and the two numbers stay comparable. ## How relevance ordering works When `pinch.relevance.enabled` and `pinch.enabled` are both true, and the conversation is actually over budget, `_relevance_order_for` makes **one batched** embeddings call (`query` + all candidates in a single request) against the Ollama endpoint, then `order_by_relevance` scores cosine similarity in **ascending order** — least relevant first. `prune_context` walks that order, compressing least-relevant candidates first until the token deficit is covered; the more-relevant (protected) candidates stay verbatim. When `relevance_order` is `None` (relevance off, too few candidates, over-budget check failed, or the embedding call failed), every candidate is compressed uniformly — byte-for-byte the historical behavior. Any failure of the embedding round-trip — `RequestException`, timeout, non-200, unparseable body, missing/malformed vectors — logs `relevance_unavailable` and returns `None`, which the pure core already treats as "compress everything". This must never be the thing that breaks a request. ## Instrumentation surface Pinch records into `route_decisions` and surfaces through `/metrics`, the admin snapshot, the admin dashboard, and the TUI. ### Schema columns Two INTEGER columns on `route_decisions`: | Column | Notes | |---|---| | `pinch_original_tokens` | Estimated tokens before pruning (includes `extra_fixed_tokens`). `NULL` when pinch is disabled or the request predates the columns. | | `pinch_final_tokens` | Estimated tokens after pruning. `NULL` when pinch is disabled or the request predates the columns. | Both columns are created in `config/schema.sql` and mirrored code-side in `ensure_route_decisions` for live DBs that predate the feature. ### `prune_context` stats keys `prune_context` returns `(pruned_messages, stats)`; on the not-pruned path: ``` pruned, original_tokens, final_tokens, tokens_saved ``` and on the pruned path: ``` pruned, original_tokens, final_tokens, tokens_saved, summarized ``` | Key | Meaning | |---|---| | `pruned` | Bool: whether anything was actually trimmed. | | `original_tokens` | Budget-guard total: messages + `extra_fixed_tokens`. | | `final_tokens` | Estimated tokens of the returned pruned list. | | `tokens_saved` | `max(original_tokens - final_tokens, 0)` — computed from what the payload actually shrunk by (whole list before vs after), never a per-message estimate. | | `summarized` | Count of tool results actually summarized/placeholder-replaced. | ### `metrics.pinch_summary(conn, cfg)` `pinch_summary` returns `None` when pinch is disabled (so callers omit the section), otherwise a 30-day aggregate: | Key | Meaning | |---|---| | `calls_30d` | Route decisions in the last 30 days. | | `pruned_calls_30d` | Decisions where `pinch_original_tokens > pinch_final_tokens`. | | `share_pruned` | `pruned_calls_30d / calls_30d`, **rounded to 4 decimals**. | | `total_tokens_saved` | Sum of `original - final` over the window. | | `median_tokens_saved` | 50th percentile of per-decision tokens saved (over pruned rows). | | `dollars_saved_usd_30d` | Token savings priced at the blended prompt rate, **rounded to 6 decimals**. | | `reset_date` / `next_reset_date` | The 30-day window start and the upcoming billing-period reset day (only when `objective.billing_reset_day` is set). | ### Where it surfaces - **`GET /metrics`** — the response includes a `"pinch"` block populated by `pinch_summary(conn, cfg)`. - **Admin snapshot** — `admin.py` embeds the same `"pinch": metrics.pinch_summary(conn, cfg)` in `/admin/api/snapshot`. It also exposes `pinch.enabled`, `pinch.prefix_probe` and `pinch.relevance.enabled` as admin runtime knobs (`pinch_enabled`, `pinch_prefix_probe`, `pinch_relevance_enabled`) and as persisted-config keys. The controls page flags `pinch.prefix_probe` in both panels, because it is the one knob there whose only effect is on measurement: turned off, routing is unchanged and the cache-loss evidence simply stops accumulating. - **Admin dashboard** — `index.html` renders a "Pinch savings" card showing **share pruned** (as a percent), **tokens saved**, **median saved**, and **30d dollars saved** (`$` with 6 decimals). When pinch is disabled the card shows "Pinch is disabled". - **TUI** — the dashboard has pinch rows (`share_pruned`, `total_tokens_saved`, `median_tokens_saved`, `dollars_saved_usd_30d`) pulled from the snapshot. ## Dollars-saved pricing The `dollars_saved_usd_30d` figure prices token savings at the **blended prompt rate** the router itself ranks on, from `routing.estimated_cost`. It reads `cfg.objective.assumed_cache_rate` and, per saved row, computes: ```python cached_price = cost_per_1m_prompt_cached or prompt_price blended = (1.0 - cache_rate) * prompt_price + cache_rate * cached_price dollars_saved += saved * blended / 1_000_000 ``` That mirrors `estimated_cost` in `src/routing.py`, so the dashboard's dollar figure agrees with the cost model the router uses to rank candidates. ### The review-caught error An earlier version priced savings **at the cached rate alone**. That was wrong in two compounding ways: 1. It **double-applied the discount** — the cached price already reflects the cache discount, so using it as the sole rate applied that discount once for the cached fraction and again across the whole amount. 2. It **dropped the fraction billed at full price** — the portion of traffic assumed to miss the cache (the `(1.0 - cache_rate)` share) was silently billed at the cheaper cached rate too, so savings were overstated. The fix prices each token at the true blend: the uncached share at full `prompt_price`, the cached share at `cached_price`, with `assumed_cache_rate` splitting the two. Only that blended rate reproduces what the router's `estimated_cost` computes, so the dashboard figure stays consistent with the cost model the router ranks on. ## Passthrough ordering constraint On the passthrough path the ordering of three steps is a **contract** — the dispatcher must NOT reorder them: 1. **prune** — `prune_context` runs first. 2. **persist** — `persist_route_decision` records the passthrough with the computed `pinch_original_tokens` / `pinch_final_tokens`. 3. **`_check_pinned_capabilities`** — raises 422 when the pinned model cannot satisfy the request. Why each constraint holds: - **Prune is hoisted above persist** so a pruned passthrough request records real pinch tokens rather than NULL. If persist ran before prune, the `pinch_stats` (computed only by the prune block) would not exist yet and the row would be written with `NULL`s, losing the measurement for the very traffic where pruning matters. - **Persist stays above `_check_pinned_capabilities`** because the check can **raise** — and a rejected passthrough must not create a partial decision row. `_check_pinned_capabilities` raises 422 for a pin that cannot possibly work (e.g. an image or JSON-mode request pinned to a model without the capability). If persist ran after the check, a failure would leave nothing behind; if persist ran after and the check raised, the raise would leave the persist unexecuted — but the constraint as written ensures that by the time the capability check can raise, no row has been written for a request that is about to be rejected. ```python if cfg.pinch.enabled: send_messages, pinch_stats = prune_context(...) persist_route_decision( "passthrough", ..., pinch_original_tokens=pinch_stats.get("original_tokens") if pinch_stats is not None else None, pinch_final_tokens=pinch_stats.get("final_tokens") if pinch_stats is not None else None, ) if caps.has_images or caps.require_json_mode: _check_pinned_capabilities(requested, caps) ``` `prune → persist → _check_pinned_capabilities` is the order the code enforces, and a future reorder would silently corrupt either the measurements or the decision-log integrity. ## Operational notes - Pinch is on by default but conservative: start with a large `budget_tokens` and watch `route_decisions` / pinch stats on real traffic before widening it. - Tool results can only be dropped safely because tool outputs are idempotent enough for a placeholder; a wrong guess here loses context, so the default leans small-reduction. - The relevance path requires an embedding model pulled on the same Ollama (`ollama pull nomic-embed-text`) and reachable at `pinch.relevance.base_url`. Like every other new-and-unproven knob, `relevance.enabled` ships guarded by `pinch.enabled` — it has no effect otherwise.