Files
6krrt/docs/pinch.md
adlee-was-taken 2a658e7294 feat(admin): pinch.prefix_probe gets a control, and a warning beside it
The knob was deliberately left off _BOOL_KNOBS when the probe landed in
fb511cd, on the grounds that a toggle whose only effect is to stop collecting
evidence is questionable UX. That is a fair reading and it loses to the
standing rule: a knob reachable only by hand-editing config.local.yaml is
invisible to whoever is actually operating the router, which is the same gap
classifier.cloud_fallback sat in while the portal's own gaming-mode text told
operators they needed it.

So it is added exactly where its two siblings already are -- _BOOL_KNOBS for
the in-memory flip, _CONFIG_ALLOWLIST plus _CONFIG_GET_ORDER for the persisted
write -- and the objection is answered in the UI instead of by omission. One
short line in the meta track of both panels, "off stops the cache-loss
measurement", with the rest of the reasoning in the title attribute and in
docs/pinch.md where it belongs. No new visual pattern: it is the same
.badge.bg-warning the zero-admit profile hint already uses, rendered into the
same track as the "file: <value>" and provenance badges.

The round-trip test is the confidence_threshold lesson applied rather than
relearned. An admin control that writes a value the config loader will later
reject or silently misread is worse than no control -- that one wrote a raw 80
meaning 80% into the overlay and was caught before the restart by luck -- so
the test reads true, writes false, reads back false, and then loads the merged
base + overlay through the real RouterConfig, asserting a genuine False rather
than the string "false" and that nothing else in the pinch block moved. It
runs entirely against a tmp_path copy; the machine-local overlay is never
opened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-13 12:05:39 -04:00

408 lines
20 KiB
Markdown

> Deep dive into the context-pruning (pinch) feature. Back to [README](../README.md).
Pinch is the router's optional **pre-dispatch context-pruning** stage. Ported
from the MIT-licensed llmrouter "pinch" module, it shrinks the provider-bound
conversation once it grows past a token budget, so a long agent session ships
fewer prompt tokens upstream — before any paid token is sent. It is **on by
default** (`pinch.enabled: true`), but is deliberately conservative.
The core is a pure, injectable module (`src/context_prune.py`): it takes
messages and limits and returns pruned messages. `dispatcher.py` owns reading
the config and deciding **when** to call it.
## Where pinch sits
The critical question for a pruning stage is *when* it runs relative to the
decisions that depend on conversation size. Pinch is hoisted **before** those
decisions on both dispatch paths.
### Routed path
On the routed path (an `auto` decision), pinch runs **before the
measured-context decision** — the window/tier/cost selection. The comment in
`dispatcher.py` is explicit: prune ONCE before the measured-context decision,
"so the window/tier/cost choice sees the size that will actually ship upstream
rather than the raw conversation."
```python
send_messages, pinch_stats = prune_context(
list(messages),
budget_tokens=cfg.pinch.budget_tokens,
keep_last_turns=cfg.pinch.keep_last_turns,
max_summarize_chars=cfg.pinch.max_summarize_chars,
relevance_order=_relevance_order_for(
messages, cfg, budget_tokens=cfg.pinch.budget_tokens,
extra_fixed_tokens=_tools_overhead,
),
extra_fixed_tokens=_tools_overhead,
)
# ... only then:
measured = estimate_prompt_tokens(send_messages, tools=body.get("tools"))
```
The pruned list (`send_messages`) is then **reused** at dispatch time — it is
never pruned twice. When pinch is disabled this whole block is a byte-for-byte
no-op and `send_messages` is the raw `messages` list, exactly as before.
Pinch trims only **tool results**, never user/assistant/system messages, so the
classification turn above (which used `_previous_context` from the **full**
untouched messages) is undisturbed — the classifier still sees the whole
history.
### Passthrough path
On the passthrough path (a pinned model id, dispatched as asked), pinch runs
**separately** with its own prune block. There the ordering matters differently,
and is a hard contract (see [Passthrough ordering constraint](#passthrough-ordering-constraint)
below): prune → persist → `_check_pinned_capabilities`.
## What pinch trims
Pinch's tailoring is deliberate and safe:
- **Only `tool` results** are trimmed or summarized. Older results outside the
protected window are the main target, but any **outsized result inside the
window** that is longer than `pinch.protected_max_chars` is also capped. That
threshold is deliberately a much higher bar than `max_summarize_chars`; recent
results are more likely to still matter, so it only catches true outliers.
Tool results are where a long agent session's tokens actually live, they are
the least likely to still be needed in full by the time a later turn is
answered, and replacing one with a short placeholder is reversible at the
semantic level — a wrong guess costs context but never breaks the request.
- **Never user / assistant / system messages.** Trimming one of those can
change what the model is being asked, so they are always kept verbatim.
- **Never the classifier input.** The classification turn uses
`_previous_context` from the full untouched messages, and pinch clamps
head+tail independently — a separate concern from the classifier's input
framing.
Messages are **never removed** (order and length are preserved): a tool result
is either summarized in place (head + tail around a `[N chars trimmed...]`
marker) or replaced with a `[name: result omitted]` placeholder, so
`len(pruned)` always equals `len(messages)` and the result still parses as a
valid conversation.
Only runs (and only mutates anything) when the estimate actually exceeds the
budget; otherwise the original list is returned untouched.
### Capping outsized results in the protected window
The protected window, controlled by `keep_last_turns`, is intentionally generous:
anything in the last few user turns stays verbatim because it is most likely to
still be needed. But that created a gap. An autonomous tool-call loop (code
execution, file reads, test runs) can produce one enormous tool result without a
new user message, and that whole result used to count as part of the protected
window. Pinch only fired when the total conversation crossed the budget, but
once it did, those outsized protected results shipped verbatim regardless of
size. In one measured case a 324k-token conversation shrank only ~8% because
nearly all of it sat inside the protected window. `protected_max_chars` closes
that gap: tool results inside the window are still preserved unless they cross
this outlier threshold, in which case they receive the same head/tail elision
applied to older candidates. The default is high enough to avoid touching normal
output; it only intervenes when a single result threatens the budget.
## Config block
All pinch configuration lives under `pinch:` in `config/config.yaml`, validated
by `PinchConfig` in `src/config.py`. Config is strict (`extra="forbid"`):
unknown keys fail at load.
The full block ships as:
```yaml
pinch:
# Relevance-based context pruning. Optional, pre-dispatch stage: when a
# conversation exceeds budget_tokens, the provider-bound messages are trimmed
# BEFORE any paid token is sent upstream.
enabled: true
budget_tokens: 50000
keep_last_turns: 4
max_summarize_chars: 4000
protected_max_chars: 20000
relevance:
enabled: true
model: "nomic-embed-text"
base_url: "http://localhost:11434/v1"
timeout_seconds: 10
min_candidates: 2
```
| Key | Default | Meaning |
|---|---|---|
| `pinch.enabled` | `true` | Master switch. Off ⇒ the whole prune block is a byte-for-byte no-op. |
| `pinch.budget_tokens` | `50000` | Above this estimated token count the conversation is pruned. Must be `> 0`. |
| `pinch.keep_last_turns` | `4` | How many recent user turns (plus their assistant replies and tool results) are protected from pruning. Must be `> 0`. |
| `pinch.max_summarize_chars` | `4000` | Tool results longer than this many chars are summarized in place; shorter ones are collapsed to a placeholder. Must be `>= 3000` — below that summarization would grow the message. |
| `pinch.protected_max_chars` | `20000` | `keep_last_turns` has no size limit inside it — an entire autonomous tool-call loop with no new user message can be one protected turn, so one outsized tool result inside it (a full verbose test run, a huge file read) used to ship verbatim regardless of size. Measured live 2026-09-06: a 324k-token conversation shrank only ~8% because nearly all of it sat inside the protected window. Any tool result inside the protected window over this many chars now gets the same head/tail elision candidates get — deliberately a much higher bar than `max_summarize_chars`, since recent results are more likely to still matter, so it only catches true outliers. Must be `>= 3000`, or `null` to disable. |
| `pinch.prefix_probe` | `true` | Record where this turn's pruned payload stopped matching the previous turn's (three integers per decision row; hashes live in process memory for one turn and never reach the database). Pure telemetry: off changes no routing and no payload, it only stops collecting the evidence for whether pruning is destroying the provider's prompt cache. |
| `pinch.relevance.enabled` | `true` | Embedding-model relevance scoring. Requires `pinch.enabled` too. Any failure reverts to uniform trimming. |
| `pinch.relevance.model` | `nomic-embed-text` | An **embedding** model — never `classifier.model` or `verification.model`. |
| `pinch.relevance.base_url` | `http://localhost:11434/v1` | OpenAI-compatible embeddings endpoint on the same local Ollama. |
| `pinch.relevance.timeout_seconds` | `10` | Embedding call timeout. Must be `> 0`. |
| `pinch.relevance.min_candidates` | `2` | Below this many trim-eligible candidates, skip the embedding round-trip entirely and fall back to uniform compression. Must be `> 0`. |
`PinchRelevanceConfig` (`src/config.py`) defaults mirror the YAML:
| Key | Default |
|---|---|
| `enabled` | `True` |
| `model` | `"nomic-embed-text"` |
| `base_url` | `"http://localhost:11434/v1"` |
| `timeout_seconds` | `10` |
| `min_candidates` | `2` |
`PinchRelevanceConfig` must point at an **embedding** model; its docstring is
explicit that pointing it at `classifier.model` or `verification.model` is
wrong, because those are chat models and the embeddings call would fail
(gracefully degrading to uniform trimming, but always failing).
## Token accounting incl. tool-definition overhead
Token estimation in pinch uses a crude characters-per-token heuristic
(`CHARS_PER_TOKEN = 3`), consistent with the dispatcher:
```python
def estimate_tokens(text): # src/context_prune.py
return len(text) // CHARS_PER_TOKEN if text else 0
```
Tool definitions are a **major omitted factor** and the reason `extra_fixed_tokens`
exists. Tool definitions ride along in `upstream_body` on every provider call and
count against the same context window, so a trivial opencode request can carry
~**32k prompt tokens of tool definitions alone**. Before commit `d7f051b`, pinch
excluded all of that from the budget comparison — every prune run undercounted
the real conversation, so the budget guard triggered far too late (or not at
all) on exactly the tool-heavy requests that newly needed pruning.
### Why `_relevance_order_for` takes `extra_fixed_tokens`
`dispatcher.py` computes the tool-definition overhead once per path:
```python
_tools_overhead = len(json.dumps(body.get("tools") or [])) // CHARS_PER_TOKEN
```
and threads it into both `prune_context` and `_relevance_order_for` as
`extra_fixed_tokens`. Inside `prune_context`:
```python
orig_tokens = sum(estimate_tokens(extract_text(m)) for m in messages) + extra_fixed_tokens
if orig_tokens <= budget_tokens:
return messages, {"pruned": False, ...}
```
So the **budget guard** compares `orig_tokens + extra_fixed_tokens` against the
budget — the same total `_relevance_order_for` computes for its skip check.
`_relevance_order_for` takes `extra_fixed_tokens` for the exact same reason.
Its skip check is:
```python
orig_tokens = sum(estimate_tokens(extract_text(m)) for m in messages) + extra_fixed_tokens
if orig_tokens <= budget_tokens:
return None
```
Without `extra_fixed_tokens`, relevance ordering silently declines to run
**on exactly the requests that newly need pruning**: those that exceed the
budget only because of tool-definition overhead. The failure mode is silent —
no error, no log — it just falls to uniform ranking instead of
least-relevant-first. The `_relevance_order_for` docstring states the contract:
"`extra_fixed_tokens` accounts for overhead that `prune_context`'s budget guard
also adds to `orig_tokens`, so the skip-budget check in here matches the same
total."
### Where the measured-context estimate fits
`estimate_prompt_tokens(send_messages, tools=body.get("tools"))` adds
`len(json.dumps(tools))` to the char count for the routed path's window/tier/cost
decision. So the tool definitions are counted **twice conceptually** — once in
pinch's budget guard (`_tools_overhead`), once in the measured-context estimate
(`tools=body.get("tools")`) — but they represent the same real overhead. Pinch's
`original_tokens`/`final_tokens` stats deliberately include `extra_fixed_tokens`
so the recorded orig matches what actually ships, and the two numbers stay
comparable.
## How relevance ordering works
When `pinch.relevance.enabled` and `pinch.enabled` are both true, and the
conversation is actually over budget, `_relevance_order_for` makes **one
batched** embeddings call (`query` + all candidates in a single request)
against the Ollama endpoint, then `order_by_relevance` scores cosine similarity
in **ascending order** — least relevant first. `prune_context` walks that order,
compressing least-relevant candidates first until the token deficit is covered;
the more-relevant (protected) candidates stay verbatim. When `relevance_order`
is `None` (relevance off, too few candidates, over-budget check failed, or the
embedding call failed), every candidate is compressed uniformly — byte-for-byte
the historical behavior.
Any failure of the embedding round-trip — `RequestException`, timeout,
non-200, unparseable body, missing/malformed vectors — logs `relevance_unavailable`
and returns `None`, which the pure core already treats as "compress everything".
This must never be the thing that breaks a request.
## Instrumentation surface
Pinch records into `route_decisions` and surfaces through `/metrics`, the admin
snapshot, the admin dashboard, and the TUI.
### Schema columns
Two INTEGER columns on `route_decisions`:
| Column | Notes |
|---|---|
| `pinch_original_tokens` | Estimated tokens before pruning (includes `extra_fixed_tokens`). `NULL` when pinch is disabled or the request predates the columns. |
| `pinch_final_tokens` | Estimated tokens after pruning. `NULL` when pinch is disabled or the request predates the columns. |
Both columns are created in `config/schema.sql` and mirrored code-side in
`ensure_route_decisions` for live DBs that predate the feature.
### `prune_context` stats keys
`prune_context` returns `(pruned_messages, stats)`; on the not-pruned path:
```
pruned, original_tokens, final_tokens, tokens_saved
```
and on the pruned path:
```
pruned, original_tokens, final_tokens, tokens_saved, summarized
```
| Key | Meaning |
|---|---|
| `pruned` | Bool: whether anything was actually trimmed. |
| `original_tokens` | Budget-guard total: messages + `extra_fixed_tokens`. |
| `final_tokens` | Estimated tokens of the returned pruned list. |
| `tokens_saved` | `max(original_tokens - final_tokens, 0)` — computed from what the payload actually shrunk by (whole list before vs after), never a per-message estimate. |
| `summarized` | Count of tool results actually summarized/placeholder-replaced. |
### `metrics.pinch_summary(conn, cfg)`
`pinch_summary` returns `None` when pinch is disabled (so callers omit the
section), otherwise a 30-day aggregate:
| Key | Meaning |
|---|---|
| `calls_30d` | Route decisions in the last 30 days. |
| `pruned_calls_30d` | Decisions where `pinch_original_tokens > pinch_final_tokens`. |
| `share_pruned` | `pruned_calls_30d / calls_30d`, **rounded to 4 decimals**. |
| `total_tokens_saved` | Sum of `original - final` over the window. |
| `median_tokens_saved` | 50th percentile of per-decision tokens saved (over pruned rows). |
| `dollars_saved_usd_30d` | Token savings priced at the blended prompt rate, **rounded to 6 decimals**. |
| `reset_date` / `next_reset_date` | The 30-day window start and the upcoming billing-period reset day (only when `objective.billing_reset_day` is set). |
### Where it surfaces
- **`GET /metrics`** — the response includes a `"pinch"` block populated by
`pinch_summary(conn, cfg)`.
- **Admin snapshot** — `admin.py` embeds the same `"pinch": metrics.pinch_summary(conn, cfg)`
in `/admin/api/snapshot`. It also exposes `pinch.enabled`,
`pinch.prefix_probe` and `pinch.relevance.enabled` as admin runtime knobs
(`pinch_enabled`, `pinch_prefix_probe`, `pinch_relevance_enabled`) and as
persisted-config keys. The controls page flags `pinch.prefix_probe` in both
panels, because it is the one knob there whose only effect is on measurement:
turned off, routing is unchanged and the cache-loss evidence simply stops
accumulating.
- **Admin dashboard** — `index.html` renders a "Pinch savings" card showing
**share pruned** (as a percent), **tokens saved**, **median saved**, and
**30d dollars saved** (`$` with 6 decimals). When pinch is disabled the card
shows "Pinch is disabled".
- **TUI** — the dashboard has pinch rows (`share_pruned`,
`total_tokens_saved`, `median_tokens_saved`, `dollars_saved_usd_30d`) pulled
from the snapshot.
## Dollars-saved pricing
The `dollars_saved_usd_30d` figure prices token savings at the **blended
prompt rate** the router itself ranks on, from `routing.estimated_cost`. It
reads `cfg.objective.assumed_cache_rate` and, per saved row, computes:
```python
cached_price = cost_per_1m_prompt_cached or prompt_price
blended = (1.0 - cache_rate) * prompt_price + cache_rate * cached_price
dollars_saved += saved * blended / 1_000_000
```
That mirrors `estimated_cost` in `src/routing.py`, so the dashboard's dollar
figure agrees with the cost model the router uses to rank candidates.
### The review-caught error
An earlier version priced savings **at the cached rate alone**. That was wrong
in two compounding ways:
1. It **double-applied the discount** — the cached price already reflects the
cache discount, so using it as the sole rate applied that discount once for
the cached fraction and again across the whole amount.
2. It **dropped the fraction billed at full price** — the portion of traffic
assumed to miss the cache (the `(1.0 - cache_rate)` share) was silently
billed at the cheaper cached rate too, so savings were overstated.
The fix prices each token at the true blend: the uncached share at full
`prompt_price`, the cached share at `cached_price`, with `assumed_cache_rate`
splitting the two. Only that blended rate reproduces what the router's
`estimated_cost` computes, so the dashboard figure stays consistent with the
cost model the router ranks on.
## Passthrough ordering constraint
On the passthrough path the ordering of three steps is a **contract** — the
dispatcher must NOT reorder them:
1. **prune** — `prune_context` runs first.
2. **persist** — `persist_route_decision` records the passthrough with the
computed `pinch_original_tokens` / `pinch_final_tokens`.
3. **`_check_pinned_capabilities`** — raises 422 when the pinned model cannot
satisfy the request.
Why each constraint holds:
- **Prune is hoisted above persist** so a pruned passthrough request records
real pinch tokens rather than NULL. If persist ran before prune, the
`pinch_stats` (computed only by the prune block) would not exist yet and the
row would be written with `NULL`s, losing the measurement for the very
traffic where pruning matters.
- **Persist stays above `_check_pinned_capabilities`** because the check can
**raise** — and a rejected passthrough must not create a partial decision
row. `_check_pinned_capabilities` raises 422 for a pin that cannot possibly
work (e.g. an image or JSON-mode request pinned to a model without the
capability). If persist ran after the check, a failure would leave nothing
behind; if persist ran after and the check raised, the raise would leave the
persist unexecuted — but the constraint as written ensures that by the time
the capability check can raise, no row has been written for a request that is
about to be rejected.
```python
if cfg.pinch.enabled:
send_messages, pinch_stats = prune_context(...)
persist_route_decision(
"passthrough", ...,
pinch_original_tokens=pinch_stats.get("original_tokens") if pinch_stats is not None else None,
pinch_final_tokens=pinch_stats.get("final_tokens") if pinch_stats is not None else None,
)
if caps.has_images or caps.require_json_mode:
_check_pinned_capabilities(requested, caps)
```
`prune → persist → _check_pinned_capabilities` is the order the code enforces,
and a future reorder would silently corrupt either the measurements or the
decision-log integrity.
## Operational notes
- Pinch is on by default but conservative: start with a large `budget_tokens`
and watch `route_decisions` / pinch stats on real traffic before widening it.
- Tool results can only be dropped safely because tool outputs are idempotent
enough for a placeholder; a wrong guess here loses context, so the default
leans small-reduction.
- The relevance path requires an embedding model pulled on the same Ollama
(`ollama pull nomic-embed-text`) and reachable at `pinch.relevance.base_url`.
Like every other new-and-unproven knob, `relevance.enabled` ships guarded by
`pinch.enabled` — it has no effect otherwise.