The knob was deliberately left off _BOOL_KNOBS when the probe landed in
fb511cd, on the grounds that a toggle whose only effect is to stop collecting
evidence is questionable UX. That is a fair reading and it loses to the
standing rule: a knob reachable only by hand-editing config.local.yaml is
invisible to whoever is actually operating the router, which is the same gap
classifier.cloud_fallback sat in while the portal's own gaming-mode text told
operators they needed it.
So it is added exactly where its two siblings already are -- _BOOL_KNOBS for
the in-memory flip, _CONFIG_ALLOWLIST plus _CONFIG_GET_ORDER for the persisted
write -- and the objection is answered in the UI instead of by omission. One
short line in the meta track of both panels, "off stops the cache-loss
measurement", with the rest of the reasoning in the title attribute and in
docs/pinch.md where it belongs. No new visual pattern: it is the same
.badge.bg-warning the zero-admit profile hint already uses, rendered into the
same track as the "file: <value>" and provenance badges.
The round-trip test is the confidence_threshold lesson applied rather than
relearned. An admin control that writes a value the config loader will later
reject or silently misread is worse than no control -- that one wrote a raw 80
meaning 80% into the overlay and was caught before the restart by luck -- so
the test reads true, writes false, reads back false, and then loads the merged
base + overlay through the real RouterConfig, asserting a genuine False rather
than the string "false" and that nothing else in the pinch block moved. It
runs entirely against a tmp_path copy; the machine-local overlay is never
opened.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
20 KiB
Deep dive into the context-pruning (pinch) feature. Back to README.
Pinch is the router's optional pre-dispatch context-pruning stage. Ported
from the MIT-licensed llmrouter "pinch" module, it shrinks the provider-bound
conversation once it grows past a token budget, so a long agent session ships
fewer prompt tokens upstream — before any paid token is sent. It is on by
default (pinch.enabled: true), but is deliberately conservative.
The core is a pure, injectable module (src/context_prune.py): it takes
messages and limits and returns pruned messages. dispatcher.py owns reading
the config and deciding when to call it.
Where pinch sits
The critical question for a pruning stage is when it runs relative to the decisions that depend on conversation size. Pinch is hoisted before those decisions on both dispatch paths.
Routed path
On the routed path (an auto decision), pinch runs before the
measured-context decision — the window/tier/cost selection. The comment in
dispatcher.py is explicit: prune ONCE before the measured-context decision,
"so the window/tier/cost choice sees the size that will actually ship upstream
rather than the raw conversation."
send_messages, pinch_stats = prune_context(
list(messages),
budget_tokens=cfg.pinch.budget_tokens,
keep_last_turns=cfg.pinch.keep_last_turns,
max_summarize_chars=cfg.pinch.max_summarize_chars,
relevance_order=_relevance_order_for(
messages, cfg, budget_tokens=cfg.pinch.budget_tokens,
extra_fixed_tokens=_tools_overhead,
),
extra_fixed_tokens=_tools_overhead,
)
# ... only then:
measured = estimate_prompt_tokens(send_messages, tools=body.get("tools"))
The pruned list (send_messages) is then reused at dispatch time — it is
never pruned twice. When pinch is disabled this whole block is a byte-for-byte
no-op and send_messages is the raw messages list, exactly as before.
Pinch trims only tool results, never user/assistant/system messages, so the
classification turn above (which used _previous_context from the full
untouched messages) is undisturbed — the classifier still sees the whole
history.
Passthrough path
On the passthrough path (a pinned model id, dispatched as asked), pinch runs
separately with its own prune block. There the ordering matters differently,
and is a hard contract (see Passthrough ordering constraint
below): prune → persist → _check_pinned_capabilities.
What pinch trims
Pinch's tailoring is deliberate and safe:
- Only
toolresults are trimmed or summarized. Older results outside the protected window are the main target, but any outsized result inside the window that is longer thanpinch.protected_max_charsis also capped. That threshold is deliberately a much higher bar thanmax_summarize_chars; recent results are more likely to still matter, so it only catches true outliers. Tool results are where a long agent session's tokens actually live, they are the least likely to still be needed in full by the time a later turn is answered, and replacing one with a short placeholder is reversible at the semantic level — a wrong guess costs context but never breaks the request. - Never user / assistant / system messages. Trimming one of those can change what the model is being asked, so they are always kept verbatim.
- Never the classifier input. The classification turn uses
_previous_contextfrom the full untouched messages, and pinch clamps head+tail independently — a separate concern from the classifier's input framing.
Messages are never removed (order and length are preserved): a tool result
is either summarized in place (head + tail around a [N chars trimmed...]
marker) or replaced with a [name: result omitted] placeholder, so
len(pruned) always equals len(messages) and the result still parses as a
valid conversation.
Only runs (and only mutates anything) when the estimate actually exceeds the budget; otherwise the original list is returned untouched.
Capping outsized results in the protected window
The protected window, controlled by keep_last_turns, is intentionally generous:
anything in the last few user turns stays verbatim because it is most likely to
still be needed. But that created a gap. An autonomous tool-call loop (code
execution, file reads, test runs) can produce one enormous tool result without a
new user message, and that whole result used to count as part of the protected
window. Pinch only fired when the total conversation crossed the budget, but
once it did, those outsized protected results shipped verbatim regardless of
size. In one measured case a 324k-token conversation shrank only ~8% because
nearly all of it sat inside the protected window. protected_max_chars closes
that gap: tool results inside the window are still preserved unless they cross
this outlier threshold, in which case they receive the same head/tail elision
applied to older candidates. The default is high enough to avoid touching normal
output; it only intervenes when a single result threatens the budget.
Config block
All pinch configuration lives under pinch: in config/config.yaml, validated
by PinchConfig in src/config.py. Config is strict (extra="forbid"):
unknown keys fail at load.
The full block ships as:
pinch:
# Relevance-based context pruning. Optional, pre-dispatch stage: when a
# conversation exceeds budget_tokens, the provider-bound messages are trimmed
# BEFORE any paid token is sent upstream.
enabled: true
budget_tokens: 50000
keep_last_turns: 4
max_summarize_chars: 4000
protected_max_chars: 20000
relevance:
enabled: true
model: "nomic-embed-text"
base_url: "http://localhost:11434/v1"
timeout_seconds: 10
min_candidates: 2
| Key | Default | Meaning |
|---|---|---|
pinch.enabled |
true |
Master switch. Off ⇒ the whole prune block is a byte-for-byte no-op. |
pinch.budget_tokens |
50000 |
Above this estimated token count the conversation is pruned. Must be > 0. |
pinch.keep_last_turns |
4 |
How many recent user turns (plus their assistant replies and tool results) are protected from pruning. Must be > 0. |
pinch.max_summarize_chars |
4000 |
Tool results longer than this many chars are summarized in place; shorter ones are collapsed to a placeholder. Must be >= 3000 — below that summarization would grow the message. |
pinch.protected_max_chars |
20000 |
keep_last_turns has no size limit inside it — an entire autonomous tool-call loop with no new user message can be one protected turn, so one outsized tool result inside it (a full verbose test run, a huge file read) used to ship verbatim regardless of size. Measured live 2026-09-06: a 324k-token conversation shrank only ~8% because nearly all of it sat inside the protected window. Any tool result inside the protected window over this many chars now gets the same head/tail elision candidates get — deliberately a much higher bar than max_summarize_chars, since recent results are more likely to still matter, so it only catches true outliers. Must be >= 3000, or null to disable. |
pinch.prefix_probe |
true |
Record where this turn's pruned payload stopped matching the previous turn's (three integers per decision row; hashes live in process memory for one turn and never reach the database). Pure telemetry: off changes no routing and no payload, it only stops collecting the evidence for whether pruning is destroying the provider's prompt cache. |
pinch.relevance.enabled |
true |
Embedding-model relevance scoring. Requires pinch.enabled too. Any failure reverts to uniform trimming. |
pinch.relevance.model |
nomic-embed-text |
An embedding model — never classifier.model or verification.model. |
pinch.relevance.base_url |
http://localhost:11434/v1 |
OpenAI-compatible embeddings endpoint on the same local Ollama. |
pinch.relevance.timeout_seconds |
10 |
Embedding call timeout. Must be > 0. |
pinch.relevance.min_candidates |
2 |
Below this many trim-eligible candidates, skip the embedding round-trip entirely and fall back to uniform compression. Must be > 0. |
PinchRelevanceConfig (src/config.py) defaults mirror the YAML:
| Key | Default |
|---|---|
enabled |
True |
model |
"nomic-embed-text" |
base_url |
"http://localhost:11434/v1" |
timeout_seconds |
10 |
min_candidates |
2 |
PinchRelevanceConfig must point at an embedding model; its docstring is
explicit that pointing it at classifier.model or verification.model is
wrong, because those are chat models and the embeddings call would fail
(gracefully degrading to uniform trimming, but always failing).
Token accounting incl. tool-definition overhead
Token estimation in pinch uses a crude characters-per-token heuristic
(CHARS_PER_TOKEN = 3), consistent with the dispatcher:
def estimate_tokens(text): # src/context_prune.py
return len(text) // CHARS_PER_TOKEN if text else 0
Tool definitions are a major omitted factor and the reason extra_fixed_tokens
exists. Tool definitions ride along in upstream_body on every provider call and
count against the same context window, so a trivial opencode request can carry
~32k prompt tokens of tool definitions alone. Before commit d7f051b, pinch
excluded all of that from the budget comparison — every prune run undercounted
the real conversation, so the budget guard triggered far too late (or not at
all) on exactly the tool-heavy requests that newly needed pruning.
Why _relevance_order_for takes extra_fixed_tokens
dispatcher.py computes the tool-definition overhead once per path:
_tools_overhead = len(json.dumps(body.get("tools") or [])) // CHARS_PER_TOKEN
and threads it into both prune_context and _relevance_order_for as
extra_fixed_tokens. Inside prune_context:
orig_tokens = sum(estimate_tokens(extract_text(m)) for m in messages) + extra_fixed_tokens
if orig_tokens <= budget_tokens:
return messages, {"pruned": False, ...}
So the budget guard compares orig_tokens + extra_fixed_tokens against the
budget — the same total _relevance_order_for computes for its skip check.
_relevance_order_for takes extra_fixed_tokens for the exact same reason.
Its skip check is:
orig_tokens = sum(estimate_tokens(extract_text(m)) for m in messages) + extra_fixed_tokens
if orig_tokens <= budget_tokens:
return None
Without extra_fixed_tokens, relevance ordering silently declines to run
on exactly the requests that newly need pruning: those that exceed the
budget only because of tool-definition overhead. The failure mode is silent —
no error, no log — it just falls to uniform ranking instead of
least-relevant-first. The _relevance_order_for docstring states the contract:
"extra_fixed_tokens accounts for overhead that prune_context's budget guard
also adds to orig_tokens, so the skip-budget check in here matches the same
total."
Where the measured-context estimate fits
estimate_prompt_tokens(send_messages, tools=body.get("tools")) adds
len(json.dumps(tools)) to the char count for the routed path's window/tier/cost
decision. So the tool definitions are counted twice conceptually — once in
pinch's budget guard (_tools_overhead), once in the measured-context estimate
(tools=body.get("tools")) — but they represent the same real overhead. Pinch's
original_tokens/final_tokens stats deliberately include extra_fixed_tokens
so the recorded orig matches what actually ships, and the two numbers stay
comparable.
How relevance ordering works
When pinch.relevance.enabled and pinch.enabled are both true, and the
conversation is actually over budget, _relevance_order_for makes one
batched embeddings call (query + all candidates in a single request)
against the Ollama endpoint, then order_by_relevance scores cosine similarity
in ascending order — least relevant first. prune_context walks that order,
compressing least-relevant candidates first until the token deficit is covered;
the more-relevant (protected) candidates stay verbatim. When relevance_order
is None (relevance off, too few candidates, over-budget check failed, or the
embedding call failed), every candidate is compressed uniformly — byte-for-byte
the historical behavior.
Any failure of the embedding round-trip — RequestException, timeout,
non-200, unparseable body, missing/malformed vectors — logs relevance_unavailable
and returns None, which the pure core already treats as "compress everything".
This must never be the thing that breaks a request.
Instrumentation surface
Pinch records into route_decisions and surfaces through /metrics, the admin
snapshot, the admin dashboard, and the TUI.
Schema columns
Two INTEGER columns on route_decisions:
| Column | Notes |
|---|---|
pinch_original_tokens |
Estimated tokens before pruning (includes extra_fixed_tokens). NULL when pinch is disabled or the request predates the columns. |
pinch_final_tokens |
Estimated tokens after pruning. NULL when pinch is disabled or the request predates the columns. |
Both columns are created in config/schema.sql and mirrored code-side in
ensure_route_decisions for live DBs that predate the feature.
prune_context stats keys
prune_context returns (pruned_messages, stats); on the not-pruned path:
pruned, original_tokens, final_tokens, tokens_saved
and on the pruned path:
pruned, original_tokens, final_tokens, tokens_saved, summarized
| Key | Meaning |
|---|---|
pruned |
Bool: whether anything was actually trimmed. |
original_tokens |
Budget-guard total: messages + extra_fixed_tokens. |
final_tokens |
Estimated tokens of the returned pruned list. |
tokens_saved |
max(original_tokens - final_tokens, 0) — computed from what the payload actually shrunk by (whole list before vs after), never a per-message estimate. |
summarized |
Count of tool results actually summarized/placeholder-replaced. |
metrics.pinch_summary(conn, cfg)
pinch_summary returns None when pinch is disabled (so callers omit the
section), otherwise a 30-day aggregate:
| Key | Meaning |
|---|---|
calls_30d |
Route decisions in the last 30 days. |
pruned_calls_30d |
Decisions where pinch_original_tokens > pinch_final_tokens. |
share_pruned |
pruned_calls_30d / calls_30d, rounded to 4 decimals. |
total_tokens_saved |
Sum of original - final over the window. |
median_tokens_saved |
50th percentile of per-decision tokens saved (over pruned rows). |
dollars_saved_usd_30d |
Token savings priced at the blended prompt rate, rounded to 6 decimals. |
reset_date / next_reset_date |
The 30-day window start and the upcoming billing-period reset day (only when objective.billing_reset_day is set). |
Where it surfaces
GET /metrics— the response includes a"pinch"block populated bypinch_summary(conn, cfg).- Admin snapshot —
admin.pyembeds the same"pinch": metrics.pinch_summary(conn, cfg)in/admin/api/snapshot. It also exposespinch.enabled,pinch.prefix_probeandpinch.relevance.enabledas admin runtime knobs (pinch_enabled,pinch_prefix_probe,pinch_relevance_enabled) and as persisted-config keys. The controls page flagspinch.prefix_probein both panels, because it is the one knob there whose only effect is on measurement: turned off, routing is unchanged and the cache-loss evidence simply stops accumulating. - Admin dashboard —
index.htmlrenders a "Pinch savings" card showing share pruned (as a percent), tokens saved, median saved, and 30d dollars saved ($with 6 decimals). When pinch is disabled the card shows "Pinch is disabled". - TUI — the dashboard has pinch rows (
share_pruned,total_tokens_saved,median_tokens_saved,dollars_saved_usd_30d) pulled from the snapshot.
Dollars-saved pricing
The dollars_saved_usd_30d figure prices token savings at the blended
prompt rate the router itself ranks on, from routing.estimated_cost. It
reads cfg.objective.assumed_cache_rate and, per saved row, computes:
cached_price = cost_per_1m_prompt_cached or prompt_price
blended = (1.0 - cache_rate) * prompt_price + cache_rate * cached_price
dollars_saved += saved * blended / 1_000_000
That mirrors estimated_cost in src/routing.py, so the dashboard's dollar
figure agrees with the cost model the router uses to rank candidates.
The review-caught error
An earlier version priced savings at the cached rate alone. That was wrong in two compounding ways:
- It double-applied the discount — the cached price already reflects the cache discount, so using it as the sole rate applied that discount once for the cached fraction and again across the whole amount.
- It dropped the fraction billed at full price — the portion of traffic
assumed to miss the cache (the
(1.0 - cache_rate)share) was silently billed at the cheaper cached rate too, so savings were overstated.
The fix prices each token at the true blend: the uncached share at full
prompt_price, the cached share at cached_price, with assumed_cache_rate
splitting the two. Only that blended rate reproduces what the router's
estimated_cost computes, so the dashboard figure stays consistent with the
cost model the router ranks on.
Passthrough ordering constraint
On the passthrough path the ordering of three steps is a contract — the dispatcher must NOT reorder them:
- prune —
prune_contextruns first. - persist —
persist_route_decisionrecords the passthrough with the computedpinch_original_tokens/pinch_final_tokens. _check_pinned_capabilities— raises 422 when the pinned model cannot satisfy the request.
Why each constraint holds:
- Prune is hoisted above persist so a pruned passthrough request records
real pinch tokens rather than NULL. If persist ran before prune, the
pinch_stats(computed only by the prune block) would not exist yet and the row would be written withNULLs, losing the measurement for the very traffic where pruning matters. - Persist stays above
_check_pinned_capabilitiesbecause the check can raise — and a rejected passthrough must not create a partial decision row._check_pinned_capabilitiesraises 422 for a pin that cannot possibly work (e.g. an image or JSON-mode request pinned to a model without the capability). If persist ran after the check, a failure would leave nothing behind; if persist ran after and the check raised, the raise would leave the persist unexecuted — but the constraint as written ensures that by the time the capability check can raise, no row has been written for a request that is about to be rejected.
if cfg.pinch.enabled:
send_messages, pinch_stats = prune_context(...)
persist_route_decision(
"passthrough", ...,
pinch_original_tokens=pinch_stats.get("original_tokens") if pinch_stats is not None else None,
pinch_final_tokens=pinch_stats.get("final_tokens") if pinch_stats is not None else None,
)
if caps.has_images or caps.require_json_mode:
_check_pinned_capabilities(requested, caps)
prune → persist → _check_pinned_capabilities is the order the code enforces,
and a future reorder would silently corrupt either the measurements or the
decision-log integrity.
Operational notes
- Pinch is on by default but conservative: start with a large
budget_tokensand watchroute_decisions/ pinch stats on real traffic before widening it. - Tool results can only be dropped safely because tool outputs are idempotent enough for a placeholder; a wrong guess here loses context, so the default leans small-reduction.
- The relevance path requires an embedding model pulled on the same Ollama
(
ollama pull nomic-embed-text) and reachable atpinch.relevance.base_url. Like every other new-and-unproven knob,relevance.enabledships guarded bypinch.enabled— it has no effect otherwise.