The knob was deliberately left off _BOOL_KNOBS when the probe landed in
fb511cd, on the grounds that a toggle whose only effect is to stop collecting
evidence is questionable UX. That is a fair reading and it loses to the
standing rule: a knob reachable only by hand-editing config.local.yaml is
invisible to whoever is actually operating the router, which is the same gap
classifier.cloud_fallback sat in while the portal's own gaming-mode text told
operators they needed it.
So it is added exactly where its two siblings already are -- _BOOL_KNOBS for
the in-memory flip, _CONFIG_ALLOWLIST plus _CONFIG_GET_ORDER for the persisted
write -- and the objection is answered in the UI instead of by omission. One
short line in the meta track of both panels, "off stops the cache-loss
measurement", with the rest of the reasoning in the title attribute and in
docs/pinch.md where it belongs. No new visual pattern: it is the same
.badge.bg-warning the zero-admit profile hint already uses, rendered into the
same track as the "file: <value>" and provenance badges.
The round-trip test is the confidence_threshold lesson applied rather than
relearned. An admin control that writes a value the config loader will later
reject or silently misread is worse than no control -- that one wrote a raw 80
meaning 80% into the overlay and was caught before the restart by luck -- so
the test reads true, writes false, reads back false, and then loads the merged
base + overlay through the real RouterConfig, asserting a genuine False rather
than the string "false" and that nothing else in the pinch block moved. It
runs entirely against a tmp_path copy; the machine-local overlay is never
opened.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
408 lines
20 KiB
Markdown
408 lines
20 KiB
Markdown
> Deep dive into the context-pruning (pinch) feature. Back to [README](../README.md).
|
|
|
|
Pinch is the router's optional **pre-dispatch context-pruning** stage. Ported
|
|
from the MIT-licensed llmrouter "pinch" module, it shrinks the provider-bound
|
|
conversation once it grows past a token budget, so a long agent session ships
|
|
fewer prompt tokens upstream — before any paid token is sent. It is **on by
|
|
default** (`pinch.enabled: true`), but is deliberately conservative.
|
|
|
|
The core is a pure, injectable module (`src/context_prune.py`): it takes
|
|
messages and limits and returns pruned messages. `dispatcher.py` owns reading
|
|
the config and deciding **when** to call it.
|
|
|
|
## Where pinch sits
|
|
|
|
The critical question for a pruning stage is *when* it runs relative to the
|
|
decisions that depend on conversation size. Pinch is hoisted **before** those
|
|
decisions on both dispatch paths.
|
|
|
|
### Routed path
|
|
|
|
On the routed path (an `auto` decision), pinch runs **before the
|
|
measured-context decision** — the window/tier/cost selection. The comment in
|
|
`dispatcher.py` is explicit: prune ONCE before the measured-context decision,
|
|
"so the window/tier/cost choice sees the size that will actually ship upstream
|
|
rather than the raw conversation."
|
|
|
|
```python
|
|
send_messages, pinch_stats = prune_context(
|
|
list(messages),
|
|
budget_tokens=cfg.pinch.budget_tokens,
|
|
keep_last_turns=cfg.pinch.keep_last_turns,
|
|
max_summarize_chars=cfg.pinch.max_summarize_chars,
|
|
relevance_order=_relevance_order_for(
|
|
messages, cfg, budget_tokens=cfg.pinch.budget_tokens,
|
|
extra_fixed_tokens=_tools_overhead,
|
|
),
|
|
extra_fixed_tokens=_tools_overhead,
|
|
)
|
|
# ... only then:
|
|
measured = estimate_prompt_tokens(send_messages, tools=body.get("tools"))
|
|
```
|
|
|
|
The pruned list (`send_messages`) is then **reused** at dispatch time — it is
|
|
never pruned twice. When pinch is disabled this whole block is a byte-for-byte
|
|
no-op and `send_messages` is the raw `messages` list, exactly as before.
|
|
|
|
Pinch trims only **tool results**, never user/assistant/system messages, so the
|
|
classification turn above (which used `_previous_context` from the **full**
|
|
untouched messages) is undisturbed — the classifier still sees the whole
|
|
history.
|
|
|
|
### Passthrough path
|
|
|
|
On the passthrough path (a pinned model id, dispatched as asked), pinch runs
|
|
**separately** with its own prune block. There the ordering matters differently,
|
|
and is a hard contract (see [Passthrough ordering constraint](#passthrough-ordering-constraint)
|
|
below): prune → persist → `_check_pinned_capabilities`.
|
|
|
|
## What pinch trims
|
|
|
|
Pinch's tailoring is deliberate and safe:
|
|
|
|
- **Only `tool` results** are trimmed or summarized. Older results outside the
|
|
protected window are the main target, but any **outsized result inside the
|
|
window** that is longer than `pinch.protected_max_chars` is also capped. That
|
|
threshold is deliberately a much higher bar than `max_summarize_chars`; recent
|
|
results are more likely to still matter, so it only catches true outliers.
|
|
Tool results are where a long agent session's tokens actually live, they are
|
|
the least likely to still be needed in full by the time a later turn is
|
|
answered, and replacing one with a short placeholder is reversible at the
|
|
semantic level — a wrong guess costs context but never breaks the request.
|
|
- **Never user / assistant / system messages.** Trimming one of those can
|
|
change what the model is being asked, so they are always kept verbatim.
|
|
- **Never the classifier input.** The classification turn uses
|
|
`_previous_context` from the full untouched messages, and pinch clamps
|
|
head+tail independently — a separate concern from the classifier's input
|
|
framing.
|
|
|
|
Messages are **never removed** (order and length are preserved): a tool result
|
|
is either summarized in place (head + tail around a `[N chars trimmed...]`
|
|
marker) or replaced with a `[name: result omitted]` placeholder, so
|
|
`len(pruned)` always equals `len(messages)` and the result still parses as a
|
|
valid conversation.
|
|
|
|
Only runs (and only mutates anything) when the estimate actually exceeds the
|
|
budget; otherwise the original list is returned untouched.
|
|
|
|
### Capping outsized results in the protected window
|
|
|
|
The protected window, controlled by `keep_last_turns`, is intentionally generous:
|
|
anything in the last few user turns stays verbatim because it is most likely to
|
|
still be needed. But that created a gap. An autonomous tool-call loop (code
|
|
execution, file reads, test runs) can produce one enormous tool result without a
|
|
new user message, and that whole result used to count as part of the protected
|
|
window. Pinch only fired when the total conversation crossed the budget, but
|
|
once it did, those outsized protected results shipped verbatim regardless of
|
|
size. In one measured case a 324k-token conversation shrank only ~8% because
|
|
nearly all of it sat inside the protected window. `protected_max_chars` closes
|
|
that gap: tool results inside the window are still preserved unless they cross
|
|
this outlier threshold, in which case they receive the same head/tail elision
|
|
applied to older candidates. The default is high enough to avoid touching normal
|
|
output; it only intervenes when a single result threatens the budget.
|
|
|
|
## Config block
|
|
|
|
All pinch configuration lives under `pinch:` in `config/config.yaml`, validated
|
|
by `PinchConfig` in `src/config.py`. Config is strict (`extra="forbid"`):
|
|
unknown keys fail at load.
|
|
|
|
The full block ships as:
|
|
|
|
```yaml
|
|
pinch:
|
|
# Relevance-based context pruning. Optional, pre-dispatch stage: when a
|
|
# conversation exceeds budget_tokens, the provider-bound messages are trimmed
|
|
# BEFORE any paid token is sent upstream.
|
|
enabled: true
|
|
budget_tokens: 50000
|
|
keep_last_turns: 4
|
|
max_summarize_chars: 4000
|
|
protected_max_chars: 20000
|
|
relevance:
|
|
enabled: true
|
|
model: "nomic-embed-text"
|
|
base_url: "http://localhost:11434/v1"
|
|
timeout_seconds: 10
|
|
min_candidates: 2
|
|
```
|
|
|
|
| Key | Default | Meaning |
|
|
|---|---|---|
|
|
| `pinch.enabled` | `true` | Master switch. Off ⇒ the whole prune block is a byte-for-byte no-op. |
|
|
| `pinch.budget_tokens` | `50000` | Above this estimated token count the conversation is pruned. Must be `> 0`. |
|
|
| `pinch.keep_last_turns` | `4` | How many recent user turns (plus their assistant replies and tool results) are protected from pruning. Must be `> 0`. |
|
|
| `pinch.max_summarize_chars` | `4000` | Tool results longer than this many chars are summarized in place; shorter ones are collapsed to a placeholder. Must be `>= 3000` — below that summarization would grow the message. |
|
|
| `pinch.protected_max_chars` | `20000` | `keep_last_turns` has no size limit inside it — an entire autonomous tool-call loop with no new user message can be one protected turn, so one outsized tool result inside it (a full verbose test run, a huge file read) used to ship verbatim regardless of size. Measured live 2026-09-06: a 324k-token conversation shrank only ~8% because nearly all of it sat inside the protected window. Any tool result inside the protected window over this many chars now gets the same head/tail elision candidates get — deliberately a much higher bar than `max_summarize_chars`, since recent results are more likely to still matter, so it only catches true outliers. Must be `>= 3000`, or `null` to disable. |
|
|
| `pinch.prefix_probe` | `true` | Record where this turn's pruned payload stopped matching the previous turn's (three integers per decision row; hashes live in process memory for one turn and never reach the database). Pure telemetry: off changes no routing and no payload, it only stops collecting the evidence for whether pruning is destroying the provider's prompt cache. |
|
|
| `pinch.relevance.enabled` | `true` | Embedding-model relevance scoring. Requires `pinch.enabled` too. Any failure reverts to uniform trimming. |
|
|
| `pinch.relevance.model` | `nomic-embed-text` | An **embedding** model — never `classifier.model` or `verification.model`. |
|
|
| `pinch.relevance.base_url` | `http://localhost:11434/v1` | OpenAI-compatible embeddings endpoint on the same local Ollama. |
|
|
| `pinch.relevance.timeout_seconds` | `10` | Embedding call timeout. Must be `> 0`. |
|
|
| `pinch.relevance.min_candidates` | `2` | Below this many trim-eligible candidates, skip the embedding round-trip entirely and fall back to uniform compression. Must be `> 0`. |
|
|
|
|
`PinchRelevanceConfig` (`src/config.py`) defaults mirror the YAML:
|
|
|
|
| Key | Default |
|
|
|---|---|
|
|
| `enabled` | `True` |
|
|
| `model` | `"nomic-embed-text"` |
|
|
| `base_url` | `"http://localhost:11434/v1"` |
|
|
| `timeout_seconds` | `10` |
|
|
| `min_candidates` | `2` |
|
|
|
|
`PinchRelevanceConfig` must point at an **embedding** model; its docstring is
|
|
explicit that pointing it at `classifier.model` or `verification.model` is
|
|
wrong, because those are chat models and the embeddings call would fail
|
|
(gracefully degrading to uniform trimming, but always failing).
|
|
|
|
## Token accounting incl. tool-definition overhead
|
|
|
|
Token estimation in pinch uses a crude characters-per-token heuristic
|
|
(`CHARS_PER_TOKEN = 3`), consistent with the dispatcher:
|
|
|
|
```python
|
|
def estimate_tokens(text): # src/context_prune.py
|
|
return len(text) // CHARS_PER_TOKEN if text else 0
|
|
```
|
|
|
|
Tool definitions are a **major omitted factor** and the reason `extra_fixed_tokens`
|
|
exists. Tool definitions ride along in `upstream_body` on every provider call and
|
|
count against the same context window, so a trivial opencode request can carry
|
|
~**32k prompt tokens of tool definitions alone**. Before commit `d7f051b`, pinch
|
|
excluded all of that from the budget comparison — every prune run undercounted
|
|
the real conversation, so the budget guard triggered far too late (or not at
|
|
all) on exactly the tool-heavy requests that newly needed pruning.
|
|
|
|
### Why `_relevance_order_for` takes `extra_fixed_tokens`
|
|
|
|
`dispatcher.py` computes the tool-definition overhead once per path:
|
|
|
|
```python
|
|
_tools_overhead = len(json.dumps(body.get("tools") or [])) // CHARS_PER_TOKEN
|
|
```
|
|
|
|
and threads it into both `prune_context` and `_relevance_order_for` as
|
|
`extra_fixed_tokens`. Inside `prune_context`:
|
|
|
|
```python
|
|
orig_tokens = sum(estimate_tokens(extract_text(m)) for m in messages) + extra_fixed_tokens
|
|
if orig_tokens <= budget_tokens:
|
|
return messages, {"pruned": False, ...}
|
|
```
|
|
|
|
So the **budget guard** compares `orig_tokens + extra_fixed_tokens` against the
|
|
budget — the same total `_relevance_order_for` computes for its skip check.
|
|
|
|
`_relevance_order_for` takes `extra_fixed_tokens` for the exact same reason.
|
|
Its skip check is:
|
|
|
|
```python
|
|
orig_tokens = sum(estimate_tokens(extract_text(m)) for m in messages) + extra_fixed_tokens
|
|
if orig_tokens <= budget_tokens:
|
|
return None
|
|
```
|
|
|
|
Without `extra_fixed_tokens`, relevance ordering silently declines to run
|
|
**on exactly the requests that newly need pruning**: those that exceed the
|
|
budget only because of tool-definition overhead. The failure mode is silent —
|
|
no error, no log — it just falls to uniform ranking instead of
|
|
least-relevant-first. The `_relevance_order_for` docstring states the contract:
|
|
"`extra_fixed_tokens` accounts for overhead that `prune_context`'s budget guard
|
|
also adds to `orig_tokens`, so the skip-budget check in here matches the same
|
|
total."
|
|
|
|
### Where the measured-context estimate fits
|
|
|
|
`estimate_prompt_tokens(send_messages, tools=body.get("tools"))` adds
|
|
`len(json.dumps(tools))` to the char count for the routed path's window/tier/cost
|
|
decision. So the tool definitions are counted **twice conceptually** — once in
|
|
pinch's budget guard (`_tools_overhead`), once in the measured-context estimate
|
|
(`tools=body.get("tools")`) — but they represent the same real overhead. Pinch's
|
|
`original_tokens`/`final_tokens` stats deliberately include `extra_fixed_tokens`
|
|
so the recorded orig matches what actually ships, and the two numbers stay
|
|
comparable.
|
|
|
|
## How relevance ordering works
|
|
|
|
When `pinch.relevance.enabled` and `pinch.enabled` are both true, and the
|
|
conversation is actually over budget, `_relevance_order_for` makes **one
|
|
batched** embeddings call (`query` + all candidates in a single request)
|
|
against the Ollama endpoint, then `order_by_relevance` scores cosine similarity
|
|
in **ascending order** — least relevant first. `prune_context` walks that order,
|
|
compressing least-relevant candidates first until the token deficit is covered;
|
|
the more-relevant (protected) candidates stay verbatim. When `relevance_order`
|
|
is `None` (relevance off, too few candidates, over-budget check failed, or the
|
|
embedding call failed), every candidate is compressed uniformly — byte-for-byte
|
|
the historical behavior.
|
|
|
|
Any failure of the embedding round-trip — `RequestException`, timeout,
|
|
non-200, unparseable body, missing/malformed vectors — logs `relevance_unavailable`
|
|
and returns `None`, which the pure core already treats as "compress everything".
|
|
This must never be the thing that breaks a request.
|
|
|
|
## Instrumentation surface
|
|
|
|
Pinch records into `route_decisions` and surfaces through `/metrics`, the admin
|
|
snapshot, the admin dashboard, and the TUI.
|
|
|
|
### Schema columns
|
|
|
|
Two INTEGER columns on `route_decisions`:
|
|
|
|
| Column | Notes |
|
|
|---|---|
|
|
| `pinch_original_tokens` | Estimated tokens before pruning (includes `extra_fixed_tokens`). `NULL` when pinch is disabled or the request predates the columns. |
|
|
| `pinch_final_tokens` | Estimated tokens after pruning. `NULL` when pinch is disabled or the request predates the columns. |
|
|
|
|
Both columns are created in `config/schema.sql` and mirrored code-side in
|
|
`ensure_route_decisions` for live DBs that predate the feature.
|
|
|
|
### `prune_context` stats keys
|
|
|
|
`prune_context` returns `(pruned_messages, stats)`; on the not-pruned path:
|
|
|
|
```
|
|
pruned, original_tokens, final_tokens, tokens_saved
|
|
```
|
|
|
|
and on the pruned path:
|
|
|
|
```
|
|
pruned, original_tokens, final_tokens, tokens_saved, summarized
|
|
```
|
|
|
|
| Key | Meaning |
|
|
|---|---|
|
|
| `pruned` | Bool: whether anything was actually trimmed. |
|
|
| `original_tokens` | Budget-guard total: messages + `extra_fixed_tokens`. |
|
|
| `final_tokens` | Estimated tokens of the returned pruned list. |
|
|
| `tokens_saved` | `max(original_tokens - final_tokens, 0)` — computed from what the payload actually shrunk by (whole list before vs after), never a per-message estimate. |
|
|
| `summarized` | Count of tool results actually summarized/placeholder-replaced. |
|
|
|
|
### `metrics.pinch_summary(conn, cfg)`
|
|
|
|
`pinch_summary` returns `None` when pinch is disabled (so callers omit the
|
|
section), otherwise a 30-day aggregate:
|
|
|
|
| Key | Meaning |
|
|
|---|---|
|
|
| `calls_30d` | Route decisions in the last 30 days. |
|
|
| `pruned_calls_30d` | Decisions where `pinch_original_tokens > pinch_final_tokens`. |
|
|
| `share_pruned` | `pruned_calls_30d / calls_30d`, **rounded to 4 decimals**. |
|
|
| `total_tokens_saved` | Sum of `original - final` over the window. |
|
|
| `median_tokens_saved` | 50th percentile of per-decision tokens saved (over pruned rows). |
|
|
| `dollars_saved_usd_30d` | Token savings priced at the blended prompt rate, **rounded to 6 decimals**. |
|
|
| `reset_date` / `next_reset_date` | The 30-day window start and the upcoming billing-period reset day (only when `objective.billing_reset_day` is set). |
|
|
|
|
### Where it surfaces
|
|
|
|
- **`GET /metrics`** — the response includes a `"pinch"` block populated by
|
|
`pinch_summary(conn, cfg)`.
|
|
- **Admin snapshot** — `admin.py` embeds the same `"pinch": metrics.pinch_summary(conn, cfg)`
|
|
in `/admin/api/snapshot`. It also exposes `pinch.enabled`,
|
|
`pinch.prefix_probe` and `pinch.relevance.enabled` as admin runtime knobs
|
|
(`pinch_enabled`, `pinch_prefix_probe`, `pinch_relevance_enabled`) and as
|
|
persisted-config keys. The controls page flags `pinch.prefix_probe` in both
|
|
panels, because it is the one knob there whose only effect is on measurement:
|
|
turned off, routing is unchanged and the cache-loss evidence simply stops
|
|
accumulating.
|
|
- **Admin dashboard** — `index.html` renders a "Pinch savings" card showing
|
|
**share pruned** (as a percent), **tokens saved**, **median saved**, and
|
|
**30d dollars saved** (`$` with 6 decimals). When pinch is disabled the card
|
|
shows "Pinch is disabled".
|
|
- **TUI** — the dashboard has pinch rows (`share_pruned`,
|
|
`total_tokens_saved`, `median_tokens_saved`, `dollars_saved_usd_30d`) pulled
|
|
from the snapshot.
|
|
|
|
## Dollars-saved pricing
|
|
|
|
The `dollars_saved_usd_30d` figure prices token savings at the **blended
|
|
prompt rate** the router itself ranks on, from `routing.estimated_cost`. It
|
|
reads `cfg.objective.assumed_cache_rate` and, per saved row, computes:
|
|
|
|
```python
|
|
cached_price = cost_per_1m_prompt_cached or prompt_price
|
|
blended = (1.0 - cache_rate) * prompt_price + cache_rate * cached_price
|
|
dollars_saved += saved * blended / 1_000_000
|
|
```
|
|
|
|
That mirrors `estimated_cost` in `src/routing.py`, so the dashboard's dollar
|
|
figure agrees with the cost model the router uses to rank candidates.
|
|
|
|
### The review-caught error
|
|
|
|
An earlier version priced savings **at the cached rate alone**. That was wrong
|
|
in two compounding ways:
|
|
|
|
1. It **double-applied the discount** — the cached price already reflects the
|
|
cache discount, so using it as the sole rate applied that discount once for
|
|
the cached fraction and again across the whole amount.
|
|
2. It **dropped the fraction billed at full price** — the portion of traffic
|
|
assumed to miss the cache (the `(1.0 - cache_rate)` share) was silently
|
|
billed at the cheaper cached rate too, so savings were overstated.
|
|
|
|
The fix prices each token at the true blend: the uncached share at full
|
|
`prompt_price`, the cached share at `cached_price`, with `assumed_cache_rate`
|
|
splitting the two. Only that blended rate reproduces what the router's
|
|
`estimated_cost` computes, so the dashboard figure stays consistent with the
|
|
cost model the router ranks on.
|
|
|
|
## Passthrough ordering constraint
|
|
|
|
On the passthrough path the ordering of three steps is a **contract** — the
|
|
dispatcher must NOT reorder them:
|
|
|
|
1. **prune** — `prune_context` runs first.
|
|
2. **persist** — `persist_route_decision` records the passthrough with the
|
|
computed `pinch_original_tokens` / `pinch_final_tokens`.
|
|
3. **`_check_pinned_capabilities`** — raises 422 when the pinned model cannot
|
|
satisfy the request.
|
|
|
|
Why each constraint holds:
|
|
|
|
- **Prune is hoisted above persist** so a pruned passthrough request records
|
|
real pinch tokens rather than NULL. If persist ran before prune, the
|
|
`pinch_stats` (computed only by the prune block) would not exist yet and the
|
|
row would be written with `NULL`s, losing the measurement for the very
|
|
traffic where pruning matters.
|
|
- **Persist stays above `_check_pinned_capabilities`** because the check can
|
|
**raise** — and a rejected passthrough must not create a partial decision
|
|
row. `_check_pinned_capabilities` raises 422 for a pin that cannot possibly
|
|
work (e.g. an image or JSON-mode request pinned to a model without the
|
|
capability). If persist ran after the check, a failure would leave nothing
|
|
behind; if persist ran after and the check raised, the raise would leave the
|
|
persist unexecuted — but the constraint as written ensures that by the time
|
|
the capability check can raise, no row has been written for a request that is
|
|
about to be rejected.
|
|
|
|
```python
|
|
if cfg.pinch.enabled:
|
|
send_messages, pinch_stats = prune_context(...)
|
|
|
|
persist_route_decision(
|
|
"passthrough", ...,
|
|
pinch_original_tokens=pinch_stats.get("original_tokens") if pinch_stats is not None else None,
|
|
pinch_final_tokens=pinch_stats.get("final_tokens") if pinch_stats is not None else None,
|
|
)
|
|
|
|
if caps.has_images or caps.require_json_mode:
|
|
_check_pinned_capabilities(requested, caps)
|
|
```
|
|
|
|
`prune → persist → _check_pinned_capabilities` is the order the code enforces,
|
|
and a future reorder would silently corrupt either the measurements or the
|
|
decision-log integrity.
|
|
|
|
## Operational notes
|
|
|
|
- Pinch is on by default but conservative: start with a large `budget_tokens`
|
|
and watch `route_decisions` / pinch stats on real traffic before widening it.
|
|
- Tool results can only be dropped safely because tool outputs are idempotent
|
|
enough for a placeholder; a wrong guess here loses context, so the default
|
|
leans small-reduction.
|
|
- The relevance path requires an embedding model pulled on the same Ollama
|
|
(`ollama pull nomic-embed-text`) and reachable at `pinch.relevance.base_url`.
|
|
Like every other new-and-unproven knob, `relevance.enabled` ships guarded by
|
|
`pinch.enabled` — it has no effect otherwise.
|