plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
11 KiB
Spec: embedding-based relevance scoring for pinch
Status: done -- pinch.relevance.enabled
Origin. context_prune.py's own docstring names this directly: "Ported
from the MIT-licensed alexrudloff/llmrouter 'pinch' module... Where llmrouter
embeds every candidate message and scores cosine relevance, this module
keeps the same SAFE invariants without requiring an embedding model on the
request path." The embedding step was cut on purpose, to keep pinch pure
and offline-testable — not because it was a bad idea. This spec is that step,
scoped to fit in without giving up the purity that cut it in the first
place. It also settles the scoping note in
magic-brainstorming-review.md
(idea #3): pinch's mechanism only ever touches old tool results, and this
spec stays inside that boundary rather than widening it.
No code changes accompany this document — this is the spec opencode builds from.
The problem, restated
Today, prune_context (context_prune.py) decides what to trim by
position only. Every tool-result message before protected_from (the
cutoff set by keep_last_turns) gets compressed — elided to head+tail if
long, replaced with a placeholder if short — regardless of whether that
result has anything to do with what the current turn is actually about. A
tool result from 3 turns ago that's central to the task being finished right
now gets flattened exactly as hard as one from 20 turns ago that's
completely irrelevant, because the only signal used is how old it is.
What changes and what must never change
Changes: which trim-eligible candidates get compressed. Instead of "all
of them, uniformly," it becomes "the least relevant ones first, stopping
once the token deficit is covered." A relevant-but-old tool result can now
survive untouched even though it's before the keep_last_turns cutoff.
Never changes — same invariants context_prune.py already documents:
- user/assistant/system messages are still always kept verbatim. The embedding step never scores or touches them; it only re-orders which tool results get compressed, using the exact same elision/placeholder mechanism that exists today for compressing them.
- No message is ever removed; order and role pairing are preserved.
- Fails closed to today's exact behavior. If relevance scoring is disabled, unavailable, times out, or errors, pinch falls back to compressing every trim-eligible candidate uniformly — precisely what it does today. This is not a degraded mode with reduced functionality; it is bit-for-bit the current, already-shipped, already-safe behavior. Nothing about pinch's risk profile gets worse by adding this — at worst, a request gets today's pinch instead of the smarter one.
This is also why an embedding model is the right tool and a generative one
isn't (see magic-brainstorming-review.md's idea #1/#2 findings on why
generative rewrites were rejected): an embedding call can fail or time out,
but it cannot hallucinate a wrong ranking that looks confident the way a
generative summary can hallucinate wrong content. The failure mode is "no
better than today," not "worse than today."
Mechanism
Pure core (testable offline, no model dependency)
A new pure function, next to prune_context in context_prune.py:
def order_by_relevance(
query_embedding: list[float],
candidate_embeddings: list[list[float]],
) -> list[int]:
"""Indexes into candidate_embeddings, LEAST relevant to query first.
Cosine similarity, ascending. The caller compresses in this order until
the token deficit is covered, so index 0 is compressed first.
"""
Trivially unit-testable with hand-built vectors (orthogonal, parallel,
near-duplicate) — no network, no model, same testing shape as
scoring.normalize_inverted.
prune_context gains one new optional parameter:
def prune_context(
messages: list[dict],
budget_tokens: int = ...,
keep_last_turns: int = ...,
max_summarize_chars: int = ...,
relevance_order: Optional[list[int]] = None, # NEW
) -> tuple[list[dict], dict]:
relevance_order indexes into the trim-eligible candidate list (tool
messages before protected_from, in the same order context_prune.py
already collects them) — same shape order_by_relevance returns, so the
dispatcher can pass its result straight through.
Cut behavior when relevance_order is provided: walk it in order,
compressing each candidate with the existing elide/placeholder logic
(unchanged), tracking cumulative tokens_saved against
orig_tokens - budget_tokens. Stop as soon as the deficit is covered —
remaining (more relevant) candidates stay verbatim. If the whole list is
exhausted before the deficit is covered, every candidate has been
compressed exactly as today, so there's no scenario where this compresses
more than the current uniform pass.
When relevance_order is None: compress every candidate, in whatever
order they're encountered — byte-for-byte the current implementation. This
is the fallback path, and it's also what happens today when the feature is
off entirely.
Impure edge (dispatcher.py owns it, matching every other model call)
A new function alongside _classify_once / _run_local_vision / the local
verification call — all of which already live in dispatcher.py per
context_prune.py's own docstring ("This module is pure... dispatcher.py
owns reading the config and deciding when to call it"):
def _embed_for_relevance(query: str, candidates: list[str], cfg) -> Optional[list[int]]:
"""Returns order_by_relevance's result, or None on any failure.
One batched embeddings call (query + all candidates in a single
request — most embedding APIs, Ollama's /api/embed included, accept a
list input) rather than N round-trips.
"""
Called only when cfg.pinch.enabled and cfg.pinch.relevance.enabled and the
trim-eligible candidate count is >= cfg.pinch.relevance.min_candidates
(below that, a network round-trip isn't worth it — ranking 1 candidate is
not a decision). Query text: the same current-turn text already computed
for classification (_last_user_text(messages) or equivalent — reuse, don't
recompute). On requests.RequestException, timeout, or an unparseable
response: logs.warning("relevance_unavailable", ...) and return None,
which the pure core already treats as "compress everything," so this can
never be the thing that breaks a request.
Config
pinch:
enabled: false
budget_tokens: 50000
keep_last_turns: 4
max_summarize_chars: 4000
relevance:
# Off by default, matching every other new-and-unproven knob in this
# project — and specifically requires pinch.enabled too, since this has
# no effect otherwise. Ship it, watch route_decisions / pinch stats on
# real traffic, then decide the default.
enabled: false
# An EMBEDDING model, not a chat model — this must not point at
# classifier.model or verification.model. Pull one on the same Ollama:
# ollama pull nomic-embed-text
model: "nomic-embed-text"
base_url: "http://localhost:11434/v1"
timeout_seconds: 10
# Below this many trim-eligible candidates, skip the embedding call
# entirely and fall back to uniform compression — a network round trip
# to rank one candidate decides nothing.
min_candidates: 2
PinchConfig gains a nested relevance: PinchRelevanceConfig = PinchRelevanceConfig() field, same pattern RouterConfig already uses for
verification/local_vision. PinchRelevanceConfig(StrictModel) with a
field_validator on timeout_seconds and min_candidates (both > 0),
matching every other positivity validator in config.py.
Before implementing: verify the endpoint live, don't assume it
This session already found two cases (classifier.num_ctx,
local_vision.keep_alive) where a request-level Ollama parameter that
should have worked, per general knowledge of the API, was silently ignored
by this specific Ollama version (0.22.0) over the OpenAI-compatible surface.
/v1/embeddings (or native /api/embed) is a standard, long-supported
endpoint rather than a vendor-extension field grafted onto chat completions,
so it's a much safer bet — but "should work" was exactly the assumption that
failed twice already this session. First implementation step: ollama pull nomic-embed-text and a direct curl against whichever endpoint
_embed_for_relevance will use, confirming it returns real, distinct
vectors for distinct inputs, before wiring it into dispatcher.py.
Interaction with the local-vision gap
local_vision's message list is never pinch-pruned at all today
(dispatcher.py:2210 passes the raw messages, not send_messages —
called out separately in magic-brainstorming-review.md's addendum). This
spec doesn't fix that; it only makes the existing pinch call sites
(chat_completions's two pinch invocations) smarter. If local_vision is
later wired through pinch, it gets this relevance scoring for free, since
it's the same prune_context function underneath.
Testing
order_by_relevance: hand-built vectors — near-identical, orthogonal, and opposite — assert the ascending-relevance ordering directly. No model, no I/O, runs in the existing offline suite.prune_contextwith an explicitrelevance_order: construct a case where an OLDER tool result is ranked more relevant than a NEWER one still inside the trim zone, and assert the older one survives verbatim while the newer, less-relevant one gets compressed — this is the actual new behavior, and it must be shown to override recency, not just coexist with it.prune_contextwithrelevance_order=None: assert byte-for-byte identical output to a call made without the parameter at all — the regression guard that the fallback path is truly a no-op, not just "close enough."_embed_for_relevancefailure paths (timeout, non-200, malformed response): assertNoneis returned and nothing raises, mirroringtests/test_config_endpoints.py's and verification's own failure-mode coverage style.
Recommendation
Worth building, and lower-risk than it might look: the pure core is a small,
fully-testable addition to a module that's already designed for exactly
this extension point (the docstring names the missing piece explicitly),
the fallback is provably identical to today's shipped behavior, and the
failure mode of an embedding call (mis-ranking) is categorically safer than
the failure mode of the generative ideas this same brainstorming session
already rejected. Verify the embeddings endpoint live first — that's a
20-minute check, not a design question — then build the pure core and its
tests before touching dispatcher.py.