plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
8.7 KiB
Spec: session-scoped classification cache with a minutes-based staleness TTL
Status: done -- session_cache.py staleness window
Origin. Settles §2 of
context-dual-use-and-classify-once-per-session.md
("Classify once per session, not once per message"), which laid out the
mechanism but explicitly deferred invalidation policy pending a decision.
That decision is now made: staleness is a config value in minutes. Also
borrows the framing (not the mechanism) from
LLMRouter's "routing memory" —
retrieving a past routing decision instead of reclassifying from scratch.
LLMRouter's version does similarity retrieval over embedded queries; this is
simpler and exact, since the thing being cached is per-session, not
per-query-similarity: same session, same category, until the TTL says
otherwise. This project already has one precedent for porting a llmrouter
idea directly — PinchConfig (config.py:304) is docstring-labeled "Port of
llmrouter's pinch." Same pattern applies here: keep what's provably useful,
drop what LLMRouter needed for its own broader (multi-provider, ML-router)
scope but this project doesn't.
No code changes accompany this document — this is the spec opencode builds from.
The problem, restated
Every turn in a long agent session pays a full classifier round-trip
(~1.7s local, ~1.0s cloud per the README's own measurements) even though a
session's category (coding_general, debugging, etc.) rarely changes
turn to turn. What changes is mostly token count, which chat_completions
already measures directly via estimate_prompt_tokens — no model call
needed for that part.
Why this is safe to cache (and what must never be cached alongside it)
Safe to cache: task_category and task_tier. These describe "what
kind of work is this," which is a property of the session's overall task,
not of any single turn's exact wording.
Never cache: the capability flags (tools_present, has_images,
require_json_mode) read by capabilities.py. These are cheap (pure
request-body reads, no model call) and can legitimately differ turn to turn
within one session — turn 5 might attach an image, turn 6 might not. Caching
them would silently misroute a request whose actual capabilities changed.
This is not a design option to weigh; it would reproduce exactly the kind of
bug this project's fail-closed capability gates exist to prevent. The
mechanism below only ever short-circuits the classifier call, never the
capability-gate reads, which already run fresh every request regardless.
Also never cached: apply_escalation. It's cheap, pure Python, and already
runs fresh on every route() call whether or not that call hit the
classification cache — no change needed there.
Mechanism
route()'s override branch (dispatcher.py, the branch used when
task_category/task_tier/required_context_tokens are all supplied)
already skips classify() entirely — this is what the measured-context
reroute already uses within a single request. A session cache is the same
mechanism applied across requests instead of just within one:
- On a cache hit (see TTL below): skip
classify()outright. Computesend_messages(pinch, if enabled) andmeasured = estimate_prompt_tokens(...)exactly as today, then callroute()once with the cachedtask_category/task_tierandrequired_context_tokens=measured. This replaces today's two-step "classify, then maybe re-route on measured context" with a single override-branch call, since there's nothing to react to — the category is already known. - On a cache miss (first turn in a session, or the entry expired):
unchanged from today — classify, then the existing measured-context
reroute if
measuredexceeds the classifier's own estimate. After the decision is made, write(task_category, task_tier, now)into the cache under this session's key.
Session identity
Reuse session_fingerprint(messages) (dispatcher.py:1645) — already
computed once per chat_completions call for observation, and its own
docstring is exactly the right property for this: stable for the life of a
session (hashes the opening system-prompt message), distinct across
sessions. No new identity mechanism needed.
Staleness (the settled decision)
New config section:
session_cache:
# Off by default, matching every other new-and-unproven knob in this
# project (pinch.enabled, routing.min_tool_proficiency): ship it, watch
# route_decisions on real traffic, then decide the right default.
enabled: false
# Minutes since the cached classification was written before it's treated
# as expired and the next turn reclassifies from scratch.
staleness_minutes: 20
SessionCacheConfig(StrictModel) in config.py, next to PinchConfig, with
a field_validator requiring staleness_minutes > 0 (same pattern as
PinchConfig.budget_positive).
On a lookup, an entry older than staleness_minutes is treated exactly like
a miss: reclassify, and overwrite the cache entry with the fresh result and
a fresh timestamp. No partial-credit or extension-on-read — a stale entry is
gone, not renewed.
Why minutes, not turns or a context-jump heuristic (the other two options §2 of the original spec raised): those need their own tuning and their own measurement pass to validate. A minutes-based TTL is one number, directly interpretable ("stop trusting this after N minutes of session inactivity-adjusted-by-nothing"), and cheap to change without a code release. If real traffic later shows category drift happening within the TTL window on active sessions (see the interaction note below), a turn-count or context-jump trigger can be layered on top — but that's a second iteration, not a blocker for this one.
Storage
An in-memory dict in a new session_cache.py, module-level, matching the
process lifetime of the dispatcher — a restart just means every active
session's next turn reclassifies once, which is a safe failure mode (same
reasoning the original spec already used for this). No new persistence
layer.
@dataclass(frozen=True)
class CachedClassification:
task_category: str
task_tier: int
cached_at: float # time.time()
def get(session_key: str, staleness_seconds: float) -> Optional[CachedClassification]: ...
def put(session_key: str, task_category: str, task_tier: int) -> None: ...
Pure functions, no I/O — same shape as scoring.py/tiering.py, testable
without a running dispatcher.
Observability
Classification.source (dispatcher.py:215,
Literal["classifier", "override", "fallback"]) gains a fourth value:
"cached". This was already anticipated by name in the original spec's
design questions. Set it at the chat_completions call site (same place
classified_src/ctx_src are already tracked locally) when a cache hit
served the decision — no changes needed in route()/tiering.py itself.
Once this Literal is updated, /metrics's existing classification-source
breakdown and the TUI's decision detail popup pick it up for free — they
already group by source.
Interaction with the classifier-input-scope finding
classifier-input-scope-check.md (this
same batch of plans) confirms the classifier currently reads only the last
user message plus one prior assistant turn, refreshed every request — which
is incidentally the thing that lets today's per-message classification
track category drift within a session for free. This cache trades that
away on purpose, on the bet that drift is rare relative to cost: a session
that's 40 turns of coding_general and then turns into docs_writing will
serve up to staleness_minutes of stale category before the TTL expires
and it corrects itself. staleness_minutes should be picked with this
trade-off in mind, not purely for latency savings — a smaller value trades
away less accuracy for less cache benefit. This is exactly the kind of
thing route_decisions.source = "cached" makes observable: watch for
"cached" decisions whose category looks wrong in hindsight (e.g. compared
against a same-session POST /outcome failure) before tuning the default
up from 20 minutes.
Recommendation
Build it. The mechanism is a small extension of code that already exists
(the override branch), the session-identity question is already answered by
existing code, and the staleness question is now answered by you. Ship
session_cache.enabled: false by default per this project's standing
pattern for new knobs, flip it on for real sessions, and watch
route_decisions.source = "cached" against outcomes before considering a
non-default staleness_minutes or a smarter invalidation trigger.