Files
6krrt/plans/session-classification-cache-ttl.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

8.7 KiB

Spec: session-scoped classification cache with a minutes-based staleness TTL

Status: done -- session_cache.py staleness window

Origin. Settles §2 of context-dual-use-and-classify-once-per-session.md ("Classify once per session, not once per message"), which laid out the mechanism but explicitly deferred invalidation policy pending a decision. That decision is now made: staleness is a config value in minutes. Also borrows the framing (not the mechanism) from LLMRouter's "routing memory" — retrieving a past routing decision instead of reclassifying from scratch. LLMRouter's version does similarity retrieval over embedded queries; this is simpler and exact, since the thing being cached is per-session, not per-query-similarity: same session, same category, until the TTL says otherwise. This project already has one precedent for porting a llmrouter idea directly — PinchConfig (config.py:304) is docstring-labeled "Port of llmrouter's pinch." Same pattern applies here: keep what's provably useful, drop what LLMRouter needed for its own broader (multi-provider, ML-router) scope but this project doesn't.

No code changes accompany this document — this is the spec opencode builds from.


The problem, restated

Every turn in a long agent session pays a full classifier round-trip (~1.7s local, ~1.0s cloud per the README's own measurements) even though a session's category (coding_general, debugging, etc.) rarely changes turn to turn. What changes is mostly token count, which chat_completions already measures directly via estimate_prompt_tokens — no model call needed for that part.

Why this is safe to cache (and what must never be cached alongside it)

Safe to cache: task_category and task_tier. These describe "what kind of work is this," which is a property of the session's overall task, not of any single turn's exact wording.

Never cache: the capability flags (tools_present, has_images, require_json_mode) read by capabilities.py. These are cheap (pure request-body reads, no model call) and can legitimately differ turn to turn within one session — turn 5 might attach an image, turn 6 might not. Caching them would silently misroute a request whose actual capabilities changed. This is not a design option to weigh; it would reproduce exactly the kind of bug this project's fail-closed capability gates exist to prevent. The mechanism below only ever short-circuits the classifier call, never the capability-gate reads, which already run fresh every request regardless.

Also never cached: apply_escalation. It's cheap, pure Python, and already runs fresh on every route() call whether or not that call hit the classification cache — no change needed there.

Mechanism

route()'s override branch (dispatcher.py, the branch used when task_category/task_tier/required_context_tokens are all supplied) already skips classify() entirely — this is what the measured-context reroute already uses within a single request. A session cache is the same mechanism applied across requests instead of just within one:

  1. On a cache hit (see TTL below): skip classify() outright. Compute send_messages (pinch, if enabled) and measured = estimate_prompt_tokens(...) exactly as today, then call route() once with the cached task_category/task_tier and required_context_tokens=measured. This replaces today's two-step "classify, then maybe re-route on measured context" with a single override-branch call, since there's nothing to react to — the category is already known.
  2. On a cache miss (first turn in a session, or the entry expired): unchanged from today — classify, then the existing measured-context reroute if measured exceeds the classifier's own estimate. After the decision is made, write (task_category, task_tier, now) into the cache under this session's key.

Session identity

Reuse session_fingerprint(messages) (dispatcher.py:1645) — already computed once per chat_completions call for observation, and its own docstring is exactly the right property for this: stable for the life of a session (hashes the opening system-prompt message), distinct across sessions. No new identity mechanism needed.

Staleness (the settled decision)

New config section:

session_cache:
  # Off by default, matching every other new-and-unproven knob in this
  # project (pinch.enabled, routing.min_tool_proficiency): ship it, watch
  # route_decisions on real traffic, then decide the right default.
  enabled: false
  # Minutes since the cached classification was written before it's treated
  # as expired and the next turn reclassifies from scratch.
  staleness_minutes: 20

SessionCacheConfig(StrictModel) in config.py, next to PinchConfig, with a field_validator requiring staleness_minutes > 0 (same pattern as PinchConfig.budget_positive).

On a lookup, an entry older than staleness_minutes is treated exactly like a miss: reclassify, and overwrite the cache entry with the fresh result and a fresh timestamp. No partial-credit or extension-on-read — a stale entry is gone, not renewed.

Why minutes, not turns or a context-jump heuristic (the other two options §2 of the original spec raised): those need their own tuning and their own measurement pass to validate. A minutes-based TTL is one number, directly interpretable ("stop trusting this after N minutes of session inactivity-adjusted-by-nothing"), and cheap to change without a code release. If real traffic later shows category drift happening within the TTL window on active sessions (see the interaction note below), a turn-count or context-jump trigger can be layered on top — but that's a second iteration, not a blocker for this one.

Storage

An in-memory dict in a new session_cache.py, module-level, matching the process lifetime of the dispatcher — a restart just means every active session's next turn reclassifies once, which is a safe failure mode (same reasoning the original spec already used for this). No new persistence layer.

@dataclass(frozen=True)
class CachedClassification:
    task_category: str
    task_tier: int
    cached_at: float  # time.time()

def get(session_key: str, staleness_seconds: float) -> Optional[CachedClassification]: ...
def put(session_key: str, task_category: str, task_tier: int) -> None: ...

Pure functions, no I/O — same shape as scoring.py/tiering.py, testable without a running dispatcher.

Observability

Classification.source (dispatcher.py:215, Literal["classifier", "override", "fallback"]) gains a fourth value: "cached". This was already anticipated by name in the original spec's design questions. Set it at the chat_completions call site (same place classified_src/ctx_src are already tracked locally) when a cache hit served the decision — no changes needed in route()/tiering.py itself. Once this Literal is updated, /metrics's existing classification-source breakdown and the TUI's decision detail popup pick it up for free — they already group by source.

Interaction with the classifier-input-scope finding

classifier-input-scope-check.md (this same batch of plans) confirms the classifier currently reads only the last user message plus one prior assistant turn, refreshed every request — which is incidentally the thing that lets today's per-message classification track category drift within a session for free. This cache trades that away on purpose, on the bet that drift is rare relative to cost: a session that's 40 turns of coding_general and then turns into docs_writing will serve up to staleness_minutes of stale category before the TTL expires and it corrects itself. staleness_minutes should be picked with this trade-off in mind, not purely for latency savings — a smaller value trades away less accuracy for less cache benefit. This is exactly the kind of thing route_decisions.source = "cached" makes observable: watch for "cached" decisions whose category looks wrong in hindsight (e.g. compared against a same-session POST /outcome failure) before tuning the default up from 20 minutes.

Recommendation

Build it. The mechanism is a small extension of code that already exists (the override branch), the session-identity question is already answered by existing code, and the staleness question is now answered by you. Ship session_cache.enabled: false by default per this project's standing pattern for new knobs, flip it on for real sessions, and watch route_decisions.source = "cached" against outcomes before considering a non-default staleness_minutes or a smarter invalidation trigger.