# Spec: session-scoped classification cache with a minutes-based staleness TTL Status: done -- session_cache.py staleness window **Origin.** Settles §2 of [`context-dual-use-and-classify-once-per-session.md`](context-dual-use-and-classify-once-per-session.md) ("Classify once per session, not once per message"), which laid out the mechanism but explicitly deferred invalidation policy pending a decision. That decision is now made: staleness is a config value in **minutes**. Also borrows the framing (not the mechanism) from [LLMRouter](https://github.com/ulab-uiuc/LLMRouter)'s "routing memory" — retrieving a past routing decision instead of reclassifying from scratch. LLMRouter's version does similarity retrieval over embedded queries; this is simpler and exact, since the thing being cached is per-session, not per-query-similarity: same session, same category, until the TTL says otherwise. This project already has one precedent for porting a llmrouter idea directly — `PinchConfig` (`config.py:304`) is docstring-labeled "Port of llmrouter's pinch." Same pattern applies here: keep what's provably useful, drop what LLMRouter needed for its own broader (multi-provider, ML-router) scope but this project doesn't. No code changes accompany this document — this is the spec opencode builds from. --- ## The problem, restated Every turn in a long agent session pays a full classifier round-trip (~1.7s local, ~1.0s cloud per the README's own measurements) even though a session's *category* (`coding_general`, `debugging`, etc.) rarely changes turn to turn. What changes is mostly token count, which `chat_completions` already measures directly via `estimate_prompt_tokens` — no model call needed for that part. ## Why this is safe to cache (and what must never be cached alongside it) **Safe to cache:** `task_category` and `task_tier`. These describe "what kind of work is this," which is a property of the session's overall task, not of any single turn's exact wording. **Never cache:** the capability flags (`tools_present`, `has_images`, `require_json_mode`) read by `capabilities.py`. These are cheap (pure request-body reads, no model call) and can legitimately differ turn to turn within one session — turn 5 might attach an image, turn 6 might not. Caching them would silently misroute a request whose actual capabilities changed. This is not a design option to weigh; it would reproduce exactly the kind of bug this project's fail-closed capability gates exist to prevent. The mechanism below only ever short-circuits the classifier call, never the capability-gate reads, which already run fresh every request regardless. Also never cached: `apply_escalation`. It's cheap, pure Python, and already runs fresh on every `route()` call whether or not that call hit the classification cache — no change needed there. ## Mechanism `route()`'s override branch (`dispatcher.py`, the branch used when `task_category`/`task_tier`/`required_context_tokens` are all supplied) already skips `classify()` entirely — this is what the measured-context reroute already uses within a single request. A session cache is the same mechanism applied across requests instead of just within one: 1. On a **cache hit** (see TTL below): skip `classify()` outright. Compute `send_messages` (pinch, if enabled) and `measured = estimate_prompt_tokens(...)` exactly as today, then call `route()` once with the cached `task_category`/`task_tier` and `required_context_tokens=measured`. This replaces today's two-step "classify, then maybe re-route on measured context" with a single override-branch call, since there's nothing to react to — the category is already known. 2. On a **cache miss** (first turn in a session, or the entry expired): unchanged from today — classify, then the existing measured-context reroute if `measured` exceeds the classifier's own estimate. After the decision is made, write `(task_category, task_tier, now)` into the cache under this session's key. ### Session identity Reuse `session_fingerprint(messages)` (`dispatcher.py:1645`) — already computed once per `chat_completions` call for observation, and its own docstring is exactly the right property for this: stable for the life of a session (hashes the opening system-prompt message), distinct across sessions. No new identity mechanism needed. ### Staleness (the settled decision) New config section: ```yaml session_cache: # Off by default, matching every other new-and-unproven knob in this # project (pinch.enabled, routing.min_tool_proficiency): ship it, watch # route_decisions on real traffic, then decide the right default. enabled: false # Minutes since the cached classification was written before it's treated # as expired and the next turn reclassifies from scratch. staleness_minutes: 20 ``` `SessionCacheConfig(StrictModel)` in `config.py`, next to `PinchConfig`, with a `field_validator` requiring `staleness_minutes > 0` (same pattern as `PinchConfig.budget_positive`). On a lookup, an entry older than `staleness_minutes` is treated exactly like a miss: reclassify, and overwrite the cache entry with the fresh result and a fresh timestamp. No partial-credit or extension-on-read — a stale entry is gone, not renewed. **Why minutes, not turns or a context-jump heuristic** (the other two options §2 of the original spec raised): those need their own tuning and their own measurement pass to validate. A minutes-based TTL is one number, directly interpretable ("stop trusting this after N minutes of session inactivity-adjusted-by-nothing"), and cheap to change without a code release. If real traffic later shows category drift happening *within* the TTL window on active sessions (see the interaction note below), a turn-count or context-jump trigger can be layered on top — but that's a second iteration, not a blocker for this one. ### Storage An in-memory dict in a new `session_cache.py`, module-level, matching the process lifetime of the dispatcher — a restart just means every active session's next turn reclassifies once, which is a safe failure mode (same reasoning the original spec already used for this). No new persistence layer. ```python @dataclass(frozen=True) class CachedClassification: task_category: str task_tier: int cached_at: float # time.time() def get(session_key: str, staleness_seconds: float) -> Optional[CachedClassification]: ... def put(session_key: str, task_category: str, task_tier: int) -> None: ... ``` Pure functions, no I/O — same shape as `scoring.py`/`tiering.py`, testable without a running dispatcher. ### Observability `Classification.source` (`dispatcher.py:215`, `Literal["classifier", "override", "fallback"]`) gains a fourth value: `"cached"`. This was already anticipated by name in the original spec's design questions. Set it at the `chat_completions` call site (same place `classified_src`/`ctx_src` are already tracked locally) when a cache hit served the decision — no changes needed in `route()`/`tiering.py` itself. Once this Literal is updated, `/metrics`'s existing classification-source breakdown and the TUI's decision detail popup pick it up for free — they already group by `source`. ## Interaction with the classifier-input-scope finding [`classifier-input-scope-check.md`](classifier-input-scope-check.md) (this same batch of plans) confirms the classifier currently reads only the last user message plus one prior assistant turn, refreshed every request — which is incidentally the thing that lets today's per-message classification track category drift within a session for free. This cache trades that away on purpose, on the bet that drift is rare relative to cost: a session that's 40 turns of `coding_general` and then turns into `docs_writing` will serve up to `staleness_minutes` of stale category before the TTL expires and it corrects itself. `staleness_minutes` should be picked with this trade-off in mind, not purely for latency savings — a smaller value trades away less accuracy for less cache benefit. This is exactly the kind of thing `route_decisions.source = "cached"` makes observable: watch for `"cached"` decisions whose category looks wrong in hindsight (e.g. compared against a same-session `POST /outcome` failure) before tuning the default up from 20 minutes. ## Recommendation Build it. The mechanism is a small extension of code that already exists (the override branch), the session-identity question is already answered by existing code, and the staleness question is now answered by you. Ship `session_cache.enabled: false` by default per this project's standing pattern for new knobs, flip it on for real sessions, and watch `route_decisions.source = "cached"` against outcomes before considering a non-default `staleness_minutes` or a smarter invalidation trigger.