plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
181 lines
8.7 KiB
Markdown
181 lines
8.7 KiB
Markdown
# Spec: session-scoped classification cache with a minutes-based staleness TTL
|
|
|
|
Status: done -- session_cache.py staleness window
|
|
|
|
**Origin.** Settles §2 of
|
|
[`context-dual-use-and-classify-once-per-session.md`](context-dual-use-and-classify-once-per-session.md)
|
|
("Classify once per session, not once per message"), which laid out the
|
|
mechanism but explicitly deferred invalidation policy pending a decision.
|
|
That decision is now made: staleness is a config value in **minutes**. Also
|
|
borrows the framing (not the mechanism) from
|
|
[LLMRouter](https://github.com/ulab-uiuc/LLMRouter)'s "routing memory" —
|
|
retrieving a past routing decision instead of reclassifying from scratch.
|
|
LLMRouter's version does similarity retrieval over embedded queries; this is
|
|
simpler and exact, since the thing being cached is per-session, not
|
|
per-query-similarity: same session, same category, until the TTL says
|
|
otherwise. This project already has one precedent for porting a llmrouter
|
|
idea directly — `PinchConfig` (`config.py:304`) is docstring-labeled "Port of
|
|
llmrouter's pinch." Same pattern applies here: keep what's provably useful,
|
|
drop what LLMRouter needed for its own broader (multi-provider, ML-router)
|
|
scope but this project doesn't.
|
|
|
|
No code changes accompany this document — this is the spec opencode builds
|
|
from.
|
|
|
|
---
|
|
|
|
## The problem, restated
|
|
|
|
Every turn in a long agent session pays a full classifier round-trip
|
|
(~1.7s local, ~1.0s cloud per the README's own measurements) even though a
|
|
session's *category* (`coding_general`, `debugging`, etc.) rarely changes
|
|
turn to turn. What changes is mostly token count, which `chat_completions`
|
|
already measures directly via `estimate_prompt_tokens` — no model call
|
|
needed for that part.
|
|
|
|
## Why this is safe to cache (and what must never be cached alongside it)
|
|
|
|
**Safe to cache:** `task_category` and `task_tier`. These describe "what
|
|
kind of work is this," which is a property of the session's overall task,
|
|
not of any single turn's exact wording.
|
|
|
|
**Never cache:** the capability flags (`tools_present`, `has_images`,
|
|
`require_json_mode`) read by `capabilities.py`. These are cheap (pure
|
|
request-body reads, no model call) and can legitimately differ turn to turn
|
|
within one session — turn 5 might attach an image, turn 6 might not. Caching
|
|
them would silently misroute a request whose actual capabilities changed.
|
|
This is not a design option to weigh; it would reproduce exactly the kind of
|
|
bug this project's fail-closed capability gates exist to prevent. The
|
|
mechanism below only ever short-circuits the classifier call, never the
|
|
capability-gate reads, which already run fresh every request regardless.
|
|
|
|
Also never cached: `apply_escalation`. It's cheap, pure Python, and already
|
|
runs fresh on every `route()` call whether or not that call hit the
|
|
classification cache — no change needed there.
|
|
|
|
## Mechanism
|
|
|
|
`route()`'s override branch (`dispatcher.py`, the branch used when
|
|
`task_category`/`task_tier`/`required_context_tokens` are all supplied)
|
|
already skips `classify()` entirely — this is what the measured-context
|
|
reroute already uses within a single request. A session cache is the same
|
|
mechanism applied across requests instead of just within one:
|
|
|
|
1. On a **cache hit** (see TTL below): skip `classify()` outright. Compute
|
|
`send_messages` (pinch, if enabled) and `measured =
|
|
estimate_prompt_tokens(...)` exactly as today, then call `route()` once
|
|
with the cached `task_category`/`task_tier` and
|
|
`required_context_tokens=measured`. This replaces today's two-step
|
|
"classify, then maybe re-route on measured context" with a single
|
|
override-branch call, since there's nothing to react to — the category
|
|
is already known.
|
|
2. On a **cache miss** (first turn in a session, or the entry expired):
|
|
unchanged from today — classify, then the existing measured-context
|
|
reroute if `measured` exceeds the classifier's own estimate. After the
|
|
decision is made, write `(task_category, task_tier, now)` into the
|
|
cache under this session's key.
|
|
|
|
### Session identity
|
|
|
|
Reuse `session_fingerprint(messages)` (`dispatcher.py:1645`) — already
|
|
computed once per `chat_completions` call for observation, and its own
|
|
docstring is exactly the right property for this: stable for the life of a
|
|
session (hashes the opening system-prompt message), distinct across
|
|
sessions. No new identity mechanism needed.
|
|
|
|
### Staleness (the settled decision)
|
|
|
|
New config section:
|
|
|
|
```yaml
|
|
session_cache:
|
|
# Off by default, matching every other new-and-unproven knob in this
|
|
# project (pinch.enabled, routing.min_tool_proficiency): ship it, watch
|
|
# route_decisions on real traffic, then decide the right default.
|
|
enabled: false
|
|
# Minutes since the cached classification was written before it's treated
|
|
# as expired and the next turn reclassifies from scratch.
|
|
staleness_minutes: 20
|
|
```
|
|
|
|
`SessionCacheConfig(StrictModel)` in `config.py`, next to `PinchConfig`, with
|
|
a `field_validator` requiring `staleness_minutes > 0` (same pattern as
|
|
`PinchConfig.budget_positive`).
|
|
|
|
On a lookup, an entry older than `staleness_minutes` is treated exactly like
|
|
a miss: reclassify, and overwrite the cache entry with the fresh result and
|
|
a fresh timestamp. No partial-credit or extension-on-read — a stale entry is
|
|
gone, not renewed.
|
|
|
|
**Why minutes, not turns or a context-jump heuristic** (the other two
|
|
options §2 of the original spec raised): those need their own tuning and
|
|
their own measurement pass to validate. A minutes-based TTL is one number,
|
|
directly interpretable ("stop trusting this after N minutes of session
|
|
inactivity-adjusted-by-nothing"), and cheap to change without a code
|
|
release. If real traffic later shows category drift happening *within* the
|
|
TTL window on active sessions (see the interaction note below), a
|
|
turn-count or context-jump trigger can be layered on top — but that's a
|
|
second iteration, not a blocker for this one.
|
|
|
|
### Storage
|
|
|
|
An in-memory dict in a new `session_cache.py`, module-level, matching the
|
|
process lifetime of the dispatcher — a restart just means every active
|
|
session's next turn reclassifies once, which is a safe failure mode (same
|
|
reasoning the original spec already used for this). No new persistence
|
|
layer.
|
|
|
|
```python
|
|
@dataclass(frozen=True)
|
|
class CachedClassification:
|
|
task_category: str
|
|
task_tier: int
|
|
cached_at: float # time.time()
|
|
|
|
def get(session_key: str, staleness_seconds: float) -> Optional[CachedClassification]: ...
|
|
def put(session_key: str, task_category: str, task_tier: int) -> None: ...
|
|
```
|
|
|
|
Pure functions, no I/O — same shape as `scoring.py`/`tiering.py`, testable
|
|
without a running dispatcher.
|
|
|
|
### Observability
|
|
|
|
`Classification.source` (`dispatcher.py:215`,
|
|
`Literal["classifier", "override", "fallback"]`) gains a fourth value:
|
|
`"cached"`. This was already anticipated by name in the original spec's
|
|
design questions. Set it at the `chat_completions` call site (same place
|
|
`classified_src`/`ctx_src` are already tracked locally) when a cache hit
|
|
served the decision — no changes needed in `route()`/`tiering.py` itself.
|
|
Once this Literal is updated, `/metrics`'s existing classification-source
|
|
breakdown and the TUI's decision detail popup pick it up for free — they
|
|
already group by `source`.
|
|
|
|
## Interaction with the classifier-input-scope finding
|
|
|
|
[`classifier-input-scope-check.md`](classifier-input-scope-check.md) (this
|
|
same batch of plans) confirms the classifier currently reads only the last
|
|
user message plus one prior assistant turn, refreshed every request — which
|
|
is incidentally the thing that lets today's per-message classification
|
|
track category drift within a session for free. This cache trades that
|
|
away on purpose, on the bet that drift is rare relative to cost: a session
|
|
that's 40 turns of `coding_general` and then turns into `docs_writing` will
|
|
serve up to `staleness_minutes` of stale category before the TTL expires
|
|
and it corrects itself. `staleness_minutes` should be picked with this
|
|
trade-off in mind, not purely for latency savings — a smaller value trades
|
|
away less accuracy for less cache benefit. This is exactly the kind of
|
|
thing `route_decisions.source = "cached"` makes observable: watch for
|
|
`"cached"` decisions whose category looks wrong in hindsight (e.g. compared
|
|
against a same-session `POST /outcome` failure) before tuning the default
|
|
up from 20 minutes.
|
|
|
|
## Recommendation
|
|
|
|
Build it. The mechanism is a small extension of code that already exists
|
|
(the override branch), the session-identity question is already answered by
|
|
existing code, and the staleness question is now answered by you. Ship
|
|
`session_cache.enabled: false` by default per this project's standing
|
|
pattern for new knobs, flip it on for real sessions, and watch
|
|
`route_decisions.source = "cached"` against outcomes before considering a
|
|
non-default `staleness_minutes` or a smarter invalidation trigger.
|