Files
6krrt/plans/classifier-input-scope-check.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

68 lines
3.4 KiB
Markdown

# Finding: classifier input scope (LLMRouter "query modes" check)
Status: done -- classifier.max_input_chars
**Origin.** LLMRouter's chat interface offers three query modes —
`current_only`, `full_context` (all history), `retrieval` (top-k similar
past queries). Before considering porting any of that, the actual question
was narrower and answerable by reading the code: does this project's
classifier already read the whole conversation, or only the tail? If only
the tail, a long session that drifts category (chat → debugging) could be
mis-tiered from a stale early read. This is a verification finding, not a
design proposal — no code changes proposed here.
---
## What the code actually does
`chat_completions` (`dispatcher.py`) builds the classifier's input from two
helpers, both scoped narrowly on purpose:
- `_last_user_text(messages)` (`dispatcher.py:1734`) — the text of the
**last** `user`-role message only. Walks backwards and returns on the
first match; never touches earlier turns.
- `_previous_context(messages)` (`dispatcher.py:1749`) — the text of the
**nearest preceding `assistant` message**, truncated to 200 characters.
Explicitly skips `system`, `user`, and `tool` roles "so that tool output
or the system prompt never contaminates the framing signal" (its own
docstring). Not the whole history — one turn of lookback, by design.
Both are recomputed **fresh on every request** — there is no session-level
memoization of the classification input today (that's exactly the gap
[`session-classification-cache-ttl.md`](session-classification-cache-ttl.md)
proposes closing, but for the *output* — the category/tier decision — not
the input).
## What this means for category drift
This is already, incidentally, LLMRouter's `current_only` mode plus a
one-turn lookback — and it's the reason category drift within a long
session is *not* currently a problem: every message gets classified from
near-term signal (the current message and the immediately preceding reply),
so a session moving from `general_chat` to `debugging` over 40 turns tracks
that drift on the very next message, for free. `full_context` (feeding the
whole conversation) would not improve this — it would make classification
slower (more tokens through `max_input_chars` clamping) and arguably noisier
(early, no-longer-relevant turns diluting the signal), for no accuracy
benefit given the categories this project tracks are about the *current*
task, not the session's history as a whole.
## The one place this now matters: the session cache
The classifier-input-scope property described here is what
`session-classification-cache-ttl.md` explicitly trades away, on purpose,
for latency. Once that cache ships, a "cached" decision no longer re-derives
category from the current message at all — it reuses whatever was true up
to `staleness_minutes` ago. That's a deliberate, bounded trade, not a
regression introduced silently; flagging it here so the connection is
on record in both documents.
## Recommendation
No action needed on classifier input scope itself — it already does the
right thing, and LLMRouter's `full_context`/`retrieval` modes don't apply
here (this project's categories are per-current-task, not
per-conversation-history, and there's no embedding/retrieval infrastructure
to reuse for `retrieval` mode even if it were wanted). Retire this as
"checked, no gap found" rather than carrying it forward as an open question.