plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
68 lines
3.4 KiB
Markdown
68 lines
3.4 KiB
Markdown
# Finding: classifier input scope (LLMRouter "query modes" check)
|
|
|
|
Status: done -- classifier.max_input_chars
|
|
|
|
**Origin.** LLMRouter's chat interface offers three query modes —
|
|
`current_only`, `full_context` (all history), `retrieval` (top-k similar
|
|
past queries). Before considering porting any of that, the actual question
|
|
was narrower and answerable by reading the code: does this project's
|
|
classifier already read the whole conversation, or only the tail? If only
|
|
the tail, a long session that drifts category (chat → debugging) could be
|
|
mis-tiered from a stale early read. This is a verification finding, not a
|
|
design proposal — no code changes proposed here.
|
|
|
|
---
|
|
|
|
## What the code actually does
|
|
|
|
`chat_completions` (`dispatcher.py`) builds the classifier's input from two
|
|
helpers, both scoped narrowly on purpose:
|
|
|
|
- `_last_user_text(messages)` (`dispatcher.py:1734`) — the text of the
|
|
**last** `user`-role message only. Walks backwards and returns on the
|
|
first match; never touches earlier turns.
|
|
- `_previous_context(messages)` (`dispatcher.py:1749`) — the text of the
|
|
**nearest preceding `assistant` message**, truncated to 200 characters.
|
|
Explicitly skips `system`, `user`, and `tool` roles "so that tool output
|
|
or the system prompt never contaminates the framing signal" (its own
|
|
docstring). Not the whole history — one turn of lookback, by design.
|
|
|
|
Both are recomputed **fresh on every request** — there is no session-level
|
|
memoization of the classification input today (that's exactly the gap
|
|
[`session-classification-cache-ttl.md`](session-classification-cache-ttl.md)
|
|
proposes closing, but for the *output* — the category/tier decision — not
|
|
the input).
|
|
|
|
## What this means for category drift
|
|
|
|
This is already, incidentally, LLMRouter's `current_only` mode plus a
|
|
one-turn lookback — and it's the reason category drift within a long
|
|
session is *not* currently a problem: every message gets classified from
|
|
near-term signal (the current message and the immediately preceding reply),
|
|
so a session moving from `general_chat` to `debugging` over 40 turns tracks
|
|
that drift on the very next message, for free. `full_context` (feeding the
|
|
whole conversation) would not improve this — it would make classification
|
|
slower (more tokens through `max_input_chars` clamping) and arguably noisier
|
|
(early, no-longer-relevant turns diluting the signal), for no accuracy
|
|
benefit given the categories this project tracks are about the *current*
|
|
task, not the session's history as a whole.
|
|
|
|
## The one place this now matters: the session cache
|
|
|
|
The classifier-input-scope property described here is what
|
|
`session-classification-cache-ttl.md` explicitly trades away, on purpose,
|
|
for latency. Once that cache ships, a "cached" decision no longer re-derives
|
|
category from the current message at all — it reuses whatever was true up
|
|
to `staleness_minutes` ago. That's a deliberate, bounded trade, not a
|
|
regression introduced silently; flagging it here so the connection is
|
|
on record in both documents.
|
|
|
|
## Recommendation
|
|
|
|
No action needed on classifier input scope itself — it already does the
|
|
right thing, and LLMRouter's `full_context`/`retrieval` modes don't apply
|
|
here (this project's categories are per-current-task, not
|
|
per-conversation-history, and there's no embedding/retrieval infrastructure
|
|
to reuse for `retrieval` mode even if it were wanted). Retire this as
|
|
"checked, no gap found" rather than carrying it forward as an open question.
|