plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
3.4 KiB
Finding: classifier input scope (LLMRouter "query modes" check)
Status: done -- classifier.max_input_chars
Origin. LLMRouter's chat interface offers three query modes —
current_only, full_context (all history), retrieval (top-k similar
past queries). Before considering porting any of that, the actual question
was narrower and answerable by reading the code: does this project's
classifier already read the whole conversation, or only the tail? If only
the tail, a long session that drifts category (chat → debugging) could be
mis-tiered from a stale early read. This is a verification finding, not a
design proposal — no code changes proposed here.
What the code actually does
chat_completions (dispatcher.py) builds the classifier's input from two
helpers, both scoped narrowly on purpose:
_last_user_text(messages)(dispatcher.py:1734) — the text of the lastuser-role message only. Walks backwards and returns on the first match; never touches earlier turns._previous_context(messages)(dispatcher.py:1749) — the text of the nearest precedingassistantmessage, truncated to 200 characters. Explicitly skipssystem,user, andtoolroles "so that tool output or the system prompt never contaminates the framing signal" (its own docstring). Not the whole history — one turn of lookback, by design.
Both are recomputed fresh on every request — there is no session-level
memoization of the classification input today (that's exactly the gap
session-classification-cache-ttl.md
proposes closing, but for the output — the category/tier decision — not
the input).
What this means for category drift
This is already, incidentally, LLMRouter's current_only mode plus a
one-turn lookback — and it's the reason category drift within a long
session is not currently a problem: every message gets classified from
near-term signal (the current message and the immediately preceding reply),
so a session moving from general_chat to debugging over 40 turns tracks
that drift on the very next message, for free. full_context (feeding the
whole conversation) would not improve this — it would make classification
slower (more tokens through max_input_chars clamping) and arguably noisier
(early, no-longer-relevant turns diluting the signal), for no accuracy
benefit given the categories this project tracks are about the current
task, not the session's history as a whole.
The one place this now matters: the session cache
The classifier-input-scope property described here is what
session-classification-cache-ttl.md explicitly trades away, on purpose,
for latency. Once that cache ships, a "cached" decision no longer re-derives
category from the current message at all — it reuses whatever was true up
to staleness_minutes ago. That's a deliberate, bounded trade, not a
regression introduced silently; flagging it here so the connection is
on record in both documents.
Recommendation
No action needed on classifier input scope itself — it already does the
right thing, and LLMRouter's full_context/retrieval modes don't apply
here (this project's categories are per-current-task, not
per-conversation-history, and there's no embedding/retrieval infrastructure
to reuse for retrieval mode even if it were wanted). Retire this as
"checked, no gap found" rather than carrying it forward as an open question.