classify_zero_shot embedded the RAW task text: a prose instruction
wrapped in realistic agent-session noise (fenced code block, "Tool
result:" line, <system-reminder> tag) mean-pooled the noise tokens
with equal weight, and the verdict scored the noise. Confirmed on the
real BAAI/bge-large-en-v1.5: 8 clean one-sentence tasks scored 8/8,
the same instructions wrapped scored 2/8 on SHORT inputs far under any
truncation limit — a second bug, distinct from the fixed 512-token
collapse; tail-biased windowing alone also measured 2/8 on long noisy
pairs, so the fit does not subsume isolation.
New _isolate_task_text() (pure, stdlib-only re, deterministic) strips
fenced code blocks, tool result/output/call lines, and closed
<system-reminder> spans before the embed pass; classify_zero_shot now
runs isolate -> fit -> prefix. Measured decisions: pure removal beats
'[code]'/'[elided]' placeholders (8/8 vs 7/8, 6/8 on long noisy pairs
— the placeholder token itself pulls toward code categories); inline
code spans stay (file names in instructions are signal); a 20%-ratio
floor guard measured harmful (3/8 — it reverts exactly the short noisy
inputs) in favor of an absolute 24-char floor that only falls back on
near-all-code inputs. Post-isolation: 8/8 on short and long noisy
pairs, clean-vs-noisy pair consistency 8/8 + 8/8; code-grounded
instructions improved 2/4 -> 3/4 (residual miss is description
similarity, not noise).
Regression: a noise tripwire (same instruction bare vs wrapped must
classify identically) fails against pre-isolation code (empirically
confirmed: docs_writing -> coding_general) and passes after; the
existing truncation tripwires still pass. CLAUDE.md open item #5
updated to reflect what is built and what is not (attention-masking
de-weighting, unclosed tags / un-fenced diff hunks, grounded-task
residual). Full suite 2170 green.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The token-budget fix (PR parent commit) addresses the truncation-on-long-inputs collapse but does not address the separate noise-isolation issue: fenced code blocks and tool-call/tool-result wrappers bias the embedding toward coding categories even when the instruction is clearly about docs_writing or other non-coding categories. Recorded as open item #5 in the What's NOT built yet section, not silently deferred.
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
classify_zero_shot's final tokenizer call passed truncation=True with no
max_length, so HuggingFace silently right-truncated to model_max_length
(512 for the BGE models) and kept the HEAD. For agent-session prompts
the head is shared boilerplate and the task-specific content lives in
the tail, so every long request scored an effectively identical prefix:
production returned (diff_checking, 0.192363) for six distinct
6.4k-6.6k-char requests for 20+ minutes.
New _fit_task_to_token_budget() selects the surviving window on real
token ids (no chars-per-token math): usable budget = model_max_length
- prefix tokens (8 for the BGE instruction) - num_special_tokens_to_add
(2); over-budget tasks keep a tail-biased head+tail slice (25/75,
rationale in module comments) spliced with the clamp_for_classifier-
style elision marker, decoded back to text and re-tokenized under an
8-token drift margin. Sentinel/missing model_max_length falls back to
512, logged once. No dispatcher changes, no tokenizer state mutation,
no new dependencies; the emergency confidence_threshold lowering stays
untouched (re-tuning is a documented Phase 2 follow-up).
Verified against the production BertTokenizerFast snapshot: an
8596-char task (1789 bare tokens) fits as 504 prefixed tokens <= 512
with head, tail and marker intact, and a 25k-char raw task likewise
fits with both ends preserved; the pre-fix path scored only head
boilerplate (last surviving token: 'and').
The regression tripwire in tests/test_local_encoder.py fails against
pre-fix code (empirically confirmed) and passes after; full suite 2160
green. Also fixes the test fakes' softmax/argmax, which read only the
first value of a (1,n) row tensor and pinned every multi-category
classify to (category_0, 1.0).
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>