classify_zero_shot embedded the RAW task text: a prose instruction wrapped in realistic agent-session noise (fenced code block, "Tool result:" line, <system-reminder> tag) mean-pooled the noise tokens with equal weight, and the verdict scored the noise. Confirmed on the real BAAI/bge-large-en-v1.5: 8 clean one-sentence tasks scored 8/8, the same instructions wrapped scored 2/8 on SHORT inputs far under any truncation limit — a second bug, distinct from the fixed 512-token collapse; tail-biased windowing alone also measured 2/8 on long noisy pairs, so the fit does not subsume isolation. New _isolate_task_text() (pure, stdlib-only re, deterministic) strips fenced code blocks, tool result/output/call lines, and closed <system-reminder> spans before the embed pass; classify_zero_shot now runs isolate -> fit -> prefix. Measured decisions: pure removal beats '[code]'/'[elided]' placeholders (8/8 vs 7/8, 6/8 on long noisy pairs — the placeholder token itself pulls toward code categories); inline code spans stay (file names in instructions are signal); a 20%-ratio floor guard measured harmful (3/8 — it reverts exactly the short noisy inputs) in favor of an absolute 24-char floor that only falls back on near-all-code inputs. Post-isolation: 8/8 on short and long noisy pairs, clean-vs-noisy pair consistency 8/8 + 8/8; code-grounded instructions improved 2/4 -> 3/4 (residual miss is description similarity, not noise). Regression: a noise tripwire (same instruction bare vs wrapped must classify identically) fails against pre-isolation code (empirically confirmed: docs_writing -> coding_general) and passes after; the existing truncation tripwires still pass. CLAUDE.md open item #5 updated to reflect what is built and what is not (attention-masking de-weighting, unclosed tags / un-fenced diff hunks, grounded-task residual). Full suite 2170 green. Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
45 KiB
45 KiB