fix(local-encoder): token-budget fit + agent-session noise isolation for classify_zero_shot #98
Reference in New Issue
Block a user
Delete Branch "fix/local-encoder-token-budget"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Summary
Two distinct bugs in
classify_zero_shot, one PR. Both confirmed on the realBAAI/bge-large-en-v1.5against the production BertTokenizerFast snapshot — measured numbers throughout, not adjectives.Bug 1 — silent HF 512-token truncation (production incident, 2026-09-18)
classify_zero_shotpassedtruncation=Truewith nomax_length, so HF silently right-truncated tomodel_max_length(512 for BERT-family BGE models). Withtruncation_side='right', only the head survived — the shared boilerplate of agent-session prompts. Production signature: six distinct 6.4k–6.6k-char requests all returning(diff_checking, 0.192363)for 20+ minutes before the emergencyconfidence_threshold=0.12was applied.Bug 2 — noise dominance (separate mechanism, short inputs too)
The same
docs_writinginstruction wrapped in realistic agent-session noise (fenced code block + "Tool result:" line +<system-reminder>tag) flips tocoding_generalon SHORT inputs far under any truncation limit. Tail-biased windowing alone measured 2/8 on long noisy pairs too — the window keeps code and instruction mixed. This is a second, distinct bug that the truncation fix does not touch; production task text is raw user_content full of tool call/result formatting, file contents, code blocks, and reminder tags.Fix 1 — token-budget fit (
_fit_task_to_token_budget)New helper selects the surviving window on real token ids (no chars-per-token math): usable budget =
model_max_length − prefix tokens − num_special_tokens_to_add(). Over-budget tasks keep a tail-biased head+tail slice (25/75, rationale: the raw task can carry its instruction at either end) spliced with the clamp_for_classifier-style elision marker, decoded back to text and re-tokenized under an 8-token drift margin. Sentinel/missingmodel_max_lengthfalls back to 512, logged once via a module-level flag. Degenerate-window path returns the task unchanged. No tokenizer state mutation (the cached tokenizer is process-wide).classify_zero_shotsignature unchanged; dispatcher untouched; the emergency confidence threshold stays at 0.12.Verification numbers (BAAI/bge-large-en-v1.5, BertTokenizerFast, prod snapshot):
model_max_lengthpassages:)num_special_tokens_to_add()'and'— pure head boilerplate; tail never reached modelFix 2 — noise isolation (
_isolate_task_text)Pure, stdlib-only (
re), deterministic helper — strip first, then fit → prefix → tokenize. Strips:'[code]'placeholder scored 7/8 and 6/8 on long pairs — the placeholder token itself pulls toward code categories and costs accuracy, so the winner is removal.^\s*tool[\s_-]*(?:result|output|call)[\s_-]*:(case-insensitive, must START with the marker — prose merely mentioning one survives, verified).<system-reminder>spans (closed-tag regex with backreference and lazy DOTALL; unclosed tags left in place — documented residual).Real-model evidence (BAAI/bge-large-en-v1.5, same snapshot, same 8-category set)
file_summarizationOrchestrator's independent real-model confirmation (ran after the commit)
docs_writing→docs_writingat 0.186482; wrapped shortdocs_writing→docs_writingat the identical 0.186482 (noise stripped to identical embedded text).docs_writingat the identical 0.186482 (isolation reduced 6248→232 chars; fit no-op — isolated text fits outright).docs_writingat the identical 0.186482.Known residuals (stated honestly, tracked in CLAUDE.md item #5)
tool_use_agentic— a description-similarity family, same mechanism as the grounded-task residual. Addressing it is attention-mask de-weighting, which is documented as not built.<system-reminder>tags and un-fenced diff hunks are left in place (only closed-tag spans, fenced blocks, and marker-prefixed lines are stripped by the current heuristics).refactor+pasted code →debuggingat 0.19 — present in baseline too (asfile_summarization), so it is description similarity, not noise dominance.Tests
Two offline fail-before/pass-after tripwires in
tests/test_local_encoder.py:SessionStyleTokenizer+ContentSignModelfaithfully emulate the HF truncation contract. Empirically fails against pre-fix code.NoiseDominantTokenizer(words in the noise fixture's vocabulary get ≥1000 ids) +NoiseDominantModel(embedding = mean of +1/−1 per token class) reproduces the mean-pool noise dominance in miniature. Empirically fails against pre-isolation module with the exactdocs_writing→coding_generalflip.[code]), floor-guard fallback on all-code tasks, idempotence, unterminated-fence leading-instruction survival.Bonus test-side fix (first commit): the test doubles'
_softmax/_argmaxread only the first value of a(1,n)row tensor — pinned every multi-category mock classify to(category_0, 1.0). Fixed via_flat_values().CLAUDE.md
Open item #5 was recorded by the operator's docs commit (
057776a) and rewritten in the noise-isolation commit (bacfd0e) to reflect what is built and the honest residuals (attention-mask de-weighting, unclosed tags / un-fenced diff hunks, grounded-task near-category miss).Follow-up flags (Phase 2)
confidence_thresholdre-tune: currently emergency-lowered from 0.4 → 0.12 inconfig.local.yamlto stop fall-through during the incident. Now that both bugs are fixed, discrimination is restored and the threshold likely needs to go back UP. Operator task via the admin portal; deliberately no number guessed here.Deploy: after merge,
systemctl --user restart llm-router.servicepicks the fix up (dispatcher code does not auto-reload).Observation: Phase 2 should watch for the documented residuals (tool-flavored prose pulling toward tool_use_agentic, unclosed tags still passing noise) on real traffic.
fix(local-encoder): fit task to the model's real token window before encodingto fix(local-encoder): token-budget fit + agent-session noise isolation for classify_zero_shot