fix(local-encoder): token-budget fit + agent-session noise isolation for classify_zero_shot #98

Merged
alee merged 3 commits from fix/local-encoder-token-budget into main 2026-09-19 15:57:15 +00:00
Owner

Summary

Two distinct bugs in classify_zero_shot, one PR. Both confirmed on the real BAAI/bge-large-en-v1.5 against the production BertTokenizerFast snapshot — measured numbers throughout, not adjectives.

Bug 1 — silent HF 512-token truncation (production incident, 2026-09-18)

classify_zero_shot passed truncation=True with no max_length, so HF silently right-truncated to model_max_length (512 for BERT-family BGE models). With truncation_side='right', only the head survived — the shared boilerplate of agent-session prompts. Production signature: six distinct 6.4k–6.6k-char requests all returning (diff_checking, 0.192363) for 20+ minutes before the emergency confidence_threshold=0.12 was applied.

Bug 2 — noise dominance (separate mechanism, short inputs too)

The same docs_writing instruction wrapped in realistic agent-session noise (fenced code block + "Tool result:" line + <system-reminder> tag) flips to coding_general on SHORT inputs far under any truncation limit. Tail-biased windowing alone measured 2/8 on long noisy pairs too — the window keeps code and instruction mixed. This is a second, distinct bug that the truncation fix does not touch; production task text is raw user_content full of tool call/result formatting, file contents, code blocks, and reminder tags.

Fix 1 — token-budget fit (_fit_task_to_token_budget)

New helper selects the surviving window on real token ids (no chars-per-token math): usable budget = model_max_length − prefix tokens − num_special_tokens_to_add(). Over-budget tasks keep a tail-biased head+tail slice (25/75, rationale: the raw task can carry its instruction at either end) spliced with the clamp_for_classifier-style elision marker, decoded back to text and re-tokenized under an 8-token drift margin. Sentinel/missing model_max_length falls back to 512, logged once via a module-level flag. Degenerate-window path returns the task unchanged. No tokenizer state mutation (the cached tokenizer is process-wide). classify_zero_shot signature unchanged; dispatcher untouched; the emergency confidence threshold stays at 0.12.

Verification numbers (BAAI/bge-large-en-v1.5, BertTokenizerFast, prod snapshot):

Measurement Value
model_max_length 512
BGE instruction prefix 8 tokens (wordpiece splits passages:)
num_special_tokens_to_add() 2
Usable content budget 502
8596-char task → 1789 bare tokens → final prefixed 504 ≤ 512 (head, tail, marker intact)
25,624-char raw task → final prefixed 504 ≤ 512 (both ends preserved)
Pre-fix last surviving token on 8596-char input 'and' — pure head boilerplate; tail never reached model

Fix 2 — noise isolation (_isolate_task_text)

Pure, stdlib-only (re), deterministic helper — strip first, then fit → prefix → tokenize. Strips:

  • Fenced code blocks stripped outright. Measured decision on the real model across 8 categories × short+long noisy pairs: pure removal scored 8/8 on both, while '[code]' placeholder scored 7/8 and 6/8 on long pairs — the placeholder token itself pulls toward code categories and costs accuracy, so the winner is removal.
  • Tool result/output/call lines line-level: ^\s*tool[\s_-]*(?:result|output|call)[\s_-]*: (case-insensitive, must START with the marker — prose merely mentioning one survives, verified).
  • Closed <system-reminder> spans (closed-tag regex with backreference and lazy DOTALL; unclosed tags left in place — documented residual).
  • Inline code spans KEPT (file names and commands in instructions are signal, not noise — verified).
  • Unterminated fence (````.*\Z` after pair removal): eats to EOS — instruction before the fence survives; instruction after + all-code → floor guard reverts to original (equals baseline, never worse).
  • Floor guard: absolute 24-char minimum kept text only — a 20%-ratio guard was measured HARMFUL (3/8: it reverts exactly the short noisy inputs isolation exists to fix).
  • Operator's explicit demand answered: tail-biased windowing ALONE does NOT fix long noisy pairs (2/8). Isolation is a genuinely separate mechanism.

Real-model evidence (BAAI/bge-large-en-v1.5, same snapshot, same 8-category set)

Variant Baseline (fit-only) Post-isolation Pair consistency
Clean (8 one-sentence tasks) 8/8 8/8 —
Noisy short (fenced code + tool result + reminder, instruction at end) 2/8 → every miss file_summarization 8/8 8/8
Noisy long (same noise + 240-line code block, 1960+ tokens) 2/8 (windowing alone insufficient) 8/8 8/8
Code-grounded (instruction + genuinely relevant pasted code) 2/4 3/4 —

Orchestrator's independent real-model confirmation (ran after the commit)

  • Operator's exact repro pair: clean docs_writing → docs_writing at 0.186482; wrapped short docs_writing → docs_writing at the identical 0.186482 (noise stripped to identical embedded text).
  • 6.2k-char prompt headed by 60 fenced JSON tool schemas plus noise → docs_writing at the identical 0.186482 (isolation reduced 6248→232 chars; fit no-op — isolated text fits outright).
  • Neutral-prose-boilerplate long variant (no code/fence noise) → docs_writing at the identical 0.186482.

Known residuals (stated honestly, tracked in CLAUDE.md item #5)

  • Repeated tool-flavored prose (e.g. "assume tools available" ×300) can still lexically pull the pooled embedding toward tool_use_agentic — a description-similarity family, same mechanism as the grounded-task residual. Addressing it is attention-mask de-weighting, which is documented as not built.
  • Unclosed <system-reminder> tags and un-fenced diff hunks are left in place (only closed-tag spans, fenced blocks, and marker-prefixed lines are stripped by the current heuristics).
  • Grounded-task near-category miss: refactor+pasted code → debugging at 0.19 — present in baseline too (as file_summarization), so it is description similarity, not noise dominance.
  • Cosmetic: transformers emits its "token indices longer than max" stderr warning when pre-fit isolated text exceeds the window (i.e. exactly when the fit is about to handle it) — log noise only; suppressing it would require mutating global transformers logging state, deliberately not done.

Tests

Two offline fail-before/pass-after tripwires in tests/test_local_encoder.py:

  • Truncation collapse: two long shared-prefix prompts differing only in tail → must NOT return identical verdicts. Fake SessionStyleTokenizer + ContentSignModel faithfully emulate the HF truncation contract. Empirically fails against pre-fix code.
  • Noise dominance: same short prose instruction, bare vs wrapped in noise shapes → must classify to the SAME category. Fake NoiseDominantTokenizer (words in the noise fixture's vocabulary get ≥1000 ids) + NoiseDominantModel (embedding = mean of +1/−1 per token class) reproduces the mean-pool noise dominance in miniature. Empirically fails against pre-isolation module with the exact docs_writing→coding_general flip.
  • Helper unit tests: fenced blocks stripped, tool lines stripped, reminder content stripped, trailing instruction preserved, inline code spans kept, removal-vs-placeholder decision locked (no [code]), floor-guard fallback on all-code tasks, idempotence, unterminated-fence leading-instruction survival.
  • Both existing truncation helper tests and the tripwire still pass.
  • Full suite: 2170 passed, 0 failed, 94.11s (PYTHONPATH=src shared venv, fully offline).

Bonus test-side fix (first commit): the test doubles' _softmax/_argmax read only the first value of a (1,n) row tensor — pinned every multi-category mock classify to (category_0, 1.0). Fixed via _flat_values().

CLAUDE.md

Open item #5 was recorded by the operator's docs commit (057776a) and rewritten in the noise-isolation commit (bacfd0e) to reflect what is built and the honest residuals (attention-mask de-weighting, unclosed tags / un-fenced diff hunks, grounded-task near-category miss).

Follow-up flags (Phase 2)

  1. confidence_threshold re-tune: currently emergency-lowered from 0.4 → 0.12 in config.local.yaml to stop fall-through during the incident. Now that both bugs are fixed, discrimination is restored and the threshold likely needs to go back UP. Operator task via the admin portal; deliberately no number guessed here.

  2. Deploy: after merge, systemctl --user restart llm-router.service picks the fix up (dispatcher code does not auto-reload).

  3. Observation: Phase 2 should watch for the documented residuals (tool-flavored prose pulling toward tool_use_agentic, unclosed tags still passing noise) on real traffic.

## Summary Two distinct bugs in `classify_zero_shot`, one PR. Both confirmed on the real `BAAI/bge-large-en-v1.5` against the production BertTokenizerFast snapshot — measured numbers throughout, not adjectives. ### Bug 1 — silent HF 512-token truncation (production incident, 2026-09-18) `classify_zero_shot` passed `truncation=True` with no `max_length`, so HF silently right-truncated to `model_max_length` (512 for BERT-family BGE models). With `truncation_side='right'`, only the head survived — the shared boilerplate of agent-session prompts. Production signature: six distinct 6.4k–6.6k-char requests all returning `(diff_checking, 0.192363)` for 20+ minutes before the emergency `confidence_threshold=0.12` was applied. ### Bug 2 — noise dominance (separate mechanism, short inputs too) The same `docs_writing` instruction wrapped in realistic agent-session noise (fenced code block + "Tool result:" line + `<system-reminder>` tag) flips to `coding_general` on SHORT inputs far under any truncation limit. Tail-biased windowing alone measured 2/8 on long noisy pairs too — the window keeps code and instruction mixed. This is a second, distinct bug that the truncation fix does not touch; production task text is raw user_content full of tool call/result formatting, file contents, code blocks, and reminder tags. ## Fix 1 — token-budget fit (`_fit_task_to_token_budget`) New helper selects the surviving window on real token ids (no chars-per-token math): usable budget = `model_max_length − prefix tokens − num_special_tokens_to_add()`. Over-budget tasks keep a tail-biased head+tail slice (25/75, rationale: the raw task can carry its instruction at either end) spliced with the clamp_for_classifier-style elision marker, decoded back to text and re-tokenized under an 8-token drift margin. Sentinel/missing `model_max_length` falls back to 512, logged once via a module-level flag. Degenerate-window path returns the task unchanged. No tokenizer state mutation (the cached tokenizer is process-wide). `classify_zero_shot` signature unchanged; dispatcher untouched; the emergency confidence threshold stays at 0.12. **Verification numbers (BAAI/bge-large-en-v1.5, BertTokenizerFast, prod snapshot):** | Measurement | Value | |---|---| | `model_max_length` | 512 | | BGE instruction prefix | 8 tokens (wordpiece splits `passages:`) | | `num_special_tokens_to_add()` | 2 | | **Usable content budget** | **502** | | 8596-char task → 1789 bare tokens → final prefixed | 504 ≤ 512 (head, tail, marker intact) | | 25,624-char raw task → final prefixed | 504 ≤ 512 (both ends preserved) | | Pre-fix last surviving token on 8596-char input | `'and'` — pure head boilerplate; tail never reached model | ## Fix 2 — noise isolation (`_isolate_task_text`) Pure, stdlib-only (`re`), deterministic helper — strip first, then fit → prefix → tokenize. Strips: - **Fenced code blocks** stripped outright. Measured decision on the real model across 8 categories × short+long noisy pairs: pure removal scored 8/8 on both, while `'[code]'` placeholder scored 7/8 and 6/8 on long pairs — the placeholder token itself pulls toward code categories and costs accuracy, so the winner is removal. - **Tool result/output/call lines** line-level: `^\s*tool[\s_-]*(?:result|output|call)[\s_-]*:` (case-insensitive, must START with the marker — prose merely mentioning one survives, verified). - **Closed `<system-reminder>` spans** (closed-tag regex with backreference and lazy DOTALL; unclosed tags left in place — documented residual). - **Inline code spans KEPT** (file names and commands in instructions are signal, not noise — verified). - **Unterminated fence** (````.*\Z` after pair removal): eats to EOS — instruction before the fence survives; instruction after + all-code → floor guard reverts to original (equals baseline, never worse). - **Floor guard**: absolute 24-char minimum kept text only — a 20%-ratio guard was measured HARMFUL (3/8: it reverts exactly the short noisy inputs isolation exists to fix). - **Operator's explicit demand answered**: tail-biased windowing ALONE does NOT fix long noisy pairs (2/8). Isolation is a genuinely separate mechanism. ### Real-model evidence (BAAI/bge-large-en-v1.5, same snapshot, same 8-category set) | Variant | Baseline (fit-only) | Post-isolation | Pair consistency | |---|---|---|---| | Clean (8 one-sentence tasks) | **8/8** | 8/8 | — | | Noisy short (fenced code + tool result + reminder, instruction at end) | **2/8** → every miss `file_summarization` | **8/8** | **8/8** | | Noisy long (same noise + 240-line code block, 1960+ tokens) | **2/8** (windowing alone insufficient) | **8/8** | **8/8** | | Code-grounded (instruction + genuinely relevant pasted code) | **2/4** | **3/4** | — | ### Orchestrator's independent real-model confirmation (ran after the commit) - Operator's exact repro pair: clean `docs_writing` → `docs_writing` at 0.186482; wrapped short `docs_writing` → `docs_writing` at the **identical** 0.186482 (noise stripped to identical embedded text). - 6.2k-char prompt headed by 60 fenced JSON tool schemas plus noise → `docs_writing` at the identical 0.186482 (isolation reduced 6248→232 chars; fit no-op — isolated text fits outright). - Neutral-prose-boilerplate long variant (no code/fence noise) → `docs_writing` at the identical 0.186482. ### Known residuals (stated honestly, tracked in CLAUDE.md item #5) - **Repeated tool-flavored prose** (e.g. "assume tools available" ×300) can still lexically pull the pooled embedding toward `tool_use_agentic` — a description-similarity family, same mechanism as the grounded-task residual. Addressing it is attention-mask de-weighting, which is documented as not built. - **Unclosed `<system-reminder>` tags** and **un-fenced diff hunks** are left in place (only closed-tag spans, fenced blocks, and marker-prefixed lines are stripped by the current heuristics). - **Grounded-task near-category miss**: `refactor`+pasted code → `debugging` at 0.19 — present in baseline too (as `file_summarization`), so it is description similarity, not noise dominance. - **Cosmetic**: transformers emits its "token indices longer than max" stderr warning when pre-fit isolated text exceeds the window (i.e. exactly when the fit is about to handle it) — log noise only; suppressing it would require mutating global transformers logging state, deliberately not done. ## Tests Two offline fail-before/pass-after tripwires in `tests/test_local_encoder.py`: - **Truncation collapse**: two long shared-prefix prompts differing only in tail → must NOT return identical verdicts. Fake `SessionStyleTokenizer` + `ContentSignModel` faithfully emulate the HF truncation contract. Empirically fails against pre-fix code. - **Noise dominance**: same short prose instruction, bare vs wrapped in noise shapes → must classify to the SAME category. Fake `NoiseDominantTokenizer` (words in the noise fixture's vocabulary get ≥1000 ids) + `NoiseDominantModel` (embedding = mean of +1/−1 per token class) reproduces the mean-pool noise dominance in miniature. Empirically fails against pre-isolation module with the exact `docs_writing→coding_general` flip. - Helper unit tests: fenced blocks stripped, tool lines stripped, reminder content stripped, trailing instruction preserved, inline code spans kept, removal-vs-placeholder decision locked (no `[code]`), floor-guard fallback on all-code tasks, idempotence, unterminated-fence leading-instruction survival. - Both existing truncation helper tests and the tripwire still pass. - **Full suite: 2170 passed, 0 failed, 94.11s** (PYTHONPATH=src shared venv, fully offline). Bonus test-side fix (first commit): the test doubles' `_softmax`/`_argmax` read only the first value of a `(1,n)` row tensor — pinned every multi-category mock classify to `(category_0, 1.0)`. Fixed via `_flat_values()`. ## CLAUDE.md Open item #5 was recorded by the operator's docs commit (`057776a`) and rewritten in the noise-isolation commit (`bacfd0e`) to reflect what is built and the honest residuals (attention-mask de-weighting, unclosed tags / un-fenced diff hunks, grounded-task near-category miss). ## Follow-up flags (Phase 2) 1. **`confidence_threshold` re-tune**: currently emergency-lowered from 0.4 → 0.12 in `config.local.yaml` to stop fall-through during the incident. Now that both bugs are fixed, discrimination is restored and the threshold likely needs to go back UP. Operator task via the admin portal; deliberately no number guessed here. 2. **Deploy**: after merge, `systemctl --user restart llm-router.service` picks the fix up (dispatcher code does not auto-reload). 3. **Observation**: Phase 2 should watch for the documented residuals (tool-flavored prose pulling toward tool_use_agentic, unclosed tags still passing noise) on real traffic.
alee added 2 commits 2026-09-19 04:24:41 +00:00
classify_zero_shot's final tokenizer call passed truncation=True with no
max_length, so HuggingFace silently right-truncated to model_max_length
(512 for the BGE models) and kept the HEAD. For agent-session prompts
the head is shared boilerplate and the task-specific content lives in
the tail, so every long request scored an effectively identical prefix:
production returned (diff_checking, 0.192363) for six distinct
6.4k-6.6k-char requests for 20+ minutes.

New _fit_task_to_token_budget() selects the surviving window on real
token ids (no chars-per-token math): usable budget = model_max_length
- prefix tokens (8 for the BGE instruction) - num_special_tokens_to_add
(2); over-budget tasks keep a tail-biased head+tail slice (25/75,
rationale in module comments) spliced with the clamp_for_classifier-
style elision marker, decoded back to text and re-tokenized under an
8-token drift margin. Sentinel/missing model_max_length falls back to
512, logged once. No dispatcher changes, no tokenizer state mutation,
no new dependencies; the emergency confidence_threshold lowering stays
untouched (re-tuning is a documented Phase 2 follow-up).

Verified against the production BertTokenizerFast snapshot: an
8596-char task (1789 bare tokens) fits as 504 prefixed tokens <= 512
with head, tail and marker intact, and a 25k-char raw task likewise
fits with both ends preserved; the pre-fix path scored only head
boilerplate (last surviving token: 'and').

The regression tripwire in tests/test_local_encoder.py fails against
pre-fix code (empirically confirmed) and passes after; full suite 2160
green. Also fixes the test fakes' softmax/argmax, which read only the
first value of a (1,n) row tensor and pinned every multi-category
classify to (category_0, 1.0).

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The token-budget fix (PR parent commit) addresses the truncation-on-long-inputs collapse but does not address the separate noise-isolation issue: fenced code blocks and tool-call/tool-result wrappers bias the embedding toward coding categories even when the instruction is clearly about docs_writing or other non-coding categories. Recorded as open item #5 in the What's NOT built yet section, not silently deferred.

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
alee added 1 commit 2026-09-19 04:59:58 +00:00
classify_zero_shot embedded the RAW task text: a prose instruction
wrapped in realistic agent-session noise (fenced code block, "Tool
result:" line, <system-reminder> tag) mean-pooled the noise tokens
with equal weight, and the verdict scored the noise. Confirmed on the
real BAAI/bge-large-en-v1.5: 8 clean one-sentence tasks scored 8/8,
the same instructions wrapped scored 2/8 on SHORT inputs far under any
truncation limit — a second bug, distinct from the fixed 512-token
collapse; tail-biased windowing alone also measured 2/8 on long noisy
pairs, so the fit does not subsume isolation.

New _isolate_task_text() (pure, stdlib-only re, deterministic) strips
fenced code blocks, tool result/output/call lines, and closed
<system-reminder> spans before the embed pass; classify_zero_shot now
runs isolate -> fit -> prefix. Measured decisions: pure removal beats
'[code]'/'[elided]' placeholders (8/8 vs 7/8, 6/8 on long noisy pairs
— the placeholder token itself pulls toward code categories); inline
code spans stay (file names in instructions are signal); a 20%-ratio
floor guard measured harmful (3/8 — it reverts exactly the short noisy
inputs) in favor of an absolute 24-char floor that only falls back on
near-all-code inputs. Post-isolation: 8/8 on short and long noisy
pairs, clean-vs-noisy pair consistency 8/8 + 8/8; code-grounded
instructions improved 2/4 -> 3/4 (residual miss is description
similarity, not noise).

Regression: a noise tripwire (same instruction bare vs wrapped must
classify identically) fails against pre-isolation code (empirically
confirmed: docs_writing -> coding_general) and passes after; the
existing truncation tripwires still pass. CLAUDE.md open item #5
updated to reflect what is built and what is not (attention-masking
de-weighting, unclosed tags / un-fenced diff hunks, grounded-task
residual). Full suite 2170 green.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
alee changed title from fix(local-encoder): fit task to the model's real token window before encoding to fix(local-encoder): token-budget fit + agent-session noise isolation for classify_zero_shot 2026-09-19 05:02:05 +00:00
alee merged commit d3dbb4a771 into main 2026-09-19 15:57:15 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#98