feat(classifier): natural-language category labels + multi_label for local_encoder #50

Merged
alee merged 1 commits from feat/encoder-natural-language-labels into main 2026-09-07 02:53:58 +00:00
Owner

HF's zero-shot pipeline scores each candidate label against the input via a hypothesis template ("This example is {}."), so the label itself needs to read as natural language for entailment scoring to work -- feeding it a raw config identifier like "tool_use_agentic" or "diff_checking" asks the model to judge "This example is tool_use_agentic.", not a sentence its NLI training ever saw.

Measured live 2026-09-06 on the same 9-prompt calibration set used to diagnose the original confidence-threshold incident (#47/#48):

before (raw labels) after (natural language + multi_label)
5/9 correct, confidence 0.15-0.90 8/9 correct, confidence 0.39-0.999

The one remaining miss (reasoning_math -> general_chat) scored 0.394 -- below the currently-configured confidence_threshold: 0.4, so it falls through to the safe fallback cascade rather than mis-routing.

Two changes:

  • _CATEGORY_DESCRIPTIONS maps each config category to a natural-language description, used as the actual candidate label; the winning description maps back to its category id. A category missing from the map falls back to its raw string rather than raising.
  • multi_label=True: the pipeline's default (single-label) normalizes every candidate's score to sum to 1, so a genuinely good match gets dragged down whenever another category is also plausible. Independent scoring lets a clear match score high on its own terms.

New regression tests confirmed failing on the unfixed code, passing after. Full suite: 1506 passed, 0 failures.

HF's zero-shot pipeline scores each candidate label against the input via a hypothesis template ("This example is {}."), so the label itself needs to read as natural language for entailment scoring to work -- feeding it a raw config identifier like "tool_use_agentic" or "diff_checking" asks the model to judge "This example is tool_use_agentic.", not a sentence its NLI training ever saw. **Measured live 2026-09-06** on the same 9-prompt calibration set used to diagnose the original confidence-threshold incident (#47/#48): | before (raw labels) | after (natural language + multi_label) | |---|---| | 5/9 correct, confidence 0.15-0.90 | 8/9 correct, confidence 0.39-0.999 | The one remaining miss (`reasoning_math` -> `general_chat`) scored 0.394 -- below the currently-configured `confidence_threshold: 0.4`, so it falls through to the safe fallback cascade rather than mis-routing. Two changes: - `_CATEGORY_DESCRIPTIONS` maps each config category to a natural-language description, used as the actual candidate label; the winning description maps back to its category id. A category missing from the map falls back to its raw string rather than raising. - `multi_label=True`: the pipeline's default (single-label) normalizes every candidate's score to sum to 1, so a genuinely good match gets dragged down whenever another category is also plausible. Independent scoring lets a clear match score high on its own terms. New regression tests confirmed failing on the unfixed code, passing after. Full suite: 1506 passed, 0 failures.
alee added 1 commit 2026-09-07 01:45:04 +00:00
HF's zero-shot pipeline scores each candidate label against the input
via a hypothesis template ("This example is {}."), so the label
itself needs to read as natural language for entailment scoring to
work -- feeding it a raw config identifier like "tool_use_agentic" or
"diff_checking" asks the model to judge "This example is
tool_use_agentic.", not a sentence its NLI training ever saw.

Measured live 2026-09-06 against bart-large-mnli with the raw labels:
4 of 9 test prompts landed on the wrong category, every miss also
scoring low (<=0.28) -- confidence and correctness tracked each other,
but the raw labels weren't giving the model enough to work with.

Two changes:
- New _CATEGORY_DESCRIPTIONS maps each config category to a natural-
  language description, used as the actual candidate label; the
  winning description maps back to its category id for the return
  value. A category missing from the map falls back to its raw
  string rather than raising, so a newly-added config category
  degrades gracefully instead of crashing.
- multi_label=True: the pipeline's default (single-label) normalizes
  every candidate's score to sum to 1, so a genuinely good match still
  gets dragged down whenever another category is also plausible.
  Scoring independently lets a clear match score high on its own
  terms.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
alee merged commit f66c8cbe66 into main 2026-09-07 02:53:58 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#50