HF's zero-shot pipeline scores each candidate label against the input
via a hypothesis template ("This example is {}."), so the label
itself needs to read as natural language for entailment scoring to
work -- feeding it a raw config identifier like "tool_use_agentic" or
"diff_checking" asks the model to judge "This example is
tool_use_agentic.", not a sentence its NLI training ever saw.
Measured live 2026-09-06 against bart-large-mnli with the raw labels:
4 of 9 test prompts landed on the wrong category, every miss also
scoring low (<=0.28) -- confidence and correctness tracked each other,
but the raw labels weren't giving the model enough to work with.
Two changes:
- New _CATEGORY_DESCRIPTIONS maps each config category to a natural-
language description, used as the actual candidate label; the
winning description maps back to its category id for the return
value. A category missing from the map falls back to its raw
string rather than raising, so a newly-added config category
degrades gracefully instead of crashing.
- multi_label=True: the pipeline's default (single-label) normalizes
every candidate's score to sum to 1, so a genuinely good match still
gets dragged down whenever another category is also plausible.
Scoring independently lets a clear match score high on its own
terms.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U