fix(classifier): local_encoder default model is gated, switch to bart-large-mnli #48
@@ -867,7 +867,7 @@ function classifierModeFieldsHtml(mode, data) {
|
|||||||
<div class="row g-2">
|
<div class="row g-2">
|
||||||
<div class="col-md-6">
|
<div class="col-md-6">
|
||||||
<input class="form-control form-control-sm" id="classifier-encoder-model"
|
<input class="form-control form-control-sm" id="classifier-encoder-model"
|
||||||
placeholder="model (default: MoritzLaurer/deberta-v3-base-zeroshot-v2)"
|
placeholder="model (default: facebook/bart-large-mnli)"
|
||||||
value="${escapeHtml(enc.model || '')}">
|
value="${escapeHtml(enc.model || '')}">
|
||||||
</div>
|
</div>
|
||||||
<div class="col-md-3">
|
<div class="col-md-3">
|
||||||
|
|||||||
@@ -707,7 +707,7 @@ classifier:
|
|||||||
# Read only when mode: local_encoder. Every field has a default, so
|
# Read only when mode: local_encoder. Every field has a default, so
|
||||||
# `encoder: {}` is enough to opt in.
|
# `encoder: {}` is enough to opt in.
|
||||||
# encoder:
|
# encoder:
|
||||||
# model: MoritzLaurer/deberta-v3-base-zeroshot-v2
|
# model: facebook/bart-large-mnli
|
||||||
# device: cpu
|
# device: cpu
|
||||||
# confidence_threshold: 0.5
|
# confidence_threshold: 0.5
|
||||||
|
|
||||||
|
|||||||
@@ -213,10 +213,24 @@ rarer each time, but never structurally impossible, because a chat-completion
|
|||||||
model can always in principle spend its budget thinking instead of
|
model can always in principle spend its budget thinking instead of
|
||||||
answering. `classifier.mode: local_encoder` sidesteps the whole failure class
|
answering. `classifier.mode: local_encoder` sidesteps the whole failure class
|
||||||
instead of picking around it: a zero-shot NLI encoder
|
instead of picking around it: a zero-shot NLI encoder
|
||||||
(`classifier.encoder.model`, default `MoritzLaurer/deberta-v3-base-
|
(`classifier.encoder.model`, default `facebook/bart-large-mnli`) scores the
|
||||||
zeroshot-v2`) scores the task directly against `proficiency.categories` and
|
task directly against `proficiency.categories` and returns a label plus a
|
||||||
returns a label plus a confidence — there is no generation step, so there is
|
confidence — there is no generation step, so there is no trace to run away.
|
||||||
no trace to run away.
|
(Was `MoritzLaurer/deberta-v3-base-zeroshot-v2` — smaller, ~184M vs ~407M
|
||||||
|
params — until that repo started returning 401 on an unauthenticated GET of
|
||||||
|
its own model page sometime after this project picked it, discovered live
|
||||||
|
2026-09-06 when it took production down: the startup check correctly
|
||||||
|
refused to boot rather than fail opaquely on the first request, but the
|
||||||
|
service still crash-looped until the default was corrected.)
|
||||||
|
|
||||||
|
`classifier.encoder.confidence_threshold` is a `[0.0, 1.0]` probability
|
||||||
|
(`classify_zero_shot`'s own output), not a percent — the admin UI takes 0-100
|
||||||
|
for a human to type and converts at the save boundary, but a config file
|
||||||
|
edit or any other caller must use the raw probability. Config load now
|
||||||
|
validates the range; a value like `80` used to be silently accepted and
|
||||||
|
would make every real confidence score read as below-threshold, since none
|
||||||
|
can exceed `1.0` (also caught live 2026-09-06, before a restart made it
|
||||||
|
active).
|
||||||
|
|
||||||
The trade is real, not free. It only produces `task_category` — no `task_tier`
|
The trade is real, not free. It only produces `task_category` — no `task_tier`
|
||||||
signal exists in a zero-shot label score, so tier falls back to
|
signal exists in a zero-shot label score, so tier falls back to
|
||||||
|
|||||||
@@ -791,7 +791,15 @@ class LocalEncoderConfig(StrictModel):
|
|||||||
learn from without a new, separate opt-in data-capture feature.
|
learn from without a new, separate opt-in data-capture feature.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
model: str = "MoritzLaurer/deberta-v3-base-zeroshot-v2"
|
# facebook/bart-large-mnli -- the reference model HuggingFace's own docs
|
||||||
|
# use for this exact pipeline. Was MoritzLaurer/deberta-v3-base-zeroshot-v2
|
||||||
|
# (smaller, ~184M vs ~407M params) until that repo started returning 401
|
||||||
|
# even on an unauthenticated GET of its model page -- gated or moved
|
||||||
|
# sometime after this project picked it. Caught live 2026-09-06: the
|
||||||
|
# startup check (ensure_available) correctly refused to boot rather than
|
||||||
|
# fail opaquely on the first request, but it still took production down
|
||||||
|
# until the default was fixed.
|
||||||
|
model: str = "facebook/bart-large-mnli"
|
||||||
device: Literal["cpu", "cuda"] = "cpu"
|
device: Literal["cpu", "cuda"] = "cpu"
|
||||||
# Below this, the classification is treated as a FAILURE, not a low-
|
# Below this, the classification is treated as a FAILURE, not a low-
|
||||||
# confidence answer -- the caller cascades exactly as it would for a
|
# confidence answer -- the caller cascades exactly as it would for a
|
||||||
|
|||||||
@@ -122,7 +122,7 @@ def test_local_encoder_with_empty_encoder_block_loads_with_defaults(raw):
|
|||||||
cfg["classifier"]["mode"] = "local_encoder"
|
cfg["classifier"]["mode"] = "local_encoder"
|
||||||
cfg["classifier"]["encoder"] = {}
|
cfg["classifier"]["encoder"] = {}
|
||||||
loaded = RouterConfig(**cfg)
|
loaded = RouterConfig(**cfg)
|
||||||
assert loaded.classifier.encoder.model == "MoritzLaurer/deberta-v3-base-zeroshot-v2"
|
assert loaded.classifier.encoder.model == "facebook/bart-large-mnli"
|
||||||
assert loaded.classifier.encoder.device == "cpu"
|
assert loaded.classifier.encoder.device == "cpu"
|
||||||
assert loaded.classifier.encoder.confidence_threshold == 0.5
|
assert loaded.classifier.encoder.confidence_threshold == 0.5
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user