Foundation for a selectable classifier primary, generalizing "which implementation answers a classification request" into config rather than always assuming the local Ollama model: - classifier.mode (default local_llm, unchanged behavior) plus cloud_primary / cloud_primary_auto for cloud_llm and an encoder block for local_encoder. Two new RouterConfig validators reject an incomplete combination at load time -- cloud_llm with neither/both primaries set, local_encoder with no encoder block -- the same model_validator(mode= "after") pattern this project already uses elsewhere. - routing.cheapest_classifier_candidate: the "auto_classifier" resolver. Reuses select_candidates + estimated_cost -- the same functions real dispatch ranking uses -- rather than a second cost model, priced for the classifier's own short-prompt/short-completion call shape (500/50 tokens, cache_rate 0) instead of the task's. required_tier=1 is a floor, not a ceiling, so a tier-3 model can still win on price -- a test pins this after an initial wrong assumption in the test itself. - local_encoder.py: zero-shot category classification via a non-generative encoder (default MoritzLaurer/deberta-v3-base-zeroshot-v2). Structurally immune to the one failure mode that has cost this project two prior classifier generations (docs/local-models.md): a generative model spending its budget on an unbounded reasoning trace. Zero-shot rather than fine-tuned, deliberately -- this router never stores raw task text anywhere, so there is no training corpus without a new, separate opt-in capture feature (scoped, not built). transformers/torch imported lazily inside the function, the same rule tui.py already follows for textual, so a deployment that never selects this mode needs neither installed. requirements-encoder.txt keeps them out of the main, pinned requirements file. Every test here was run against unmodified main first and observed to fail for the right reason before this commit made it pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
15 lines
725 B
Plaintext
15 lines
725 B
Plaintext
# Optional: only needed when classifier.mode: local_encoder is configured.
|
|
# Not part of requirements.txt deliberately -- this project's dependency
|
|
# tree is pinned and bumped deliberately, and torch is large enough (and
|
|
# CPU/CUDA-wheel-specific enough) to warrant staying out of every install
|
|
# rather than every deployment paying for it whether or not the mode is used.
|
|
#
|
|
# CPU install (the classifier.encoder.device: cpu default):
|
|
# pip install -r requirements-encoder.txt
|
|
#
|
|
# CUDA install: replace the torch line with the CUDA wheel index per
|
|
# https://pytorch.org/get-started/locally/ -- the exact index URL is
|
|
# CUDA-version-specific and changes upstream, so it is not pinned here.
|
|
transformers==4.57.1
|
|
torch==2.9.1
|