ConnectionError from unreachable Ollama was propagating as 500 instead of
degrading to the fallback cascade. Add requests.RequestException to the
except tuple in classify() so any requests-based exception (including
ConnectionError, Timeout, etc.) gracefully degrades to the configured
fallback classification.
Replace all 3 classify_choice() calls with classify_category() so the
dispatcher receives category names (not letters) that match
_DECISION_DESCRIPTIONS keys — previously coverage was always 0.
In the metering code path, run the tier classification call in parallel
with the category call via ThreadPoolExecutor so both classifier round-
trips happen simultaneously instead of sequentially, cutting latency.
Update all tests: replace classify_choice mocks with classify_category,
and add a timing test that verifies the tier call completes in parallel
(~0.1s) rather than sequentially (~0.2s).
Replace the manual letter→category mapping in eval_classifier.py with a
new classify_category() function in local_decision.py that handles
letter assignment and mapping internally. eval_classifier.py now calls
classify_category() directly, which returns a category name.
Tests in test_eval_classifier.py updated to mock classify_category instead
of classify_choice.
qwen3.5:4b is a reasoning model that generates "Thinking" tokens even
with think: false. With num_predict=1 (the code default) it outputs
only the first thinking token and never reaches an option letter,
causing coverage=0 and a cascading fallback.
Testing showed num_predict=5/20/80/256 all fail — the model always
gets stuck in verbose thinking. Switch to mistral-nemo:12b (the same
non-reasoning model the main classifier uses), where num_predict=1
works reliably.
Update the corresponding config test that asserts the default value.
At 32k the 14b is 16.5-17.8 GB resident and the local_decision classifier
(qwen3.5:4b, 5.9 GB at num_ctx 8192) evicts it; at 16k it is 12.26 GB and the
pair sits at ~19.9 of 24 GB (plans/local-decision-classifier-results.md,
sections 7 and 9). context_window follows the tag, which the operator rebuilds
at 16384 after this merges; the config must shrink first.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
Copy plans/local-decision-classifier-heldout.yaml (30 tasks) to
evals/heldout.yaml unchanged, and add a loader test asserting the
file loads 30 scoreable tasks with the expected categories.
The held-out set is weak: single author, short prompts, no true
distributional shift from the training set. Treat 100% as a ceiling,
not a forecast -- it cannot measure generalization.