feat(classifier): add classifier.mode local_decision (first-token logprob classifier on local Ollama) #106

Merged
alee merged 41 commits from feat/local-decision-classifier into main 2026-10-03 03:45:22 +00:00

41 Commits

Author SHA1 Message Date
adlee-was-taken
46e6da89af fix(dispatcher): do not probe Ollama at startup while local compute is off 2026-10-02 20:38:17 -04:00
adlee-was-taken
23b39bd64c fix(dispatcher): open the classifier backoff when local_decision cannot reach Ollama 2026-10-02 20:34:45 -04:00
adlee-was-taken
891f176c87 fix(eval): resolve basedpyright Optional[dict] narrowing errors 2026-09-29 23:41:11 -04:00
adlee-was-taken
f2ddcae4ec test(dispatcher): unreachable Ollama in local_decision mode walks the cascade 2026-09-29 23:02:16 -04:00
adlee-was-taken
3559eb02c7 test(dispatcher): snapshot local_decision energy at log time 2026-09-29 22:56:17 -04:00
adlee-was-taken
558f869ad7 test(local-decision): exercise letter mapping through the HTTP boundary 2026-09-29 22:55:44 -04:00
adlee-was-taken
bd5aa57566 fix(dispatcher): log local_decision energy after the measurement closes 2026-09-29 22:47:20 -04:00
adlee-was-taken
c53bb20568 fix(dispatcher): run local_decision tier call in parallel; fix metering log timing 2026-09-29 22:31:03 -04:00
adlee-was-taken
c06bc36dba fix(dispatcher): catch requests.RequestException in classify() cascade handler
ConnectionError from unreachable Ollama was propagating as 500 instead of
degrading to the fallback cascade. Add requests.RequestException to the
except tuple in classify() so any requests-based exception (including
ConnectionError, Timeout, etc.) gracefully degrades to the configured
fallback classification.
2026-09-29 18:42:18 -04:00
adlee-was-taken
66404f53c1 feat(dispatcher): fix letter→category mapping + run tier call in parallel inside metering
Replace all 3 classify_choice() calls with classify_category() so the
dispatcher receives category names (not letters) that match
_DECISION_DESCRIPTIONS keys — previously coverage was always 0.

In the metering code path, run the tier classification call in parallel
with the category call via ThreadPoolExecutor so both classifier round-
trips happen simultaneously instead of sequentially, cutting latency.

Update all tests: replace classify_choice mocks with classify_category,
and add a timing test that verifies the tier call completes in parallel
(~0.1s) rather than sequentially (~0.2s).
2026-09-29 18:18:06 -04:00
adlee-was-taken
3a52307a59 fix(local-decision): add classify_category() and move letter mapping out of eval harness
Replace the manual letter→category mapping in eval_classifier.py with a
new classify_category() function in local_decision.py that handles
letter assignment and mapping internally. eval_classifier.py now calls
classify_category() directly, which returns a category name.

Tests in test_eval_classifier.py updated to mock classify_category instead
of classify_choice.
2026-09-29 18:10:59 -04:00
adlee-was-taken
c946606e65 Revert "fix(config): switch local_decision default model to mistral-nemo:12b to avoid qwen3.5 reasoning model"
This reverts commit 2387111478.
2026-09-29 18:00:51 -04:00
adlee-was-taken
2387111478 fix(config): switch local_decision default model to mistral-nemo:12b to avoid qwen3.5 reasoning model
qwen3.5:4b is a reasoning model that generates "Thinking" tokens even
with think: false. With num_predict=1 (the code default) it outputs
only the first thinking token and never reaches an option letter,
causing coverage=0 and a cascading fallback.

Testing showed num_predict=5/20/80/256 all fail — the model always
gets stuck in verbose thinking. Switch to mistral-nemo:12b (the same
non-reasoning model the main classifier uses), where num_predict=1
works reliably.

Update the corresponding config test that asserts the default value.
2026-09-29 01:18:50 -04:00
adlee-was-taken
2dd1cdf082 fix(local-decision): fix Optional typing and remove type ignore 2026-09-29 01:05:11 -04:00
adlee-was-taken
0cf6b5df29 docs(admin-portal): document local_decision admin fields 2026-09-29 00:26:59 -04:00
adlee-was-taken
6d6e346358 docs(claude): update classifier mode descriptions, fix stale encoder device note 2026-09-29 00:26:48 -04:00
adlee-was-taken
de5d9ce969 docs(local-models): document local_decision mode and VRAM requirements 2026-09-29 00:26:36 -04:00
adlee-was-taken
7d221440f3 feat(dispatcher): add tier classification with parallel execution behind tier_enabled flag 2026-09-29 00:24:35 -04:00
adlee-was-taken
45c1493c94 feat(admin): include decision block in classifier-config GET response 2026-09-29 00:19:06 -04:00
adlee-was-taken
09a4dc992e feat(admin): wire local_decision fields through collectClassifierConfigBody + loadClassifierConfig 2026-09-29 00:13:53 -04:00
adlee-was-taken
a62056d83d feat(admin): add local_decision mode fields to classifier card UI 2026-09-29 00:04:54 -04:00
adlee-was-taken
58e6b81e57 feat(admin): add local_decision to _ClassifierConfigBody mode Literal 2026-09-28 23:57:46 -04:00
adlee-was-taken
6c7602265d fix(config): correct LocalDecisionConfig.coverage_min comment 2026-09-28 23:48:33 -04:00
adlee-was-taken
7ac4d4da01 feat(dispatcher): add startup readiness check for local_decision mode 2026-09-28 23:48:31 -04:00
adlee-was-taken
9213a0870a feat(dispatcher): wire local_decision in _classify_via_configured_mode 2026-09-28 23:27:25 -04:00
adlee-was-taken
b9de4c0530 feat(dispatcher): add _classify_via_local_decision with circuit breaker + metering 2026-09-28 23:26:56 -04:00
adlee-was-taken
1a6c82f4b2 fix(config): meter local_decision via base_url loopback check, not encoder path 2026-09-28 23:22:22 -04:00
adlee-was-taken
3010d98f89 feat(config): add local_decision to mode Literal 2026-09-28 23:17:42 -04:00
adlee-was-taken
0bb6f564cf feat(config): add classifier.decision field and cross-mode validator 2026-09-28 23:13:45 -04:00
adlee-was-taken
03d19a64b7 feat(config): add LocalDecisionConfig with validators 2026-09-28 23:08:05 -04:00
adlee-was-taken
cc5ac2357e feat(admin): local_vision.enabled gets a persisted control 2026-09-28 22:57:28 -04:00
adlee-was-taken
a7d4c699f1 config: run the local-dispatch 14b at num_ctx 16384 to make room for local_decision
At 32k the 14b is 16.5-17.8 GB resident and the local_decision classifier
(qwen3.5:4b, 5.9 GB at num_ctx 8192) evicts it; at 16k it is 12.26 GB and the
pair sits at ~19.9 of 24 GB (plans/local-decision-classifier-results.md,
sections 7 and 9). context_window follows the tag, which the operator rebuilds
at 16384 after this merges; the config must shrink first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-28 22:43:33 -04:00
adlee-was-taken
af9be29b15 docs(local-decision): re-measure VRAM with the 14b at 16k (gate G passes)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-28 22:24:27 -04:00
adlee-was-taken
979a11b850 docs(local-decision): gate G measurement results and description tuning
Gate G results:
- Eval-set accuracy: 97.8% (PASS >= 85%)
- Held-out accuracy: 90.0% (PASS >= 90%)
- p95 latency: 502ms (PASS <= 800ms)
- 14b router resident after 100 calls: EVICTED (FAIL)

VRAM constraint: 14b router (17.76GB) + qwen3.5:4b (>=6.15GB) =
25.1GB > 24GB RTX 6000. No num_ctx allows coexistence.

STOP per plan: todos 4-17 NOT started.
2026-09-28 21:11:36 -04:00
adlee-was-taken
e66387cd9f feat(local-decision): add _DECISION_DESCRIPTIONS with confusion-tuned wording 2026-09-28 20:16:19 -04:00
adlee-was-taken
48923caa47 feat(local-decision): add classify_choice with Ollama /api/chat 2026-09-28 20:12:38 -04:00
adlee-was-taken
cf3b685545 fix(local-decision): fix logprobs read path to top-level Ollama key + fixture test 2026-09-28 20:09:46 -04:00
adlee-was-taken
c8224aec7d feat(eval): add --backend decision and --head-off flags 2026-09-28 20:02:37 -04:00
adlee-was-taken
1b9af75576 feat(local-decision): add parse_logprobs with unit tests 2026-09-28 19:57:09 -04:00
adlee-was-taken
0b85b6366c feat(eval): promote held-out set to evals/heldout.yaml
Copy plans/local-decision-classifier-heldout.yaml (30 tasks) to
evals/heldout.yaml unchanged, and add a loader test asserting the
file loads 30 scoreable tasks with the expected categories.

The held-out set is weak: single author, short prompts, no true
distributional shift from the training set. Treat 100% as a ceiling,
not a forecast -- it cannot measure generalization.
2026-09-28 19:56:55 -04:00
adlee-was-taken
e5f0be633b plans: local_decision classifier brief, prototype and held-out prompts
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
2026-09-28 19:46:56 -04:00