feat(classifier): add classifier.mode local_decision (first-token logprob classifier on local Ollama) #106

Merged
alee merged 41 commits from feat/local-decision-classifier into main 2026-10-03 03:45:22 +00:00
Owner

What

Adds a fourth classifier.mode, local_decision. A small local LLM (qwen3.5:4b on Ollama's native /api/chat) is asked one multiple-choice question, and the answer is read from the first output token's logprobs, normalised over the option letters that were offered. There is no text generation, JSON or reasoning trace, so the runaway-reasoning failure that retired local_llm cannot happen, and confidence is calibrated (the encoder's zero-shot confidence sits at 0.2-0.3 whether it is right or wrong).

Opt-in only. The config.yaml default stays local_llm, and the live overlay still selects local_encoder. Nothing changes for anyone until they switch the mode. A failure still walks the unchanged classifier cascade, and tool_use_agentic stays out of the option set.

Brief: plans/local-decision-classifier.md. Measurements: plans/local-decision-classifier-results.md.

Evidence (RTX 6000 24 GB, 2026-09-28)

result bar
evals/tasks.yaml (46 tasks x 3 noise levels) 97.8% after description tuning (76.1% before) >= 85%
evals/heldout.yaml (30 prompts x 3 noise levels) 90.0% (96.7% before tuning) >= 90%
p50 / p95 latency, 14b router resident 321 ms / 474 ms p95 <= 800 ms
Confidence when right / wrong about 0.80 / 0.35-0.41 calibrated in the right direction

Noise isolation was measured and rejected for this backend (it cost 3.3 points on the eval set and 11.7 on held-out, because the code and tool content it strips is the signal), so raw text is sent. Tier uses separate calls rather than one joint prompt (the joint read produced valid mass at both positions on only 5 of 30 prompts).

Gate G first came back STOP on VRAM: at the 14b's production num_ctx 32768 the two models do not reliably coexist. It passes with the 14b at num_ctx 16384 and local vision off (12.26 + 5.89 GB, about 19.9 of 24 GB used). That is why config.yaml changes below.

What changes

  • src/local_decision.py (new): parse_logprobs, classify_choice, classify_category, and a separate _DECISION_DESCRIPTIONS map tuned against the measured confusions. local_encoder._CATEGORY_DESCRIPTIONS is untouched, so the encoder and its head are not shifted.
  • src/config.py: "local_decision" in the mode Literal, a LocalDecisionConfig block (classifier.decision: base_url, model, num_ctx, timeout_s, confidence_min, coverage_min, tier_enabled), a validator that requires the block when the mode is selected, and metering via the HTTP base_url loopback check (like local_llm, not like the in-process encoder).
  • src/dispatcher.py: _classify_via_local_decision, the branch in _classify_via_configured_mode, the startup readiness check, requests.RequestException added to classify()'s cascade handler, and an optional tier call behind tier_enabled (default false, fired in parallel with the category call).
  • Admin (src/admin.py, admin/frontend/controls.html): the Classifier card renders the local_decision fields, and the classifier-config GET includes the decision block. local_vision.enabled also gets a persisted control, so the operator steps below need no config edit. Knob coverage (tests/test_admin_knob_coverage.py) is updated.
  • Eval harness (src/eval_classifier.py): --backend decision, a flag to run the encoder with its trained head off (so the leakage stays visible), and the held-out set promoted to evals/heldout.yaml.
  • Docs: CLAUDE.md (classifier mode descriptions; also fixes the stale "encoder is device: cuda" note, the overlay says cpu), docs/local-models.md, docs/admin-portal.md.
  • config/config.yaml: the local-dispatch 14b context_window goes from 32768 to 16384, with the measurement in the comment. See "Decision for the reviewer" below.

27 files, +3348 / -52, 41 commits.

Fixes found in review (last two commits)

  1. 23b39bd, a hung Ollama never opened the classifier backoff. _classify_via_local_decision read the circuit through _local_classifier_skip_reason but never wrote it (only _classify_via_local_llm called _record_failure()), so every request would wait the full timeout_s (10s) indefinitely. It now records on requests.RequestException only, at both the metered and unmetered sites. Deliberately not recorded: below-confidence_min verdicts and parse_logprobs errors (the endpoint answered, and recording would skip a healthy classifier for cooldown_seconds), tier-call failures (already swallowed to fallback_tier), and skips (which must never re-stamp the clock).
  2. 46e6da8, the startup probe ignored gaming mode. _ensure_classifier_mode_ready runs at import and made a live Ollama call, loading the 4b model onto the GPU even with local_compute.enabled: false. It now skips the probe and logs classifier_mode_probe_skipped. It checks cfg.local_compute.enabled directly, because _local_classifier_skip_reason is defined later in the module than the import-time call.

Before you enable it (merging does not do these)

  1. Retag the router model with PARAMETER num_ctx 16384 on the Ollama host. The live qwen2.5-coder-router:14b was still num_ctx 32768 on 2026-10-02, while config.yaml here says 16384.
  2. Turn off local_vision. config.yaml still has enabled: true, and the VRAM pass assumes it is off (the 5.6 GB vision model does not fit beside both).
  3. Switch classifier.mode from the admin Classifier card (runtime plus persisted overlay), after backing up config/config.local.yaml.

Decision for the reviewer

The 14b context_window: 16384 lives in the shared config.yaml, so it lowers the local-dispatch context ceiling for every deployment, not only this one (a request above 16k can no longer be admitted to the local row by the hard context filter). It was chosen as option A in the VRAM discussion. If it should be per-host instead, move it to the overlay before merging.

Verification

  • Full suite: 2508 passed, 1 failed. The one failure, tests/test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity, fails identically on a clean export of main, so it is not from this branch.
  • The new backoff and probe tests were checked against the pre-fix dispatcher in a scratch export: 3 fail there. Each of the two failure-recording sites was also neutered separately, and each is caught by its own test.
  • The branch merges cleanly into current main (20 commits behind at the time of writing).

Known gaps and not verified

  • Production input shape was never measured. The accuracy numbers come from raw task text through the eval harness. In production, chat_completions passes the previous turn as context, so the classifier sees Context: <prev> --- Message: <task> plus a follow-up sentence on most agent turns.
  • The held-out 90.0% is thinner than it looks. It sits exactly on the threshold, was re-measured after every description change (so it is no longer truly held out), and is 30 prompts from one author.
  • Tier is unmeasured under concurrency. "Adds no wall-clock time" is computed as the max of two separate runs. Ollama serialises same-model requests unless OLLAMA_NUM_PARALLEL is above 1. Tier is off by default.
  • Live QA on the 8081 sandbox is not independently verified. An agent run reported it as passing, but nothing records it.
  • Two test weaknesses. The unmetered backoff test asserts _last_classifier_failure > 0 after the skipped second call rather than "unchanged", so a skip that re-stamped the clock would pass. The "low confidence" test actually trips the coverage floor (coverage 0.0116 below minimum 0.3), not confidence_min.
  • Cleanup debt, not fixed here. The metered and unmetered paths of _classify_via_local_decision duplicate about 60 lines, the local_decision.py module docstring says four options (A-D) where there are ten, and _TOP_LOGBPROBS is a typo.

🤖 Generated with Claude Code

https://claude.ai/code/session_01KkCGRantZsSwmcFpet6FTa

## What Adds a fourth `classifier.mode`, `local_decision`. A small local LLM (`qwen3.5:4b` on Ollama's native `/api/chat`) is asked one multiple-choice question, and the answer is read from the **first output token's logprobs**, normalised over the option letters that were offered. There is no text generation, JSON or reasoning trace, so the runaway-reasoning failure that retired `local_llm` cannot happen, and confidence is calibrated (the encoder's zero-shot confidence sits at 0.2-0.3 whether it is right or wrong). **Opt-in only.** The `config.yaml` default stays `local_llm`, and the live overlay still selects `local_encoder`. Nothing changes for anyone until they switch the mode. A failure still walks the unchanged classifier cascade, and `tool_use_agentic` stays out of the option set. Brief: `plans/local-decision-classifier.md`. Measurements: `plans/local-decision-classifier-results.md`. ## Evidence (RTX 6000 24 GB, 2026-09-28) | | result | bar | |---|---|---| | `evals/tasks.yaml` (46 tasks x 3 noise levels) | 97.8% after description tuning (76.1% before) | >= 85% | | `evals/heldout.yaml` (30 prompts x 3 noise levels) | 90.0% (96.7% before tuning) | >= 90% | | p50 / p95 latency, 14b router resident | 321 ms / 474 ms | p95 <= 800 ms | | Confidence when right / wrong | about 0.80 / 0.35-0.41 | calibrated in the right direction | Noise isolation was measured and **rejected** for this backend (it cost 3.3 points on the eval set and 11.7 on held-out, because the code and tool content it strips is the signal), so raw text is sent. Tier uses separate calls rather than one joint prompt (the joint read produced valid mass at both positions on only 5 of 30 prompts). Gate G first came back STOP on VRAM: at the 14b's production `num_ctx 32768` the two models do not reliably coexist. It passes with the 14b at `num_ctx 16384` and local vision off (12.26 + 5.89 GB, about 19.9 of 24 GB used). That is why `config.yaml` changes below. ## What changes - **`src/local_decision.py`** (new): `parse_logprobs`, `classify_choice`, `classify_category`, and a separate `_DECISION_DESCRIPTIONS` map tuned against the measured confusions. `local_encoder._CATEGORY_DESCRIPTIONS` is untouched, so the encoder and its head are not shifted. - **`src/config.py`**: `"local_decision"` in the `mode` Literal, a `LocalDecisionConfig` block (`classifier.decision`: `base_url`, `model`, `num_ctx`, `timeout_s`, `confidence_min`, `coverage_min`, `tier_enabled`), a validator that requires the block when the mode is selected, and metering via the HTTP `base_url` loopback check (like `local_llm`, not like the in-process encoder). - **`src/dispatcher.py`**: `_classify_via_local_decision`, the branch in `_classify_via_configured_mode`, the startup readiness check, `requests.RequestException` added to `classify()`'s cascade handler, and an optional tier call behind `tier_enabled` (default false, fired in parallel with the category call). - **Admin** (`src/admin.py`, `admin/frontend/controls.html`): the Classifier card renders the `local_decision` fields, and the classifier-config GET includes the `decision` block. `local_vision.enabled` also gets a persisted control, so the operator steps below need no config edit. Knob coverage (`tests/test_admin_knob_coverage.py`) is updated. - **Eval harness** (`src/eval_classifier.py`): `--backend decision`, a flag to run the encoder with its trained head off (so the leakage stays visible), and the held-out set promoted to `evals/heldout.yaml`. - **Docs**: `CLAUDE.md` (classifier mode descriptions; also fixes the stale "encoder is `device: cuda`" note, the overlay says `cpu`), `docs/local-models.md`, `docs/admin-portal.md`. - **`config/config.yaml`**: the local-dispatch 14b `context_window` goes from 32768 to 16384, with the measurement in the comment. See "Decision for the reviewer" below. 27 files, +3348 / -52, 41 commits. ## Fixes found in review (last two commits) 1. **`23b39bd`, a hung Ollama never opened the classifier backoff.** `_classify_via_local_decision` read the circuit through `_local_classifier_skip_reason` but never wrote it (only `_classify_via_local_llm` called `_record_failure()`), so every request would wait the full `timeout_s` (10s) indefinitely. It now records on `requests.RequestException` only, at both the metered and unmetered sites. Deliberately **not** recorded: below-`confidence_min` verdicts and `parse_logprobs` errors (the endpoint answered, and recording would skip a healthy classifier for `cooldown_seconds`), tier-call failures (already swallowed to `fallback_tier`), and skips (which must never re-stamp the clock). 2. **`46e6da8`, the startup probe ignored gaming mode.** `_ensure_classifier_mode_ready` runs at import and made a live Ollama call, loading the 4b model onto the GPU even with `local_compute.enabled: false`. It now skips the probe and logs `classifier_mode_probe_skipped`. It checks `cfg.local_compute.enabled` directly, because `_local_classifier_skip_reason` is defined later in the module than the import-time call. ## Before you enable it (merging does not do these) 1. Retag the router model with `PARAMETER num_ctx 16384` on the Ollama host. The live `qwen2.5-coder-router:14b` was still `num_ctx 32768` on 2026-10-02, while `config.yaml` here says 16384. 2. Turn off `local_vision`. `config.yaml` still has `enabled: true`, and the VRAM pass assumes it is off (the 5.6 GB vision model does not fit beside both). 3. Switch `classifier.mode` from the admin Classifier card (runtime plus persisted overlay), after backing up `config/config.local.yaml`. ## Decision for the reviewer The 14b `context_window: 16384` lives in the shared `config.yaml`, so it lowers the local-dispatch context ceiling for every deployment, not only this one (a request above 16k can no longer be admitted to the local row by the hard context filter). It was chosen as option A in the VRAM discussion. If it should be per-host instead, move it to the overlay before merging. ## Verification - Full suite: **2508 passed, 1 failed**. The one failure, `tests/test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity`, fails identically on a clean export of main, so it is not from this branch. - The new backoff and probe tests were checked against the pre-fix dispatcher in a scratch export: 3 fail there. Each of the two failure-recording sites was also neutered separately, and each is caught by its own test. - The branch merges cleanly into current main (20 commits behind at the time of writing). ## Known gaps and not verified - **Production input shape was never measured.** The accuracy numbers come from raw task text through the eval harness. In production, `chat_completions` passes the previous turn as context, so the classifier sees `Context: <prev> --- Message: <task>` plus a follow-up sentence on most agent turns. - **The held-out 90.0% is thinner than it looks.** It sits exactly on the threshold, was re-measured after every description change (so it is no longer truly held out), and is 30 prompts from one author. - **Tier is unmeasured under concurrency.** "Adds no wall-clock time" is computed as the max of two separate runs. Ollama serialises same-model requests unless `OLLAMA_NUM_PARALLEL` is above 1. Tier is off by default. - **Live QA on the 8081 sandbox is not independently verified.** An agent run reported it as passing, but nothing records it. - **Two test weaknesses.** The unmetered backoff test asserts `_last_classifier_failure > 0` after the skipped second call rather than "unchanged", so a skip that re-stamped the clock would pass. The "low confidence" test actually trips the coverage floor (`coverage 0.0116 below minimum 0.3`), not `confidence_min`. - **Cleanup debt, not fixed here.** The metered and unmetered paths of `_classify_via_local_decision` duplicate about 60 lines, the `local_decision.py` module docstring says four options (A-D) where there are ten, and `_TOP_LOGBPROBS` is a typo. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01KkCGRantZsSwmcFpet6FTa
alee added 41 commits 2026-10-03 01:04:02 +00:00
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
Copy plans/local-decision-classifier-heldout.yaml (30 tasks) to
evals/heldout.yaml unchanged, and add a loader test asserting the
file loads 30 scoreable tasks with the expected categories.

The held-out set is weak: single author, short prompts, no true
distributional shift from the training set. Treat 100% as a ceiling,
not a forecast -- it cannot measure generalization.
Gate G results:
- Eval-set accuracy: 97.8% (PASS >= 85%)
- Held-out accuracy: 90.0% (PASS >= 90%)
- p95 latency: 502ms (PASS <= 800ms)
- 14b router resident after 100 calls: EVICTED (FAIL)

VRAM constraint: 14b router (17.76GB) + qwen3.5:4b (>=6.15GB) =
25.1GB > 24GB RTX 6000. No num_ctx allows coexistence.

STOP per plan: todos 4-17 NOT started.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
At 32k the 14b is 16.5-17.8 GB resident and the local_decision classifier
(qwen3.5:4b, 5.9 GB at num_ctx 8192) evicts it; at 16k it is 12.26 GB and the
pair sits at ~19.9 of 24 GB (plans/local-decision-classifier-results.md,
sections 7 and 9). context_window follows the tag, which the operator rebuilds
at 16384 after this merges; the config must shrink first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
qwen3.5:4b is a reasoning model that generates "Thinking" tokens even
with think: false. With num_predict=1 (the code default) it outputs
only the first thinking token and never reaches an option letter,
causing coverage=0 and a cascading fallback.

Testing showed num_predict=5/20/80/256 all fail — the model always
gets stuck in verbose thinking. Switch to mistral-nemo:12b (the same
non-reasoning model the main classifier uses), where num_predict=1
works reliably.

Update the corresponding config test that asserts the default value.
This reverts commit 2387111478.
Replace the manual letter→category mapping in eval_classifier.py with a
new classify_category() function in local_decision.py that handles
letter assignment and mapping internally. eval_classifier.py now calls
classify_category() directly, which returns a category name.

Tests in test_eval_classifier.py updated to mock classify_category instead
of classify_choice.
Replace all 3 classify_choice() calls with classify_category() so the
dispatcher receives category names (not letters) that match
_DECISION_DESCRIPTIONS keys — previously coverage was always 0.

In the metering code path, run the tier classification call in parallel
with the category call via ThreadPoolExecutor so both classifier round-
trips happen simultaneously instead of sequentially, cutting latency.

Update all tests: replace classify_choice mocks with classify_category,
and add a timing test that verifies the tier call completes in parallel
(~0.1s) rather than sequentially (~0.2s).
ConnectionError from unreachable Ollama was propagating as 500 instead of
degrading to the fallback cascade. Add requests.RequestException to the
except tuple in classify() so any requests-based exception (including
ConnectionError, Timeout, etc.) gracefully degrades to the configured
fallback classification.
alee changed title from fix(dispatcher): classifier circuit breaker and startup probe skip to feat(classifier): add classifier.mode local_decision (first-token logprob classifier on local Ollama) 2026-10-03 02:54:31 +00:00
alee merged commit cc72f205f4 into main 2026-10-03 03:45:22 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#106