feat(classifier): add classifier.mode local_decision (first-token logprob classifier on local Ollama) #106
Reference in New Issue
Block a user
Delete Branch "feat/local-decision-classifier"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What
Adds a fourth
classifier.mode,local_decision. A small local LLM (qwen3.5:4bon Ollama's native/api/chat) is asked one multiple-choice question, and the answer is read from the first output token's logprobs, normalised over the option letters that were offered. There is no text generation, JSON or reasoning trace, so the runaway-reasoning failure that retiredlocal_llmcannot happen, and confidence is calibrated (the encoder's zero-shot confidence sits at 0.2-0.3 whether it is right or wrong).Opt-in only. The
config.yamldefault stayslocal_llm, and the live overlay still selectslocal_encoder. Nothing changes for anyone until they switch the mode. A failure still walks the unchanged classifier cascade, andtool_use_agenticstays out of the option set.Brief:
plans/local-decision-classifier.md. Measurements:plans/local-decision-classifier-results.md.Evidence (RTX 6000 24 GB, 2026-09-28)
evals/tasks.yaml(46 tasks x 3 noise levels)evals/heldout.yaml(30 prompts x 3 noise levels)Noise isolation was measured and rejected for this backend (it cost 3.3 points on the eval set and 11.7 on held-out, because the code and tool content it strips is the signal), so raw text is sent. Tier uses separate calls rather than one joint prompt (the joint read produced valid mass at both positions on only 5 of 30 prompts).
Gate G first came back STOP on VRAM: at the 14b's production
num_ctx 32768the two models do not reliably coexist. It passes with the 14b atnum_ctx 16384and local vision off (12.26 + 5.89 GB, about 19.9 of 24 GB used). That is whyconfig.yamlchanges below.What changes
src/local_decision.py(new):parse_logprobs,classify_choice,classify_category, and a separate_DECISION_DESCRIPTIONSmap tuned against the measured confusions.local_encoder._CATEGORY_DESCRIPTIONSis untouched, so the encoder and its head are not shifted.src/config.py:"local_decision"in themodeLiteral, aLocalDecisionConfigblock (classifier.decision:base_url,model,num_ctx,timeout_s,confidence_min,coverage_min,tier_enabled), a validator that requires the block when the mode is selected, and metering via the HTTPbase_urlloopback check (likelocal_llm, not like the in-process encoder).src/dispatcher.py:_classify_via_local_decision, the branch in_classify_via_configured_mode, the startup readiness check,requests.RequestExceptionadded toclassify()'s cascade handler, and an optional tier call behindtier_enabled(default false, fired in parallel with the category call).src/admin.py,admin/frontend/controls.html): the Classifier card renders thelocal_decisionfields, and the classifier-config GET includes thedecisionblock.local_vision.enabledalso gets a persisted control, so the operator steps below need no config edit. Knob coverage (tests/test_admin_knob_coverage.py) is updated.src/eval_classifier.py):--backend decision, a flag to run the encoder with its trained head off (so the leakage stays visible), and the held-out set promoted toevals/heldout.yaml.CLAUDE.md(classifier mode descriptions; also fixes the stale "encoder isdevice: cuda" note, the overlay sayscpu),docs/local-models.md,docs/admin-portal.md.config/config.yaml: the local-dispatch 14bcontext_windowgoes from 32768 to 16384, with the measurement in the comment. See "Decision for the reviewer" below.27 files, +3348 / -52, 41 commits.
Fixes found in review (last two commits)
23b39bd, a hung Ollama never opened the classifier backoff._classify_via_local_decisionread the circuit through_local_classifier_skip_reasonbut never wrote it (only_classify_via_local_llmcalled_record_failure()), so every request would wait the fulltimeout_s(10s) indefinitely. It now records onrequests.RequestExceptiononly, at both the metered and unmetered sites. Deliberately not recorded: below-confidence_minverdicts andparse_logprobserrors (the endpoint answered, and recording would skip a healthy classifier forcooldown_seconds), tier-call failures (already swallowed tofallback_tier), and skips (which must never re-stamp the clock).46e6da8, the startup probe ignored gaming mode._ensure_classifier_mode_readyruns at import and made a live Ollama call, loading the 4b model onto the GPU even withlocal_compute.enabled: false. It now skips the probe and logsclassifier_mode_probe_skipped. It checkscfg.local_compute.enableddirectly, because_local_classifier_skip_reasonis defined later in the module than the import-time call.Before you enable it (merging does not do these)
PARAMETER num_ctx 16384on the Ollama host. The liveqwen2.5-coder-router:14bwas stillnum_ctx 32768on 2026-10-02, whileconfig.yamlhere says 16384.local_vision.config.yamlstill hasenabled: true, and the VRAM pass assumes it is off (the 5.6 GB vision model does not fit beside both).classifier.modefrom the admin Classifier card (runtime plus persisted overlay), after backing upconfig/config.local.yaml.Decision for the reviewer
The 14b
context_window: 16384lives in the sharedconfig.yaml, so it lowers the local-dispatch context ceiling for every deployment, not only this one (a request above 16k can no longer be admitted to the local row by the hard context filter). It was chosen as option A in the VRAM discussion. If it should be per-host instead, move it to the overlay before merging.Verification
tests/test_incumbent_routing.py::TestDebugLog::test_debug_log_emits_incumbent_identity, fails identically on a clean export of main, so it is not from this branch.Known gaps and not verified
chat_completionspasses the previous turn as context, so the classifier seesContext: <prev> --- Message: <task>plus a follow-up sentence on most agent turns.OLLAMA_NUM_PARALLELis above 1. Tier is off by default._last_classifier_failure > 0after the skipped second call rather than "unchanged", so a skip that re-stamped the clock would pass. The "low confidence" test actually trips the coverage floor (coverage 0.0116 below minimum 0.3), notconfidence_min._classify_via_local_decisionduplicate about 60 lines, thelocal_decision.pymodule docstring says four options (A-D) where there are ten, and_TOP_LOGBPROBSis a typo.🤖 Generated with Claude Code
https://claude.ai/code/session_01KkCGRantZsSwmcFpet6FTa
fix(dispatcher): classifier circuit breaker and startup probe skipto feat(classifier): add classifier.mode local_decision (first-token logprob classifier on local Ollama)