# `classifier.mode: local_decision` — a Jev-style classifier on the local GPU Status: planned -- plan reviewed 2026-09-28; gate G (tuning + measurement) decides whether the wiring proceeds Date: 2026-09-27 Prototype: `plans/local-decision-classifier-prototype.py` (a benchmark you can run, NOT the implementation) Held-out prompts: `plans/local-decision-classifier-heldout.yaml` Background: TypeSafe's hosted "Jev" decision model and the OpenJev/SemIf open copies. We are copying the **technique**, not using either product. ## Goal Add a fourth `classifier.mode`, `local_decision`. It asks a small local LLM (`qwen3.5:4b` on Ollama) one multiple-choice question and reads the answer from the **first output token's logprobs**, normalised over the option letters that were offered. There is no text generation, JSON or reasoning trace, so the runaway-reasoning failure that retired `local_llm` can't happen, and it is calibrated far better than `local_encoder`. ## Evidence (2026-09-27, RTX 6000 24 GB, measured with the prototype) 46 eval tasks and 30 held-out tasks, each at 3 noise levels (`_wrap_agent_noise`): | | `evals/tasks.yaml` | held-out | p50 / p95 ms | |---|---|---|---| | `local_decision` (`qwen3.5:4b`, GPU) | 71% | 100% | 340-385 / 510-610 | | `local_encoder` + trained head (live, CPU) | 100%\* | 47% | 80-115 / 100-360 | | `local_encoder` zero-shot (CPU) | 39% | 60% | 80-115 / 100-350 | \* The head was fitted on a synthetic corpus derived from `evals/tasks.yaml`, so this is leakage. On held-out prompts the head is **worse** than zero-shot. - **Calibration:** `local_decision` averages 0.96 confidence when right and 0.71 when wrong. The encoder zero-shot sits at 0.2-0.3 on both. - **Eval-set misses:** 18 were `coding_general`→`reasoning_math` (algorithm tasks), 18 were `file_summarization`→`debugging`/`reasoning_math`/`docs_writing`. Both look like option-description problems (component 2). - **VRAM:** `qwen3.5:4b` holds about 7 GB at `num_ctx` 8192 next to `qwen2.5-coder-router:14b`. The card sat at 23.4/24 GB. It's tight, so measure again when planning. - **Weak held-out set:** a single author wrote all 30 prompts and they are short and clean. Treat 100% as a ceiling, not a forecast. ## Mechanism (proven by the prototype) Ollama 0.22 native `POST /api/chat`: `think: false`, `logprobs: true`, `top_logprobs: 20`, `options: {num_predict: 1, temperature: 0, num_ctx: N}`. The system prompt says "answer with the letter only". The options are `A.`–`J.` followed by a natural-language description. Parse `logprobs[0].top_logprobs`: - strip each token and remove a trailing `.` - sum the probability mass for each option letter (`"B"` and `" B"` both count) - confidence = winner's mass ÷ total option mass - coverage = total option mass Tokens like `` show up in the top 20 and must be ignored. ## Components ### 1. `src/local_decision.py` (stdlib + `requests`, no new dependency) - **`classify_choice(text, options, *, base_url, model, num_ctx, timeout)`** returns `(label, confidence, coverage)`. It is generic over the option list, so tier reuses it. - **Coverage floor:** below a configurable `coverage_min` (the model didn't answer with a letter), raise. Don't guess. - **Pure parse function:** split the logprob parsing into its own function so tests can drive it with recorded responses and no Ollama. - **Noise:** decide whether to run `local_encoder._isolate_task_text` first. The prototype sent the raw noisy text and noise cost about 2 points; measure both. ### 2. Its own option descriptions Give it a separate `_DECISION_DESCRIPTIONS` map. **Do not edit** `local_encoder._CATEGORY_DESCRIPTIONS`, because that would shift the encoder and invalidate its head. Tune it against the two confusions above, and re-measure after every change on **both** sets so nobody overfits `evals/tasks.yaml` the way the head did. ### 3. Wiring (follow how `local_encoder` was added) - **`src/config.py`:** add `"local_decision"` to the `mode` Literal (line ~1187). Add a `classifier.decision` block (`base_url`, `model`, `num_ctx`, `timeout_s`, `confidence_min`, `coverage_min`, `tier_enabled`) and a validator that requires it when the mode is selected. Metering (line ~1712): this mode is an HTTP call to Ollama, so it meters like `local_llm` via its own `base_url`, not like the in-process encoder. - **`src/dispatcher.py`:** - add `_classify_via_local_decision` and a branch in `_classify_via_configured_mode` - add a startup readiness check in `_ensure_classifier_mode_ready` (model is pulled and logprobs come back) - on success record `source="classifier"` - treat below-threshold confidence as a failure that walks the **unchanged** cascade - **Gating:** it uses the GPU, so it obeys `_local_classifier_skip_reason()` (`local_compute.enabled` plus the circuit breaker), the same as `local_llm`. - **Admin:** every new scalar knob gets a control on the Classifier card (North Star #1). `tests/test_admin_knob_coverage.py` enforces this. ### 4. Tier (behind `tier_enabled`, default false) - **How:** a second `classify_choice` call with three options (cheap/simple, general, frontier/high-stakes), fired **in parallel** with the category call so it adds no wall-clock time. - **Why it defaults off:** no gold tier labels exist yet. Until they do, tier keeps falling back to `fallback_tier`, exactly as `local_encoder` does now. - **Output:** a gold tier set of about 30 labelled prompts, used to decide the default. ### 5. Evaluation harness - **Backend flag:** add a `--backend encoder|decision` flag to `src/eval_classifier.py` (latency is reported per call). - **Held-out set:** promote the held-out prompts to `evals/` and add a `--tasks` default for them. - **Head flag:** add a flag that runs the encoder with the head disabled, so the leakage stays visible. - **Report:** accuracy for both sets, calibration (confidence correct vs wrong), and p50/p95. ### 6. Docs - **CLAUDE.md:** the "Which implementation is PRIMARY" section, which also still says the live encoder is `device: cuda` while the overlay says `cpu`. Fix that. - **Also:** `docs/local-models.md` and `docs/admin-portal.md`. ## Acceptance - **Held-out accuracy:** at least 90% on the held-out set. - **Eval-set accuracy:** at least 85% on `evals/tasks.yaml` after component 2. **Stop and report** if that needs more than description tuning. - **Latency:** p95 at or below 800 ms on the RTX 6000 with the 14b router model resident. The router model must not be evicted: check `/api/ps` before and after a 100-call run. - **Failure paths:** Ollama down, model missing, coverage below the floor and low confidence each walk the cascade, covered by tests with a fake Ollama. - **Tests:** the full suite is green, including `test_admin_knob_coverage.py`. ## Out of scope - **Replacing gaming mode** with an automatic Ollama reachability check. It's a separate lift, and when it lands this mode inherits it for free. - **What to do about the trained encoder head.** It's a separate decision, but the evidence above is a strong argument for turning it off. - **Other routes:** hosted Jev (cloud-only, bans benchmarking, sends raw task text off-box), OpenJev 27B (no FP8 on Turing, CC BY-NC, evicts the router model), and fine-tuning. - **Defaults:** making `local_decision` the `config.yaml` default. It's an overlay opt-in until it has run on live traffic. ## Guardrails - **Worktree:** do the work in its own worktree (North Star #3). There are 15 open ones, and `feat/local-encoder-accuracy-rebuild` touches classifier code. Commit by explicit path. - **Live router:** exercise the admin portal on the 8081 sandbox, never on production 8080, and don't edit `config.local.yaml` (see the `6krrt-ops` skill). - **Live database:** open `router.db` read-only. ## Open questions for planning 1. Isolate noise before calling or not (component 1)? Decide from a measurement. 2. Should tier's second call be a separate request, or both questions in one prompt read from two token positions? Separate is simpler; measure the latency before choosing the clever option. 3. What `num_ctx` gives the best VRAM trade-off against the 8000-char `max_input_chars` clamp?