Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
8.1 KiB
classifier.mode: local_decision — a Jev-style classifier on the local GPU
Status: planned -- plan reviewed 2026-09-28; gate G (tuning + measurement) decides whether the wiring proceeds
Date: 2026-09-27
Prototype: plans/local-decision-classifier-prototype.py (a benchmark you can run, NOT the
implementation)
Held-out prompts: plans/local-decision-classifier-heldout.yaml
Background: TypeSafe's hosted "Jev" decision model and the OpenJev/SemIf open
copies. We are copying the technique, not using either product.
Goal
Add a fourth classifier.mode, local_decision. It asks a small local LLM
(qwen3.5:4b on Ollama) one multiple-choice question and reads the answer
from the first output token's logprobs, normalised over the option letters
that were offered. There is no text generation, JSON or reasoning trace, so the
runaway-reasoning failure that retired local_llm can't happen, and it is
calibrated far better than local_encoder.
Evidence (2026-09-27, RTX 6000 24 GB, measured with the prototype)
46 eval tasks and 30 held-out tasks, each at 3 noise levels (_wrap_agent_noise):
evals/tasks.yaml |
held-out | p50 / p95 ms | |
|---|---|---|---|
local_decision (qwen3.5:4b, GPU) |
71% | 100% | 340-385 / 510-610 |
local_encoder + trained head (live, CPU) |
100%* | 47% | 80-115 / 100-360 |
local_encoder zero-shot (CPU) |
39% | 60% | 80-115 / 100-350 |
* The head was fitted on a synthetic corpus derived from evals/tasks.yaml, so
this is leakage. On held-out prompts the head is worse than zero-shot.
- Calibration:
local_decisionaverages 0.96 confidence when right and 0.71 when wrong. The encoder zero-shot sits at 0.2-0.3 on both. - Eval-set misses: 18 were
coding_general→reasoning_math(algorithm tasks), 18 werefile_summarization→debugging/reasoning_math/docs_writing. Both look like option-description problems (component 2). - VRAM:
qwen3.5:4bholds about 7 GB atnum_ctx8192 next toqwen2.5-coder-router:14b. The card sat at 23.4/24 GB. It's tight, so measure again when planning. - Weak held-out set: a single author wrote all 30 prompts and they are short and clean. Treat 100% as a ceiling, not a forecast.
Mechanism (proven by the prototype)
Ollama 0.22 native POST /api/chat:
think: false, logprobs: true, top_logprobs: 20,
options: {num_predict: 1, temperature: 0, num_ctx: N}. The system prompt says
"answer with the letter only". The options are A.–J. followed by a natural-language description.
Parse logprobs[0].top_logprobs:
- strip each token and remove a trailing
. - sum the probability mass for each option letter (
"B"and" B"both count) - confidence = winner's mass ÷ total option mass
- coverage = total option mass
Tokens like <think> show up in the top 20 and must be ignored.
Components
1. src/local_decision.py (stdlib + requests, no new dependency)
classify_choice(text, options, *, base_url, model, num_ctx, timeout)returns(label, confidence, coverage). It is generic over the option list, so tier reuses it.- Coverage floor: below a configurable
coverage_min(the model didn't answer with a letter), raise. Don't guess. - Pure parse function: split the logprob parsing into its own function so tests can drive it with recorded responses and no Ollama.
- Noise: decide whether to run
local_encoder._isolate_task_textfirst. The prototype sent the raw noisy text and noise cost about 2 points; measure both.
2. Its own option descriptions
Give it a separate _DECISION_DESCRIPTIONS map. Do not edit
local_encoder._CATEGORY_DESCRIPTIONS, because that would shift the encoder and
invalidate its head. Tune it against the two confusions above, and re-measure after
every change on both sets so nobody overfits evals/tasks.yaml the way the head did.
3. Wiring (follow how local_encoder was added)
src/config.py: add"local_decision"to themodeLiteral (line ~1187). Add aclassifier.decisionblock (base_url,model,num_ctx,timeout_s,confidence_min,coverage_min,tier_enabled) and a validator that requires it when the mode is selected. Metering (line ~1712): this mode is an HTTP call to Ollama, so it meters likelocal_llmvia its ownbase_url, not like the in-process encoder.src/dispatcher.py:- add
_classify_via_local_decisionand a branch in_classify_via_configured_mode - add a startup readiness check in
_ensure_classifier_mode_ready(model is pulled and logprobs come back) - on success record
source="classifier" - treat below-threshold confidence as a failure that walks the unchanged cascade
- add
- Gating: it uses the GPU, so it obeys
_local_classifier_skip_reason()(local_compute.enabledplus the circuit breaker), the same aslocal_llm. - Admin: every new scalar knob gets a control on the Classifier card (North Star #1).
tests/test_admin_knob_coverage.pyenforces this.
4. Tier (behind tier_enabled, default false)
- How: a second
classify_choicecall with three options (cheap/simple, general, frontier/high-stakes), fired in parallel with the category call so it adds no wall-clock time. - Why it defaults off: no gold tier labels exist yet. Until they do, tier keeps
falling back to
fallback_tier, exactly aslocal_encoderdoes now. - Output: a gold tier set of about 30 labelled prompts, used to decide the default.
5. Evaluation harness
- Backend flag: add a
--backend encoder|decisionflag tosrc/eval_classifier.py(latency is reported per call). - Held-out set: promote the held-out prompts to
evals/and add a--tasksdefault for them. - Head flag: add a flag that runs the encoder with the head disabled, so the leakage stays visible.
- Report: accuracy for both sets, calibration (confidence correct vs wrong), and p50/p95.
6. Docs
- CLAUDE.md: the "Which implementation is PRIMARY" section, which also still says
the live encoder is
device: cudawhile the overlay sayscpu. Fix that. - Also:
docs/local-models.mdanddocs/admin-portal.md.
Acceptance
- Held-out accuracy: at least 90% on the held-out set.
- Eval-set accuracy: at least 85% on
evals/tasks.yamlafter component 2. Stop and report if that needs more than description tuning. - Latency: p95 at or below 800 ms on the RTX 6000 with the 14b router model resident.
The router model must not be evicted: check
/api/psbefore and after a 100-call run. - Failure paths: Ollama down, model missing, coverage below the floor and low confidence each walk the cascade, covered by tests with a fake Ollama.
- Tests: the full suite is green, including
test_admin_knob_coverage.py.
Out of scope
- Replacing gaming mode with an automatic Ollama reachability check. It's a separate lift, and when it lands this mode inherits it for free.
- What to do about the trained encoder head. It's a separate decision, but the evidence above is a strong argument for turning it off.
- Other routes: hosted Jev (cloud-only, bans benchmarking, sends raw task text off-box), OpenJev 27B (no FP8 on Turing, CC BY-NC, evicts the router model), and fine-tuning.
- Defaults: making
local_decisiontheconfig.yamldefault. It's an overlay opt-in until it has run on live traffic.
Guardrails
- Worktree: do the work in its own worktree (North Star #3). There are 15 open ones,
and
feat/local-encoder-accuracy-rebuildtouches classifier code. Commit by explicit path. - Live router: exercise the admin portal on the 8081 sandbox, never on production 8080,
and don't edit
config.local.yaml(see the6krrt-opsskill). - Live database: open
router.dbread-only.
Open questions for planning
- Isolate noise before calling or not (component 1)? Decide from a measurement.
- Should tier's second call be a separate request, or both questions in one prompt read from two token positions? Separate is simpler; measure the latency before choosing the clever option.
- What
num_ctxgives the best VRAM trade-off against the 8000-charmax_input_charsclamp?