Files
6krrt/plans/local-decision-classifier.md

8.1 KiB
Raw Permalink Blame History

classifier.mode: local_decision — a Jev-style classifier on the local GPU

Status: planned -- plan reviewed 2026-09-28; gate G (tuning + measurement) decides whether the wiring proceeds Date: 2026-09-27 Prototype: plans/local-decision-classifier-prototype.py (a benchmark you can run, NOT the implementation) Held-out prompts: plans/local-decision-classifier-heldout.yaml Background: TypeSafe's hosted "Jev" decision model and the OpenJev/SemIf open copies. We are copying the technique, not using either product.

Goal

Add a fourth classifier.mode, local_decision. It asks a small local LLM (qwen3.5:4b on Ollama) one multiple-choice question and reads the answer from the first output token's logprobs, normalised over the option letters that were offered. There is no text generation, JSON or reasoning trace, so the runaway-reasoning failure that retired local_llm can't happen, and it is calibrated far better than local_encoder.

Evidence (2026-09-27, RTX 6000 24 GB, measured with the prototype)

46 eval tasks and 30 held-out tasks, each at 3 noise levels (_wrap_agent_noise):

evals/tasks.yaml held-out p50 / p95 ms
local_decision (qwen3.5:4b, GPU) 71% 100% 340-385 / 510-610
local_encoder + trained head (live, CPU) 100%* 47% 80-115 / 100-360
local_encoder zero-shot (CPU) 39% 60% 80-115 / 100-350

* The head was fitted on a synthetic corpus derived from evals/tasks.yaml, so this is leakage. On held-out prompts the head is worse than zero-shot.

  • Calibration: local_decision averages 0.96 confidence when right and 0.71 when wrong. The encoder zero-shot sits at 0.2-0.3 on both.
  • Eval-set misses: 18 were coding_general→reasoning_math (algorithm tasks), 18 were file_summarization→debugging/reasoning_math/docs_writing. Both look like option-description problems (component 2).
  • VRAM: qwen3.5:4b holds about 7 GB at num_ctx 8192 next to qwen2.5-coder-router:14b. The card sat at 23.4/24 GB. It's tight, so measure again when planning.
  • Weak held-out set: a single author wrote all 30 prompts and they are short and clean. Treat 100% as a ceiling, not a forecast.

Mechanism (proven by the prototype)

Ollama 0.22 native POST /api/chat: think: false, logprobs: true, top_logprobs: 20, options: {num_predict: 1, temperature: 0, num_ctx: N}. The system prompt says "answer with the letter only". The options are A.–J. followed by a natural-language description. Parse logprobs[0].top_logprobs:

  • strip each token and remove a trailing .
  • sum the probability mass for each option letter ("B" and " B" both count)
  • confidence = winner's mass ÷ total option mass
  • coverage = total option mass

Tokens like <think> show up in the top 20 and must be ignored.

Components

1. src/local_decision.py (stdlib + requests, no new dependency)

  • classify_choice(text, options, *, base_url, model, num_ctx, timeout) returns (label, confidence, coverage). It is generic over the option list, so tier reuses it.
  • Coverage floor: below a configurable coverage_min (the model didn't answer with a letter), raise. Don't guess.
  • Pure parse function: split the logprob parsing into its own function so tests can drive it with recorded responses and no Ollama.
  • Noise: decide whether to run local_encoder._isolate_task_text first. The prototype sent the raw noisy text and noise cost about 2 points; measure both.

2. Its own option descriptions

Give it a separate _DECISION_DESCRIPTIONS map. Do not edit local_encoder._CATEGORY_DESCRIPTIONS, because that would shift the encoder and invalidate its head. Tune it against the two confusions above, and re-measure after every change on both sets so nobody overfits evals/tasks.yaml the way the head did.

3. Wiring (follow how local_encoder was added)

  • src/config.py: add "local_decision" to the mode Literal (line ~1187). Add a classifier.decision block (base_url, model, num_ctx, timeout_s, confidence_min, coverage_min, tier_enabled) and a validator that requires it when the mode is selected. Metering (line ~1712): this mode is an HTTP call to Ollama, so it meters like local_llm via its own base_url, not like the in-process encoder.
  • src/dispatcher.py:
    • add _classify_via_local_decision and a branch in _classify_via_configured_mode
    • add a startup readiness check in _ensure_classifier_mode_ready (model is pulled and logprobs come back)
    • on success record source="classifier"
    • treat below-threshold confidence as a failure that walks the unchanged cascade
  • Gating: it uses the GPU, so it obeys _local_classifier_skip_reason() (local_compute.enabled plus the circuit breaker), the same as local_llm.
  • Admin: every new scalar knob gets a control on the Classifier card (North Star #1). tests/test_admin_knob_coverage.py enforces this.

4. Tier (behind tier_enabled, default false)

  • How: a second classify_choice call with three options (cheap/simple, general, frontier/high-stakes), fired in parallel with the category call so it adds no wall-clock time.
  • Why it defaults off: no gold tier labels exist yet. Until they do, tier keeps falling back to fallback_tier, exactly as local_encoder does now.
  • Output: a gold tier set of about 30 labelled prompts, used to decide the default.

5. Evaluation harness

  • Backend flag: add a --backend encoder|decision flag to src/eval_classifier.py (latency is reported per call).
  • Held-out set: promote the held-out prompts to evals/ and add a --tasks default for them.
  • Head flag: add a flag that runs the encoder with the head disabled, so the leakage stays visible.
  • Report: accuracy for both sets, calibration (confidence correct vs wrong), and p50/p95.

6. Docs

  • CLAUDE.md: the "Which implementation is PRIMARY" section, which also still says the live encoder is device: cuda while the overlay says cpu. Fix that.
  • Also: docs/local-models.md and docs/admin-portal.md.

Acceptance

  • Held-out accuracy: at least 90% on the held-out set.
  • Eval-set accuracy: at least 85% on evals/tasks.yaml after component 2. Stop and report if that needs more than description tuning.
  • Latency: p95 at or below 800 ms on the RTX 6000 with the 14b router model resident. The router model must not be evicted: check /api/ps before and after a 100-call run.
  • Failure paths: Ollama down, model missing, coverage below the floor and low confidence each walk the cascade, covered by tests with a fake Ollama.
  • Tests: the full suite is green, including test_admin_knob_coverage.py.

Out of scope

  • Replacing gaming mode with an automatic Ollama reachability check. It's a separate lift, and when it lands this mode inherits it for free.
  • What to do about the trained encoder head. It's a separate decision, but the evidence above is a strong argument for turning it off.
  • Other routes: hosted Jev (cloud-only, bans benchmarking, sends raw task text off-box), OpenJev 27B (no FP8 on Turing, CC BY-NC, evicts the router model), and fine-tuning.
  • Defaults: making local_decision the config.yaml default. It's an overlay opt-in until it has run on live traffic.

Guardrails

  • Worktree: do the work in its own worktree (North Star #3). There are 15 open ones, and feat/local-encoder-accuracy-rebuild touches classifier code. Commit by explicit path.
  • Live router: exercise the admin portal on the 8081 sandbox, never on production 8080, and don't edit config.local.yaml (see the 6krrt-ops skill).
  • Live database: open router.db read-only.

Open questions for planning

  1. Isolate noise before calling or not (component 1)? Decide from a measurement.
  2. Should tier's second call be a separate request, or both questions in one prompt read from two token positions? Separate is simpler; measure the latency before choosing the clever option.
  3. What num_ctx gives the best VRAM trade-off against the 8000-char max_input_chars clamp?