Files
6krrt/plans/local-decision-classifier.md

158 lines
8.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# `classifier.mode: local_decision` — a Jev-style classifier on the local GPU
Status: planned -- plan reviewed 2026-09-28; gate G (tuning + measurement) decides whether the wiring proceeds
Date: 2026-09-27
Prototype: `plans/local-decision-classifier-prototype.py` (a benchmark you can run, NOT the
implementation)
Held-out prompts: `plans/local-decision-classifier-heldout.yaml`
Background: TypeSafe's hosted "Jev" decision model and the OpenJev/SemIf open
copies. We are copying the **technique**, not using either product.
## Goal
Add a fourth `classifier.mode`, `local_decision`. It asks a small local LLM
(`qwen3.5:4b` on Ollama) one multiple-choice question and reads the answer
from the **first output token's logprobs**, normalised over the option letters
that were offered. There is no text generation, JSON or reasoning trace, so the
runaway-reasoning failure that retired `local_llm` can't happen, and it is
calibrated far better than `local_encoder`.
## Evidence (2026-09-27, RTX 6000 24 GB, measured with the prototype)
46 eval tasks and 30 held-out tasks, each at 3 noise levels (`_wrap_agent_noise`):
| | `evals/tasks.yaml` | held-out | p50 / p95 ms |
|---|---|---|---|
| `local_decision` (`qwen3.5:4b`, GPU) | 71% | 100% | 340-385 / 510-610 |
| `local_encoder` + trained head (live, CPU) | 100%\* | 47% | 80-115 / 100-360 |
| `local_encoder` zero-shot (CPU) | 39% | 60% | 80-115 / 100-350 |
\* The head was fitted on a synthetic corpus derived from `evals/tasks.yaml`, so
this is leakage. On held-out prompts the head is **worse** than zero-shot.
- **Calibration:** `local_decision` averages 0.96 confidence when right and 0.71 when wrong. The encoder
zero-shot sits at 0.2-0.3 on both.
- **Eval-set misses:** 18 were `coding_general`→`reasoning_math` (algorithm tasks), 18 were
`file_summarization`→`debugging`/`reasoning_math`/`docs_writing`. Both look like
option-description problems (component 2).
- **VRAM:** `qwen3.5:4b` holds about 7 GB at `num_ctx` 8192 next to `qwen2.5-coder-router:14b`.
The card sat at 23.4/24 GB. It's tight, so measure again when planning.
- **Weak held-out set:** a single author wrote all 30 prompts and they are short and clean. Treat
100% as a ceiling, not a forecast.
## Mechanism (proven by the prototype)
Ollama 0.22 native `POST /api/chat`:
`think: false`, `logprobs: true`, `top_logprobs: 20`,
`options: {num_predict: 1, temperature: 0, num_ctx: N}`. The system prompt says
"answer with the letter only". The options are `A.`–`J.` followed by a natural-language description.
Parse `logprobs[0].top_logprobs`:
- strip each token and remove a trailing `.`
- sum the probability mass for each option letter (`"B"` and `" B"` both count)
- confidence = winner's mass ÷ total option mass
- coverage = total option mass
Tokens like `<think>` show up in the top 20 and must be ignored.
## Components
### 1. `src/local_decision.py` (stdlib + `requests`, no new dependency)
- **`classify_choice(text, options, *, base_url, model, num_ctx, timeout)`** returns
`(label, confidence, coverage)`. It is generic over the option list, so tier
reuses it.
- **Coverage floor:** below a configurable `coverage_min` (the model didn't answer
with a letter), raise. Don't guess.
- **Pure parse function:** split the logprob parsing into its own function so tests
can drive it with recorded responses and no Ollama.
- **Noise:** decide whether to run `local_encoder._isolate_task_text` first. The prototype
sent the raw noisy text and noise cost about 2 points; measure both.
### 2. Its own option descriptions
Give it a separate `_DECISION_DESCRIPTIONS` map. **Do not edit**
`local_encoder._CATEGORY_DESCRIPTIONS`, because that would shift the encoder and
invalidate its head. Tune it against the two confusions above, and re-measure after
every change on **both** sets so nobody overfits `evals/tasks.yaml` the way the head did.
### 3. Wiring (follow how `local_encoder` was added)
- **`src/config.py`:** add `"local_decision"` to the `mode` Literal (line ~1187). Add a
`classifier.decision` block (`base_url`, `model`, `num_ctx`, `timeout_s`,
`confidence_min`, `coverage_min`, `tier_enabled`) and a validator that requires
it when the mode is selected. Metering (line ~1712): this mode is an HTTP call to
Ollama, so it meters like `local_llm` via its own `base_url`, not like the in-process encoder.
- **`src/dispatcher.py`:**
- add `_classify_via_local_decision` and a branch in `_classify_via_configured_mode`
- add a startup readiness check in `_ensure_classifier_mode_ready` (model is pulled and
logprobs come back)
- on success record `source="classifier"`
- treat below-threshold confidence as a failure that walks the **unchanged** cascade
- **Gating:** it uses the GPU, so it obeys `_local_classifier_skip_reason()`
(`local_compute.enabled` plus the circuit breaker), the same as `local_llm`.
- **Admin:** every new scalar knob gets a control on the Classifier card (North Star #1).
`tests/test_admin_knob_coverage.py` enforces this.
### 4. Tier (behind `tier_enabled`, default false)
- **How:** a second `classify_choice` call with three options (cheap/simple, general,
frontier/high-stakes), fired **in parallel** with the category call so it adds no
wall-clock time.
- **Why it defaults off:** no gold tier labels exist yet. Until they do, tier keeps
falling back to `fallback_tier`, exactly as `local_encoder` does now.
- **Output:** a gold tier set of about 30 labelled prompts, used to decide the default.
### 5. Evaluation harness
- **Backend flag:** add a `--backend encoder|decision` flag to `src/eval_classifier.py`
(latency is reported per call).
- **Held-out set:** promote the held-out prompts to `evals/` and add a `--tasks` default for them.
- **Head flag:** add a flag that runs the encoder with the head disabled, so the leakage
stays visible.
- **Report:** accuracy for both sets, calibration (confidence correct vs wrong), and p50/p95.
### 6. Docs
- **CLAUDE.md:** the "Which implementation is PRIMARY" section, which also still says
the live encoder is `device: cuda` while the overlay says `cpu`. Fix that.
- **Also:** `docs/local-models.md` and `docs/admin-portal.md`.
## Acceptance
- **Held-out accuracy:** at least 90% on the held-out set.
- **Eval-set accuracy:** at least 85% on `evals/tasks.yaml` after component 2. **Stop and
report** if that needs more than description tuning.
- **Latency:** p95 at or below 800 ms on the RTX 6000 with the 14b router model resident.
The router model must not be evicted: check `/api/ps` before and after a 100-call run.
- **Failure paths:** Ollama down, model missing, coverage below the floor and low
confidence each walk the cascade, covered by tests with a fake Ollama.
- **Tests:** the full suite is green, including `test_admin_knob_coverage.py`.
## Out of scope
- **Replacing gaming mode** with an automatic Ollama reachability check. It's a separate
lift, and when it lands this mode inherits it for free.
- **What to do about the trained encoder head.** It's a separate decision, but the
evidence above is a strong argument for turning it off.
- **Other routes:** hosted Jev (cloud-only, bans benchmarking, sends raw task text off-box),
OpenJev 27B (no FP8 on Turing, CC BY-NC, evicts the router model), and fine-tuning.
- **Defaults:** making `local_decision` the `config.yaml` default. It's an overlay opt-in
until it has run on live traffic.
## Guardrails
- **Worktree:** do the work in its own worktree (North Star #3). There are 15 open ones,
and `feat/local-encoder-accuracy-rebuild` touches classifier code. Commit by explicit path.
- **Live router:** exercise the admin portal on the 8081 sandbox, never on production 8080,
and don't edit `config.local.yaml` (see the `6krrt-ops` skill).
- **Live database:** open `router.db` read-only.
## Open questions for planning
1. Isolate noise before calling or not (component 1)? Decide from a measurement.
2. Should tier's second call be a separate request, or both questions in one prompt
read from two token positions? Separate is simpler; measure the latency before
choosing the clever option.
3. What `num_ctx` gives the best VRAM trade-off against the 8000-char
`max_input_chars` clamp?