Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
158 lines
8.1 KiB
Markdown
158 lines
8.1 KiB
Markdown
# `classifier.mode: local_decision` — a Jev-style classifier on the local GPU
|
||
|
||
Status: planned -- plan reviewed 2026-09-28; gate G (tuning + measurement) decides whether the wiring proceeds
|
||
Date: 2026-09-27
|
||
Prototype: `plans/local-decision-classifier-prototype.py` (a benchmark you can run, NOT the
|
||
implementation)
|
||
Held-out prompts: `plans/local-decision-classifier-heldout.yaml`
|
||
Background: TypeSafe's hosted "Jev" decision model and the OpenJev/SemIf open
|
||
copies. We are copying the **technique**, not using either product.
|
||
|
||
## Goal
|
||
|
||
Add a fourth `classifier.mode`, `local_decision`. It asks a small local LLM
|
||
(`qwen3.5:4b` on Ollama) one multiple-choice question and reads the answer
|
||
from the **first output token's logprobs**, normalised over the option letters
|
||
that were offered. There is no text generation, JSON or reasoning trace, so the
|
||
runaway-reasoning failure that retired `local_llm` can't happen, and it is
|
||
calibrated far better than `local_encoder`.
|
||
|
||
## Evidence (2026-09-27, RTX 6000 24 GB, measured with the prototype)
|
||
|
||
46 eval tasks and 30 held-out tasks, each at 3 noise levels (`_wrap_agent_noise`):
|
||
|
||
| | `evals/tasks.yaml` | held-out | p50 / p95 ms |
|
||
|---|---|---|---|
|
||
| `local_decision` (`qwen3.5:4b`, GPU) | 71% | 100% | 340-385 / 510-610 |
|
||
| `local_encoder` + trained head (live, CPU) | 100%\* | 47% | 80-115 / 100-360 |
|
||
| `local_encoder` zero-shot (CPU) | 39% | 60% | 80-115 / 100-350 |
|
||
|
||
\* The head was fitted on a synthetic corpus derived from `evals/tasks.yaml`, so
|
||
this is leakage. On held-out prompts the head is **worse** than zero-shot.
|
||
|
||
- **Calibration:** `local_decision` averages 0.96 confidence when right and 0.71 when wrong. The encoder
|
||
zero-shot sits at 0.2-0.3 on both.
|
||
- **Eval-set misses:** 18 were `coding_general`→`reasoning_math` (algorithm tasks), 18 were
|
||
`file_summarization`→`debugging`/`reasoning_math`/`docs_writing`. Both look like
|
||
option-description problems (component 2).
|
||
- **VRAM:** `qwen3.5:4b` holds about 7 GB at `num_ctx` 8192 next to `qwen2.5-coder-router:14b`.
|
||
The card sat at 23.4/24 GB. It's tight, so measure again when planning.
|
||
- **Weak held-out set:** a single author wrote all 30 prompts and they are short and clean. Treat
|
||
100% as a ceiling, not a forecast.
|
||
|
||
## Mechanism (proven by the prototype)
|
||
|
||
Ollama 0.22 native `POST /api/chat`:
|
||
`think: false`, `logprobs: true`, `top_logprobs: 20`,
|
||
`options: {num_predict: 1, temperature: 0, num_ctx: N}`. The system prompt says
|
||
"answer with the letter only". The options are `A.`–`J.` followed by a natural-language description.
|
||
Parse `logprobs[0].top_logprobs`:
|
||
- strip each token and remove a trailing `.`
|
||
- sum the probability mass for each option letter (`"B"` and `" B"` both count)
|
||
- confidence = winner's mass ÷ total option mass
|
||
- coverage = total option mass
|
||
|
||
Tokens like `<think>` show up in the top 20 and must be ignored.
|
||
|
||
## Components
|
||
|
||
### 1. `src/local_decision.py` (stdlib + `requests`, no new dependency)
|
||
|
||
- **`classify_choice(text, options, *, base_url, model, num_ctx, timeout)`** returns
|
||
`(label, confidence, coverage)`. It is generic over the option list, so tier
|
||
reuses it.
|
||
- **Coverage floor:** below a configurable `coverage_min` (the model didn't answer
|
||
with a letter), raise. Don't guess.
|
||
- **Pure parse function:** split the logprob parsing into its own function so tests
|
||
can drive it with recorded responses and no Ollama.
|
||
- **Noise:** decide whether to run `local_encoder._isolate_task_text` first. The prototype
|
||
sent the raw noisy text and noise cost about 2 points; measure both.
|
||
|
||
### 2. Its own option descriptions
|
||
|
||
Give it a separate `_DECISION_DESCRIPTIONS` map. **Do not edit**
|
||
`local_encoder._CATEGORY_DESCRIPTIONS`, because that would shift the encoder and
|
||
invalidate its head. Tune it against the two confusions above, and re-measure after
|
||
every change on **both** sets so nobody overfits `evals/tasks.yaml` the way the head did.
|
||
|
||
### 3. Wiring (follow how `local_encoder` was added)
|
||
|
||
- **`src/config.py`:** add `"local_decision"` to the `mode` Literal (line ~1187). Add a
|
||
`classifier.decision` block (`base_url`, `model`, `num_ctx`, `timeout_s`,
|
||
`confidence_min`, `coverage_min`, `tier_enabled`) and a validator that requires
|
||
it when the mode is selected. Metering (line ~1712): this mode is an HTTP call to
|
||
Ollama, so it meters like `local_llm` via its own `base_url`, not like the in-process encoder.
|
||
- **`src/dispatcher.py`:**
|
||
- add `_classify_via_local_decision` and a branch in `_classify_via_configured_mode`
|
||
- add a startup readiness check in `_ensure_classifier_mode_ready` (model is pulled and
|
||
logprobs come back)
|
||
- on success record `source="classifier"`
|
||
- treat below-threshold confidence as a failure that walks the **unchanged** cascade
|
||
- **Gating:** it uses the GPU, so it obeys `_local_classifier_skip_reason()`
|
||
(`local_compute.enabled` plus the circuit breaker), the same as `local_llm`.
|
||
- **Admin:** every new scalar knob gets a control on the Classifier card (North Star #1).
|
||
`tests/test_admin_knob_coverage.py` enforces this.
|
||
|
||
### 4. Tier (behind `tier_enabled`, default false)
|
||
|
||
- **How:** a second `classify_choice` call with three options (cheap/simple, general,
|
||
frontier/high-stakes), fired **in parallel** with the category call so it adds no
|
||
wall-clock time.
|
||
- **Why it defaults off:** no gold tier labels exist yet. Until they do, tier keeps
|
||
falling back to `fallback_tier`, exactly as `local_encoder` does now.
|
||
- **Output:** a gold tier set of about 30 labelled prompts, used to decide the default.
|
||
|
||
### 5. Evaluation harness
|
||
|
||
- **Backend flag:** add a `--backend encoder|decision` flag to `src/eval_classifier.py`
|
||
(latency is reported per call).
|
||
- **Held-out set:** promote the held-out prompts to `evals/` and add a `--tasks` default for them.
|
||
- **Head flag:** add a flag that runs the encoder with the head disabled, so the leakage
|
||
stays visible.
|
||
- **Report:** accuracy for both sets, calibration (confidence correct vs wrong), and p50/p95.
|
||
|
||
### 6. Docs
|
||
|
||
- **CLAUDE.md:** the "Which implementation is PRIMARY" section, which also still says
|
||
the live encoder is `device: cuda` while the overlay says `cpu`. Fix that.
|
||
- **Also:** `docs/local-models.md` and `docs/admin-portal.md`.
|
||
|
||
## Acceptance
|
||
|
||
- **Held-out accuracy:** at least 90% on the held-out set.
|
||
- **Eval-set accuracy:** at least 85% on `evals/tasks.yaml` after component 2. **Stop and
|
||
report** if that needs more than description tuning.
|
||
- **Latency:** p95 at or below 800 ms on the RTX 6000 with the 14b router model resident.
|
||
The router model must not be evicted: check `/api/ps` before and after a 100-call run.
|
||
- **Failure paths:** Ollama down, model missing, coverage below the floor and low
|
||
confidence each walk the cascade, covered by tests with a fake Ollama.
|
||
- **Tests:** the full suite is green, including `test_admin_knob_coverage.py`.
|
||
|
||
## Out of scope
|
||
|
||
- **Replacing gaming mode** with an automatic Ollama reachability check. It's a separate
|
||
lift, and when it lands this mode inherits it for free.
|
||
- **What to do about the trained encoder head.** It's a separate decision, but the
|
||
evidence above is a strong argument for turning it off.
|
||
- **Other routes:** hosted Jev (cloud-only, bans benchmarking, sends raw task text off-box),
|
||
OpenJev 27B (no FP8 on Turing, CC BY-NC, evicts the router model), and fine-tuning.
|
||
- **Defaults:** making `local_decision` the `config.yaml` default. It's an overlay opt-in
|
||
until it has run on live traffic.
|
||
|
||
## Guardrails
|
||
|
||
- **Worktree:** do the work in its own worktree (North Star #3). There are 15 open ones,
|
||
and `feat/local-encoder-accuracy-rebuild` touches classifier code. Commit by explicit path.
|
||
- **Live router:** exercise the admin portal on the 8081 sandbox, never on production 8080,
|
||
and don't edit `config.local.yaml` (see the `6krrt-ops` skill).
|
||
- **Live database:** open `router.db` read-only.
|
||
|
||
## Open questions for planning
|
||
|
||
1. Isolate noise before calling or not (component 1)? Decide from a measurement.
|
||
2. Should tier's second call be a separate request, or both questions in one prompt
|
||
read from two token positions? Separate is simpler; measure the latency before
|
||
choosing the clever option.
|
||
3. What `num_ctx` gives the best VRAM trade-off against the 8000-char
|
||
`max_input_chars` clamp?
|