Files
6krrt/plans/local-decision-classifier-results.md

211 lines
9.6 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Gate G — local_decision measurement results
Status: done -- gate G measured; VRAM passes with the 14b at 16k and local vision off (section 9); accuracy and latency pass
Date: 2026-09-28
Plan: `plans/local-decision-classifier.md`
Hardware: Quadro RTX 6000, 24 GB (24576 MiB total)
Models: `qwen3.5:4b` (classifier), `qwen2.5-coder-router:14b` (router, resident at 32768 ctx = 17.76 GB)
Ollama: `http://localhost:11434`
## Headline
| Criterion | Measured | Threshold | Verdict |
|---|---|---|---|
| Eval-set accuracy (`evals/tasks.yaml`) | **97.8%** (tuned) | ≥ 85% | PASS |
| Held-out accuracy (`evals/heldout.yaml`) | **90.0%** (tuned) | ≥ 90% | PASS (at threshold) |
| p95 latency (14b resident) | **502 ms** | ≤ 800 ms | PASS |
| 14b router resident after 100 calls | **EVICTED** | must stay resident | **FAIL** |
**Gate verdict: STOP.** The 14b router model is evicted whenever `qwen3.5:4b` is loaded —
on this 24 GB card there is no `num_ctx` at which the two coexist. Per the plan, "STOP AND
REPORT after G if ... the 14b router model was evicted. Do NOT start todos 4-17 in that case."
Todos 4-17 are NOT started.
---
## 1. Accuracy per set
Measured with the eval harness (`--backend decision --decision-base-url http://localhost:11434`),
majority-vote overall accuracy over the clean/short/long noise variants of each task.
### Baseline (original `_DECISION_DESCRIPTIONS`)
| Set | clean | short | long | overall |
|---|---|---|---|---|
| `evals/tasks.yaml` (46) | 0.739 | 0.783 | 0.761 | **0.761** |
| `evals/heldout.yaml` (30) | 0.800 | 0.933 | 0.900 | **0.967** |
Baseline confusion (eval-set, dominant cells): `coding_general → debugging` (14) and
`→ reasoning_math` (4); `file_summarization → debugging` (9). The `debugging` description
was attracting both code-tracing and code-summary tasks.
### After description tuning
`_DECISION_DESCRIPTIONS` changes: widen `coding_general` to include "tracing what existing
code returns" and "implementing a feature or algorithm"; narrow `debugging` to "fixing a bug";
narrow `reasoning_math` to "pure math or logic word problem (no code)"; soften
`file_summarization` to "summarizing or explaining".
| Set | clean | short | long | overall |
|---|---|---|---|---|
| `evals/tasks.yaml` (46) | 0.978 | 0.978 | 0.978 | **0.978** |
| `evals/heldout.yaml` (30) | 0.800 | 0.900 | 0.833 | **0.900** |
Tuning lifted the eval-set from 76.1% → 97.8% (the confusion cells are gone: eval-set
`coding_general` 27/27, `file_summarization` 18/18, `debugging` 18/18). It cost 6.7 points on
held-out (96.7% → 90.0%), landing exactly on the 90% threshold. Held-out residual misses:
`h2` coding_general→file_summarization, `h30` general_chat→file_summarization, `h9`
debugging→general_chat — all low-confidence (0.33-0.49) edge cases the confidence cascade
would absorb. Further tuning to chase these risks overfitting `evals/tasks.yaml` (the exact
failure the plan warns about), so tuning stops here.
## 2. Calibration (tuned)
Mean confidence when correct vs incorrect (per-noise-level rows).
| Set | correct mean | incorrect mean |
|---|---|---|
| `evals/tasks.yaml` | 0.798 | 0.350 |
| `evals/heldout.yaml` | 0.802 | 0.413 |
Calibrated in the right direction on both sets: the model is roughly twice as confident when
right as when wrong. (Baseline was 0.833/0.438 eval-set, 0.815/0.300 heldout — the tuning
moved both closer together but kept correct ≫ incorrect.)
## 3. Latency (p50 / p95)
Per-call latency of a single `classify_choice` over the eval-set, `qwen3.5:4b`, `num_ctx 8192`,
with the 14b router resident (the production steady state).
```
single category call: n=46 p50=363ms p95=502ms p99=598ms mean=379ms
```
**p95 502 ms ≤ 800 ms — PASS.**
## 4. Open question (a) — noise isolation
Compare `classify_choice` on raw wrapped (noisy) text vs text passed through
`local_encoder._isolate_task_text` first.
| Set | no-isolation (short+long) | isolated (short+long) | clean baseline |
|---|---|---|---|
| `evals/tasks.yaml` | 0.772 | 0.739 | 0.739 |
| `evals/heldout.yaml` | 0.917 | 0.800 | 0.800 |
**Isolation HURTS** on both sets (−3.3 pt eval, −11.7 pt heldout). For `local_decision` the
fenced-code/tool content `_isolate_task_text` strips is itself the signal (code-tracing,
file-summary tasks). **Do NOT run `_isolate_task_text` before `classify_choice`; send raw text.**
This overrides the prototype's earlier "noise cost about 2 points" note — for this backend
isolation is strictly worse.
## 5. Open question (b) — tier call strategy
Measure separate-request latency vs a joint prompt reading two token positions.
| strategy | p50 | p95 | reliability |
|---|---|---|---|
| category call (10 options) | 361 ms | 595 ms | — |
| tier call (3 options, separate) | 326 ms | 553 ms | — |
| separate, sequential total | 689 ms | 1154 ms | 100% |
| joint prompt, two token positions | 434 ms | 635 ms | **5/30 = 17%** |
The joint two-position read is faster (p50 434 vs 689 ms) but **unreliable**: only 5/30 (17%)
prompts produced valid letter mass (>0.5) at BOTH token positions — the model does not
reliably answer two questions with two letters in one generation.
**Recommendation: two separate calls fired in parallel** (category + tier). Parallel wall-clock
= max(category, tier) ≈ p95 ~595 ms, fully reliable, and matches the plan's component-4 design
("fired in parallel ... so it adds no wall-clock time"). The joint cleverness is rejected on
reliability, not speed.
## 6. Open question (c) — num_ctx trade-off
100-call run at each `num_ctx`, `/api/ps` before/after, `qwen3.5:4b` footprint and 14b residency.
| num_ctx | p50 | p95 | qwen3.5:4b VRAM | 14b router resident? |
|---|---|---|---|---|
| 4096 | 361 ms | 588 ms | 6.15 GB | **NO** |
| 8192 | 364 ms | 594 ms | 6.32 GB | **NO** |
| 16384 | 364 ms | 591 ms | 6.65 GB | **NO** |
Latency is essentially flat across `num_ctx` (the classification prompts are short), and VRAM
grows 6.15 → 6.65 GB. If coexistence were possible, `num_ctx 4096` would be the best
trade-off. It is not possible: see below.
## 7. VRAM / router-model eviction (HARD CONSTRAINT — FAILS)
`/api/ps` before and after a 100-call run at `num_ctx 8192`:
```
BEFORE: {'qwen2.5-coder-router:14b': 17.76}
AFTER : {'qwen3.5:4b': 6.32} # 14b router GONE
14b router resident after 100 calls: False
```
Definitive coexistence check — 14b loaded at its production context (32768 = 17.76 GB), then a
single `qwen3.5:4b` call at each `num_ctx`:
```
after qwen3.5@4096 : {'qwen3.5:4b': 6.15} coexist? False
after qwen3.5@8192 : {'qwen3.5:4b': 6.32} coexist? False
after qwen3.5@16384: {'qwen3.5:4b': 6.65} coexist? False
```
Why: the RTX 6000 has 24 GB. The 14b router occupies 17.76 GB (Q4_K_M @ 32768 ctx) plus ~1.2 GB
of other GPU processes → ~18.4 GB used, ~5.5 GB free. `qwen3.5:4b` needs ≥ 6.15 GB at any
`num_ctx`, so loading it always exceeds free VRAM and Ollama evicts the 14b router. 17.76 + 1.2 +
6.15 = 25.1 GB > 24 GB — it cannot physically fit at `num_ctx 4096`, let alone 8192.
(The earlier plan note "the card sat at 23.4/24 GB with both resident at num_ctx 8192" was not
reproducible — the 14b router here is 17.76 GB at its production 32768 context, and the two do
not coexist.)
## 8. Gate summary and next action
- Eval-set 97.8% ≥ 85% — PASS
- Held-out 90.0% ≥ 90% — PASS
- p95 502 ms ≤ 800 ms — PASS
- 14b router resident after 100 calls — **FAIL (evicted)**
**The local_decision approach fails gate G on VRAM and must NOT proceed to todos 4-17.**
The accuracy and latency bars are met and the description tuning is complete, but the classifier
cannot run on this 24 GB card without evicting the 14b router model that local dispatch depends
on. Options to revisit before re-running gate G (out of scope here): a smaller classifier model
that coexists in the remaining ~5.5 GB (e.g. a ~3-4 GB 2-3B model at low num_ctx), or a model
that is small enough that 14b + classifier + other procs fit in 24 GB.
## 9. Re-measure with the 14b at 16k (reviewer, 2026-09-28, owner option A)
The owner chose option A: drop the router 14b to `num_ctx 16384` and re-measure.
Local vision fallback (`qwen3-vl-router:4b`, 5.6 GB) is being turned off: it has
never fired (0 route decisions have ever selected a `qwen3-vl` model).
Measured with a separate test tag (`FROM qwen2.5-coder:14b`, `PARAMETER num_ctx
16384`), so the production tag was untouched. The tag was deleted afterwards.
| state | `/api/ps` size_vram | nvidia-smi used |
|---|---|---|
| 14b @ 16k alone | 12.26 GB | 14.2 GB |
| + `qwen3.5:4b` @ 4096 | 12.26 + 5.73 GB | 19.7 GB |
| + `qwen3.5:4b` @ 8192 | 12.26 + 5.89 GB | 19.9 GB (about 4 GB spare) |
**At 16k the two coexist, with about 4 GB of headroom.** The vision model (5.6 GB)
would not also fit, which is why it is being turned off.
**Gate G's "evicted at every num_ctx" is not stable.** During the latency run,
production traffic reloaded its own 14b at 32k, and it stayed resident alongside
the 4b: 16.54 + 5.89 GB, 22.9 of 24 GB used, about 1 GB spare. Ollama's eviction
follows its own memory estimate, which varies with load order and other GPU
processes (gate G saw the same 14b as 17.76 GB). At 32k the pair fits only by
about 1 GB and can flip back to eviction; 16k is the robust setting.
Latency, 100 calls of `classify_choice` (`num_ctx 8192`) with both models resident:
p50 321 ms, p95 474 ms, max 2545 ms (the first cold call). The 14b stayed resident
throughout. **p95 474 ms, PASS.**
**Revised gate verdict:** PASS on every criterion **provided** the 14b router tag
runs at `num_ctx 16384` and local vision is off. Both are production changes and
the owner's call; todos 4-17 may proceed once they are made.