# Gate G — local_decision measurement results Status: done -- gate G measured; VRAM passes with the 14b at 16k and local vision off (section 9); accuracy and latency pass Date: 2026-09-28 Plan: `plans/local-decision-classifier.md` Hardware: Quadro RTX 6000, 24 GB (24576 MiB total) Models: `qwen3.5:4b` (classifier), `qwen2.5-coder-router:14b` (router, resident at 32768 ctx = 17.76 GB) Ollama: `http://localhost:11434` ## Headline | Criterion | Measured | Threshold | Verdict | |---|---|---|---| | Eval-set accuracy (`evals/tasks.yaml`) | **97.8%** (tuned) | ≥ 85% | PASS | | Held-out accuracy (`evals/heldout.yaml`) | **90.0%** (tuned) | ≥ 90% | PASS (at threshold) | | p95 latency (14b resident) | **502 ms** | ≤ 800 ms | PASS | | 14b router resident after 100 calls | **EVICTED** | must stay resident | **FAIL** | **Gate verdict: STOP.** The 14b router model is evicted whenever `qwen3.5:4b` is loaded — on this 24 GB card there is no `num_ctx` at which the two coexist. Per the plan, "STOP AND REPORT after G if ... the 14b router model was evicted. Do NOT start todos 4-17 in that case." Todos 4-17 are NOT started. --- ## 1. Accuracy per set Measured with the eval harness (`--backend decision --decision-base-url http://localhost:11434`), majority-vote overall accuracy over the clean/short/long noise variants of each task. ### Baseline (original `_DECISION_DESCRIPTIONS`) | Set | clean | short | long | overall | |---|---|---|---|---| | `evals/tasks.yaml` (46) | 0.739 | 0.783 | 0.761 | **0.761** | | `evals/heldout.yaml` (30) | 0.800 | 0.933 | 0.900 | **0.967** | Baseline confusion (eval-set, dominant cells): `coding_general → debugging` (14) and `→ reasoning_math` (4); `file_summarization → debugging` (9). The `debugging` description was attracting both code-tracing and code-summary tasks. ### After description tuning `_DECISION_DESCRIPTIONS` changes: widen `coding_general` to include "tracing what existing code returns" and "implementing a feature or algorithm"; narrow `debugging` to "fixing a bug"; narrow `reasoning_math` to "pure math or logic word problem (no code)"; soften `file_summarization` to "summarizing or explaining". | Set | clean | short | long | overall | |---|---|---|---|---| | `evals/tasks.yaml` (46) | 0.978 | 0.978 | 0.978 | **0.978** | | `evals/heldout.yaml` (30) | 0.800 | 0.900 | 0.833 | **0.900** | Tuning lifted the eval-set from 76.1% → 97.8% (the confusion cells are gone: eval-set `coding_general` 27/27, `file_summarization` 18/18, `debugging` 18/18). It cost 6.7 points on held-out (96.7% → 90.0%), landing exactly on the 90% threshold. Held-out residual misses: `h2` coding_general→file_summarization, `h30` general_chat→file_summarization, `h9` debugging→general_chat — all low-confidence (0.33-0.49) edge cases the confidence cascade would absorb. Further tuning to chase these risks overfitting `evals/tasks.yaml` (the exact failure the plan warns about), so tuning stops here. ## 2. Calibration (tuned) Mean confidence when correct vs incorrect (per-noise-level rows). | Set | correct mean | incorrect mean | |---|---|---| | `evals/tasks.yaml` | 0.798 | 0.350 | | `evals/heldout.yaml` | 0.802 | 0.413 | Calibrated in the right direction on both sets: the model is roughly twice as confident when right as when wrong. (Baseline was 0.833/0.438 eval-set, 0.815/0.300 heldout — the tuning moved both closer together but kept correct ≫ incorrect.) ## 3. Latency (p50 / p95) Per-call latency of a single `classify_choice` over the eval-set, `qwen3.5:4b`, `num_ctx 8192`, with the 14b router resident (the production steady state). ``` single category call: n=46 p50=363ms p95=502ms p99=598ms mean=379ms ``` **p95 502 ms ≤ 800 ms — PASS.** ## 4. Open question (a) — noise isolation Compare `classify_choice` on raw wrapped (noisy) text vs text passed through `local_encoder._isolate_task_text` first. | Set | no-isolation (short+long) | isolated (short+long) | clean baseline | |---|---|---|---| | `evals/tasks.yaml` | 0.772 | 0.739 | 0.739 | | `evals/heldout.yaml` | 0.917 | 0.800 | 0.800 | **Isolation HURTS** on both sets (−3.3 pt eval, −11.7 pt heldout). For `local_decision` the fenced-code/tool content `_isolate_task_text` strips is itself the signal (code-tracing, file-summary tasks). **Do NOT run `_isolate_task_text` before `classify_choice`; send raw text.** This overrides the prototype's earlier "noise cost about 2 points" note — for this backend isolation is strictly worse. ## 5. Open question (b) — tier call strategy Measure separate-request latency vs a joint prompt reading two token positions. | strategy | p50 | p95 | reliability | |---|---|---|---| | category call (10 options) | 361 ms | 595 ms | — | | tier call (3 options, separate) | 326 ms | 553 ms | — | | separate, sequential total | 689 ms | 1154 ms | 100% | | joint prompt, two token positions | 434 ms | 635 ms | **5/30 = 17%** | The joint two-position read is faster (p50 434 vs 689 ms) but **unreliable**: only 5/30 (17%) prompts produced valid letter mass (>0.5) at BOTH token positions — the model does not reliably answer two questions with two letters in one generation. **Recommendation: two separate calls fired in parallel** (category + tier). Parallel wall-clock = max(category, tier) ≈ p95 ~595 ms, fully reliable, and matches the plan's component-4 design ("fired in parallel ... so it adds no wall-clock time"). The joint cleverness is rejected on reliability, not speed. ## 6. Open question (c) — num_ctx trade-off 100-call run at each `num_ctx`, `/api/ps` before/after, `qwen3.5:4b` footprint and 14b residency. | num_ctx | p50 | p95 | qwen3.5:4b VRAM | 14b router resident? | |---|---|---|---|---| | 4096 | 361 ms | 588 ms | 6.15 GB | **NO** | | 8192 | 364 ms | 594 ms | 6.32 GB | **NO** | | 16384 | 364 ms | 591 ms | 6.65 GB | **NO** | Latency is essentially flat across `num_ctx` (the classification prompts are short), and VRAM grows 6.15 → 6.65 GB. If coexistence were possible, `num_ctx 4096` would be the best trade-off. It is not possible: see below. ## 7. VRAM / router-model eviction (HARD CONSTRAINT — FAILS) `/api/ps` before and after a 100-call run at `num_ctx 8192`: ``` BEFORE: {'qwen2.5-coder-router:14b': 17.76} AFTER : {'qwen3.5:4b': 6.32} # 14b router GONE 14b router resident after 100 calls: False ``` Definitive coexistence check — 14b loaded at its production context (32768 = 17.76 GB), then a single `qwen3.5:4b` call at each `num_ctx`: ``` after qwen3.5@4096 : {'qwen3.5:4b': 6.15} coexist? False after qwen3.5@8192 : {'qwen3.5:4b': 6.32} coexist? False after qwen3.5@16384: {'qwen3.5:4b': 6.65} coexist? False ``` Why: the RTX 6000 has 24 GB. The 14b router occupies 17.76 GB (Q4_K_M @ 32768 ctx) plus ~1.2 GB of other GPU processes → ~18.4 GB used, ~5.5 GB free. `qwen3.5:4b` needs ≥ 6.15 GB at any `num_ctx`, so loading it always exceeds free VRAM and Ollama evicts the 14b router. 17.76 + 1.2 + 6.15 = 25.1 GB > 24 GB — it cannot physically fit at `num_ctx 4096`, let alone 8192. (The earlier plan note "the card sat at 23.4/24 GB with both resident at num_ctx 8192" was not reproducible — the 14b router here is 17.76 GB at its production 32768 context, and the two do not coexist.) ## 8. Gate summary and next action - Eval-set 97.8% ≥ 85% — PASS - Held-out 90.0% ≥ 90% — PASS - p95 502 ms ≤ 800 ms — PASS - 14b router resident after 100 calls — **FAIL (evicted)** **The local_decision approach fails gate G on VRAM and must NOT proceed to todos 4-17.** The accuracy and latency bars are met and the description tuning is complete, but the classifier cannot run on this 24 GB card without evicting the 14b router model that local dispatch depends on. Options to revisit before re-running gate G (out of scope here): a smaller classifier model that coexists in the remaining ~5.5 GB (e.g. a ~3-4 GB 2-3B model at low num_ctx), or a model that is small enough that 14b + classifier + other procs fit in 24 GB. ## 9. Re-measure with the 14b at 16k (reviewer, 2026-09-28, owner option A) The owner chose option A: drop the router 14b to `num_ctx 16384` and re-measure. Local vision fallback (`qwen3-vl-router:4b`, 5.6 GB) is being turned off: it has never fired (0 route decisions have ever selected a `qwen3-vl` model). Measured with a separate test tag (`FROM qwen2.5-coder:14b`, `PARAMETER num_ctx 16384`), so the production tag was untouched. The tag was deleted afterwards. | state | `/api/ps` size_vram | nvidia-smi used | |---|---|---| | 14b @ 16k alone | 12.26 GB | 14.2 GB | | + `qwen3.5:4b` @ 4096 | 12.26 + 5.73 GB | 19.7 GB | | + `qwen3.5:4b` @ 8192 | 12.26 + 5.89 GB | 19.9 GB (about 4 GB spare) | **At 16k the two coexist, with about 4 GB of headroom.** The vision model (5.6 GB) would not also fit, which is why it is being turned off. **Gate G's "evicted at every num_ctx" is not stable.** During the latency run, production traffic reloaded its own 14b at 32k, and it stayed resident alongside the 4b: 16.54 + 5.89 GB, 22.9 of 24 GB used, about 1 GB spare. Ollama's eviction follows its own memory estimate, which varies with load order and other GPU processes (gate G saw the same 14b as 17.76 GB). At 32k the pair fits only by about 1 GB and can flip back to eviction; 16k is the robust setting. Latency, 100 calls of `classify_choice` (`num_ctx 8192`) with both models resident: p50 321 ms, p95 474 ms, max 2545 ms (the first cold call). The 14b stayed resident throughout. **p95 474 ms, PASS.** **Revised gate verdict:** PASS on every criterion **provided** the 14b router tag runs at `num_ctx 16384` and local vision is off. Both are production changes and the owner's call; todos 4-17 may proceed once they are made.