Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
9.6 KiB
Gate G — local_decision measurement results
Status: done -- gate G measured; VRAM passes with the 14b at 16k and local vision off (section 9); accuracy and latency pass
Date: 2026-09-28
Plan: plans/local-decision-classifier.md
Hardware: Quadro RTX 6000, 24 GB (24576 MiB total)
Models: qwen3.5:4b (classifier), qwen2.5-coder-router:14b (router, resident at 32768 ctx = 17.76 GB)
Ollama: http://localhost:11434
Headline
| Criterion | Measured | Threshold | Verdict |
|---|---|---|---|
Eval-set accuracy (evals/tasks.yaml) |
97.8% (tuned) | ≥ 85% | PASS |
Held-out accuracy (evals/heldout.yaml) |
90.0% (tuned) | ≥ 90% | PASS (at threshold) |
| p95 latency (14b resident) | 502 ms | ≤ 800 ms | PASS |
| 14b router resident after 100 calls | EVICTED | must stay resident | FAIL |
Gate verdict: STOP. The 14b router model is evicted whenever qwen3.5:4b is loaded —
on this 24 GB card there is no num_ctx at which the two coexist. Per the plan, "STOP AND
REPORT after G if ... the 14b router model was evicted. Do NOT start todos 4-17 in that case."
Todos 4-17 are NOT started.
1. Accuracy per set
Measured with the eval harness (--backend decision --decision-base-url http://localhost:11434),
majority-vote overall accuracy over the clean/short/long noise variants of each task.
Baseline (original _DECISION_DESCRIPTIONS)
| Set | clean | short | long | overall |
|---|---|---|---|---|
evals/tasks.yaml (46) |
0.739 | 0.783 | 0.761 | 0.761 |
evals/heldout.yaml (30) |
0.800 | 0.933 | 0.900 | 0.967 |
Baseline confusion (eval-set, dominant cells): coding_general → debugging (14) and
→ reasoning_math (4); file_summarization → debugging (9). The debugging description
was attracting both code-tracing and code-summary tasks.
After description tuning
_DECISION_DESCRIPTIONS changes: widen coding_general to include "tracing what existing
code returns" and "implementing a feature or algorithm"; narrow debugging to "fixing a bug";
narrow reasoning_math to "pure math or logic word problem (no code)"; soften
file_summarization to "summarizing or explaining".
| Set | clean | short | long | overall |
|---|---|---|---|---|
evals/tasks.yaml (46) |
0.978 | 0.978 | 0.978 | 0.978 |
evals/heldout.yaml (30) |
0.800 | 0.900 | 0.833 | 0.900 |
Tuning lifted the eval-set from 76.1% → 97.8% (the confusion cells are gone: eval-set
coding_general 27/27, file_summarization 18/18, debugging 18/18). It cost 6.7 points on
held-out (96.7% → 90.0%), landing exactly on the 90% threshold. Held-out residual misses:
h2 coding_general→file_summarization, h30 general_chat→file_summarization, h9
debugging→general_chat — all low-confidence (0.33-0.49) edge cases the confidence cascade
would absorb. Further tuning to chase these risks overfitting evals/tasks.yaml (the exact
failure the plan warns about), so tuning stops here.
2. Calibration (tuned)
Mean confidence when correct vs incorrect (per-noise-level rows).
| Set | correct mean | incorrect mean |
|---|---|---|
evals/tasks.yaml |
0.798 | 0.350 |
evals/heldout.yaml |
0.802 | 0.413 |
Calibrated in the right direction on both sets: the model is roughly twice as confident when right as when wrong. (Baseline was 0.833/0.438 eval-set, 0.815/0.300 heldout — the tuning moved both closer together but kept correct ≫ incorrect.)
3. Latency (p50 / p95)
Per-call latency of a single classify_choice over the eval-set, qwen3.5:4b, num_ctx 8192,
with the 14b router resident (the production steady state).
single category call: n=46 p50=363ms p95=502ms p99=598ms mean=379ms
p95 502 ms ≤ 800 ms — PASS.
4. Open question (a) — noise isolation
Compare classify_choice on raw wrapped (noisy) text vs text passed through
local_encoder._isolate_task_text first.
| Set | no-isolation (short+long) | isolated (short+long) | clean baseline |
|---|---|---|---|
evals/tasks.yaml |
0.772 | 0.739 | 0.739 |
evals/heldout.yaml |
0.917 | 0.800 | 0.800 |
Isolation HURTS on both sets (−3.3 pt eval, −11.7 pt heldout). For local_decision the
fenced-code/tool content _isolate_task_text strips is itself the signal (code-tracing,
file-summary tasks). Do NOT run _isolate_task_text before classify_choice; send raw text.
This overrides the prototype's earlier "noise cost about 2 points" note — for this backend
isolation is strictly worse.
5. Open question (b) — tier call strategy
Measure separate-request latency vs a joint prompt reading two token positions.
| strategy | p50 | p95 | reliability |
|---|---|---|---|
| category call (10 options) | 361 ms | 595 ms | — |
| tier call (3 options, separate) | 326 ms | 553 ms | — |
| separate, sequential total | 689 ms | 1154 ms | 100% |
| joint prompt, two token positions | 434 ms | 635 ms | 5/30 = 17% |
The joint two-position read is faster (p50 434 vs 689 ms) but unreliable: only 5/30 (17%) prompts produced valid letter mass (>0.5) at BOTH token positions — the model does not reliably answer two questions with two letters in one generation.
Recommendation: two separate calls fired in parallel (category + tier). Parallel wall-clock = max(category, tier) ≈ p95 ~595 ms, fully reliable, and matches the plan's component-4 design ("fired in parallel ... so it adds no wall-clock time"). The joint cleverness is rejected on reliability, not speed.
6. Open question (c) — num_ctx trade-off
100-call run at each num_ctx, /api/ps before/after, qwen3.5:4b footprint and 14b residency.
| num_ctx | p50 | p95 | qwen3.5:4b VRAM | 14b router resident? |
|---|---|---|---|---|
| 4096 | 361 ms | 588 ms | 6.15 GB | NO |
| 8192 | 364 ms | 594 ms | 6.32 GB | NO |
| 16384 | 364 ms | 591 ms | 6.65 GB | NO |
Latency is essentially flat across num_ctx (the classification prompts are short), and VRAM
grows 6.15 → 6.65 GB. If coexistence were possible, num_ctx 4096 would be the best
trade-off. It is not possible: see below.
7. VRAM / router-model eviction (HARD CONSTRAINT — FAILS)
/api/ps before and after a 100-call run at num_ctx 8192:
BEFORE: {'qwen2.5-coder-router:14b': 17.76}
AFTER : {'qwen3.5:4b': 6.32} # 14b router GONE
14b router resident after 100 calls: False
Definitive coexistence check — 14b loaded at its production context (32768 = 17.76 GB), then a
single qwen3.5:4b call at each num_ctx:
after qwen3.5@4096 : {'qwen3.5:4b': 6.15} coexist? False
after qwen3.5@8192 : {'qwen3.5:4b': 6.32} coexist? False
after qwen3.5@16384: {'qwen3.5:4b': 6.65} coexist? False
Why: the RTX 6000 has 24 GB. The 14b router occupies 17.76 GB (Q4_K_M @ 32768 ctx) plus ~1.2 GB
of other GPU processes → ~18.4 GB used, ~5.5 GB free. qwen3.5:4b needs ≥ 6.15 GB at any
num_ctx, so loading it always exceeds free VRAM and Ollama evicts the 14b router. 17.76 + 1.2 +
6.15 = 25.1 GB > 24 GB — it cannot physically fit at num_ctx 4096, let alone 8192.
(The earlier plan note "the card sat at 23.4/24 GB with both resident at num_ctx 8192" was not reproducible — the 14b router here is 17.76 GB at its production 32768 context, and the two do not coexist.)
8. Gate summary and next action
- Eval-set 97.8% ≥ 85% — PASS
- Held-out 90.0% ≥ 90% — PASS
- p95 502 ms ≤ 800 ms — PASS
- 14b router resident after 100 calls — FAIL (evicted)
The local_decision approach fails gate G on VRAM and must NOT proceed to todos 4-17. The accuracy and latency bars are met and the description tuning is complete, but the classifier cannot run on this 24 GB card without evicting the 14b router model that local dispatch depends on. Options to revisit before re-running gate G (out of scope here): a smaller classifier model that coexists in the remaining ~5.5 GB (e.g. a ~3-4 GB 2-3B model at low num_ctx), or a model that is small enough that 14b + classifier + other procs fit in 24 GB.
9. Re-measure with the 14b at 16k (reviewer, 2026-09-28, owner option A)
The owner chose option A: drop the router 14b to num_ctx 16384 and re-measure.
Local vision fallback (qwen3-vl-router:4b, 5.6 GB) is being turned off: it has
never fired (0 route decisions have ever selected a qwen3-vl model).
Measured with a separate test tag (FROM qwen2.5-coder:14b, PARAMETER num_ctx 16384), so the production tag was untouched. The tag was deleted afterwards.
| state | /api/ps size_vram |
nvidia-smi used |
|---|---|---|
| 14b @ 16k alone | 12.26 GB | 14.2 GB |
+ qwen3.5:4b @ 4096 |
12.26 + 5.73 GB | 19.7 GB |
+ qwen3.5:4b @ 8192 |
12.26 + 5.89 GB | 19.9 GB (about 4 GB spare) |
At 16k the two coexist, with about 4 GB of headroom. The vision model (5.6 GB) would not also fit, which is why it is being turned off.
Gate G's "evicted at every num_ctx" is not stable. During the latency run, production traffic reloaded its own 14b at 32k, and it stayed resident alongside the 4b: 16.54 + 5.89 GB, 22.9 of 24 GB used, about 1 GB spare. Ollama's eviction follows its own memory estimate, which varies with load order and other GPU processes (gate G saw the same 14b as 17.76 GB). At 32k the pair fits only by about 1 GB and can flip back to eviction; 16k is the robust setting.
Latency, 100 calls of classify_choice (num_ctx 8192) with both models resident:
p50 321 ms, p95 474 ms, max 2545 ms (the first cold call). The 14b stayed resident
throughout. p95 474 ms, PASS.
Revised gate verdict: PASS on every criterion provided the 14b router tag
runs at num_ctx 16384 and local vision is off. Both are production changes and
the owner's call; todos 4-17 may proceed once they are made.