Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N9biTbFC63yDfYfUsZmhgd
211 lines
9.6 KiB
Markdown
211 lines
9.6 KiB
Markdown
# Gate G — local_decision measurement results
|
||
|
||
Status: done -- gate G measured; VRAM passes with the 14b at 16k and local vision off (section 9); accuracy and latency pass
|
||
|
||
Date: 2026-09-28
|
||
Plan: `plans/local-decision-classifier.md`
|
||
Hardware: Quadro RTX 6000, 24 GB (24576 MiB total)
|
||
Models: `qwen3.5:4b` (classifier), `qwen2.5-coder-router:14b` (router, resident at 32768 ctx = 17.76 GB)
|
||
Ollama: `http://localhost:11434`
|
||
|
||
## Headline
|
||
|
||
| Criterion | Measured | Threshold | Verdict |
|
||
|---|---|---|---|
|
||
| Eval-set accuracy (`evals/tasks.yaml`) | **97.8%** (tuned) | ≥ 85% | PASS |
|
||
| Held-out accuracy (`evals/heldout.yaml`) | **90.0%** (tuned) | ≥ 90% | PASS (at threshold) |
|
||
| p95 latency (14b resident) | **502 ms** | ≤ 800 ms | PASS |
|
||
| 14b router resident after 100 calls | **EVICTED** | must stay resident | **FAIL** |
|
||
|
||
**Gate verdict: STOP.** The 14b router model is evicted whenever `qwen3.5:4b` is loaded —
|
||
on this 24 GB card there is no `num_ctx` at which the two coexist. Per the plan, "STOP AND
|
||
REPORT after G if ... the 14b router model was evicted. Do NOT start todos 4-17 in that case."
|
||
Todos 4-17 are NOT started.
|
||
|
||
---
|
||
|
||
## 1. Accuracy per set
|
||
|
||
Measured with the eval harness (`--backend decision --decision-base-url http://localhost:11434`),
|
||
majority-vote overall accuracy over the clean/short/long noise variants of each task.
|
||
|
||
### Baseline (original `_DECISION_DESCRIPTIONS`)
|
||
|
||
| Set | clean | short | long | overall |
|
||
|---|---|---|---|---|
|
||
| `evals/tasks.yaml` (46) | 0.739 | 0.783 | 0.761 | **0.761** |
|
||
| `evals/heldout.yaml` (30) | 0.800 | 0.933 | 0.900 | **0.967** |
|
||
|
||
Baseline confusion (eval-set, dominant cells): `coding_general → debugging` (14) and
|
||
`→ reasoning_math` (4); `file_summarization → debugging` (9). The `debugging` description
|
||
was attracting both code-tracing and code-summary tasks.
|
||
|
||
### After description tuning
|
||
|
||
`_DECISION_DESCRIPTIONS` changes: widen `coding_general` to include "tracing what existing
|
||
code returns" and "implementing a feature or algorithm"; narrow `debugging` to "fixing a bug";
|
||
narrow `reasoning_math` to "pure math or logic word problem (no code)"; soften
|
||
`file_summarization` to "summarizing or explaining".
|
||
|
||
| Set | clean | short | long | overall |
|
||
|---|---|---|---|---|
|
||
| `evals/tasks.yaml` (46) | 0.978 | 0.978 | 0.978 | **0.978** |
|
||
| `evals/heldout.yaml` (30) | 0.800 | 0.900 | 0.833 | **0.900** |
|
||
|
||
Tuning lifted the eval-set from 76.1% → 97.8% (the confusion cells are gone: eval-set
|
||
`coding_general` 27/27, `file_summarization` 18/18, `debugging` 18/18). It cost 6.7 points on
|
||
held-out (96.7% → 90.0%), landing exactly on the 90% threshold. Held-out residual misses:
|
||
`h2` coding_general→file_summarization, `h30` general_chat→file_summarization, `h9`
|
||
debugging→general_chat — all low-confidence (0.33-0.49) edge cases the confidence cascade
|
||
would absorb. Further tuning to chase these risks overfitting `evals/tasks.yaml` (the exact
|
||
failure the plan warns about), so tuning stops here.
|
||
|
||
## 2. Calibration (tuned)
|
||
|
||
Mean confidence when correct vs incorrect (per-noise-level rows).
|
||
|
||
| Set | correct mean | incorrect mean |
|
||
|---|---|---|
|
||
| `evals/tasks.yaml` | 0.798 | 0.350 |
|
||
| `evals/heldout.yaml` | 0.802 | 0.413 |
|
||
|
||
Calibrated in the right direction on both sets: the model is roughly twice as confident when
|
||
right as when wrong. (Baseline was 0.833/0.438 eval-set, 0.815/0.300 heldout — the tuning
|
||
moved both closer together but kept correct ≫ incorrect.)
|
||
|
||
## 3. Latency (p50 / p95)
|
||
|
||
Per-call latency of a single `classify_choice` over the eval-set, `qwen3.5:4b`, `num_ctx 8192`,
|
||
with the 14b router resident (the production steady state).
|
||
|
||
```
|
||
single category call: n=46 p50=363ms p95=502ms p99=598ms mean=379ms
|
||
```
|
||
|
||
**p95 502 ms ≤ 800 ms — PASS.**
|
||
|
||
## 4. Open question (a) — noise isolation
|
||
|
||
Compare `classify_choice` on raw wrapped (noisy) text vs text passed through
|
||
`local_encoder._isolate_task_text` first.
|
||
|
||
| Set | no-isolation (short+long) | isolated (short+long) | clean baseline |
|
||
|---|---|---|---|
|
||
| `evals/tasks.yaml` | 0.772 | 0.739 | 0.739 |
|
||
| `evals/heldout.yaml` | 0.917 | 0.800 | 0.800 |
|
||
|
||
**Isolation HURTS** on both sets (−3.3 pt eval, −11.7 pt heldout). For `local_decision` the
|
||
fenced-code/tool content `_isolate_task_text` strips is itself the signal (code-tracing,
|
||
file-summary tasks). **Do NOT run `_isolate_task_text` before `classify_choice`; send raw text.**
|
||
This overrides the prototype's earlier "noise cost about 2 points" note — for this backend
|
||
isolation is strictly worse.
|
||
|
||
## 5. Open question (b) — tier call strategy
|
||
|
||
Measure separate-request latency vs a joint prompt reading two token positions.
|
||
|
||
| strategy | p50 | p95 | reliability |
|
||
|---|---|---|---|
|
||
| category call (10 options) | 361 ms | 595 ms | — |
|
||
| tier call (3 options, separate) | 326 ms | 553 ms | — |
|
||
| separate, sequential total | 689 ms | 1154 ms | 100% |
|
||
| joint prompt, two token positions | 434 ms | 635 ms | **5/30 = 17%** |
|
||
|
||
The joint two-position read is faster (p50 434 vs 689 ms) but **unreliable**: only 5/30 (17%)
|
||
prompts produced valid letter mass (>0.5) at BOTH token positions — the model does not
|
||
reliably answer two questions with two letters in one generation.
|
||
|
||
**Recommendation: two separate calls fired in parallel** (category + tier). Parallel wall-clock
|
||
= max(category, tier) ≈ p95 ~595 ms, fully reliable, and matches the plan's component-4 design
|
||
("fired in parallel ... so it adds no wall-clock time"). The joint cleverness is rejected on
|
||
reliability, not speed.
|
||
|
||
## 6. Open question (c) — num_ctx trade-off
|
||
|
||
100-call run at each `num_ctx`, `/api/ps` before/after, `qwen3.5:4b` footprint and 14b residency.
|
||
|
||
| num_ctx | p50 | p95 | qwen3.5:4b VRAM | 14b router resident? |
|
||
|---|---|---|---|---|
|
||
| 4096 | 361 ms | 588 ms | 6.15 GB | **NO** |
|
||
| 8192 | 364 ms | 594 ms | 6.32 GB | **NO** |
|
||
| 16384 | 364 ms | 591 ms | 6.65 GB | **NO** |
|
||
|
||
Latency is essentially flat across `num_ctx` (the classification prompts are short), and VRAM
|
||
grows 6.15 → 6.65 GB. If coexistence were possible, `num_ctx 4096` would be the best
|
||
trade-off. It is not possible: see below.
|
||
|
||
## 7. VRAM / router-model eviction (HARD CONSTRAINT — FAILS)
|
||
|
||
`/api/ps` before and after a 100-call run at `num_ctx 8192`:
|
||
|
||
```
|
||
BEFORE: {'qwen2.5-coder-router:14b': 17.76}
|
||
AFTER : {'qwen3.5:4b': 6.32} # 14b router GONE
|
||
14b router resident after 100 calls: False
|
||
```
|
||
|
||
Definitive coexistence check — 14b loaded at its production context (32768 = 17.76 GB), then a
|
||
single `qwen3.5:4b` call at each `num_ctx`:
|
||
|
||
```
|
||
after qwen3.5@4096 : {'qwen3.5:4b': 6.15} coexist? False
|
||
after qwen3.5@8192 : {'qwen3.5:4b': 6.32} coexist? False
|
||
after qwen3.5@16384: {'qwen3.5:4b': 6.65} coexist? False
|
||
```
|
||
|
||
Why: the RTX 6000 has 24 GB. The 14b router occupies 17.76 GB (Q4_K_M @ 32768 ctx) plus ~1.2 GB
|
||
of other GPU processes → ~18.4 GB used, ~5.5 GB free. `qwen3.5:4b` needs ≥ 6.15 GB at any
|
||
`num_ctx`, so loading it always exceeds free VRAM and Ollama evicts the 14b router. 17.76 + 1.2 +
|
||
6.15 = 25.1 GB > 24 GB — it cannot physically fit at `num_ctx 4096`, let alone 8192.
|
||
|
||
(The earlier plan note "the card sat at 23.4/24 GB with both resident at num_ctx 8192" was not
|
||
reproducible — the 14b router here is 17.76 GB at its production 32768 context, and the two do
|
||
not coexist.)
|
||
|
||
## 8. Gate summary and next action
|
||
|
||
- Eval-set 97.8% ≥ 85% — PASS
|
||
- Held-out 90.0% ≥ 90% — PASS
|
||
- p95 502 ms ≤ 800 ms — PASS
|
||
- 14b router resident after 100 calls — **FAIL (evicted)**
|
||
|
||
**The local_decision approach fails gate G on VRAM and must NOT proceed to todos 4-17.**
|
||
The accuracy and latency bars are met and the description tuning is complete, but the classifier
|
||
cannot run on this 24 GB card without evicting the 14b router model that local dispatch depends
|
||
on. Options to revisit before re-running gate G (out of scope here): a smaller classifier model
|
||
that coexists in the remaining ~5.5 GB (e.g. a ~3-4 GB 2-3B model at low num_ctx), or a model
|
||
that is small enough that 14b + classifier + other procs fit in 24 GB.
|
||
|
||
## 9. Re-measure with the 14b at 16k (reviewer, 2026-09-28, owner option A)
|
||
|
||
The owner chose option A: drop the router 14b to `num_ctx 16384` and re-measure.
|
||
Local vision fallback (`qwen3-vl-router:4b`, 5.6 GB) is being turned off: it has
|
||
never fired (0 route decisions have ever selected a `qwen3-vl` model).
|
||
|
||
Measured with a separate test tag (`FROM qwen2.5-coder:14b`, `PARAMETER num_ctx
|
||
16384`), so the production tag was untouched. The tag was deleted afterwards.
|
||
|
||
| state | `/api/ps` size_vram | nvidia-smi used |
|
||
|---|---|---|
|
||
| 14b @ 16k alone | 12.26 GB | 14.2 GB |
|
||
| + `qwen3.5:4b` @ 4096 | 12.26 + 5.73 GB | 19.7 GB |
|
||
| + `qwen3.5:4b` @ 8192 | 12.26 + 5.89 GB | 19.9 GB (about 4 GB spare) |
|
||
|
||
**At 16k the two coexist, with about 4 GB of headroom.** The vision model (5.6 GB)
|
||
would not also fit, which is why it is being turned off.
|
||
|
||
**Gate G's "evicted at every num_ctx" is not stable.** During the latency run,
|
||
production traffic reloaded its own 14b at 32k, and it stayed resident alongside
|
||
the 4b: 16.54 + 5.89 GB, 22.9 of 24 GB used, about 1 GB spare. Ollama's eviction
|
||
follows its own memory estimate, which varies with load order and other GPU
|
||
processes (gate G saw the same 14b as 17.76 GB). At 32k the pair fits only by
|
||
about 1 GB and can flip back to eviction; 16k is the robust setting.
|
||
|
||
Latency, 100 calls of `classify_choice` (`num_ctx 8192`) with both models resident:
|
||
p50 321 ms, p95 474 ms, max 2545 ms (the first cold call). The 14b stayed resident
|
||
throughout. **p95 474 ms, PASS.**
|
||
|
||
**Revised gate verdict:** PASS on every criterion **provided** the 14b router tag
|
||
runs at `num_ctx 16384` and local vision is off. Both are production changes and
|
||
the owner's call; todos 4-17 may proceed once they are made.
|