Files
6krrt/docs/local-models.md
adlee-was-taken 92031dbd0b feat(classifier): record the classifier's own attempt on route_decisions
A declined answer left one number behind and it was in a log line, so
confidence_min could not be tuned from data. route_decisions.confidence cannot
say it either: the chat path re-routes through the override branch, which
hard-codes 1.0, and 43,804 of 43,856 live rows hold exactly that.

Measured on 2026-10-04 under local_decision (qwen3.5:4b): 769 of 2,557 turns
(30%) had no fresh classification, against 0% for local_llm and local_encoder.
The cause was recorded in only 4 of them.

- classifier_confidence, classifier_coverage and classifier_reject on
  route_decisions, filled from a ClassifierAttempt carried on Classification.
  Both accepted and declined answers carry one, so the two distributions can be
  compared around the floor. Reason codes are listed in docs/data-model.md.
- ClassifierRejected (a RuntimeError subclass, messages unchanged) replaces the
  plain RuntimeErrors at the six floor-miss sites and the two local_decision
  refusals, so the number and reason travel out of the raise site.
- The chat path captures the classifier's verdict before the re-route and passes
  it to persist_route_decision (attempt_of), like it already does for source.
- Admin decisions page: the source badge's tooltip shows the attempt, and
  session_history / session_stale get an amber badge instead of neutral grey.
  TUI detail popup and the live event carry the same three keys.
- The /metrics degraded-share warning lists the recorded reasons instead of
  claiming the classifier "has been failing", which was wrong for a classifier
  that answers and is declined.
- degraded_warn_threshold must be in (0, 1] and degraded_warn_min at least 1,
  refused at load: a value above 1 could never fire.

The three columns arrive by ALTER and are NULL on every earlier row; metrics
selects them only when present, so the live DB reads NULL until its restart.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KkCGRantZsSwmcFpet6FTa
2026-10-04 22:06:57 -04:00

506 lines
28 KiB
Markdown

> Sizing local models for your GPU. Back to [README](../README.md).
## Fitting local models into 24GB VRAM
These defaults are measured on a 24GB Quadro RTX 6000 (Turing, compute 7.5)
and are a sound starting point for any 24-32GB card. They are **defaults, not
prescriptions** — see [Swapping in your own model](#swapping-in-your-own-model)
below, which is the part that matters if you run different hardware or simply
want a different model.
Ollama loads a model at the context size baked into its tag. You cannot set it
per request: the OpenAI-compatible `/v1/chat/completions` endpoint in Ollama
0.22.0 was verified live against every plausible field shape, and in every case
returned 200 while silently keeping the loaded context unchanged. So context
size has to go into the tag via a `Modelfile`:
```bash
# Classification + verification + local dispatch (ONE resident model)
printf 'FROM qwen2.5-coder:14b\nPARAMETER num_ctx 32768\n' > Modelfile.local-dispatch
ollama create qwen2.5-coder-router:14b -f Modelfile.local-dispatch
# Vision fallback only
printf 'FROM qwen3-vl:4b\nPARAMETER num_ctx 8192\n' > Modelfile.vision
ollama create qwen3-vl-router:4b -f Modelfile.vision
```
Measured co-residency, both fully on GPU:
| model | role | resident |
|---|---|---|
| `qwen2.5-coder-router:14b` (Q4_K_M, 32k) | classify + verify + dispatch | 17.0GB |
| `qwen3-vl-router:4b` (8k) | vision fallback | 5.6GB |
| | **total** | **23.2GB of 24GB** |
That leaves ~800MB spare, which is thin but works. The single most important
thing to understand is **where the memory actually goes**.
### KV cache is the cheapest GB to reclaim
Weights are fixed; KV cache scales with `num_ctx` and is often the larger
lever. Measured on this hardware:
| model | on disk | resident @ num_ctx |
|---|---|---|
| `qwen2.5-coder:14b` Q4_K_M | 9.0GB | 13GB @ 16k, **17GB @ 32k** |
| `qwen3-vl:4b` | 3.3GB | 5.6GB @ 8k, **6.9GB @ 16k** |
Halving the vision model's context bought 1.3GB — exactly what lets the
dispatch model run at 32k with vision still resident. Without it, loading
vision **evicts the dispatch model entirely**, and the next classification
pays a ~6s cold reload on the interactive latency floor.
If you are short on VRAM, cut `num_ctx` before you cut anything else. The
classifier in particular never needs a large window: `classifier.max_input_chars`
clamps input to 8000 chars (~2.7k tokens), so even 8192 leaves 2x headroom.
### Bigger quantization is usually the wrong trade
Tempting and mostly counterproductive. Measured, same model, same tasks:
| variant | resident | placement | classify accuracy | classify mean |
|---|---|---|---|---|
| Q4_K_M @ 16k | 13GB | 100% GPU | 11/14 | **1.07s** |
| Q4_K_M @ 32k | 17GB | 100% GPU | 11/14 | 1.14s |
| Q8_0 @ 16k | 19GB | 100% GPU | 11/14 | 1.62s |
| Q8_0 @ 32k | 24GB | **83% GPU** | 11/14 | 3.79s |
Accuracy never moved — only speed did. Q8 costs ~1.5x latency even when it
fits (Turing is memory-bandwidth bound here, and Q8 moves twice the weight
bytes per token), and `83% GPU` in the placement column means Ollama spilled
the rest to system RAM for a 3.3x latency hit — a config that "fits" on paper
can silently land there. Going up in *parameters* beat going up in *bits*: a
32B model at Q4 scored higher on quality than this 14B at Q8, though it needs
22GB and evicts the vision model entirely. Watch `ollama ps`'s placement
column, not just the resident-GB math.
## Swapping in your own model
Nothing here is load-bearing on the specific models. Swap freely — the defaults
are a starting point, and an abliterated or fine-tuned model of your choosing is
a perfectly legitimate replacement. What follows is the procedure so you can
make the change with numbers instead of vibes.
**1. Build a tagged variant.** Always tag with an explicit `num_ctx`; the
library default is usually far larger than you want and costs GB of KV cache.
```bash
printf 'FROM <your-model>\nPARAMETER num_ctx 32768\n' > Modelfile.mine
ollama create my-router-model:tag -f Modelfile.mine
```
**2. Point the config at it.** For classification and verification, set both —
they must name the **same tag**, or Ollama reloads the base model at a
different context every time a request alternates between the two:
```yaml
classifier: { model: "my-router-model:tag" }
verification: { model: "my-router-model:tag" }
```
For local dispatch you also need three things that are easy to miss:
- an entry under `local_dispatch_models` (with `context_window` matching the
tag's `num_ctx`),
- the categories it may serve in its `eligible_categories`, which must also
exist in `proficiency.categories`,
- **a pin under `tiering.model_tiers` keyed by the new model id** — `src/tier.py`
overwrites `models.tier` on every run, so without the pin the model silently
loses its tier and drops out of routing.
**3. Measure it, in this order.** Cheap and decisive first:
- **VRAM and placement**: load it, then `ollama ps`. Confirm `100%` — anything
less means a CPU spill and a large latency penalty.
- **Classification** (free, local): accuracy, **hard failures**, and the
latency **tail**. The tail is what has disqualified two classifiers in this
project's history, not the mean; a model that is usually fast but sometimes
spends 6s producing no parseable JSON degrades that request to
`source: "fallback"`.
- **Quality on the categories it will serve**:
`PYTHONPATH=src python -m eval_proficiency --models <id> --categories <cats>`.
Always bench a **cloud model on the same tasks** as a control — without it you
cannot tell "this model is weak" from "these tasks are unfair", and this
project has been wrong about that more than once.
- **Cost**: `PYTHONPATH=src python -m seed_local_dispatch_energy --samples 5`,
which measures real GPU draw and prices it at your `local_energy` tariff.
**4. Beware the sample size.** The bundled task sets are small —
`file_summarization` is 6 tasks (0.167 per task) and `diff_checking` is 8 tasks
over only 4 scenarios (0.125 per task). Differences under ~0.15 are noise. At
`temperature: 0` re-running is byte-identical, so more confidence means more
*tasks*, not more runs.
**5. Do not widen `eligible_categories` without measuring the new category.**
The hard filter is what keeps a summarization-grade model away from code. A
confidently wrong answer is worse than an honest error.
### A caution about small models
Small does not degrade gracefully; it fabricates. Two measured examples:
- `nemotron-mini:4b` answered "no" to **8/8** diff-checking tasks and to 3/3
blatant probes including `return a + b` -> `return a - b` — a detector that
never fires — and scored 0.25 on summarization while inventing behaviour that
was not in the source.
- `moondream` (1B) described a dense dashboard screenshot as "a spreadsheet
... possibly related to business decisions or financial analysis", reading no
title, no column and no value. `qwen3-vl:4b` on the same image read the page
title, quoted the subtitle verbatim and recovered the row count.
Test with **your real workload** — for vision here that means screenshots, not
a synthetic shape. A toy test will pass a model that cannot do the job.
## The classifier is the latency floor
Every routed request pays a full local classification round-trip before a
single upstream token is requested, so the classifier model is the single
biggest lever on interactive latency.
### Three generations, and the one lesson that outlived all of them
**The classifier default is `qwen2.5-coder-router:14b`** (see the sizing
section above), the second replacement in this project's history. Each swap
was measured, not assumed — same prompts, same system prompt,
`temperature: 0`, cold load excluded:
| | `qwen3.5` (original) | `mistral-nemo:12b` | `qwen2.5-coder:14b` (current) |
|---|---|---|---|
| category correct | 10/14 | 9/14 or 10/14¹ | 11/14 |
| **hard failures** | **4 of 19 calls (21%)** | 1/14 | **0/14** |
| latency mean | 6.6s | 1.25s | **1.07s** |
| **latency MAX** | **15.6s** | **6.10s** | **1.14s** |
¹ `mistral-nemo`'s own category-correct score reads differently depending on
which of the two original comparison runs it is pulled from (9/14 against
`qwen3.5`, 10/14 against `qwen2.5-coder`) — flagged rather than silently
picked, since this project has been burned before by treating a single run
as ground truth. Neither run changes the conclusion: the failure column and
the tail decided both swaps, not the category-accuracy column.
Category accuracy barely moves across all three. The failure column and the
tail are what actually decided each swap, and both keep shrinking: every
`qwen3.5` failure and most of `mistral-nemo`'s tail is the same
runaway-thinking-trace mode — the model spends its budget on an unbounded
chain of thought and returns no parseable JSON, which degrades that request
to `source: "fallback"` (tier 2, `general_chat`). End-to-end `/route` went
~10s (`qwen3.5`) → ~1.7s (`mistral-nemo`) → ~1.1s (`qwen2.5-coder`).
Consolidating onto one tag also let it serve classification, verification
AND local dispatch, so only one model stays resident.
**The lesson that outlived every model named above: pick a non-reasoning
model, and judge it by the tail, not the mean.** The classifier's entire
output is ~45 tokens of JSON — a model that reasons before answering is
solving the wrong problem, and a model that is usually fast but occasionally
spends 6-15s producing nothing is worse than one that is uniformly mediocre,
because that tail sits directly on the interactive latency floor.
### The four settings, and why each exists
| setting | value | fixes |
|---|---|---|
| `max_retries` | `0` | The OpenAI SDK retries twice by default, so `timeout_seconds` silently became a 3x wall-clock bound — a request hung past 250s on a 120s setting and logged nothing. |
| `max_output_tokens` | `1024` | Bounds a reasoning model's chain of thought. Unbounded, it cascades: Ollama keeps generating after the client gives up *and* serializes per model, so one runaway request queues every later request behind it. 256 was tried first and was too tight — the trace consumed the whole budget and the model was truncated before emitting any JSON. Does not bind for a non-reasoning model; kept because it costs nothing unused and is the only guard if a reasoning model is swapped back in. |
| `max_input_chars` | `8000` | The classifier decides a category and a tier; it does not need the document. Feeding it one is harmful, not wasteful — on a ~20k-token prompt, two different models spent 28.7s and 41.8s respectively without producing a usable answer, both landing on `source: "fallback"`. Clamped to head + tail (the instruction sits at one end or the other, never the middle), the same prompts classify in ~2.2s. Nothing is lost: `chat_completions` measures the real conversation with `estimate_prompt_tokens` and takes the larger value. |
| `fallback_tier` / `fallback_category` | mid tier / `general_chat` | A classifier that times out, errors, or returns garbage degrades to a configured mid tier flagged `source: "fallback"` instead of 502/503 — a coding agent would rather have a mid-tier answer than an error. Escalation deliberately skips fallbacks, so an unavailable local model does not silently promote every request to the frontier tier. |
| `temperature` | `0` | At the default, the same prompt classified tier 2 then tier 1 on consecutive calls and routed to two different models. |
### `local_encoder`: a different KIND of answer to the same failure mode
Every swap above (`qwen3.5` → `mistral-nemo` → `qwen2.5-coder`) picked a
*better-behaved generative model* — the runaway-reasoning-trace failure got
rarer each time, but never structurally impossible, because a chat-completion
model can always in principle spend its budget thinking instead of
answering. `classifier.mode: local_encoder` sidesteps the whole failure class
instead of picking around it: a zero-shot embedding classifier
(`classifier.encoder.model`, default `BAAI/bge-large-en-v1.5`) scores the
task directly against `proficiency.categories` and returns a label plus a
confidence — there is no generation step, so there is no trace to run away.
Originally a cross-encoder NLI pipeline (`facebook/bart-large-mnli`, itself a
fallback after `MoritzLaurer/deberta-v3-base-zeroshot-v2` — smaller, ~184M vs
~407M params — started returning 401 on an unauthenticated GET of its own
model page sometime after this project picked it, discovered live
2026-09-06 when it took production down: the startup check correctly
refused to boot rather than fail opaquely on the first request, but the
service still crash-looped until the default was corrected). Replaced with a
sentence-embedding + nearest-centroid architecture for Wave 5.1 of
`plans/token-waste-waves.md`: NLI cross-encoding scores the input against
*every* candidate label separately (one forward pass per category), while an
embedding model does one forward pass for the input and compares it by cosine
similarity against precomputed category-description embeddings — cost
independent of category count, which also affords a stronger backbone
(`bge-large-en-v1.5`, ~1.3GB, CPU-viable) on the same compute budget.
The trade is real, not free. It only produces `task_category` — no `task_tier`
signal exists in a zero-shot label score, so tier falls back to
`classifier.fallback_tier` for this mode, a documented limitation rather than
a second heuristic bolted on to guess one. And it is **zero-shot, not
fine-tuned**, deliberately: this router never stores raw task text anywhere —
`docs/operations.md` states it plainly and a test enforces it — so there is no
labeled corpus of your own traffic to train against without a new, separate,
opt-in capture feature (scoped in `CLAUDE.md`'s "What's NOT built yet", not
built). A below-`confidence_min` result is treated as a failure and
walks the same cascade a local-LLM parse failure would, unchanged.
`confidence_min` is a **minimum similarity score in [0.0, 1.0]**, not a
percentage. `classify_zero_shot` returns a similarity score between 0 and 1,
so a value of `0.5` means the score must be at least 0.5 to avoid being
treated as a failure. The shipped default is `0.5`. The validator in
`src/config.py` rejects anything outside [0.0, 1.0] with an error. The knob
is *not* a probability — with nearest-centroid (pre-head) scoring it is a
softmax-amplified cosine similarity that has no probabilistic meaning, so it
is named `confidence_min`, not `confidence_threshold`. Only when the
trainable logistic-regression head is present (below) does the returned score
carry real `P(correct)` meaning via Platt calibration.
The admin UI lets you type 0-100 and scales to the probability before saving,
so typing `80` in the UI writes `0.8` to the overlay. The config file and any
other direct writer must supply the raw probability. That boundary conversion
matters: a raw `80` in the YAML is two orders of magnitude above the maximum
and would make every real score read as below threshold, since no probability
exceeds 1.0.
That validator exists because of a **production near-miss** (issue #47). The
admin UI used to save the threshold as written; an `80` was entered, the UI
wrote `80.0` straight into the overlay, and it was only caught because the
service had not yet restarted and loaded the bad value. Once restarted,
classification would have failed on *every* request. The UI now converts
0-100 to 0.0-1.0 before saving, and the validator is the fail-closed last line
of defense for every other caller.
**Measuring accuracy: the `eval_classifier.py` harness.** The included
`eval_classifier.py` harness (run via `PYTHONPATH=src python
src/eval_classifier.py`) measures accuracy on the 46-task eval set —
`evals/tasks.yaml` minus the 11 `tool_use_agentic` rows the encoder cannot
emit — with three noise variants per prompt (clean, short-noise-wrap,
long-noise-wrap) and a confusion-matrix output showing per-category
precision/recall and the `blended_score` gap per cell. `--dry-run` plans
the run without loading the model. This is what makes further accuracy
regression measurable instead of a vibe check, and it stays standalone
(`requirements-encoder.txt`) so the router's dispatch path never imports
transformers/torch.
**A fitted trainable head.** A fitted logistic-regression head
(`_TrainableHead`) replaces nearest-centroid when
`evals/synthetic/encoder-head-coefficients.json` is present, improving
accuracy and providing a calibrated confidence score. Trained by
`scripts/train_encoder_head.py` on the committed synthetic corpus
(`evals/synthetic/encoder-training.jsonl`), it runs a convex
`LogisticRegression(solver="lbfgs")` over the same frozen embeddings with
per-class Platt (sigmoid) calibration, so `predict_proba` genuinely means
`P(correct)` — the confidence gate reads as a real abstention signal rather
than a monotonic score. The artifact is keyed to a specific backbone and
candidate set; if the file is absent or the model/categories don't match, it
degrades to the unchanged nearest-centroid path with a logged warning.
**Candidate labels are natural-language descriptions, never the raw config
identifier.** HF's zero-shot pipeline scores a candidate label against the
input via a hypothesis template (`"This example is {}."`), so the label
itself has to read as a sentence for entailment scoring to work — feeding it
`tool_use_agentic` or `diff_checking` verbatim asks the model to judge
`"This example is tool_use_agentic."`, not a sentence its NLI training ever
saw. `local_encoder.py`'s `_CATEGORY_DESCRIPTIONS` maps each category to a
real description before scoring, and `multi_label=True` scores each
candidate independently rather than normalizing every score to sum to 1 (the
pipeline's single-label default, which drags a genuinely good match down
whenever another category is also plausible). Measured live 2026-09-06 on the
same 9-prompt set: raw identifiers scored 5/9 correct at confidence
0.15-0.90; natural-language descriptions scored 8/9 correct at 0.39-0.999. A
category missing from the map falls back to its own raw string rather than
raising, so a newly-added config category degrades gracefully instead of
crashing.
`local_encoder` runs on **cpu by default** in the shipped example; setting it
to `cuda` requires a `config.local.yaml` overlay and a host that actually has
the hardware. The project does not ship a settled default device. Match the
device to the machine the router is running on, otherwise `torch` will fail
fast with a clear error at first classification. CPU inference is the safe,
portable starting point; CUDA is host-dependent and should be chosen only
where the router process has a card available.
`transformers`/`torch` are optional (`requirements-encoder.txt`, not the main
`requirements.txt` — pinned-dependency discipline applies here too) and
imported lazily inside `local_encoder.py`, the same rule `tui.py` follows for
`textual`: a deployment that never selects this mode needs neither installed.
CPU install:
```bash
pip install -r requirements-encoder.txt
```
### `local_decision`: read the first-token logprobs, never parse a generation
`classifier.mode: local_decision` is a third answer to the same
runaway-reasoning problem, and the most direct one. Instead of asking a small
generative model to emit JSON and hoping it stops, it asks for a single
lettered choice and reads the decision out of the **logprobs on the first
output token**.
Mechanism, a Jev-style first-token-logprob classifier:
- The task text and a lettered list of candidate categories go into one prompt;
the model is asked to answer with one letter.
- Ollama returns `logprobs` for that first position. `local_decision.parse_logprobs`
sums `exp(logprob)` per option letter (A through J), ignoring non-option
tokens and Ollama's `.` separator tokens, then normalizes the accumulated mass
into a confidence. The letter with the most mass is the predicted category.
- There is **no JSON parse, no reasoning trace, and no multi-token generation
to run away**. `num_predict` is 1 and `think` is false, so a model that would
normally spend its budget thinking has nowhere to spend it. The whole failure
class that took `qwen3.5` down on 4 of 19 calls simply does not exist here.
- Confidence is the chosen letter's share of total option mass. `coverage_min`
(default `0.3`) is the minimum total option mass for a call to be accepted at
all; below it, or when no logprobs come back, the call raises and walks the
same cascade a `local_llm` parse failure would. `confidence_min` (default
`0.5`) treats a below-threshold verdict as a failure rather than a
low-confidence answer, mirroring the encoder's semantics.
- As with `local_encoder`, it produces only `task_category` unless
`classifier.decision.tier_enabled` is set, in which case it may also choose
`task_tier`; otherwise tier falls back to `classifier.fallback_tier`.
- A declined answer is recorded, not just logged: `route_decisions.classifier_confidence`,
`classifier_coverage` and `classifier_reject` hold the numbers and the reason
(see [data-model](data-model.md#what-the-classifier-did-and-why-its-answer-was-not-used)).
On the 4b the decline is by design, not an outage: it ran at 30% of turns on
2026-10-04 (8% the day before), concentrated in long agent sessions.
Pull and enable it:
```bash
ollama pull qwen3.5:4b
```
```yaml
# config.local.yaml
classifier:
mode: "local_decision"
decision: {}
```
Every field in `classifier.decision` has a default, so `decision: {}` is enough
to opt in. The one that matters for hardware is **`classifier.decision.num_ctx`**:
it sets the context window the classifier loads at (default `8192`) and scales
the KV cache directly, so it is the cheapest GB to reclaim when the pair below
does not fit. Model, `base_url`, timeout, and the confidence gates all live
under the same block.
VRAM, and why this is tight on a 24 GB card:
| model | role | num_ctx | resident |
|---|---|---|---|
| `qwen2.5-coder-router:14b` | classify + verify + dispatch | 16384 | 12.3 GB |
| `qwen3.5:4b` | `local_decision` | 8192 | ~7 GB |
| | | **total** | **~19.9 GB of 24 GB** |
The pair fits, with about 4 GB spare, but **only because the 14b router runs at
`num_ctx 16384`, not the 32768 it was originally tagged with**. At 32k the
router alone measured 16.5 to 17.8 GB resident, and the two models then fit by
about 1 GB or get evicted outright depending on load order. Cutting the router
to 16k is what makes the coexistence robust; that is a change to the
`qwen2.5-coder-router:14b` Modelfile, not to this config block. Local vision is
off in this profile because its 5.6 GB model no longer fits alongside both.
### The latency tax, and what has and hasn't addressed it
~10s of local overhead per message was the original number on a reasoning
classifier; the current `qwen2.5-coder-router:14b` brings the round-trip
itself down to ~1.1s (see the table above), so the *per-call* cost of that
tax is mostly gone. What is still unaddressed is the *per-message* repetition
of it: `session_cache` (off by default, `config/config.yaml` ->
`session_cache.enabled`) now exists specifically to classify once per
session rather than once per message, but nobody has measured its effect on
a real interactive session yet. Caching by prompt hash and skipping
classification for short prompts remain unbuilt. `classifier.mode:
local_encoder` above addresses a different cost entirely (which model
answers) and does not touch this one (how often it is asked).
Separately, a **cold reload** is its own latency spike — ~6s when loading a
second model evicts the resident one (see the KV-cache section above) — and
is not the same phenomenon as the per-message round-trip tax: cold load
happens once per eviction, the round-trip tax happens on every single
message even with everything warm.
### "Local" just means reachable, not necessarily on this machine
"Local" here means the model runs on hardware **you** control and don't pay a
cloud provider for — not that it has to share a motherboard with the process
asking it questions. All three of these count as "local" for this router's
purposes, and the config looks identical for each:
- **localhost** — Ollama and the router on the same machine.
- **local network** — Ollama on another box on your LAN.
- **a VPN tunnel** — Ollama on a home-lab workstation, router and editor on a
laptop, connected over WireGuard or similar. This is the common shape for
anyone whose GPU isn't in their daily-driver machine.
`classifier.base_url` / `api_key_env` / `model` take any OpenAI-compatible
endpoint, and switching between the three above is just changing that URL —
verified end to end against a non-loopback address, classifier and verifier
both.
Ollama binds `127.0.0.1` by default, so the VPN case fails with
connection-refused until the serving host applies
`deploy/ollama-over-vpn.conf`. Bind it to the VPN address, not `0.0.0.0`:
Ollama has no auth of any kind, so anything that can reach the port can run
inference and enumerate your models, and `0.0.0.0` publishes it on whatever
network the serving machine happens to be sitting on.
#### Local vs cloud classifier, measured (historical — against the superseded `qwen3.5`)
**This comparison predates the swap to `qwen2.5-coder-router:14b` above** and
has not been re-run against the current classifier. It measured the
*original* `qwen3.5` classifier against a cloud one, five prompts, same
system prompt, `temperature: 0`:
| | local `qwen3.5` (RTX 6000, superseded) | NeuralWatt `deepseek-v4-flash` |
|---|---|---|
| mean latency | 11.58s (4.95-15.76) | 1.02s |
| categories agreed with the label | 2 of 4 | 5 of 5 |
| hard failures | 1 of 5 (empty after 15.76s) | 0 |
| energy per call | ~7e-05 kWh, on your meter | 1.17e-05 kWh attributed |
| cost per 1,000 calls | electricity + 6.6GB resident | $0.093 (0.19% of quota) |
At the time, the cloud side was faster, more accurate, and less attributed
energy — the one hard failure is the same runaway-thinking-trace mode
described above. The general lesson still holds regardless of which local
model is current: local classification is not free, it is just unbilled, and
that is a substitution the cost axis had to unlearn independently (see
`CLAUDE.md`'s billing section). The shipped default is still local Ollama,
because switching spends quota and that is a deployment choice, not a code
one.
**Worth re-running against `qwen2.5-coder-router:14b`** before treating the
11x figure as current — the classifier's own latency dropped roughly 10x
since this table was measured (11.58s mean here vs ~1.1s in the three-way
table above), which would substantially narrow or erase the local/cloud gap
this table reports. Flagging this rather than re-deriving a number with no
live hardware to measure against.
### The verifier follows, but only to another Ollama
The local LLM check speaks Ollama's **native** `/api/chat` (the only way to
set `think: False`), so it follows the classifier across a VPN but not to a
cloud provider. It used to derive its URL by stripping `/v1` off
`classifier.base_url`, which meant moving the classifier at all would have
pointed it at `<that host>/api/chat`. It now has its own
`verification.base_url` and `verification.model`.
`verification.model` may be null only while both run on one host. Config load
**refuses** the null once the hostnames differ, because the failure is silent:
observed directly, with the classifier on NeuralWatt the verifier POSTed
`deepseek-v4-flash` to `localhost:11434`, 404'd, caught it, logged "local
verification unavailable" and recorded no sample. Verification would have
looked enabled while producing nothing.
Structural verification needs no model at all — it is pure Python — so
`verification.local_llm_enabled: false` leaves a host with no local inference
fully functional, minus the refusal/incoherence class of failure.
### A note on prompt-token estimation
opencode sends ~32K prompt tokens of system prompt and tool definitions on a
trivial request, so the measured-size floor in `estimate_prompt_tokens` does
real work — the classifier's own estimate for that request was two orders of
magnitude low. This is the same phenomenon pinch confronts on the provider
side; see [docs/pinch.md](pinch.md).