Files
6krrt/docs/local-models.md
adlee-was-taken 5ef411496b docs: dedupe confidence_threshold section, extend with the #50 accuracy fix
The merge from origin/main brought in #48's own confidence_threshold
paragraph, which duplicated this branch's own independent write-up of
the same incident. Kept the more detailed version, removed the
redundant one, and added the #50 natural-language-labels accuracy fix
(5/9 -> 8/9 correct, confidence 0.15-0.90 -> 0.39-0.999) since it's the
direct continuation of the same incident thread.
2026-09-06 22:15:05 -04:00

22 KiB

Sizing local models for your GPU. Back to README.

Fitting local models into 24GB VRAM

These defaults are measured on a 24GB Quadro RTX 6000 (Turing, compute 7.5) and are a sound starting point for any 24-32GB card. They are defaults, not prescriptions — see Swapping in your own model below, which is the part that matters if you run different hardware or simply want a different model.

Ollama loads a model at the context size baked into its tag. You cannot set it per request: the OpenAI-compatible /v1/chat/completions endpoint in Ollama 0.22.0 was verified live against every plausible field shape, and in every case returned 200 while silently keeping the loaded context unchanged. So context size has to go into the tag via a Modelfile:

# Classification + verification + local dispatch (ONE resident model)
printf 'FROM qwen2.5-coder:14b\nPARAMETER num_ctx 32768\n' > Modelfile.local-dispatch
ollama create qwen2.5-coder-router:14b -f Modelfile.local-dispatch

# Vision fallback only
printf 'FROM qwen3-vl:4b\nPARAMETER num_ctx 8192\n' > Modelfile.vision
ollama create qwen3-vl-router:4b -f Modelfile.vision

Measured co-residency, both fully on GPU:

model role resident
qwen2.5-coder-router:14b (Q4_K_M, 32k) classify + verify + dispatch 17.0GB
qwen3-vl-router:4b (8k) vision fallback 5.6GB
total 23.2GB of 24GB

That leaves ~800MB spare, which is thin but works. The single most important thing to understand is where the memory actually goes.

KV cache is the cheapest GB to reclaim

Weights are fixed; KV cache scales with num_ctx and is often the larger lever. Measured on this hardware:

model on disk resident @ num_ctx
qwen2.5-coder:14b Q4_K_M 9.0GB 13GB @ 16k, 17GB @ 32k
qwen3-vl:4b 3.3GB 5.6GB @ 8k, 6.9GB @ 16k

Halving the vision model's context bought 1.3GB — exactly what lets the dispatch model run at 32k with vision still resident. Without it, loading vision evicts the dispatch model entirely, and the next classification pays a ~6s cold reload on the interactive latency floor.

If you are short on VRAM, cut num_ctx before you cut anything else. The classifier in particular never needs a large window: classifier.max_input_chars clamps input to 8000 chars (~2.7k tokens), so even 8192 leaves 2x headroom.

Bigger quantization is usually the wrong trade

Tempting and mostly counterproductive. Measured, same model, same tasks:

variant resident placement classify accuracy classify mean
Q4_K_M @ 16k 13GB 100% GPU 11/14 1.07s
Q4_K_M @ 32k 17GB 100% GPU 11/14 1.14s
Q8_0 @ 16k 19GB 100% GPU 11/14 1.62s
Q8_0 @ 32k 24GB 83% GPU 11/14 3.79s

Accuracy never moved — only speed did. Q8 costs ~1.5x latency even when it fits (Turing is memory-bandwidth bound here, and Q8 moves twice the weight bytes per token), and 83% GPU in the placement column means Ollama spilled the rest to system RAM for a 3.3x latency hit — a config that "fits" on paper can silently land there. Going up in parameters beat going up in bits: a 32B model at Q4 scored higher on quality than this 14B at Q8, though it needs 22GB and evicts the vision model entirely. Watch ollama ps's placement column, not just the resident-GB math.

Swapping in your own model

Nothing here is load-bearing on the specific models. Swap freely — the defaults are a starting point, and an abliterated or fine-tuned model of your choosing is a perfectly legitimate replacement. What follows is the procedure so you can make the change with numbers instead of vibes.

1. Build a tagged variant. Always tag with an explicit num_ctx; the library default is usually far larger than you want and costs GB of KV cache.

printf 'FROM <your-model>\nPARAMETER num_ctx 32768\n' > Modelfile.mine
ollama create my-router-model:tag -f Modelfile.mine

2. Point the config at it. For classification and verification, set both — they must name the same tag, or Ollama reloads the base model at a different context every time a request alternates between the two:

classifier:   { model: "my-router-model:tag" }
verification: { model: "my-router-model:tag" }

For local dispatch you also need three things that are easy to miss:

  • an entry under local_dispatch_models (with context_window matching the tag's num_ctx),
  • the categories it may serve in its eligible_categories, which must also exist in proficiency.categories,
  • a pin under tiering.model_tiers keyed by the new model id — src/tier.py overwrites models.tier on every run, so without the pin the model silently loses its tier and drops out of routing.

3. Measure it, in this order. Cheap and decisive first:

  • VRAM and placement: load it, then ollama ps. Confirm 100% — anything less means a CPU spill and a large latency penalty.
  • Classification (free, local): accuracy, hard failures, and the latency tail. The tail is what has disqualified two classifiers in this project's history, not the mean; a model that is usually fast but sometimes spends 6s producing no parseable JSON degrades that request to source: "fallback".
  • Quality on the categories it will serve: PYTHONPATH=src python -m eval_proficiency --models <id> --categories <cats>. Always bench a cloud model on the same tasks as a control — without it you cannot tell "this model is weak" from "these tasks are unfair", and this project has been wrong about that more than once.
  • Cost: PYTHONPATH=src python -m seed_local_dispatch_energy --samples 5, which measures real GPU draw and prices it at your local_energy tariff.

4. Beware the sample size. The bundled task sets are small — file_summarization is 6 tasks (0.167 per task) and diff_checking is 8 tasks over only 4 scenarios (0.125 per task). Differences under ~0.15 are noise. At temperature: 0 re-running is byte-identical, so more confidence means more tasks, not more runs.

5. Do not widen eligible_categories without measuring the new category. The hard filter is what keeps a summarization-grade model away from code. A confidently wrong answer is worse than an honest error.

A caution about small models

Small does not degrade gracefully; it fabricates. Two measured examples:

  • nemotron-mini:4b answered "no" to 8/8 diff-checking tasks and to 3/3 blatant probes including return a + b -> return a - b — a detector that never fires — and scored 0.25 on summarization while inventing behaviour that was not in the source.
  • moondream (1B) described a dense dashboard screenshot as "a spreadsheet ... possibly related to business decisions or financial analysis", reading no title, no column and no value. qwen3-vl:4b on the same image read the page title, quoted the subtitle verbatim and recovered the row count.

Test with your real workload — for vision here that means screenshots, not a synthetic shape. A toy test will pass a model that cannot do the job.

The classifier is the latency floor

Every routed request pays a full local classification round-trip before a single upstream token is requested, so the classifier model is the single biggest lever on interactive latency.

Three generations, and the one lesson that outlived all of them

The classifier default is qwen2.5-coder-router:14b (see the sizing section above), the second replacement in this project's history. Each swap was measured, not assumed — same prompts, same system prompt, temperature: 0, cold load excluded:

qwen3.5 (original) mistral-nemo:12b qwen2.5-coder:14b (current)
category correct 10/14 9/14 or 10/14¹ 11/14
hard failures 4 of 19 calls (21%) 1/14 0/14
latency mean 6.6s 1.25s 1.07s
latency MAX 15.6s 6.10s 1.14s

¹ mistral-nemo's own category-correct score reads differently depending on which of the two original comparison runs it is pulled from (9/14 against qwen3.5, 10/14 against qwen2.5-coder) — flagged rather than silently picked, since this project has been burned before by treating a single run as ground truth. Neither run changes the conclusion: the failure column and the tail decided both swaps, not the category-accuracy column.

Category accuracy barely moves across all three. The failure column and the tail are what actually decided each swap, and both keep shrinking: every qwen3.5 failure and most of mistral-nemo's tail is the same runaway-thinking-trace mode — the model spends its budget on an unbounded chain of thought and returns no parseable JSON, which degrades that request to source: "fallback" (tier 2, general_chat). End-to-end /route went ~10s (qwen3.5) → ~1.7s (mistral-nemo) → ~1.1s (qwen2.5-coder). Consolidating onto one tag also let it serve classification, verification AND local dispatch, so only one model stays resident.

The lesson that outlived every model named above: pick a non-reasoning model, and judge it by the tail, not the mean. The classifier's entire output is ~45 tokens of JSON — a model that reasons before answering is solving the wrong problem, and a model that is usually fast but occasionally spends 6-15s producing nothing is worse than one that is uniformly mediocre, because that tail sits directly on the interactive latency floor.

The four settings, and why each exists

setting value fixes
max_retries 0 The OpenAI SDK retries twice by default, so timeout_seconds silently became a 3x wall-clock bound — a request hung past 250s on a 120s setting and logged nothing.
max_output_tokens 1024 Bounds a reasoning model's chain of thought. Unbounded, it cascades: Ollama keeps generating after the client gives up and serializes per model, so one runaway request queues every later request behind it. 256 was tried first and was too tight — the trace consumed the whole budget and the model was truncated before emitting any JSON. Does not bind for a non-reasoning model; kept because it costs nothing unused and is the only guard if a reasoning model is swapped back in.
max_input_chars 8000 The classifier decides a category and a tier; it does not need the document. Feeding it one is harmful, not wasteful — on a ~20k-token prompt, two different models spent 28.7s and 41.8s respectively without producing a usable answer, both landing on source: "fallback". Clamped to head + tail (the instruction sits at one end or the other, never the middle), the same prompts classify in ~2.2s. Nothing is lost: chat_completions measures the real conversation with estimate_prompt_tokens and takes the larger value.
fallback_tier / fallback_category mid tier / general_chat A classifier that times out, errors, or returns garbage degrades to a configured mid tier flagged source: "fallback" instead of 502/503 — a coding agent would rather have a mid-tier answer than an error. Escalation deliberately skips fallbacks, so an unavailable local model does not silently promote every request to the frontier tier.
temperature 0 At the default, the same prompt classified tier 2 then tier 1 on consecutive calls and routed to two different models.

local_encoder: a different KIND of answer to the same failure mode

Every swap above (qwen3.5 → mistral-nemo → qwen2.5-coder) picked a better-behaved generative model — the runaway-reasoning-trace failure got rarer each time, but never structurally impossible, because a chat-completion model can always in principle spend its budget thinking instead of answering. classifier.mode: local_encoder sidesteps the whole failure class instead of picking around it: a zero-shot NLI encoder (classifier.encoder.model, default facebook/bart-large-mnli) scores the task directly against proficiency.categories and returns a label plus a confidence — there is no generation step, so there is no trace to run away. (Was MoritzLaurer/deberta-v3-base-zeroshot-v2 — smaller, ~184M vs ~407M params — until that repo started returning 401 on an unauthenticated GET of its own model page sometime after this project picked it, discovered live 2026-09-06 when it took production down: the startup check correctly refused to boot rather than fail opaquely on the first request, but the service still crash-looped until the default was corrected.)

The trade is real, not free. It only produces task_category — no task_tier signal exists in a zero-shot label score, so tier falls back to classifier.fallback_tier for this mode, a documented limitation rather than a second heuristic bolted on to guess one. And it is zero-shot, not fine-tuned, deliberately: this router never stores raw task text anywhere — docs/operations.md states it plainly and a test enforces it — so there is no labeled corpus of your own traffic to train against without a new, separate, opt-in capture feature (scoped in CLAUDE.md's "What's NOT built yet", not built). A below-confidence_threshold result is treated as a failure and walks the same cascade a local-LLM parse failure would, unchanged.

confidence_threshold is a probability in [0.0, 1.0], not a percentage. classify_zero_shot returns a confidence score between 0 and 1, so a value of 0.5 means 50% confidence. The shipped default is 0.5. The validator in src/config.py rejects anything outside [0.0, 1.0] with an error that says exactly this: it is a probability, not a percent.

The admin UI lets you type 0-100 and scales to the probability before saving, so typing 80 in the UI writes 0.8 to the overlay. The config file and any other direct writer must supply the raw probability. That boundary conversion matters: a raw 80 in the YAML is two orders of magnitude above the maximum and would make every real score read as below threshold, since no probability exceeds 1.0.

That validator exists because of a production near-miss (issue #47). The admin UI used to save the threshold as written; an 80 was entered, the UI wrote 80.0 straight into the overlay, and it was only caught because the service had not yet restarted and loaded the bad value. Once restarted, classification would have failed on every request. The UI now converts 0-100 to 0.0-1.0 before saving, and the validator is the fail-closed last line of defense for every other caller.

Candidate labels are natural-language descriptions, never the raw config identifier. HF's zero-shot pipeline scores a candidate label against the input via a hypothesis template ("This example is {}."), so the label itself has to read as a sentence for entailment scoring to work — feeding it tool_use_agentic or diff_checking verbatim asks the model to judge "This example is tool_use_agentic.", not a sentence its NLI training ever saw. local_encoder.py's _CATEGORY_DESCRIPTIONS maps each category to a real description before scoring, and multi_label=True scores each candidate independently rather than normalizing every score to sum to 1 (the pipeline's single-label default, which drags a genuinely good match down whenever another category is also plausible). Measured live 2026-09-06 on the same 9-prompt set: raw identifiers scored 5/9 correct at confidence 0.15-0.90; natural-language descriptions scored 8/9 correct at 0.39-0.999. A category missing from the map falls back to its own raw string rather than raising, so a newly-added config category degrades gracefully instead of crashing.

local_encoder runs on cpu by default in the shipped example; setting it to cuda requires a config.local.yaml overlay and a host that actually has the hardware. The project does not ship a settled default device. Match the device to the machine the router is running on, otherwise torch will fail fast with a clear error at first classification. CPU inference is the safe, portable starting point; CUDA is host-dependent and should be chosen only where the router process has a card available.

transformers/torch are optional (requirements-encoder.txt, not the main requirements.txt — pinned-dependency discipline applies here too) and imported lazily inside local_encoder.py, the same rule tui.py follows for textual: a deployment that never selects this mode needs neither installed. CPU install:

pip install -r requirements-encoder.txt

The latency tax, and what has and hasn't addressed it

~10s of local overhead per message was the original number on a reasoning classifier; the current qwen2.5-coder-router:14b brings the round-trip itself down to ~1.1s (see the table above), so the per-call cost of that tax is mostly gone. What is still unaddressed is the per-message repetition of it: session_cache (off by default, config/config.yaml -> session_cache.enabled) now exists specifically to classify once per session rather than once per message, but nobody has measured its effect on a real interactive session yet. Caching by prompt hash and skipping classification for short prompts remain unbuilt. classifier.mode: local_encoder above addresses a different cost entirely (which model answers) and does not touch this one (how often it is asked).

Separately, a cold reload is its own latency spike — ~6s when loading a second model evicts the resident one (see the KV-cache section above) — and is not the same phenomenon as the per-message round-trip tax: cold load happens once per eviction, the round-trip tax happens on every single message even with everything warm.

"Local" just means reachable, not necessarily on this machine

"Local" here means the model runs on hardware you control and don't pay a cloud provider for — not that it has to share a motherboard with the process asking it questions. All three of these count as "local" for this router's purposes, and the config looks identical for each:

  • localhost — Ollama and the router on the same machine.
  • local network — Ollama on another box on your LAN.
  • a VPN tunnel — Ollama on a home-lab workstation, router and editor on a laptop, connected over WireGuard or similar. This is the common shape for anyone whose GPU isn't in their daily-driver machine.

classifier.base_url / api_key_env / model take any OpenAI-compatible endpoint, and switching between the three above is just changing that URL — verified end to end against a non-loopback address, classifier and verifier both.

Ollama binds 127.0.0.1 by default, so the VPN case fails with connection-refused until the serving host applies deploy/ollama-over-vpn.conf. Bind it to the VPN address, not 0.0.0.0: Ollama has no auth of any kind, so anything that can reach the port can run inference and enumerate your models, and 0.0.0.0 publishes it on whatever network the serving machine happens to be sitting on.

Local vs cloud classifier, measured (historical — against the superseded qwen3.5)

This comparison predates the swap to qwen2.5-coder-router:14b above and has not been re-run against the current classifier. It measured the original qwen3.5 classifier against a cloud one, five prompts, same system prompt, temperature: 0:

local qwen3.5 (RTX 6000, superseded) NeuralWatt deepseek-v4-flash
mean latency 11.58s (4.95-15.76) 1.02s
categories agreed with the label 2 of 4 5 of 5
hard failures 1 of 5 (empty after 15.76s) 0
energy per call ~7e-05 kWh, on your meter 1.17e-05 kWh attributed
cost per 1,000 calls electricity + 6.6GB resident $0.093 (0.19% of quota)

At the time, the cloud side was faster, more accurate, and less attributed energy — the one hard failure is the same runaway-thinking-trace mode described above. The general lesson still holds regardless of which local model is current: local classification is not free, it is just unbilled, and that is a substitution the cost axis had to unlearn independently (see CLAUDE.md's billing section). The shipped default is still local Ollama, because switching spends quota and that is a deployment choice, not a code one.

Worth re-running against qwen2.5-coder-router:14b before treating the 11x figure as current — the classifier's own latency dropped roughly 10x since this table was measured (11.58s mean here vs ~1.1s in the three-way table above), which would substantially narrow or erase the local/cloud gap this table reports. Flagging this rather than re-deriving a number with no live hardware to measure against.

The verifier follows, but only to another Ollama

The local LLM check speaks Ollama's native /api/chat (the only way to set think: False), so it follows the classifier across a VPN but not to a cloud provider. It used to derive its URL by stripping /v1 off classifier.base_url, which meant moving the classifier at all would have pointed it at <that host>/api/chat. It now has its own verification.base_url and verification.model.

verification.model may be null only while both run on one host. Config load refuses the null once the hostnames differ, because the failure is silent: observed directly, with the classifier on NeuralWatt the verifier POSTed deepseek-v4-flash to localhost:11434, 404'd, caught it, logged "local verification unavailable" and recorded no sample. Verification would have looked enabled while producing nothing.

Structural verification needs no model at all — it is pure Python — so verification.local_llm_enabled: false leaves a host with no local inference fully functional, minus the refusal/incoherence class of failure.

A note on prompt-token estimation

opencode sends ~32K prompt tokens of system prompt and tool definitions on a trivial request, so the measured-size floor in estimate_prompt_tokens does real work — the classifier's own estimate for that request was two orders of magnitude low. This is the same phenomenon pinch confronts on the provider side; see docs/pinch.md.