> Sizing local models for your GPU. Back to [README](../README.md). ## Fitting local models into 24GB VRAM These defaults are measured on a 24GB Quadro RTX 6000 (Turing, compute 7.5) and are a sound starting point for any 24-32GB card. They are **defaults, not prescriptions** — see [Swapping in your own model](#swapping-in-your-own-model) below, which is the part that matters if you run different hardware or simply want a different model. Ollama loads a model at the context size baked into its tag. You cannot set it per request: the OpenAI-compatible `/v1/chat/completions` endpoint in Ollama 0.22.0 was verified live against every plausible field shape, and in every case returned 200 while silently keeping the loaded context unchanged. So context size has to go into the tag via a `Modelfile`: ```bash # Classification + verification + local dispatch (ONE resident model) printf 'FROM qwen2.5-coder:14b\nPARAMETER num_ctx 32768\n' > Modelfile.local-dispatch ollama create qwen2.5-coder-router:14b -f Modelfile.local-dispatch # Vision fallback only printf 'FROM qwen3-vl:4b\nPARAMETER num_ctx 8192\n' > Modelfile.vision ollama create qwen3-vl-router:4b -f Modelfile.vision ``` Measured co-residency, both fully on GPU: | model | role | resident | |---|---|---| | `qwen2.5-coder-router:14b` (Q4_K_M, 32k) | classify + verify + dispatch | 17.0GB | | `qwen3-vl-router:4b` (8k) | vision fallback | 5.6GB | | | **total** | **23.2GB of 24GB** | That leaves ~800MB spare, which is thin but works. The single most important thing to understand is **where the memory actually goes**. ### KV cache is the cheapest GB to reclaim Weights are fixed; KV cache scales with `num_ctx` and is often the larger lever. Measured on this hardware: | model | on disk | resident @ num_ctx | |---|---|---| | `qwen2.5-coder:14b` Q4_K_M | 9.0GB | 13GB @ 16k, **17GB @ 32k** | | `qwen3-vl:4b` | 3.3GB | 5.6GB @ 8k, **6.9GB @ 16k** | Halving the vision model's context bought 1.3GB — exactly what lets the dispatch model run at 32k with vision still resident. Without it, loading vision **evicts the dispatch model entirely**, and the next classification pays a ~6s cold reload on the interactive latency floor. If you are short on VRAM, cut `num_ctx` before you cut anything else. The classifier in particular never needs a large window: `classifier.max_input_chars` clamps input to 8000 chars (~2.7k tokens), so even 8192 leaves 2x headroom. ### Bigger quantization is usually the wrong trade Tempting and mostly counterproductive. Measured, same model, same tasks: | variant | resident | placement | classify accuracy | classify mean | |---|---|---|---|---| | Q4_K_M @ 16k | 13GB | 100% GPU | 11/14 | **1.07s** | | Q4_K_M @ 32k | 17GB | 100% GPU | 11/14 | 1.14s | | Q8_0 @ 16k | 19GB | 100% GPU | 11/14 | 1.62s | | Q8_0 @ 32k | 24GB | **83% GPU** | 11/14 | 3.79s | Accuracy never moved — only speed did. Q8 costs ~1.5x latency even when it fits (Turing is memory-bandwidth bound here, and Q8 moves twice the weight bytes per token), and `83% GPU` in the placement column means Ollama spilled the rest to system RAM for a 3.3x latency hit — a config that "fits" on paper can silently land there. Going up in *parameters* beat going up in *bits*: a 32B model at Q4 scored higher on quality than this 14B at Q8, though it needs 22GB and evicts the vision model entirely. Watch `ollama ps`'s placement column, not just the resident-GB math. ## Swapping in your own model Nothing here is load-bearing on the specific models. Swap freely — the defaults are a starting point, and an abliterated or fine-tuned model of your choosing is a perfectly legitimate replacement. What follows is the procedure so you can make the change with numbers instead of vibes. **1. Build a tagged variant.** Always tag with an explicit `num_ctx`; the library default is usually far larger than you want and costs GB of KV cache. ```bash printf 'FROM \nPARAMETER num_ctx 32768\n' > Modelfile.mine ollama create my-router-model:tag -f Modelfile.mine ``` **2. Point the config at it.** For classification and verification, set both — they must name the **same tag**, or Ollama reloads the base model at a different context every time a request alternates between the two: ```yaml classifier: { model: "my-router-model:tag" } verification: { model: "my-router-model:tag" } ``` For local dispatch you also need three things that are easy to miss: - an entry under `local_dispatch_models` (with `context_window` matching the tag's `num_ctx`), - the categories it may serve in its `eligible_categories`, which must also exist in `proficiency.categories`, - **a pin under `tiering.model_tiers` keyed by the new model id** — `src/tier.py` overwrites `models.tier` on every run, so without the pin the model silently loses its tier and drops out of routing. **3. Measure it, in this order.** Cheap and decisive first: - **VRAM and placement**: load it, then `ollama ps`. Confirm `100%` — anything less means a CPU spill and a large latency penalty. - **Classification** (free, local): accuracy, **hard failures**, and the latency **tail**. The tail is what has disqualified two classifiers in this project's history, not the mean; a model that is usually fast but sometimes spends 6s producing no parseable JSON degrades that request to `source: "fallback"`. - **Quality on the categories it will serve**: `PYTHONPATH=src python -m eval_proficiency --models --categories `. Always bench a **cloud model on the same tasks** as a control — without it you cannot tell "this model is weak" from "these tasks are unfair", and this project has been wrong about that more than once. - **Cost**: `PYTHONPATH=src python -m seed_local_dispatch_energy --samples 5`, which measures real GPU draw and prices it at your `local_energy` tariff. **4. Beware the sample size.** The bundled task sets are small — `file_summarization` is 6 tasks (0.167 per task) and `diff_checking` is 8 tasks over only 4 scenarios (0.125 per task). Differences under ~0.15 are noise. At `temperature: 0` re-running is byte-identical, so more confidence means more *tasks*, not more runs. **5. Do not widen `eligible_categories` without measuring the new category.** The hard filter is what keeps a summarization-grade model away from code. A confidently wrong answer is worse than an honest error. ### A caution about small models Small does not degrade gracefully; it fabricates. Two measured examples: - `nemotron-mini:4b` answered "no" to **8/8** diff-checking tasks and to 3/3 blatant probes including `return a + b` -> `return a - b` — a detector that never fires — and scored 0.25 on summarization while inventing behaviour that was not in the source. - `moondream` (1B) described a dense dashboard screenshot as "a spreadsheet ... possibly related to business decisions or financial analysis", reading no title, no column and no value. `qwen3-vl:4b` on the same image read the page title, quoted the subtitle verbatim and recovered the row count. Test with **your real workload** — for vision here that means screenshots, not a synthetic shape. A toy test will pass a model that cannot do the job. ## The classifier is the latency floor Every routed request pays a full local classification round-trip before a single upstream token is requested, so the classifier model is the single biggest lever on interactive latency. ### Three generations, and the one lesson that outlived all of them **The classifier default is `qwen2.5-coder-router:14b`** (see the sizing section above), the second replacement in this project's history. Each swap was measured, not assumed — same prompts, same system prompt, `temperature: 0`, cold load excluded: | | `qwen3.5` (original) | `mistral-nemo:12b` | `qwen2.5-coder:14b` (current) | |---|---|---|---| | category correct | 10/14 | 9/14 or 10/14¹ | 11/14 | | **hard failures** | **4 of 19 calls (21%)** | 1/14 | **0/14** | | latency mean | 6.6s | 1.25s | **1.07s** | | **latency MAX** | **15.6s** | **6.10s** | **1.14s** | ¹ `mistral-nemo`'s own category-correct score reads differently depending on which of the two original comparison runs it is pulled from (9/14 against `qwen3.5`, 10/14 against `qwen2.5-coder`) — flagged rather than silently picked, since this project has been burned before by treating a single run as ground truth. Neither run changes the conclusion: the failure column and the tail decided both swaps, not the category-accuracy column. Category accuracy barely moves across all three. The failure column and the tail are what actually decided each swap, and both keep shrinking: every `qwen3.5` failure and most of `mistral-nemo`'s tail is the same runaway-thinking-trace mode — the model spends its budget on an unbounded chain of thought and returns no parseable JSON, which degrades that request to `source: "fallback"` (tier 2, `general_chat`). End-to-end `/route` went ~10s (`qwen3.5`) → ~1.7s (`mistral-nemo`) → ~1.1s (`qwen2.5-coder`). Consolidating onto one tag also let it serve classification, verification AND local dispatch, so only one model stays resident. **The lesson that outlived every model named above: pick a non-reasoning model, and judge it by the tail, not the mean.** The classifier's entire output is ~45 tokens of JSON — a model that reasons before answering is solving the wrong problem, and a model that is usually fast but occasionally spends 6-15s producing nothing is worse than one that is uniformly mediocre, because that tail sits directly on the interactive latency floor. ### The four settings, and why each exists | setting | value | fixes | |---|---|---| | `max_retries` | `0` | The OpenAI SDK retries twice by default, so `timeout_seconds` silently became a 3x wall-clock bound — a request hung past 250s on a 120s setting and logged nothing. | | `max_output_tokens` | `1024` | Bounds a reasoning model's chain of thought. Unbounded, it cascades: Ollama keeps generating after the client gives up *and* serializes per model, so one runaway request queues every later request behind it. 256 was tried first and was too tight — the trace consumed the whole budget and the model was truncated before emitting any JSON. Does not bind for a non-reasoning model; kept because it costs nothing unused and is the only guard if a reasoning model is swapped back in. | | `max_input_chars` | `8000` | The classifier decides a category and a tier; it does not need the document. Feeding it one is harmful, not wasteful — on a ~20k-token prompt, two different models spent 28.7s and 41.8s respectively without producing a usable answer, both landing on `source: "fallback"`. Clamped to head + tail (the instruction sits at one end or the other, never the middle), the same prompts classify in ~2.2s. Nothing is lost: `chat_completions` measures the real conversation with `estimate_prompt_tokens` and takes the larger value. | | `fallback_tier` / `fallback_category` | mid tier / `general_chat` | A classifier that times out, errors, or returns garbage degrades to a configured mid tier flagged `source: "fallback"` instead of 502/503 — a coding agent would rather have a mid-tier answer than an error. Escalation deliberately skips fallbacks, so an unavailable local model does not silently promote every request to the frontier tier. | | `temperature` | `0` | At the default, the same prompt classified tier 2 then tier 1 on consecutive calls and routed to two different models. | ### `local_encoder`: a different KIND of answer to the same failure mode Every swap above (`qwen3.5` → `mistral-nemo` → `qwen2.5-coder`) picked a *better-behaved generative model* — the runaway-reasoning-trace failure got rarer each time, but never structurally impossible, because a chat-completion model can always in principle spend its budget thinking instead of answering. `classifier.mode: local_encoder` sidesteps the whole failure class instead of picking around it: a zero-shot embedding classifier (`classifier.encoder.model`, default `BAAI/bge-large-en-v1.5`) scores the task directly against `proficiency.categories` and returns a label plus a confidence — there is no generation step, so there is no trace to run away. Originally a cross-encoder NLI pipeline (`facebook/bart-large-mnli`, itself a fallback after `MoritzLaurer/deberta-v3-base-zeroshot-v2` — smaller, ~184M vs ~407M params — started returning 401 on an unauthenticated GET of its own model page sometime after this project picked it, discovered live 2026-09-06 when it took production down: the startup check correctly refused to boot rather than fail opaquely on the first request, but the service still crash-looped until the default was corrected). Replaced with a sentence-embedding + nearest-centroid architecture for Wave 5.1 of `plans/token-waste-waves.md`: NLI cross-encoding scores the input against *every* candidate label separately (one forward pass per category), while an embedding model does one forward pass for the input and compares it by cosine similarity against precomputed category-description embeddings — cost independent of category count, which also affords a stronger backbone (`bge-large-en-v1.5`, ~1.3GB, CPU-viable) on the same compute budget. The trade is real, not free. It only produces `task_category` — no `task_tier` signal exists in a zero-shot label score, so tier falls back to `classifier.fallback_tier` for this mode, a documented limitation rather than a second heuristic bolted on to guess one. And it is **zero-shot, not fine-tuned**, deliberately: this router never stores raw task text anywhere — `docs/operations.md` states it plainly and a test enforces it — so there is no labeled corpus of your own traffic to train against without a new, separate, opt-in capture feature (scoped in `CLAUDE.md`'s "What's NOT built yet", not built). A below-`confidence_min` result is treated as a failure and walks the same cascade a local-LLM parse failure would, unchanged. `confidence_min` is a **minimum similarity score in [0.0, 1.0]**, not a percentage. `classify_zero_shot` returns a similarity score between 0 and 1, so a value of `0.5` means the score must be at least 0.5 to avoid being treated as a failure. The shipped default is `0.5`. The validator in `src/config.py` rejects anything outside [0.0, 1.0] with an error. The knob is *not* a probability — with nearest-centroid (pre-head) scoring it is a softmax-amplified cosine similarity that has no probabilistic meaning, so it is named `confidence_min`, not `confidence_threshold`. Only when the trainable logistic-regression head is present (below) does the returned score carry real `P(correct)` meaning via Platt calibration. The admin UI lets you type 0-100 and scales to the probability before saving, so typing `80` in the UI writes `0.8` to the overlay. The config file and any other direct writer must supply the raw probability. That boundary conversion matters: a raw `80` in the YAML is two orders of magnitude above the maximum and would make every real score read as below threshold, since no probability exceeds 1.0. That validator exists because of a **production near-miss** (issue #47). The admin UI used to save the threshold as written; an `80` was entered, the UI wrote `80.0` straight into the overlay, and it was only caught because the service had not yet restarted and loaded the bad value. Once restarted, classification would have failed on *every* request. The UI now converts 0-100 to 0.0-1.0 before saving, and the validator is the fail-closed last line of defense for every other caller. **Measuring accuracy: the `eval_classifier.py` harness.** The included `eval_classifier.py` harness (run via `PYTHONPATH=src python src/eval_classifier.py`) measures accuracy on the 46-task eval set — `evals/tasks.yaml` minus the 11 `tool_use_agentic` rows the encoder cannot emit — with three noise variants per prompt (clean, short-noise-wrap, long-noise-wrap) and a confusion-matrix output showing per-category precision/recall and the `blended_score` gap per cell. `--dry-run` plans the run without loading the model. This is what makes further accuracy regression measurable instead of a vibe check, and it stays standalone (`requirements-encoder.txt`) so the router's dispatch path never imports transformers/torch. **A fitted trainable head.** A fitted logistic-regression head (`_TrainableHead`) replaces nearest-centroid when `evals/synthetic/encoder-head-coefficients.json` is present, improving accuracy and providing a calibrated confidence score. Trained by `scripts/train_encoder_head.py` on the committed synthetic corpus (`evals/synthetic/encoder-training.jsonl`), it runs a convex `LogisticRegression(solver="lbfgs")` over the same frozen embeddings with per-class Platt (sigmoid) calibration, so `predict_proba` genuinely means `P(correct)` — the confidence gate reads as a real abstention signal rather than a monotonic score. The artifact is keyed to a specific backbone and candidate set; if the file is absent or the model/categories don't match, it degrades to the unchanged nearest-centroid path with a logged warning. **Candidate labels are natural-language descriptions, never the raw config identifier.** HF's zero-shot pipeline scores a candidate label against the input via a hypothesis template (`"This example is {}."`), so the label itself has to read as a sentence for entailment scoring to work — feeding it `tool_use_agentic` or `diff_checking` verbatim asks the model to judge `"This example is tool_use_agentic."`, not a sentence its NLI training ever saw. `local_encoder.py`'s `_CATEGORY_DESCRIPTIONS` maps each category to a real description before scoring, and `multi_label=True` scores each candidate independently rather than normalizing every score to sum to 1 (the pipeline's single-label default, which drags a genuinely good match down whenever another category is also plausible). Measured live 2026-09-06 on the same 9-prompt set: raw identifiers scored 5/9 correct at confidence 0.15-0.90; natural-language descriptions scored 8/9 correct at 0.39-0.999. A category missing from the map falls back to its own raw string rather than raising, so a newly-added config category degrades gracefully instead of crashing. `local_encoder` runs on **cpu by default** in the shipped example; setting it to `cuda` requires a `config.local.yaml` overlay and a host that actually has the hardware. The project does not ship a settled default device. Match the device to the machine the router is running on, otherwise `torch` will fail fast with a clear error at first classification. CPU inference is the safe, portable starting point; CUDA is host-dependent and should be chosen only where the router process has a card available. `transformers`/`torch` are optional (`requirements-encoder.txt`, not the main `requirements.txt` — pinned-dependency discipline applies here too) and imported lazily inside `local_encoder.py`, the same rule `tui.py` follows for `textual`: a deployment that never selects this mode needs neither installed. CPU install: ```bash pip install -r requirements-encoder.txt ``` ### `local_decision`: read the first-token logprobs, never parse a generation `classifier.mode: local_decision` is a third answer to the same runaway-reasoning problem, and the most direct one. Instead of asking a small generative model to emit JSON and hoping it stops, it asks for a single lettered choice and reads the decision out of the **logprobs on the first output token**. Mechanism, a Jev-style first-token-logprob classifier: - The task text and a lettered list of candidate categories go into one prompt; the model is asked to answer with one letter. - Ollama returns `logprobs` for that first position. `local_decision.parse_logprobs` sums `exp(logprob)` per option letter (A through J), ignoring non-option tokens and Ollama's `.` separator tokens, then normalizes the accumulated mass into a confidence. The letter with the most mass is the predicted category. - There is **no JSON parse, no reasoning trace, and no multi-token generation to run away**. `num_predict` is 1 and `think` is false, so a model that would normally spend its budget thinking has nowhere to spend it. The whole failure class that took `qwen3.5` down on 4 of 19 calls simply does not exist here. - Confidence is the chosen letter's share of total option mass. `coverage_min` (default `0.3`) is the minimum total option mass for a call to be accepted at all; below it, or when no logprobs come back, the call raises and walks the same cascade a `local_llm` parse failure would. `confidence_min` (default `0.5`) treats a below-threshold verdict as a failure rather than a low-confidence answer, mirroring the encoder's semantics. - As with `local_encoder`, it produces only `task_category` unless `classifier.decision.tier_enabled` is set, in which case it may also choose `task_tier`; otherwise tier falls back to `classifier.fallback_tier`. - A declined answer is recorded, not just logged: `route_decisions.classifier_confidence`, `classifier_coverage` and `classifier_reject` hold the numbers and the reason (see [data-model](data-model.md#what-the-classifier-did-and-why-its-answer-was-not-used)). On the 4b the decline is by design, not an outage: it ran at 30% of turns on 2026-10-04 (8% the day before), concentrated in long agent sessions. Pull and enable it: ```bash ollama pull qwen3.5:4b ``` ```yaml # config.local.yaml classifier: mode: "local_decision" decision: {} ``` Every field in `classifier.decision` has a default, so `decision: {}` is enough to opt in. The one that matters for hardware is **`classifier.decision.num_ctx`**: it sets the context window the classifier loads at (default `8192`) and scales the KV cache directly, so it is the cheapest GB to reclaim when the pair below does not fit. Model, `base_url`, timeout, and the confidence gates all live under the same block. VRAM, and why this is tight on a 24 GB card: | model | role | num_ctx | resident | |---|---|---|---| | `qwen2.5-coder-router:14b` | classify + verify + dispatch | 16384 | 12.3 GB | | `qwen3.5:4b` | `local_decision` | 8192 | ~7 GB | | | | **total** | **~19.9 GB of 24 GB** | The pair fits, with about 4 GB spare, but **only because the 14b router runs at `num_ctx 16384`, not the 32768 it was originally tagged with**. At 32k the router alone measured 16.5 to 17.8 GB resident, and the two models then fit by about 1 GB or get evicted outright depending on load order. Cutting the router to 16k is what makes the coexistence robust; that is a change to the `qwen2.5-coder-router:14b` Modelfile, not to this config block. Local vision is off in this profile because its 5.6 GB model no longer fits alongside both. ### The latency tax, and what has and hasn't addressed it ~10s of local overhead per message was the original number on a reasoning classifier; the current `qwen2.5-coder-router:14b` brings the round-trip itself down to ~1.1s (see the table above), so the *per-call* cost of that tax is mostly gone. What is still unaddressed is the *per-message* repetition of it: `session_cache` (off by default, `config/config.yaml` -> `session_cache.enabled`) now exists specifically to classify once per session rather than once per message, but nobody has measured its effect on a real interactive session yet. Caching by prompt hash and skipping classification for short prompts remain unbuilt. `classifier.mode: local_encoder` above addresses a different cost entirely (which model answers) and does not touch this one (how often it is asked). Separately, a **cold reload** is its own latency spike — ~6s when loading a second model evicts the resident one (see the KV-cache section above) — and is not the same phenomenon as the per-message round-trip tax: cold load happens once per eviction, the round-trip tax happens on every single message even with everything warm. ### "Local" just means reachable, not necessarily on this machine "Local" here means the model runs on hardware **you** control and don't pay a cloud provider for — not that it has to share a motherboard with the process asking it questions. All three of these count as "local" for this router's purposes, and the config looks identical for each: - **localhost** — Ollama and the router on the same machine. - **local network** — Ollama on another box on your LAN. - **a VPN tunnel** — Ollama on a home-lab workstation, router and editor on a laptop, connected over WireGuard or similar. This is the common shape for anyone whose GPU isn't in their daily-driver machine. `classifier.base_url` / `api_key_env` / `model` take any OpenAI-compatible endpoint, and switching between the three above is just changing that URL — verified end to end against a non-loopback address, classifier and verifier both. Ollama binds `127.0.0.1` by default, so the VPN case fails with connection-refused until the serving host applies `deploy/ollama-over-vpn.conf`. Bind it to the VPN address, not `0.0.0.0`: Ollama has no auth of any kind, so anything that can reach the port can run inference and enumerate your models, and `0.0.0.0` publishes it on whatever network the serving machine happens to be sitting on. #### Local vs cloud classifier, measured (historical — against the superseded `qwen3.5`) **This comparison predates the swap to `qwen2.5-coder-router:14b` above** and has not been re-run against the current classifier. It measured the *original* `qwen3.5` classifier against a cloud one, five prompts, same system prompt, `temperature: 0`: | | local `qwen3.5` (RTX 6000, superseded) | NeuralWatt `deepseek-v4-flash` | |---|---|---| | mean latency | 11.58s (4.95-15.76) | 1.02s | | categories agreed with the label | 2 of 4 | 5 of 5 | | hard failures | 1 of 5 (empty after 15.76s) | 0 | | energy per call | ~7e-05 kWh, on your meter | 1.17e-05 kWh attributed | | cost per 1,000 calls | electricity + 6.6GB resident | $0.093 (0.19% of quota) | At the time, the cloud side was faster, more accurate, and less attributed energy — the one hard failure is the same runaway-thinking-trace mode described above. The general lesson still holds regardless of which local model is current: local classification is not free, it is just unbilled, and that is a substitution the cost axis had to unlearn independently (see `CLAUDE.md`'s billing section). The shipped default is still local Ollama, because switching spends quota and that is a deployment choice, not a code one. **Worth re-running against `qwen2.5-coder-router:14b`** before treating the 11x figure as current — the classifier's own latency dropped roughly 10x since this table was measured (11.58s mean here vs ~1.1s in the three-way table above), which would substantially narrow or erase the local/cloud gap this table reports. Flagging this rather than re-deriving a number with no live hardware to measure against. ### The verifier follows, but only to another Ollama The local LLM check speaks Ollama's **native** `/api/chat` (the only way to set `think: False`), so it follows the classifier across a VPN but not to a cloud provider. It used to derive its URL by stripping `/v1` off `classifier.base_url`, which meant moving the classifier at all would have pointed it at `/api/chat`. It now has its own `verification.base_url` and `verification.model`. `verification.model` may be null only while both run on one host. Config load **refuses** the null once the hostnames differ, because the failure is silent: observed directly, with the classifier on NeuralWatt the verifier POSTed `deepseek-v4-flash` to `localhost:11434`, 404'd, caught it, logged "local verification unavailable" and recorded no sample. Verification would have looked enabled while producing nothing. Structural verification needs no model at all — it is pure Python — so `verification.local_llm_enabled: false` leaves a host with no local inference fully functional, minus the refusal/incoherence class of failure. ### A note on prompt-token estimation opencode sends ~32K prompt tokens of system prompt and tool definitions on a trivial request, so the measured-size floor in `estimate_prompt_tokens` does real work — the classifier's own estimate for that request was two orders of magnitude low. This is the same phenomenon pinch confronts on the provider side; see [docs/pinch.md](pinch.md).