A declined answer left one number behind and it was in a log line, so confidence_min could not be tuned from data. route_decisions.confidence cannot say it either: the chat path re-routes through the override branch, which hard-codes 1.0, and 43,804 of 43,856 live rows hold exactly that. Measured on 2026-10-04 under local_decision (qwen3.5:4b): 769 of 2,557 turns (30%) had no fresh classification, against 0% for local_llm and local_encoder. The cause was recorded in only 4 of them. - classifier_confidence, classifier_coverage and classifier_reject on route_decisions, filled from a ClassifierAttempt carried on Classification. Both accepted and declined answers carry one, so the two distributions can be compared around the floor. Reason codes are listed in docs/data-model.md. - ClassifierRejected (a RuntimeError subclass, messages unchanged) replaces the plain RuntimeErrors at the six floor-miss sites and the two local_decision refusals, so the number and reason travel out of the raise site. - The chat path captures the classifier's verdict before the re-route and passes it to persist_route_decision (attempt_of), like it already does for source. - Admin decisions page: the source badge's tooltip shows the attempt, and session_history / session_stale get an amber badge instead of neutral grey. TUI detail popup and the live event carry the same three keys. - The /metrics degraded-share warning lists the recorded reasons instead of claiming the classifier "has been failing", which was wrong for a classifier that answers and is declined. - degraded_warn_threshold must be in (0, 1] and degraded_warn_min at least 1, refused at load: a value above 1 could never fire. The three columns arrive by ALTER and are NULL on every earlier row; metrics selects them only when present, so the live DB reads NULL until its restart. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KkCGRantZsSwmcFpet6FTa
506 lines
28 KiB
Markdown
506 lines
28 KiB
Markdown
> Sizing local models for your GPU. Back to [README](../README.md).
|
|
|
|
## Fitting local models into 24GB VRAM
|
|
|
|
These defaults are measured on a 24GB Quadro RTX 6000 (Turing, compute 7.5)
|
|
and are a sound starting point for any 24-32GB card. They are **defaults, not
|
|
prescriptions** — see [Swapping in your own model](#swapping-in-your-own-model)
|
|
below, which is the part that matters if you run different hardware or simply
|
|
want a different model.
|
|
|
|
Ollama loads a model at the context size baked into its tag. You cannot set it
|
|
per request: the OpenAI-compatible `/v1/chat/completions` endpoint in Ollama
|
|
0.22.0 was verified live against every plausible field shape, and in every case
|
|
returned 200 while silently keeping the loaded context unchanged. So context
|
|
size has to go into the tag via a `Modelfile`:
|
|
|
|
```bash
|
|
# Classification + verification + local dispatch (ONE resident model)
|
|
printf 'FROM qwen2.5-coder:14b\nPARAMETER num_ctx 32768\n' > Modelfile.local-dispatch
|
|
ollama create qwen2.5-coder-router:14b -f Modelfile.local-dispatch
|
|
|
|
# Vision fallback only
|
|
printf 'FROM qwen3-vl:4b\nPARAMETER num_ctx 8192\n' > Modelfile.vision
|
|
ollama create qwen3-vl-router:4b -f Modelfile.vision
|
|
```
|
|
|
|
Measured co-residency, both fully on GPU:
|
|
|
|
| model | role | resident |
|
|
|---|---|---|
|
|
| `qwen2.5-coder-router:14b` (Q4_K_M, 32k) | classify + verify + dispatch | 17.0GB |
|
|
| `qwen3-vl-router:4b` (8k) | vision fallback | 5.6GB |
|
|
| | **total** | **23.2GB of 24GB** |
|
|
|
|
That leaves ~800MB spare, which is thin but works. The single most important
|
|
thing to understand is **where the memory actually goes**.
|
|
|
|
### KV cache is the cheapest GB to reclaim
|
|
|
|
Weights are fixed; KV cache scales with `num_ctx` and is often the larger
|
|
lever. Measured on this hardware:
|
|
|
|
| model | on disk | resident @ num_ctx |
|
|
|---|---|---|
|
|
| `qwen2.5-coder:14b` Q4_K_M | 9.0GB | 13GB @ 16k, **17GB @ 32k** |
|
|
| `qwen3-vl:4b` | 3.3GB | 5.6GB @ 8k, **6.9GB @ 16k** |
|
|
|
|
Halving the vision model's context bought 1.3GB — exactly what lets the
|
|
dispatch model run at 32k with vision still resident. Without it, loading
|
|
vision **evicts the dispatch model entirely**, and the next classification
|
|
pays a ~6s cold reload on the interactive latency floor.
|
|
|
|
If you are short on VRAM, cut `num_ctx` before you cut anything else. The
|
|
classifier in particular never needs a large window: `classifier.max_input_chars`
|
|
clamps input to 8000 chars (~2.7k tokens), so even 8192 leaves 2x headroom.
|
|
|
|
### Bigger quantization is usually the wrong trade
|
|
|
|
Tempting and mostly counterproductive. Measured, same model, same tasks:
|
|
|
|
| variant | resident | placement | classify accuracy | classify mean |
|
|
|---|---|---|---|---|
|
|
| Q4_K_M @ 16k | 13GB | 100% GPU | 11/14 | **1.07s** |
|
|
| Q4_K_M @ 32k | 17GB | 100% GPU | 11/14 | 1.14s |
|
|
| Q8_0 @ 16k | 19GB | 100% GPU | 11/14 | 1.62s |
|
|
| Q8_0 @ 32k | 24GB | **83% GPU** | 11/14 | 3.79s |
|
|
|
|
Accuracy never moved — only speed did. Q8 costs ~1.5x latency even when it
|
|
fits (Turing is memory-bandwidth bound here, and Q8 moves twice the weight
|
|
bytes per token), and `83% GPU` in the placement column means Ollama spilled
|
|
the rest to system RAM for a 3.3x latency hit — a config that "fits" on paper
|
|
can silently land there. Going up in *parameters* beat going up in *bits*: a
|
|
32B model at Q4 scored higher on quality than this 14B at Q8, though it needs
|
|
22GB and evicts the vision model entirely. Watch `ollama ps`'s placement
|
|
column, not just the resident-GB math.
|
|
|
|
## Swapping in your own model
|
|
|
|
Nothing here is load-bearing on the specific models. Swap freely — the defaults
|
|
are a starting point, and an abliterated or fine-tuned model of your choosing is
|
|
a perfectly legitimate replacement. What follows is the procedure so you can
|
|
make the change with numbers instead of vibes.
|
|
|
|
**1. Build a tagged variant.** Always tag with an explicit `num_ctx`; the
|
|
library default is usually far larger than you want and costs GB of KV cache.
|
|
|
|
```bash
|
|
printf 'FROM <your-model>\nPARAMETER num_ctx 32768\n' > Modelfile.mine
|
|
ollama create my-router-model:tag -f Modelfile.mine
|
|
```
|
|
|
|
**2. Point the config at it.** For classification and verification, set both —
|
|
they must name the **same tag**, or Ollama reloads the base model at a
|
|
different context every time a request alternates between the two:
|
|
|
|
```yaml
|
|
classifier: { model: "my-router-model:tag" }
|
|
verification: { model: "my-router-model:tag" }
|
|
```
|
|
|
|
For local dispatch you also need three things that are easy to miss:
|
|
|
|
- an entry under `local_dispatch_models` (with `context_window` matching the
|
|
tag's `num_ctx`),
|
|
- the categories it may serve in its `eligible_categories`, which must also
|
|
exist in `proficiency.categories`,
|
|
- **a pin under `tiering.model_tiers` keyed by the new model id** — `src/tier.py`
|
|
overwrites `models.tier` on every run, so without the pin the model silently
|
|
loses its tier and drops out of routing.
|
|
|
|
**3. Measure it, in this order.** Cheap and decisive first:
|
|
|
|
- **VRAM and placement**: load it, then `ollama ps`. Confirm `100%` — anything
|
|
less means a CPU spill and a large latency penalty.
|
|
- **Classification** (free, local): accuracy, **hard failures**, and the
|
|
latency **tail**. The tail is what has disqualified two classifiers in this
|
|
project's history, not the mean; a model that is usually fast but sometimes
|
|
spends 6s producing no parseable JSON degrades that request to
|
|
`source: "fallback"`.
|
|
- **Quality on the categories it will serve**:
|
|
`PYTHONPATH=src python -m eval_proficiency --models <id> --categories <cats>`.
|
|
Always bench a **cloud model on the same tasks** as a control — without it you
|
|
cannot tell "this model is weak" from "these tasks are unfair", and this
|
|
project has been wrong about that more than once.
|
|
- **Cost**: `PYTHONPATH=src python -m seed_local_dispatch_energy --samples 5`,
|
|
which measures real GPU draw and prices it at your `local_energy` tariff.
|
|
|
|
**4. Beware the sample size.** The bundled task sets are small —
|
|
`file_summarization` is 6 tasks (0.167 per task) and `diff_checking` is 8 tasks
|
|
over only 4 scenarios (0.125 per task). Differences under ~0.15 are noise. At
|
|
`temperature: 0` re-running is byte-identical, so more confidence means more
|
|
*tasks*, not more runs.
|
|
|
|
**5. Do not widen `eligible_categories` without measuring the new category.**
|
|
The hard filter is what keeps a summarization-grade model away from code. A
|
|
confidently wrong answer is worse than an honest error.
|
|
|
|
### A caution about small models
|
|
|
|
Small does not degrade gracefully; it fabricates. Two measured examples:
|
|
|
|
- `nemotron-mini:4b` answered "no" to **8/8** diff-checking tasks and to 3/3
|
|
blatant probes including `return a + b` -> `return a - b` — a detector that
|
|
never fires — and scored 0.25 on summarization while inventing behaviour that
|
|
was not in the source.
|
|
- `moondream` (1B) described a dense dashboard screenshot as "a spreadsheet
|
|
... possibly related to business decisions or financial analysis", reading no
|
|
title, no column and no value. `qwen3-vl:4b` on the same image read the page
|
|
title, quoted the subtitle verbatim and recovered the row count.
|
|
|
|
Test with **your real workload** — for vision here that means screenshots, not
|
|
a synthetic shape. A toy test will pass a model that cannot do the job.
|
|
|
|
## The classifier is the latency floor
|
|
|
|
Every routed request pays a full local classification round-trip before a
|
|
single upstream token is requested, so the classifier model is the single
|
|
biggest lever on interactive latency.
|
|
|
|
### Three generations, and the one lesson that outlived all of them
|
|
|
|
**The classifier default is `qwen2.5-coder-router:14b`** (see the sizing
|
|
section above), the second replacement in this project's history. Each swap
|
|
was measured, not assumed — same prompts, same system prompt,
|
|
`temperature: 0`, cold load excluded:
|
|
|
|
| | `qwen3.5` (original) | `mistral-nemo:12b` | `qwen2.5-coder:14b` (current) |
|
|
|---|---|---|---|
|
|
| category correct | 10/14 | 9/14 or 10/14¹ | 11/14 |
|
|
| **hard failures** | **4 of 19 calls (21%)** | 1/14 | **0/14** |
|
|
| latency mean | 6.6s | 1.25s | **1.07s** |
|
|
| **latency MAX** | **15.6s** | **6.10s** | **1.14s** |
|
|
|
|
¹ `mistral-nemo`'s own category-correct score reads differently depending on
|
|
which of the two original comparison runs it is pulled from (9/14 against
|
|
`qwen3.5`, 10/14 against `qwen2.5-coder`) — flagged rather than silently
|
|
picked, since this project has been burned before by treating a single run
|
|
as ground truth. Neither run changes the conclusion: the failure column and
|
|
the tail decided both swaps, not the category-accuracy column.
|
|
|
|
Category accuracy barely moves across all three. The failure column and the
|
|
tail are what actually decided each swap, and both keep shrinking: every
|
|
`qwen3.5` failure and most of `mistral-nemo`'s tail is the same
|
|
runaway-thinking-trace mode — the model spends its budget on an unbounded
|
|
chain of thought and returns no parseable JSON, which degrades that request
|
|
to `source: "fallback"` (tier 2, `general_chat`). End-to-end `/route` went
|
|
~10s (`qwen3.5`) → ~1.7s (`mistral-nemo`) → ~1.1s (`qwen2.5-coder`).
|
|
Consolidating onto one tag also let it serve classification, verification
|
|
AND local dispatch, so only one model stays resident.
|
|
|
|
**The lesson that outlived every model named above: pick a non-reasoning
|
|
model, and judge it by the tail, not the mean.** The classifier's entire
|
|
output is ~45 tokens of JSON — a model that reasons before answering is
|
|
solving the wrong problem, and a model that is usually fast but occasionally
|
|
spends 6-15s producing nothing is worse than one that is uniformly mediocre,
|
|
because that tail sits directly on the interactive latency floor.
|
|
|
|
### The four settings, and why each exists
|
|
|
|
| setting | value | fixes |
|
|
|---|---|---|
|
|
| `max_retries` | `0` | The OpenAI SDK retries twice by default, so `timeout_seconds` silently became a 3x wall-clock bound — a request hung past 250s on a 120s setting and logged nothing. |
|
|
| `max_output_tokens` | `1024` | Bounds a reasoning model's chain of thought. Unbounded, it cascades: Ollama keeps generating after the client gives up *and* serializes per model, so one runaway request queues every later request behind it. 256 was tried first and was too tight — the trace consumed the whole budget and the model was truncated before emitting any JSON. Does not bind for a non-reasoning model; kept because it costs nothing unused and is the only guard if a reasoning model is swapped back in. |
|
|
| `max_input_chars` | `8000` | The classifier decides a category and a tier; it does not need the document. Feeding it one is harmful, not wasteful — on a ~20k-token prompt, two different models spent 28.7s and 41.8s respectively without producing a usable answer, both landing on `source: "fallback"`. Clamped to head + tail (the instruction sits at one end or the other, never the middle), the same prompts classify in ~2.2s. Nothing is lost: `chat_completions` measures the real conversation with `estimate_prompt_tokens` and takes the larger value. |
|
|
| `fallback_tier` / `fallback_category` | mid tier / `general_chat` | A classifier that times out, errors, or returns garbage degrades to a configured mid tier flagged `source: "fallback"` instead of 502/503 — a coding agent would rather have a mid-tier answer than an error. Escalation deliberately skips fallbacks, so an unavailable local model does not silently promote every request to the frontier tier. |
|
|
| `temperature` | `0` | At the default, the same prompt classified tier 2 then tier 1 on consecutive calls and routed to two different models. |
|
|
|
|
### `local_encoder`: a different KIND of answer to the same failure mode
|
|
|
|
Every swap above (`qwen3.5` → `mistral-nemo` → `qwen2.5-coder`) picked a
|
|
*better-behaved generative model* — the runaway-reasoning-trace failure got
|
|
rarer each time, but never structurally impossible, because a chat-completion
|
|
model can always in principle spend its budget thinking instead of
|
|
answering. `classifier.mode: local_encoder` sidesteps the whole failure class
|
|
instead of picking around it: a zero-shot embedding classifier
|
|
(`classifier.encoder.model`, default `BAAI/bge-large-en-v1.5`) scores the
|
|
task directly against `proficiency.categories` and returns a label plus a
|
|
confidence — there is no generation step, so there is no trace to run away.
|
|
|
|
Originally a cross-encoder NLI pipeline (`facebook/bart-large-mnli`, itself a
|
|
fallback after `MoritzLaurer/deberta-v3-base-zeroshot-v2` — smaller, ~184M vs
|
|
~407M params — started returning 401 on an unauthenticated GET of its own
|
|
model page sometime after this project picked it, discovered live
|
|
2026-09-06 when it took production down: the startup check correctly
|
|
refused to boot rather than fail opaquely on the first request, but the
|
|
service still crash-looped until the default was corrected). Replaced with a
|
|
sentence-embedding + nearest-centroid architecture for Wave 5.1 of
|
|
`plans/token-waste-waves.md`: NLI cross-encoding scores the input against
|
|
*every* candidate label separately (one forward pass per category), while an
|
|
embedding model does one forward pass for the input and compares it by cosine
|
|
similarity against precomputed category-description embeddings — cost
|
|
independent of category count, which also affords a stronger backbone
|
|
(`bge-large-en-v1.5`, ~1.3GB, CPU-viable) on the same compute budget.
|
|
|
|
The trade is real, not free. It only produces `task_category` — no `task_tier`
|
|
signal exists in a zero-shot label score, so tier falls back to
|
|
`classifier.fallback_tier` for this mode, a documented limitation rather than
|
|
a second heuristic bolted on to guess one. And it is **zero-shot, not
|
|
fine-tuned**, deliberately: this router never stores raw task text anywhere —
|
|
`docs/operations.md` states it plainly and a test enforces it — so there is no
|
|
labeled corpus of your own traffic to train against without a new, separate,
|
|
opt-in capture feature (scoped in `CLAUDE.md`'s "What's NOT built yet", not
|
|
built). A below-`confidence_min` result is treated as a failure and
|
|
walks the same cascade a local-LLM parse failure would, unchanged.
|
|
|
|
`confidence_min` is a **minimum similarity score in [0.0, 1.0]**, not a
|
|
percentage. `classify_zero_shot` returns a similarity score between 0 and 1,
|
|
so a value of `0.5` means the score must be at least 0.5 to avoid being
|
|
treated as a failure. The shipped default is `0.5`. The validator in
|
|
`src/config.py` rejects anything outside [0.0, 1.0] with an error. The knob
|
|
is *not* a probability — with nearest-centroid (pre-head) scoring it is a
|
|
softmax-amplified cosine similarity that has no probabilistic meaning, so it
|
|
is named `confidence_min`, not `confidence_threshold`. Only when the
|
|
trainable logistic-regression head is present (below) does the returned score
|
|
carry real `P(correct)` meaning via Platt calibration.
|
|
|
|
The admin UI lets you type 0-100 and scales to the probability before saving,
|
|
so typing `80` in the UI writes `0.8` to the overlay. The config file and any
|
|
other direct writer must supply the raw probability. That boundary conversion
|
|
matters: a raw `80` in the YAML is two orders of magnitude above the maximum
|
|
and would make every real score read as below threshold, since no probability
|
|
exceeds 1.0.
|
|
|
|
That validator exists because of a **production near-miss** (issue #47). The
|
|
admin UI used to save the threshold as written; an `80` was entered, the UI
|
|
wrote `80.0` straight into the overlay, and it was only caught because the
|
|
service had not yet restarted and loaded the bad value. Once restarted,
|
|
classification would have failed on *every* request. The UI now converts
|
|
0-100 to 0.0-1.0 before saving, and the validator is the fail-closed last line
|
|
of defense for every other caller.
|
|
|
|
**Measuring accuracy: the `eval_classifier.py` harness.** The included
|
|
`eval_classifier.py` harness (run via `PYTHONPATH=src python
|
|
src/eval_classifier.py`) measures accuracy on the 46-task eval set —
|
|
`evals/tasks.yaml` minus the 11 `tool_use_agentic` rows the encoder cannot
|
|
emit — with three noise variants per prompt (clean, short-noise-wrap,
|
|
long-noise-wrap) and a confusion-matrix output showing per-category
|
|
precision/recall and the `blended_score` gap per cell. `--dry-run` plans
|
|
the run without loading the model. This is what makes further accuracy
|
|
regression measurable instead of a vibe check, and it stays standalone
|
|
(`requirements-encoder.txt`) so the router's dispatch path never imports
|
|
transformers/torch.
|
|
|
|
**A fitted trainable head.** A fitted logistic-regression head
|
|
(`_TrainableHead`) replaces nearest-centroid when
|
|
`evals/synthetic/encoder-head-coefficients.json` is present, improving
|
|
accuracy and providing a calibrated confidence score. Trained by
|
|
`scripts/train_encoder_head.py` on the committed synthetic corpus
|
|
(`evals/synthetic/encoder-training.jsonl`), it runs a convex
|
|
`LogisticRegression(solver="lbfgs")` over the same frozen embeddings with
|
|
per-class Platt (sigmoid) calibration, so `predict_proba` genuinely means
|
|
`P(correct)` — the confidence gate reads as a real abstention signal rather
|
|
than a monotonic score. The artifact is keyed to a specific backbone and
|
|
candidate set; if the file is absent or the model/categories don't match, it
|
|
degrades to the unchanged nearest-centroid path with a logged warning.
|
|
|
|
**Candidate labels are natural-language descriptions, never the raw config
|
|
identifier.** HF's zero-shot pipeline scores a candidate label against the
|
|
input via a hypothesis template (`"This example is {}."`), so the label
|
|
itself has to read as a sentence for entailment scoring to work — feeding it
|
|
`tool_use_agentic` or `diff_checking` verbatim asks the model to judge
|
|
`"This example is tool_use_agentic."`, not a sentence its NLI training ever
|
|
saw. `local_encoder.py`'s `_CATEGORY_DESCRIPTIONS` maps each category to a
|
|
real description before scoring, and `multi_label=True` scores each
|
|
candidate independently rather than normalizing every score to sum to 1 (the
|
|
pipeline's single-label default, which drags a genuinely good match down
|
|
whenever another category is also plausible). Measured live 2026-09-06 on the
|
|
same 9-prompt set: raw identifiers scored 5/9 correct at confidence
|
|
0.15-0.90; natural-language descriptions scored 8/9 correct at 0.39-0.999. A
|
|
category missing from the map falls back to its own raw string rather than
|
|
raising, so a newly-added config category degrades gracefully instead of
|
|
crashing.
|
|
|
|
`local_encoder` runs on **cpu by default** in the shipped example; setting it
|
|
to `cuda` requires a `config.local.yaml` overlay and a host that actually has
|
|
the hardware. The project does not ship a settled default device. Match the
|
|
device to the machine the router is running on, otherwise `torch` will fail
|
|
fast with a clear error at first classification. CPU inference is the safe,
|
|
portable starting point; CUDA is host-dependent and should be chosen only
|
|
where the router process has a card available.
|
|
|
|
`transformers`/`torch` are optional (`requirements-encoder.txt`, not the main
|
|
`requirements.txt` — pinned-dependency discipline applies here too) and
|
|
imported lazily inside `local_encoder.py`, the same rule `tui.py` follows for
|
|
`textual`: a deployment that never selects this mode needs neither installed.
|
|
CPU install:
|
|
|
|
```bash
|
|
pip install -r requirements-encoder.txt
|
|
```
|
|
|
|
### `local_decision`: read the first-token logprobs, never parse a generation
|
|
|
|
`classifier.mode: local_decision` is a third answer to the same
|
|
runaway-reasoning problem, and the most direct one. Instead of asking a small
|
|
generative model to emit JSON and hoping it stops, it asks for a single
|
|
lettered choice and reads the decision out of the **logprobs on the first
|
|
output token**.
|
|
|
|
Mechanism, a Jev-style first-token-logprob classifier:
|
|
|
|
- The task text and a lettered list of candidate categories go into one prompt;
|
|
the model is asked to answer with one letter.
|
|
- Ollama returns `logprobs` for that first position. `local_decision.parse_logprobs`
|
|
sums `exp(logprob)` per option letter (A through J), ignoring non-option
|
|
tokens and Ollama's `.` separator tokens, then normalizes the accumulated mass
|
|
into a confidence. The letter with the most mass is the predicted category.
|
|
- There is **no JSON parse, no reasoning trace, and no multi-token generation
|
|
to run away**. `num_predict` is 1 and `think` is false, so a model that would
|
|
normally spend its budget thinking has nowhere to spend it. The whole failure
|
|
class that took `qwen3.5` down on 4 of 19 calls simply does not exist here.
|
|
- Confidence is the chosen letter's share of total option mass. `coverage_min`
|
|
(default `0.3`) is the minimum total option mass for a call to be accepted at
|
|
all; below it, or when no logprobs come back, the call raises and walks the
|
|
same cascade a `local_llm` parse failure would. `confidence_min` (default
|
|
`0.5`) treats a below-threshold verdict as a failure rather than a
|
|
low-confidence answer, mirroring the encoder's semantics.
|
|
- As with `local_encoder`, it produces only `task_category` unless
|
|
`classifier.decision.tier_enabled` is set, in which case it may also choose
|
|
`task_tier`; otherwise tier falls back to `classifier.fallback_tier`.
|
|
- A declined answer is recorded, not just logged: `route_decisions.classifier_confidence`,
|
|
`classifier_coverage` and `classifier_reject` hold the numbers and the reason
|
|
(see [data-model](data-model.md#what-the-classifier-did-and-why-its-answer-was-not-used)).
|
|
On the 4b the decline is by design, not an outage: it ran at 30% of turns on
|
|
2026-10-04 (8% the day before), concentrated in long agent sessions.
|
|
|
|
Pull and enable it:
|
|
|
|
```bash
|
|
ollama pull qwen3.5:4b
|
|
```
|
|
|
|
```yaml
|
|
# config.local.yaml
|
|
classifier:
|
|
mode: "local_decision"
|
|
decision: {}
|
|
```
|
|
|
|
Every field in `classifier.decision` has a default, so `decision: {}` is enough
|
|
to opt in. The one that matters for hardware is **`classifier.decision.num_ctx`**:
|
|
it sets the context window the classifier loads at (default `8192`) and scales
|
|
the KV cache directly, so it is the cheapest GB to reclaim when the pair below
|
|
does not fit. Model, `base_url`, timeout, and the confidence gates all live
|
|
under the same block.
|
|
|
|
VRAM, and why this is tight on a 24 GB card:
|
|
|
|
| model | role | num_ctx | resident |
|
|
|---|---|---|---|
|
|
| `qwen2.5-coder-router:14b` | classify + verify + dispatch | 16384 | 12.3 GB |
|
|
| `qwen3.5:4b` | `local_decision` | 8192 | ~7 GB |
|
|
| | | **total** | **~19.9 GB of 24 GB** |
|
|
|
|
The pair fits, with about 4 GB spare, but **only because the 14b router runs at
|
|
`num_ctx 16384`, not the 32768 it was originally tagged with**. At 32k the
|
|
router alone measured 16.5 to 17.8 GB resident, and the two models then fit by
|
|
about 1 GB or get evicted outright depending on load order. Cutting the router
|
|
to 16k is what makes the coexistence robust; that is a change to the
|
|
`qwen2.5-coder-router:14b` Modelfile, not to this config block. Local vision is
|
|
off in this profile because its 5.6 GB model no longer fits alongside both.
|
|
|
|
### The latency tax, and what has and hasn't addressed it
|
|
|
|
~10s of local overhead per message was the original number on a reasoning
|
|
classifier; the current `qwen2.5-coder-router:14b` brings the round-trip
|
|
itself down to ~1.1s (see the table above), so the *per-call* cost of that
|
|
tax is mostly gone. What is still unaddressed is the *per-message* repetition
|
|
of it: `session_cache` (off by default, `config/config.yaml` ->
|
|
`session_cache.enabled`) now exists specifically to classify once per
|
|
session rather than once per message, but nobody has measured its effect on
|
|
a real interactive session yet. Caching by prompt hash and skipping
|
|
classification for short prompts remain unbuilt. `classifier.mode:
|
|
local_encoder` above addresses a different cost entirely (which model
|
|
answers) and does not touch this one (how often it is asked).
|
|
|
|
Separately, a **cold reload** is its own latency spike — ~6s when loading a
|
|
second model evicts the resident one (see the KV-cache section above) — and
|
|
is not the same phenomenon as the per-message round-trip tax: cold load
|
|
happens once per eviction, the round-trip tax happens on every single
|
|
message even with everything warm.
|
|
|
|
### "Local" just means reachable, not necessarily on this machine
|
|
|
|
"Local" here means the model runs on hardware **you** control and don't pay a
|
|
cloud provider for — not that it has to share a motherboard with the process
|
|
asking it questions. All three of these count as "local" for this router's
|
|
purposes, and the config looks identical for each:
|
|
|
|
- **localhost** — Ollama and the router on the same machine.
|
|
- **local network** — Ollama on another box on your LAN.
|
|
- **a VPN tunnel** — Ollama on a home-lab workstation, router and editor on a
|
|
laptop, connected over WireGuard or similar. This is the common shape for
|
|
anyone whose GPU isn't in their daily-driver machine.
|
|
|
|
`classifier.base_url` / `api_key_env` / `model` take any OpenAI-compatible
|
|
endpoint, and switching between the three above is just changing that URL —
|
|
verified end to end against a non-loopback address, classifier and verifier
|
|
both.
|
|
|
|
Ollama binds `127.0.0.1` by default, so the VPN case fails with
|
|
connection-refused until the serving host applies
|
|
`deploy/ollama-over-vpn.conf`. Bind it to the VPN address, not `0.0.0.0`:
|
|
Ollama has no auth of any kind, so anything that can reach the port can run
|
|
inference and enumerate your models, and `0.0.0.0` publishes it on whatever
|
|
network the serving machine happens to be sitting on.
|
|
|
|
#### Local vs cloud classifier, measured (historical — against the superseded `qwen3.5`)
|
|
|
|
**This comparison predates the swap to `qwen2.5-coder-router:14b` above** and
|
|
has not been re-run against the current classifier. It measured the
|
|
*original* `qwen3.5` classifier against a cloud one, five prompts, same
|
|
system prompt, `temperature: 0`:
|
|
|
|
| | local `qwen3.5` (RTX 6000, superseded) | NeuralWatt `deepseek-v4-flash` |
|
|
|---|---|---|
|
|
| mean latency | 11.58s (4.95-15.76) | 1.02s |
|
|
| categories agreed with the label | 2 of 4 | 5 of 5 |
|
|
| hard failures | 1 of 5 (empty after 15.76s) | 0 |
|
|
| energy per call | ~7e-05 kWh, on your meter | 1.17e-05 kWh attributed |
|
|
| cost per 1,000 calls | electricity + 6.6GB resident | $0.093 (0.19% of quota) |
|
|
|
|
At the time, the cloud side was faster, more accurate, and less attributed
|
|
energy — the one hard failure is the same runaway-thinking-trace mode
|
|
described above. The general lesson still holds regardless of which local
|
|
model is current: local classification is not free, it is just unbilled, and
|
|
that is a substitution the cost axis had to unlearn independently (see
|
|
`CLAUDE.md`'s billing section). The shipped default is still local Ollama,
|
|
because switching spends quota and that is a deployment choice, not a code
|
|
one.
|
|
|
|
**Worth re-running against `qwen2.5-coder-router:14b`** before treating the
|
|
11x figure as current — the classifier's own latency dropped roughly 10x
|
|
since this table was measured (11.58s mean here vs ~1.1s in the three-way
|
|
table above), which would substantially narrow or erase the local/cloud gap
|
|
this table reports. Flagging this rather than re-deriving a number with no
|
|
live hardware to measure against.
|
|
|
|
### The verifier follows, but only to another Ollama
|
|
|
|
The local LLM check speaks Ollama's **native** `/api/chat` (the only way to
|
|
set `think: False`), so it follows the classifier across a VPN but not to a
|
|
cloud provider. It used to derive its URL by stripping `/v1` off
|
|
`classifier.base_url`, which meant moving the classifier at all would have
|
|
pointed it at `<that host>/api/chat`. It now has its own
|
|
`verification.base_url` and `verification.model`.
|
|
|
|
`verification.model` may be null only while both run on one host. Config load
|
|
**refuses** the null once the hostnames differ, because the failure is silent:
|
|
observed directly, with the classifier on NeuralWatt the verifier POSTed
|
|
`deepseek-v4-flash` to `localhost:11434`, 404'd, caught it, logged "local
|
|
verification unavailable" and recorded no sample. Verification would have
|
|
looked enabled while producing nothing.
|
|
|
|
Structural verification needs no model at all — it is pure Python — so
|
|
`verification.local_llm_enabled: false` leaves a host with no local inference
|
|
fully functional, minus the refusal/incoherence class of failure.
|
|
|
|
### A note on prompt-token estimation
|
|
|
|
opencode sends ~32K prompt tokens of system prompt and tool definitions on a
|
|
trivial request, so the measured-size floor in `estimate_prompt_tokens` does
|
|
real work — the classifier's own estimate for that request was two orders of
|
|
magnitude low. This is the same phenomenon pinch confronts on the provider
|
|
side; see [docs/pinch.md](pinch.md).
|