# Local dispatch model: file summarization + diff checking Status: done -- ollama-local rows with eligible_categories ## Context Every routable model today comes from one place: `poller.py` fetching NeuralWatt's `/models` endpoint. The classifier, verifier, and vision fallback already run small local Ollama models, but only as *router-internal utility calls* — none of them is a candidate `chat_completions` can actually select and dispatch a real task to. The model: **`nemotron-mini:4b`** (2.70GB, confirmed against Ollama's registry manifest — layer size 2.70e9 bytes). Settled after the first working name, "Nemotron 3.5 Lightning", failed verification: it's a real pullable tag, but a 30B MoE with 25GB of Q4 weights resident regardless of its 3B *active* parameter count — "3B active" describes compute per token, not VRAM footprint, and it would have evicted the classifier/verifier/ vision models to run at all. `nemotron-mini:4b` is the real small member of the Nemotron family and comfortably fits the actual budget: measured worst-case combined resident footprint (classifier + verifier + vision) is 15.9GB/24GB, ~8GB free, and this uses 2.70GB of it — a third, with room to spare even under concurrent load (classify + verify + vision + this dispatch model all resident at once is the real worst case, not just the two-model figure originally measured). One model for both task categories, not two specialized ones — considered and deliberately declined for v1 despite both fitting easily on cost alone (a code-tuned model + a prose-tuned model together would still total well under 4GB): keeps the catalog/config/eval footprint to one row instead of two for the first pass, revisit only if real eval data shows one task category suffering for want of a more specialized model. Rationale for doing this at all: there is real spare VRAM, and the `(model_id, provider)` catalog key was explicitly kept around for "a second provider... without a migration." This is that extension, scoped to local only — full multi-cloud-provider support is a separate, larger effort the user explicitly deferred, not in scope here. **Hard constraint carried over from the local-energy-ledger work**: this must not repeat the "local is free" mistake CLAUDE.md already had to unlearn once (the classifier/verifier comparison). A local dispatch model's real cost is electricity, and it must compete in ranking on a real number, not a None/zero that makes it look artificially free. **Sequencing dependency, not just a nice-to-have**: this plan reuses the local-energy metering machinery (`local_energy.py`, `LocalEnergyConfig`, `_log_local_energy`, `cfg.local_energy_call_sites`) that Atlas is currently building from `plans/local-energy-ledger-and-session-attribution.md`. Verified live in the current (in-progress) `src/dispatcher.py` and `src/config.py`: `classify()` and `_run_local_vision()` already call `local_energy.measure(...)` / `_log_local_energy(...)`, and `RouterConfig.local_energy_call_sites` is a cached dict computed once at config load (`classify`/`verify`/`local_vision` keys, each gated on the corresponding call site's `base_url` being loopback). **Execution of this plan should not start until that work has landed** — both to avoid two concurrent streams editing `dispatcher.py`/`config.py` at once, and because this plan extends that exact machinery rather than re-implementing it. Verified against the current repo before writing this: - `poller.upsert()` (`src/poller.py:216`) is `INSERT ... ON CONFLICT(model_id, provider) DO UPDATE ...` — already keyed exactly right for a second provider value (e.g. `ollama-local`) to coexist with NeuralWatt rows with zero schema change to the key itself. - `routing.rejection_reason()` / `is_eligible()` (`src/routing.py:62-155`) is a straight chain of independent checks against `dict` rows, each returning early with a reason string — a new "is this category allowed for this row" check is a one-function addition in the same shape as the existing tool-proficiency gate. - `dispatcher.py` already has a non-NeuralWatt provider value in production use: the vision fallback sets `selected_provider="local"` (confirmed at `dispatcher.py:2516`) and the main dispatch path already reads `decision.selected_provider` to decide behavior. A new local-dispatch branch is the same shape, not new territory. - `scoring.normalize_inverted()` gives a candidate with `cost=None` a **neutral 0.5**, not a winning 1.0 — so leaving a local row's cost unset would not make it falsely look cheapest, but it also means no real evidence is competing, which is the wrong default given the point of this feature. Seed it with a real number (below). - `proficiency.categories` (`config/config.yaml`) flows directly into the classifier's system prompt (`classify()`, `src/dispatcher.py`) — adding new categories there is sufficient; no separate classifier-prompt wiring needed. - `evals/tasks.yaml` categories use four scoring kinds (`code`/`exact`/`tool`/`judge`), and the project's own stated preference is "objective wherever the category allows" — worth pushing for `exact` where possible on `diff_checking` rather than defaulting straight to `judge`. --- ## Scope for this pass 1. Add local Ollama models as a new, distinct provider in the `models` catalog, restricted to specific task categories. 2. Add two new proficiency categories: `file_summarization`, `diff_checking`. 3. Add eval tasks for both, scored as objectively as each admits. 4. Wire real dispatch: when routing selects a local-provider row, call local Ollama directly instead of NeuralWatt, reusing the local-energy metering pattern already established for `local_vision`. 5. Seed real (not fabricated, not free) cost figures for the local row via measurement, so it competes fairly in ranking. Explicitly **not** in scope: tool-calling support for the local model, vision, multi-cloud-provider work, or opening the local model up to any category beyond the two above (start narrow; broaden only once real eval/ outcome data justifies it — same discipline this project already applies to leaderboard priors). --- ## 1. Catalog: local models as a new provider **New config section**, `config/config.yaml` (empty list by default — no behavior change until populated, same opt-in posture as `local_vision` before its model is pulled): ```yaml local_dispatch_models: # Each entry becomes one `models` row with provider='ollama-local'. # eligible_categories is a hard filter: this row is never a candidate # for any other category, however good its cost/proficiency might look. - model_id: "nemotron-mini-router:4b" # Modelfile-tagged from nemotron-mini:4b — see setup below context_window: 16384 # match the tag's PARAMETER num_ctx; confirm against nemotron-mini's real base context at execution time max_output_tokens: 2048 tier: 1 eligible_categories: [file_summarization, diff_checking] ``` Setup mirrors the classifier/vision Modelfile-tagging convention already established in this repo: ```bash ollama pull nemotron-mini:4b printf 'FROM nemotron-mini:4b\nPARAMETER num_ctx 16384\n' > Modelfile.local-dispatch ollama create nemotron-mini-router:4b -f Modelfile.local-dispatch ``` `num_ctx` here is a starting estimate, not a measured value — confirm against real file/diff sizes and nemotron-mini's actual base context window during implementation, same as the classifier/vision tags were tuned from measurement rather than guessed. **Schema**: `config/schema.sql` — add `eligible_categories TEXT` (nullable, comma-joined category list) to `models`. NULL for every existing NeuralWatt row (no restriction, unchanged behavior). Additive migration, same pattern as `proficiency_store.ensure_columns` / `_ensure_route_decisions_table`: a live `router.db` predating this column needs a code-side `ALTER TABLE ... ADD COLUMN` guarded by `IF NOT EXISTS`-equivalent handling (SQLite has no `ADD COLUMN IF NOT EXISTS`; catch the "duplicate column" `OperationalError` the way the rest of this codebase already does for additive migrations). **Seeding the rows**: extend `poller.py` with a small `upsert_local_dispatch_models(conn, cfg)` run right after the NeuralWatt `upsert()` in `main()` — same `INSERT ... ON CONFLICT(model_id, provider)` shape, `provider='ollama-local'`, `base_model_id=model_id` (no `-fast`/`-flex`/`-short` suffixes to parse for a local tag), `access_level='public'`, `availability='active'`, `cost_per_1m_prompt`/`cost_per_1m_completion` filled from the seeding step below (not left NULL — see §5), `supports_tools=0`, `supports_vision=0`, `supports_json_mode=0` unless a specific entry says otherwise, `eligible_categories` joined from config. Runs from static config each poll, so it never goes stale the way a NeuralWatt row would from a missed catalog fetch — nothing to reconcile against an upstream source of truth. --- ## 2 & 3. New categories + eval tasks `config/config.yaml`: add `file_summarization` and `diff_checking` to `proficiency.categories`. `evals/tasks.yaml`: - **`file_summarization`** (kind: `judge` — inherently prose, same treatment as `docs_writing`/`summarization`). A handful of small real-shaped files (a config file with a non-obvious gotcha, a function with a subtle edge case) with a rubric that checks the summary states the gotcha, not just the happy path — same lesson `docs_writing`'s rubric already learned about one-line-decides-the-score risk. Don't let this be 2-sample thin; budget for enough passes to cross `self_eval_min_samples` (10), the same trap `docs_writing` fell into at n=2 before six more passes spread it out. - **`diff_checking`** — push for `exact` or `tool` scoring where the task can be framed that way: "does this diff introduce bug X, and on which line" has a discrete right answer and shouldn't default to `judge` just because "diff review" sounds like a prose task. Model a few tasks on the same shape as the `debugging` category's bowling/dominoes pairs (one diff that's correct, one with a deliberately planted regression, model must tell them apart) — reuse that pattern rather than inventing a new one. `tests/test_task_set.py` already asserts every `code` task's reference solution passes and every `exact` answer is independently recomputed — extend it to cover the new tasks the same way, not a separate mechanism. --- ## 4. Dispatch: routing to a local-provider row **`src/routing.py`**: add `eligible_categories: frozenset[str] | None` handling to `rejection_reason()`/`is_eligible()`/`select_candidates()` — same shape as the existing tool-proficiency gate: a row with a non-null `eligible_categories` that doesn't contain the request's `task_category` is rejected with reason `category_ineligible`. A row with `eligible_categories` unset (every NeuralWatt row) is never affected — this only ever restricts, never expands, a candidate set. **`src/dispatcher.py`**: when `route()`'s winning candidate has `provider == 'ollama-local'`, the main dispatch branch in `chat_completions()` needs a local-call path instead of the NeuralWatt upstream client. Model this directly on `_run_local_vision()` (`dispatcher.py:2083`): POST to the local model's OpenAI-compatible `/chat/completions`, wrap the call in `local_energy.measure(...)` / `_log_local_energy(...)` exactly as that function already does, and reuse whatever response-shaping helper `_local_vision_response` uses to return a normal OpenAI-shaped completion (including the streamed variant) — don't reinvent that shape. Log a `route_decisions` row the same as any other dispatch (`selected_provider` already has precedent for a non-NeuralWatt value here). Add a new call-site key to `RouterConfig.local_energy_call_sites` (`model_post_init`, `src/config.py:621`) — e.g. `"local_dispatch"` — checked against the local model's own `base_url` for loopback, same as the other three sites. Local dispatch energy should land in the **same** `local_energy_observations` table Atlas is building (not `energy_observations` — that separation is the whole point of the other plan), with `call_type` either `"local_dispatch"` or the actual task category (`"file_summarization"` / `"diff_checking"`) — prefer the actual category: these are real answered tasks, worth breaking out individually in the ledger, unlike the fixed internal roles `classify`/`verify` play. Update that table's `call_type` doc comment (`config/schema.sql`) to note it's not a closed enum. --- ## 5. Real cost, not free and not fabricated Mirror `seed_energy.py`'s approach, scaled down: a small `seed_local_dispatch_energy.py` (or a mode of the existing script — decide at implementation time which reads cleaner) that runs a handful of reference file-summarization/diff-checking prompts against the local model, metering real watts via `local_energy.py`'s sampler, and derives an effective `$/1M prompt` / `$/1M completion` rate from `tariff_usd_per_kwh × measured energy / tokens`. Write those into the `models` row's `cost_per_1m_prompt`/`cost_per_1m_completion` columns. This is the key design choice of this plan: it means `routing.estimated_cost()` and `rank_candidates()` need **zero changes** to handle a local row correctly — they already work purely off catalog price columns scaled to request shape. A local model with a real, measured, tariff-priced cost competes on the same axis as every cloud row, honestly. --- ## Verification - `tests/test_task_set.py` extended and passing for the two new categories. - `python leaderboard.py --check` still reports correctly (local models have no leaderboard family — they start on self-eval alone, same as any unlisted family; confirm this doesn't error on a provider it's never seen). - `PYTHONPATH=src python -m poller` (or the new local-seed step) inserts the `ollama-local` row(s); confirm via `sqlite3 router.db "SELECT model_id, provider, eligible_categories, cost_per_1m_prompt FROM models WHERE provider='ollama-local'"`. - `/route` with `task_category: file_summarization` returns the local model as a candidate; `/route` with `task_category: coding_general` (or any category outside the eligible set) does **not**, even if cost/proficiency would otherwise favor it — confirms the hard filter, not just a soft preference. - A real `/dispatch` call against `file_summarization` produces a row in `local_energy_observations` (not `energy_observations`) with a sane `cost_usd` — confirms the metering/logging path and the separation constraint both hold. - Full `pytest` stays green.