Files
6krrt/plans/local-model-dispatch-file-summarization-diff-checking.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

14 KiB
Raw Permalink Blame History

Local dispatch model: file summarization + diff checking

Status: done -- ollama-local rows with eligible_categories

Context

Every routable model today comes from one place: poller.py fetching NeuralWatt's /models endpoint. The classifier, verifier, and vision fallback already run small local Ollama models, but only as router-internal utility calls — none of them is a candidate chat_completions can actually select and dispatch a real task to.

The model: nemotron-mini:4b (2.70GB, confirmed against Ollama's registry manifest — layer size 2.70e9 bytes). Settled after the first working name, "Nemotron 3.5 Lightning", failed verification: it's a real pullable tag, but a 30B MoE with 25GB of Q4 weights resident regardless of its 3B active parameter count — "3B active" describes compute per token, not VRAM footprint, and it would have evicted the classifier/verifier/ vision models to run at all. nemotron-mini:4b is the real small member of the Nemotron family and comfortably fits the actual budget: measured worst-case combined resident footprint (classifier + verifier + vision) is 15.9GB/24GB, ~8GB free, and this uses 2.70GB of it — a third, with room to spare even under concurrent load (classify + verify + vision + this dispatch model all resident at once is the real worst case, not just the two-model figure originally measured).

One model for both task categories, not two specialized ones — considered and deliberately declined for v1 despite both fitting easily on cost alone (a code-tuned model + a prose-tuned model together would still total well under 4GB): keeps the catalog/config/eval footprint to one row instead of two for the first pass, revisit only if real eval data shows one task category suffering for want of a more specialized model.

Rationale for doing this at all: there is real spare VRAM, and the (model_id, provider) catalog key was explicitly kept around for "a second provider... without a migration." This is that extension, scoped to local only — full multi-cloud-provider support is a separate, larger effort the user explicitly deferred, not in scope here.

Hard constraint carried over from the local-energy-ledger work: this must not repeat the "local is free" mistake CLAUDE.md already had to unlearn once (the classifier/verifier comparison). A local dispatch model's real cost is electricity, and it must compete in ranking on a real number, not a None/zero that makes it look artificially free.

Sequencing dependency, not just a nice-to-have: this plan reuses the local-energy metering machinery (local_energy.py, LocalEnergyConfig, _log_local_energy, cfg.local_energy_call_sites) that Atlas is currently building from plans/local-energy-ledger-and-session-attribution.md. Verified live in the current (in-progress) src/dispatcher.py and src/config.py: classify() and _run_local_vision() already call local_energy.measure(...) / _log_local_energy(...), and RouterConfig.local_energy_call_sites is a cached dict computed once at config load (classify/verify/local_vision keys, each gated on the corresponding call site's base_url being loopback). Execution of this plan should not start until that work has landed — both to avoid two concurrent streams editing dispatcher.py/config.py at once, and because this plan extends that exact machinery rather than re-implementing it.

Verified against the current repo before writing this:

  • poller.upsert() (src/poller.py:216) is INSERT ... ON CONFLICT(model_id, provider) DO UPDATE ... — already keyed exactly right for a second provider value (e.g. ollama-local) to coexist with NeuralWatt rows with zero schema change to the key itself.
  • routing.rejection_reason() / is_eligible() (src/routing.py:62-155) is a straight chain of independent checks against dict rows, each returning early with a reason string — a new "is this category allowed for this row" check is a one-function addition in the same shape as the existing tool-proficiency gate.
  • dispatcher.py already has a non-NeuralWatt provider value in production use: the vision fallback sets selected_provider="local" (confirmed at dispatcher.py:2516) and the main dispatch path already reads decision.selected_provider to decide behavior. A new local-dispatch branch is the same shape, not new territory.
  • scoring.normalize_inverted() gives a candidate with cost=None a neutral 0.5, not a winning 1.0 — so leaving a local row's cost unset would not make it falsely look cheapest, but it also means no real evidence is competing, which is the wrong default given the point of this feature. Seed it with a real number (below).
  • proficiency.categories (config/config.yaml) flows directly into the classifier's system prompt (classify(), src/dispatcher.py) — adding new categories there is sufficient; no separate classifier-prompt wiring needed.
  • evals/tasks.yaml categories use four scoring kinds (code/exact/tool/judge), and the project's own stated preference is "objective wherever the category allows" — worth pushing for exact where possible on diff_checking rather than defaulting straight to judge.

Scope for this pass

  1. Add local Ollama models as a new, distinct provider in the models catalog, restricted to specific task categories.
  2. Add two new proficiency categories: file_summarization, diff_checking.
  3. Add eval tasks for both, scored as objectively as each admits.
  4. Wire real dispatch: when routing selects a local-provider row, call local Ollama directly instead of NeuralWatt, reusing the local-energy metering pattern already established for local_vision.
  5. Seed real (not fabricated, not free) cost figures for the local row via measurement, so it competes fairly in ranking.

Explicitly not in scope: tool-calling support for the local model, vision, multi-cloud-provider work, or opening the local model up to any category beyond the two above (start narrow; broaden only once real eval/ outcome data justifies it — same discipline this project already applies to leaderboard priors).


1. Catalog: local models as a new provider

New config section, config/config.yaml (empty list by default — no behavior change until populated, same opt-in posture as local_vision before its model is pulled):

local_dispatch_models:
  # Each entry becomes one `models` row with provider='ollama-local'.
  # eligible_categories is a hard filter: this row is never a candidate
  # for any other category, however good its cost/proficiency might look.
  - model_id: "nemotron-mini-router:4b"   # Modelfile-tagged from nemotron-mini:4b — see setup below
    context_window: 16384                 # match the tag's PARAMETER num_ctx; confirm against nemotron-mini's real base context at execution time
    max_output_tokens: 2048
    tier: 1
    eligible_categories: [file_summarization, diff_checking]

Setup mirrors the classifier/vision Modelfile-tagging convention already established in this repo:

ollama pull nemotron-mini:4b
printf 'FROM nemotron-mini:4b\nPARAMETER num_ctx 16384\n' > Modelfile.local-dispatch
ollama create nemotron-mini-router:4b -f Modelfile.local-dispatch

num_ctx here is a starting estimate, not a measured value — confirm against real file/diff sizes and nemotron-mini's actual base context window during implementation, same as the classifier/vision tags were tuned from measurement rather than guessed.

Schema: config/schema.sql — add eligible_categories TEXT (nullable, comma-joined category list) to models. NULL for every existing NeuralWatt row (no restriction, unchanged behavior). Additive migration, same pattern as proficiency_store.ensure_columns / _ensure_route_decisions_table: a live router.db predating this column needs a code-side ALTER TABLE ... ADD COLUMN guarded by IF NOT EXISTS-equivalent handling (SQLite has no ADD COLUMN IF NOT EXISTS; catch the "duplicate column" OperationalError the way the rest of this codebase already does for additive migrations).

Seeding the rows: extend poller.py with a small upsert_local_dispatch_models(conn, cfg) run right after the NeuralWatt upsert() in main() — same INSERT ... ON CONFLICT(model_id, provider) shape, provider='ollama-local', base_model_id=model_id (no -fast/-flex/-short suffixes to parse for a local tag), access_level='public', availability='active', cost_per_1m_prompt/cost_per_1m_completion filled from the seeding step below (not left NULL — see §5), supports_tools=0, supports_vision=0, supports_json_mode=0 unless a specific entry says otherwise, eligible_categories joined from config. Runs from static config each poll, so it never goes stale the way a NeuralWatt row would from a missed catalog fetch — nothing to reconcile against an upstream source of truth.


2 & 3. New categories + eval tasks

config/config.yaml: add file_summarization and diff_checking to proficiency.categories.

evals/tasks.yaml:

  • file_summarization (kind: judge — inherently prose, same treatment as docs_writing/summarization). A handful of small real-shaped files (a config file with a non-obvious gotcha, a function with a subtle edge case) with a rubric that checks the summary states the gotcha, not just the happy path — same lesson docs_writing's rubric already learned about one-line-decides-the-score risk. Don't let this be 2-sample thin; budget for enough passes to cross self_eval_min_samples (10), the same trap docs_writing fell into at n=2 before six more passes spread it out.
  • diff_checking — push for exact or tool scoring where the task can be framed that way: "does this diff introduce bug X, and on which line" has a discrete right answer and shouldn't default to judge just because "diff review" sounds like a prose task. Model a few tasks on the same shape as the debugging category's bowling/dominoes pairs (one diff that's correct, one with a deliberately planted regression, model must tell them apart) — reuse that pattern rather than inventing a new one.

tests/test_task_set.py already asserts every code task's reference solution passes and every exact answer is independently recomputed — extend it to cover the new tasks the same way, not a separate mechanism.


4. Dispatch: routing to a local-provider row

src/routing.py: add eligible_categories: frozenset[str] | None handling to rejection_reason()/is_eligible()/select_candidates() — same shape as the existing tool-proficiency gate: a row with a non-null eligible_categories that doesn't contain the request's task_category is rejected with reason category_ineligible. A row with eligible_categories unset (every NeuralWatt row) is never affected — this only ever restricts, never expands, a candidate set.

src/dispatcher.py: when route()'s winning candidate has provider == 'ollama-local', the main dispatch branch in chat_completions() needs a local-call path instead of the NeuralWatt upstream client. Model this directly on _run_local_vision() (dispatcher.py:2083): POST to the local model's OpenAI-compatible /chat/completions, wrap the call in local_energy.measure(...) / _log_local_energy(...) exactly as that function already does, and reuse whatever response-shaping helper _local_vision_response uses to return a normal OpenAI-shaped completion (including the streamed variant) — don't reinvent that shape. Log a route_decisions row the same as any other dispatch (selected_provider already has precedent for a non-NeuralWatt value here).

Add a new call-site key to RouterConfig.local_energy_call_sites (model_post_init, src/config.py:621) — e.g. "local_dispatch" — checked against the local model's own base_url for loopback, same as the other three sites. Local dispatch energy should land in the same local_energy_observations table Atlas is building (not energy_observations — that separation is the whole point of the other plan), with call_type either "local_dispatch" or the actual task category ("file_summarization" / "diff_checking") — prefer the actual category: these are real answered tasks, worth breaking out individually in the ledger, unlike the fixed internal roles classify/verify play. Update that table's call_type doc comment (config/schema.sql) to note it's not a closed enum.


5. Real cost, not free and not fabricated

Mirror seed_energy.py's approach, scaled down: a small seed_local_dispatch_energy.py (or a mode of the existing script — decide at implementation time which reads cleaner) that runs a handful of reference file-summarization/diff-checking prompts against the local model, metering real watts via local_energy.py's sampler, and derives an effective $/1M prompt / $/1M completion rate from tariff_usd_per_kwh × measured energy / tokens. Write those into the models row's cost_per_1m_prompt/cost_per_1m_completion columns.

This is the key design choice of this plan: it means routing.estimated_cost() and rank_candidates() need zero changes to handle a local row correctly — they already work purely off catalog price columns scaled to request shape. A local model with a real, measured, tariff-priced cost competes on the same axis as every cloud row, honestly.


Verification

  • tests/test_task_set.py extended and passing for the two new categories.
  • python leaderboard.py --check still reports correctly (local models have no leaderboard family — they start on self-eval alone, same as any unlisted family; confirm this doesn't error on a provider it's never seen).
  • PYTHONPATH=src python -m poller (or the new local-seed step) inserts the ollama-local row(s); confirm via sqlite3 router.db "SELECT model_id, provider, eligible_categories, cost_per_1m_prompt FROM models WHERE provider='ollama-local'".
  • /route with task_category: file_summarization returns the local model as a candidate; /route with task_category: coding_general (or any category outside the eligible set) does not, even if cost/proficiency would otherwise favor it — confirms the hard filter, not just a soft preference.
  • A real /dispatch call against file_summarization produces a row in local_energy_observations (not energy_observations) with a sane cost_usd — confirms the metering/logging path and the separation constraint both hold.
  • Full pytest stays green.