Files
6krrt/plans/local-model-dispatch-file-summarization-diff-checking.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

250 lines
14 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Local dispatch model: file summarization + diff checking
Status: done -- ollama-local rows with eligible_categories
## Context
Every routable model today comes from one place: `poller.py` fetching
NeuralWatt's `/models` endpoint. The classifier, verifier, and vision
fallback already run small local Ollama models, but only as *router-internal
utility calls* — none of them is a candidate `chat_completions` can actually
select and dispatch a real task to.
The model: **`nemotron-mini:4b`** (2.70GB, confirmed against Ollama's
registry manifest — layer size 2.70e9 bytes). Settled after the first
working name, "Nemotron 3.5 Lightning", failed verification: it's a real
pullable tag, but a 30B MoE with 25GB of Q4 weights resident regardless of
its 3B *active* parameter count — "3B active" describes compute per token,
not VRAM footprint, and it would have evicted the classifier/verifier/
vision models to run at all. `nemotron-mini:4b` is the real small member of
the Nemotron family and comfortably fits the actual budget: measured
worst-case combined resident footprint (classifier + verifier + vision) is
15.9GB/24GB, ~8GB free, and this uses 2.70GB of it — a third, with room to
spare even under concurrent load (classify + verify + vision + this dispatch
model all resident at once is the real worst case, not just the two-model
figure originally measured).
One model for both task categories, not two specialized ones — considered
and deliberately declined for v1 despite both fitting easily on cost alone
(a code-tuned model + a prose-tuned model together would still total well
under 4GB): keeps the catalog/config/eval footprint to one row instead of
two for the first pass, revisit only if real eval data shows one task
category suffering for want of a more specialized model.
Rationale for doing this at all: there is real spare VRAM, and the
`(model_id, provider)` catalog key was explicitly kept around for "a second
provider... without a migration." This is that extension, scoped to local
only — full multi-cloud-provider support is a separate, larger effort the
user explicitly deferred, not in scope here.
**Hard constraint carried over from the local-energy-ledger work**: this
must not repeat the "local is free" mistake CLAUDE.md already had to
unlearn once (the classifier/verifier comparison). A local dispatch model's
real cost is electricity, and it must compete in ranking on a real number,
not a None/zero that makes it look artificially free.
**Sequencing dependency, not just a nice-to-have**: this plan reuses the
local-energy metering machinery (`local_energy.py`, `LocalEnergyConfig`,
`_log_local_energy`, `cfg.local_energy_call_sites`) that Atlas is currently
building from `plans/local-energy-ledger-and-session-attribution.md`.
Verified live in the current (in-progress) `src/dispatcher.py` and
`src/config.py`: `classify()` and `_run_local_vision()` already call
`local_energy.measure(...)` / `_log_local_energy(...)`, and
`RouterConfig.local_energy_call_sites` is a cached dict computed once at
config load (`classify`/`verify`/`local_vision` keys, each gated on the
corresponding call site's `base_url` being loopback). **Execution of this
plan should not start until that work has landed** — both to avoid two
concurrent streams editing `dispatcher.py`/`config.py` at once, and because
this plan extends that exact machinery rather than re-implementing it.
Verified against the current repo before writing this:
- `poller.upsert()` (`src/poller.py:216`) is `INSERT ... ON CONFLICT(model_id, provider) DO UPDATE ...` — already keyed exactly right for a second provider value (e.g. `ollama-local`) to coexist with NeuralWatt rows with zero schema change to the key itself.
- `routing.rejection_reason()` / `is_eligible()` (`src/routing.py:62-155`) is a straight chain of independent checks against `dict` rows, each returning early with a reason string — a new "is this category allowed for this row" check is a one-function addition in the same shape as the existing tool-proficiency gate.
- `dispatcher.py` already has a non-NeuralWatt provider value in production use: the vision fallback sets `selected_provider="local"` (confirmed at `dispatcher.py:2516`) and the main dispatch path already reads `decision.selected_provider` to decide behavior. A new local-dispatch branch is the same shape, not new territory.
- `scoring.normalize_inverted()` gives a candidate with `cost=None` a **neutral 0.5**, not a winning 1.0 — so leaving a local row's cost unset would not make it falsely look cheapest, but it also means no real evidence is competing, which is the wrong default given the point of this feature. Seed it with a real number (below).
- `proficiency.categories` (`config/config.yaml`) flows directly into the classifier's system prompt (`classify()`, `src/dispatcher.py`) — adding new categories there is sufficient; no separate classifier-prompt wiring needed.
- `evals/tasks.yaml` categories use four scoring kinds (`code`/`exact`/`tool`/`judge`), and the project's own stated preference is "objective wherever the category allows" — worth pushing for `exact` where possible on `diff_checking` rather than defaulting straight to `judge`.
---
## Scope for this pass
1. Add local Ollama models as a new, distinct provider in the `models`
catalog, restricted to specific task categories.
2. Add two new proficiency categories: `file_summarization`,
`diff_checking`.
3. Add eval tasks for both, scored as objectively as each admits.
4. Wire real dispatch: when routing selects a local-provider row, call
local Ollama directly instead of NeuralWatt, reusing the local-energy
metering pattern already established for `local_vision`.
5. Seed real (not fabricated, not free) cost figures for the local row via
measurement, so it competes fairly in ranking.
Explicitly **not** in scope: tool-calling support for the local model,
vision, multi-cloud-provider work, or opening the local model up to any
category beyond the two above (start narrow; broaden only once real eval/
outcome data justifies it — same discipline this project already applies
to leaderboard priors).
---
## 1. Catalog: local models as a new provider
**New config section**, `config/config.yaml` (empty list by default — no
behavior change until populated, same opt-in posture as `local_vision`
before its model is pulled):
```yaml
local_dispatch_models:
# Each entry becomes one `models` row with provider='ollama-local'.
# eligible_categories is a hard filter: this row is never a candidate
# for any other category, however good its cost/proficiency might look.
- model_id: "nemotron-mini-router:4b" # Modelfile-tagged from nemotron-mini:4b — see setup below
context_window: 16384 # match the tag's PARAMETER num_ctx; confirm against nemotron-mini's real base context at execution time
max_output_tokens: 2048
tier: 1
eligible_categories: [file_summarization, diff_checking]
```
Setup mirrors the classifier/vision Modelfile-tagging convention already
established in this repo:
```bash
ollama pull nemotron-mini:4b
printf 'FROM nemotron-mini:4b\nPARAMETER num_ctx 16384\n' > Modelfile.local-dispatch
ollama create nemotron-mini-router:4b -f Modelfile.local-dispatch
```
`num_ctx` here is a starting estimate, not a measured value — confirm
against real file/diff sizes and nemotron-mini's actual base context window
during implementation, same as the classifier/vision tags were tuned from
measurement rather than guessed.
**Schema**: `config/schema.sql` — add `eligible_categories TEXT` (nullable,
comma-joined category list) to `models`. NULL for every existing NeuralWatt
row (no restriction, unchanged behavior). Additive migration, same pattern
as `proficiency_store.ensure_columns` / `_ensure_route_decisions_table`: a
live `router.db` predating this column needs a code-side
`ALTER TABLE ... ADD COLUMN` guarded by `IF NOT EXISTS`-equivalent handling
(SQLite has no `ADD COLUMN IF NOT EXISTS`; catch the "duplicate column"
`OperationalError` the way the rest of this codebase already does for
additive migrations).
**Seeding the rows**: extend `poller.py` with a small
`upsert_local_dispatch_models(conn, cfg)` run right after the NeuralWatt
`upsert()` in `main()` — same `INSERT ... ON CONFLICT(model_id, provider)`
shape, `provider='ollama-local'`, `base_model_id=model_id` (no
`-fast`/`-flex`/`-short` suffixes to parse for a local tag),
`access_level='public'`, `availability='active'`,
`cost_per_1m_prompt`/`cost_per_1m_completion` filled from the seeding step
below (not left NULL — see §5), `supports_tools=0`, `supports_vision=0`,
`supports_json_mode=0` unless a specific entry says otherwise,
`eligible_categories` joined from config. Runs from static config each poll,
so it never goes stale the way a NeuralWatt row would from a missed catalog
fetch — nothing to reconcile against an upstream source of truth.
---
## 2 & 3. New categories + eval tasks
`config/config.yaml`: add `file_summarization` and `diff_checking` to
`proficiency.categories`.
`evals/tasks.yaml`:
- **`file_summarization`** (kind: `judge` — inherently prose, same
treatment as `docs_writing`/`summarization`). A handful of small
real-shaped files (a config file with a non-obvious gotcha, a function
with a subtle edge case) with a rubric that checks the summary states the
gotcha, not just the happy path — same lesson `docs_writing`'s rubric
already learned about one-line-decides-the-score risk. Don't let this be
2-sample thin; budget for enough passes to cross
`self_eval_min_samples` (10), the same trap `docs_writing` fell into at
n=2 before six more passes spread it out.
- **`diff_checking`** — push for `exact` or `tool` scoring where the task
can be framed that way: "does this diff introduce bug X, and on which
line" has a discrete right answer and shouldn't default to `judge` just
because "diff review" sounds like a prose task. Model a few tasks on the
same shape as the `debugging` category's bowling/dominoes pairs (one diff
that's correct, one with a deliberately planted regression, model must
tell them apart) — reuse that pattern rather than inventing a new one.
`tests/test_task_set.py` already asserts every `code` task's reference
solution passes and every `exact` answer is independently recomputed —
extend it to cover the new tasks the same way, not a separate mechanism.
---
## 4. Dispatch: routing to a local-provider row
**`src/routing.py`**: add `eligible_categories: frozenset[str] | None`
handling to `rejection_reason()`/`is_eligible()`/`select_candidates()` —
same shape as the existing tool-proficiency gate: a row with a non-null
`eligible_categories` that doesn't contain the request's `task_category` is
rejected with reason `category_ineligible`. A row with `eligible_categories`
unset (every NeuralWatt row) is never affected — this only ever restricts,
never expands, a candidate set.
**`src/dispatcher.py`**: when `route()`'s winning candidate has
`provider == 'ollama-local'`, the main dispatch branch in `chat_completions()`
needs a local-call path instead of the NeuralWatt upstream client. Model
this directly on `_run_local_vision()` (`dispatcher.py:2083`): POST to the
local model's OpenAI-compatible `/chat/completions`, wrap the call in
`local_energy.measure(...)` / `_log_local_energy(...)` exactly as that
function already does, and reuse whatever response-shaping helper
`_local_vision_response` uses to return a normal OpenAI-shaped completion
(including the streamed variant) — don't reinvent that shape. Log a
`route_decisions` row the same as any other dispatch (`selected_provider`
already has precedent for a non-NeuralWatt value here).
Add a new call-site key to `RouterConfig.local_energy_call_sites`
(`model_post_init`, `src/config.py:621`) — e.g. `"local_dispatch"` — checked
against the local model's own `base_url` for loopback, same as the other
three sites. Local dispatch energy should land in the **same**
`local_energy_observations` table Atlas is building (not `energy_observations`
— that separation is the whole point of the other plan), with `call_type`
either `"local_dispatch"` or the actual task category
(`"file_summarization"` / `"diff_checking"`) — prefer the actual category:
these are real answered tasks, worth breaking out individually in the
ledger, unlike the fixed internal roles `classify`/`verify` play. Update
that table's `call_type` doc comment (`config/schema.sql`) to note it's not
a closed enum.
---
## 5. Real cost, not free and not fabricated
Mirror `seed_energy.py`'s approach, scaled down: a small
`seed_local_dispatch_energy.py` (or a mode of the existing script — decide
at implementation time which reads cleaner) that runs a handful of
reference file-summarization/diff-checking prompts against the local model,
metering real watts via `local_energy.py`'s sampler, and derives an
effective `$/1M prompt` / `$/1M completion` rate from
`tariff_usd_per_kwh × measured energy / tokens`. Write those into the
`models` row's `cost_per_1m_prompt`/`cost_per_1m_completion` columns.
This is the key design choice of this plan: it means `routing.estimated_cost()`
and `rank_candidates()` need **zero changes** to handle a local row
correctly — they already work purely off catalog price columns scaled to
request shape. A local model with a real, measured, tariff-priced cost
competes on the same axis as every cloud row, honestly.
---
## Verification
- `tests/test_task_set.py` extended and passing for the two new categories.
- `python leaderboard.py --check` still reports correctly (local models have
no leaderboard family — they start on self-eval alone, same as any
unlisted family; confirm this doesn't error on a provider it's never seen).
- `PYTHONPATH=src python -m poller` (or the new local-seed step) inserts the
`ollama-local` row(s); confirm via
`sqlite3 router.db "SELECT model_id, provider, eligible_categories, cost_per_1m_prompt FROM models WHERE provider='ollama-local'"`.
- `/route` with `task_category: file_summarization` returns the local model
as a candidate; `/route` with `task_category: coding_general` (or any
category outside the eligible set) does **not**, even if cost/proficiency
would otherwise favor it — confirms the hard filter, not just a soft
preference.
- A real `/dispatch` call against `file_summarization` produces a row in
`local_energy_observations` (not `energy_observations`) with a sane
`cost_usd` — confirms the metering/logging path and the separation
constraint both hold.
- Full `pytest` stays green.