plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
250 lines
14 KiB
Markdown
250 lines
14 KiB
Markdown
# Local dispatch model: file summarization + diff checking
|
||
|
||
Status: done -- ollama-local rows with eligible_categories
|
||
|
||
## Context
|
||
|
||
Every routable model today comes from one place: `poller.py` fetching
|
||
NeuralWatt's `/models` endpoint. The classifier, verifier, and vision
|
||
fallback already run small local Ollama models, but only as *router-internal
|
||
utility calls* — none of them is a candidate `chat_completions` can actually
|
||
select and dispatch a real task to.
|
||
|
||
The model: **`nemotron-mini:4b`** (2.70GB, confirmed against Ollama's
|
||
registry manifest — layer size 2.70e9 bytes). Settled after the first
|
||
working name, "Nemotron 3.5 Lightning", failed verification: it's a real
|
||
pullable tag, but a 30B MoE with 25GB of Q4 weights resident regardless of
|
||
its 3B *active* parameter count — "3B active" describes compute per token,
|
||
not VRAM footprint, and it would have evicted the classifier/verifier/
|
||
vision models to run at all. `nemotron-mini:4b` is the real small member of
|
||
the Nemotron family and comfortably fits the actual budget: measured
|
||
worst-case combined resident footprint (classifier + verifier + vision) is
|
||
15.9GB/24GB, ~8GB free, and this uses 2.70GB of it — a third, with room to
|
||
spare even under concurrent load (classify + verify + vision + this dispatch
|
||
model all resident at once is the real worst case, not just the two-model
|
||
figure originally measured).
|
||
|
||
One model for both task categories, not two specialized ones — considered
|
||
and deliberately declined for v1 despite both fitting easily on cost alone
|
||
(a code-tuned model + a prose-tuned model together would still total well
|
||
under 4GB): keeps the catalog/config/eval footprint to one row instead of
|
||
two for the first pass, revisit only if real eval data shows one task
|
||
category suffering for want of a more specialized model.
|
||
|
||
Rationale for doing this at all: there is real spare VRAM, and the
|
||
`(model_id, provider)` catalog key was explicitly kept around for "a second
|
||
provider... without a migration." This is that extension, scoped to local
|
||
only — full multi-cloud-provider support is a separate, larger effort the
|
||
user explicitly deferred, not in scope here.
|
||
|
||
**Hard constraint carried over from the local-energy-ledger work**: this
|
||
must not repeat the "local is free" mistake CLAUDE.md already had to
|
||
unlearn once (the classifier/verifier comparison). A local dispatch model's
|
||
real cost is electricity, and it must compete in ranking on a real number,
|
||
not a None/zero that makes it look artificially free.
|
||
|
||
**Sequencing dependency, not just a nice-to-have**: this plan reuses the
|
||
local-energy metering machinery (`local_energy.py`, `LocalEnergyConfig`,
|
||
`_log_local_energy`, `cfg.local_energy_call_sites`) that Atlas is currently
|
||
building from `plans/local-energy-ledger-and-session-attribution.md`.
|
||
Verified live in the current (in-progress) `src/dispatcher.py` and
|
||
`src/config.py`: `classify()` and `_run_local_vision()` already call
|
||
`local_energy.measure(...)` / `_log_local_energy(...)`, and
|
||
`RouterConfig.local_energy_call_sites` is a cached dict computed once at
|
||
config load (`classify`/`verify`/`local_vision` keys, each gated on the
|
||
corresponding call site's `base_url` being loopback). **Execution of this
|
||
plan should not start until that work has landed** — both to avoid two
|
||
concurrent streams editing `dispatcher.py`/`config.py` at once, and because
|
||
this plan extends that exact machinery rather than re-implementing it.
|
||
|
||
Verified against the current repo before writing this:
|
||
- `poller.upsert()` (`src/poller.py:216`) is `INSERT ... ON CONFLICT(model_id, provider) DO UPDATE ...` — already keyed exactly right for a second provider value (e.g. `ollama-local`) to coexist with NeuralWatt rows with zero schema change to the key itself.
|
||
- `routing.rejection_reason()` / `is_eligible()` (`src/routing.py:62-155`) is a straight chain of independent checks against `dict` rows, each returning early with a reason string — a new "is this category allowed for this row" check is a one-function addition in the same shape as the existing tool-proficiency gate.
|
||
- `dispatcher.py` already has a non-NeuralWatt provider value in production use: the vision fallback sets `selected_provider="local"` (confirmed at `dispatcher.py:2516`) and the main dispatch path already reads `decision.selected_provider` to decide behavior. A new local-dispatch branch is the same shape, not new territory.
|
||
- `scoring.normalize_inverted()` gives a candidate with `cost=None` a **neutral 0.5**, not a winning 1.0 — so leaving a local row's cost unset would not make it falsely look cheapest, but it also means no real evidence is competing, which is the wrong default given the point of this feature. Seed it with a real number (below).
|
||
- `proficiency.categories` (`config/config.yaml`) flows directly into the classifier's system prompt (`classify()`, `src/dispatcher.py`) — adding new categories there is sufficient; no separate classifier-prompt wiring needed.
|
||
- `evals/tasks.yaml` categories use four scoring kinds (`code`/`exact`/`tool`/`judge`), and the project's own stated preference is "objective wherever the category allows" — worth pushing for `exact` where possible on `diff_checking` rather than defaulting straight to `judge`.
|
||
|
||
---
|
||
|
||
## Scope for this pass
|
||
|
||
1. Add local Ollama models as a new, distinct provider in the `models`
|
||
catalog, restricted to specific task categories.
|
||
2. Add two new proficiency categories: `file_summarization`,
|
||
`diff_checking`.
|
||
3. Add eval tasks for both, scored as objectively as each admits.
|
||
4. Wire real dispatch: when routing selects a local-provider row, call
|
||
local Ollama directly instead of NeuralWatt, reusing the local-energy
|
||
metering pattern already established for `local_vision`.
|
||
5. Seed real (not fabricated, not free) cost figures for the local row via
|
||
measurement, so it competes fairly in ranking.
|
||
|
||
Explicitly **not** in scope: tool-calling support for the local model,
|
||
vision, multi-cloud-provider work, or opening the local model up to any
|
||
category beyond the two above (start narrow; broaden only once real eval/
|
||
outcome data justifies it — same discipline this project already applies
|
||
to leaderboard priors).
|
||
|
||
---
|
||
|
||
## 1. Catalog: local models as a new provider
|
||
|
||
**New config section**, `config/config.yaml` (empty list by default — no
|
||
behavior change until populated, same opt-in posture as `local_vision`
|
||
before its model is pulled):
|
||
|
||
```yaml
|
||
local_dispatch_models:
|
||
# Each entry becomes one `models` row with provider='ollama-local'.
|
||
# eligible_categories is a hard filter: this row is never a candidate
|
||
# for any other category, however good its cost/proficiency might look.
|
||
- model_id: "nemotron-mini-router:4b" # Modelfile-tagged from nemotron-mini:4b — see setup below
|
||
context_window: 16384 # match the tag's PARAMETER num_ctx; confirm against nemotron-mini's real base context at execution time
|
||
max_output_tokens: 2048
|
||
tier: 1
|
||
eligible_categories: [file_summarization, diff_checking]
|
||
```
|
||
|
||
Setup mirrors the classifier/vision Modelfile-tagging convention already
|
||
established in this repo:
|
||
```bash
|
||
ollama pull nemotron-mini:4b
|
||
printf 'FROM nemotron-mini:4b\nPARAMETER num_ctx 16384\n' > Modelfile.local-dispatch
|
||
ollama create nemotron-mini-router:4b -f Modelfile.local-dispatch
|
||
```
|
||
`num_ctx` here is a starting estimate, not a measured value — confirm
|
||
against real file/diff sizes and nemotron-mini's actual base context window
|
||
during implementation, same as the classifier/vision tags were tuned from
|
||
measurement rather than guessed.
|
||
|
||
**Schema**: `config/schema.sql` — add `eligible_categories TEXT` (nullable,
|
||
comma-joined category list) to `models`. NULL for every existing NeuralWatt
|
||
row (no restriction, unchanged behavior). Additive migration, same pattern
|
||
as `proficiency_store.ensure_columns` / `_ensure_route_decisions_table`: a
|
||
live `router.db` predating this column needs a code-side
|
||
`ALTER TABLE ... ADD COLUMN` guarded by `IF NOT EXISTS`-equivalent handling
|
||
(SQLite has no `ADD COLUMN IF NOT EXISTS`; catch the "duplicate column"
|
||
`OperationalError` the way the rest of this codebase already does for
|
||
additive migrations).
|
||
|
||
**Seeding the rows**: extend `poller.py` with a small
|
||
`upsert_local_dispatch_models(conn, cfg)` run right after the NeuralWatt
|
||
`upsert()` in `main()` — same `INSERT ... ON CONFLICT(model_id, provider)`
|
||
shape, `provider='ollama-local'`, `base_model_id=model_id` (no
|
||
`-fast`/`-flex`/`-short` suffixes to parse for a local tag),
|
||
`access_level='public'`, `availability='active'`,
|
||
`cost_per_1m_prompt`/`cost_per_1m_completion` filled from the seeding step
|
||
below (not left NULL — see §5), `supports_tools=0`, `supports_vision=0`,
|
||
`supports_json_mode=0` unless a specific entry says otherwise,
|
||
`eligible_categories` joined from config. Runs from static config each poll,
|
||
so it never goes stale the way a NeuralWatt row would from a missed catalog
|
||
fetch — nothing to reconcile against an upstream source of truth.
|
||
|
||
---
|
||
|
||
## 2 & 3. New categories + eval tasks
|
||
|
||
`config/config.yaml`: add `file_summarization` and `diff_checking` to
|
||
`proficiency.categories`.
|
||
|
||
`evals/tasks.yaml`:
|
||
- **`file_summarization`** (kind: `judge` — inherently prose, same
|
||
treatment as `docs_writing`/`summarization`). A handful of small
|
||
real-shaped files (a config file with a non-obvious gotcha, a function
|
||
with a subtle edge case) with a rubric that checks the summary states the
|
||
gotcha, not just the happy path — same lesson `docs_writing`'s rubric
|
||
already learned about one-line-decides-the-score risk. Don't let this be
|
||
2-sample thin; budget for enough passes to cross
|
||
`self_eval_min_samples` (10), the same trap `docs_writing` fell into at
|
||
n=2 before six more passes spread it out.
|
||
- **`diff_checking`** — push for `exact` or `tool` scoring where the task
|
||
can be framed that way: "does this diff introduce bug X, and on which
|
||
line" has a discrete right answer and shouldn't default to `judge` just
|
||
because "diff review" sounds like a prose task. Model a few tasks on the
|
||
same shape as the `debugging` category's bowling/dominoes pairs (one diff
|
||
that's correct, one with a deliberately planted regression, model must
|
||
tell them apart) — reuse that pattern rather than inventing a new one.
|
||
|
||
`tests/test_task_set.py` already asserts every `code` task's reference
|
||
solution passes and every `exact` answer is independently recomputed —
|
||
extend it to cover the new tasks the same way, not a separate mechanism.
|
||
|
||
---
|
||
|
||
## 4. Dispatch: routing to a local-provider row
|
||
|
||
**`src/routing.py`**: add `eligible_categories: frozenset[str] | None`
|
||
handling to `rejection_reason()`/`is_eligible()`/`select_candidates()` —
|
||
same shape as the existing tool-proficiency gate: a row with a non-null
|
||
`eligible_categories` that doesn't contain the request's `task_category` is
|
||
rejected with reason `category_ineligible`. A row with `eligible_categories`
|
||
unset (every NeuralWatt row) is never affected — this only ever restricts,
|
||
never expands, a candidate set.
|
||
|
||
**`src/dispatcher.py`**: when `route()`'s winning candidate has
|
||
`provider == 'ollama-local'`, the main dispatch branch in `chat_completions()`
|
||
needs a local-call path instead of the NeuralWatt upstream client. Model
|
||
this directly on `_run_local_vision()` (`dispatcher.py:2083`): POST to the
|
||
local model's OpenAI-compatible `/chat/completions`, wrap the call in
|
||
`local_energy.measure(...)` / `_log_local_energy(...)` exactly as that
|
||
function already does, and reuse whatever response-shaping helper
|
||
`_local_vision_response` uses to return a normal OpenAI-shaped completion
|
||
(including the streamed variant) — don't reinvent that shape. Log a
|
||
`route_decisions` row the same as any other dispatch (`selected_provider`
|
||
already has precedent for a non-NeuralWatt value here).
|
||
|
||
Add a new call-site key to `RouterConfig.local_energy_call_sites`
|
||
(`model_post_init`, `src/config.py:621`) — e.g. `"local_dispatch"` — checked
|
||
against the local model's own `base_url` for loopback, same as the other
|
||
three sites. Local dispatch energy should land in the **same**
|
||
`local_energy_observations` table Atlas is building (not `energy_observations`
|
||
— that separation is the whole point of the other plan), with `call_type`
|
||
either `"local_dispatch"` or the actual task category
|
||
(`"file_summarization"` / `"diff_checking"`) — prefer the actual category:
|
||
these are real answered tasks, worth breaking out individually in the
|
||
ledger, unlike the fixed internal roles `classify`/`verify` play. Update
|
||
that table's `call_type` doc comment (`config/schema.sql`) to note it's not
|
||
a closed enum.
|
||
|
||
---
|
||
|
||
## 5. Real cost, not free and not fabricated
|
||
|
||
Mirror `seed_energy.py`'s approach, scaled down: a small
|
||
`seed_local_dispatch_energy.py` (or a mode of the existing script — decide
|
||
at implementation time which reads cleaner) that runs a handful of
|
||
reference file-summarization/diff-checking prompts against the local model,
|
||
metering real watts via `local_energy.py`'s sampler, and derives an
|
||
effective `$/1M prompt` / `$/1M completion` rate from
|
||
`tariff_usd_per_kwh × measured energy / tokens`. Write those into the
|
||
`models` row's `cost_per_1m_prompt`/`cost_per_1m_completion` columns.
|
||
|
||
This is the key design choice of this plan: it means `routing.estimated_cost()`
|
||
and `rank_candidates()` need **zero changes** to handle a local row
|
||
correctly — they already work purely off catalog price columns scaled to
|
||
request shape. A local model with a real, measured, tariff-priced cost
|
||
competes on the same axis as every cloud row, honestly.
|
||
|
||
---
|
||
|
||
## Verification
|
||
|
||
- `tests/test_task_set.py` extended and passing for the two new categories.
|
||
- `python leaderboard.py --check` still reports correctly (local models have
|
||
no leaderboard family — they start on self-eval alone, same as any
|
||
unlisted family; confirm this doesn't error on a provider it's never seen).
|
||
- `PYTHONPATH=src python -m poller` (or the new local-seed step) inserts the
|
||
`ollama-local` row(s); confirm via
|
||
`sqlite3 router.db "SELECT model_id, provider, eligible_categories, cost_per_1m_prompt FROM models WHERE provider='ollama-local'"`.
|
||
- `/route` with `task_category: file_summarization` returns the local model
|
||
as a candidate; `/route` with `task_category: coding_general` (or any
|
||
category outside the eligible set) does **not**, even if cost/proficiency
|
||
would otherwise favor it — confirms the hard filter, not just a soft
|
||
preference.
|
||
- A real `/dispatch` call against `file_summarization` produces a row in
|
||
`local_energy_observations` (not `energy_observations`) with a sane
|
||
`cost_usd` — confirms the metering/logging path and the separation
|
||
constraint both hold.
|
||
- Full `pytest` stays green.
|