plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
14 KiB
Local dispatch model: file summarization + diff checking
Status: done -- ollama-local rows with eligible_categories
Context
Every routable model today comes from one place: poller.py fetching
NeuralWatt's /models endpoint. The classifier, verifier, and vision
fallback already run small local Ollama models, but only as router-internal
utility calls — none of them is a candidate chat_completions can actually
select and dispatch a real task to.
The model: nemotron-mini:4b (2.70GB, confirmed against Ollama's
registry manifest — layer size 2.70e9 bytes). Settled after the first
working name, "Nemotron 3.5 Lightning", failed verification: it's a real
pullable tag, but a 30B MoE with 25GB of Q4 weights resident regardless of
its 3B active parameter count — "3B active" describes compute per token,
not VRAM footprint, and it would have evicted the classifier/verifier/
vision models to run at all. nemotron-mini:4b is the real small member of
the Nemotron family and comfortably fits the actual budget: measured
worst-case combined resident footprint (classifier + verifier + vision) is
15.9GB/24GB, ~8GB free, and this uses 2.70GB of it — a third, with room to
spare even under concurrent load (classify + verify + vision + this dispatch
model all resident at once is the real worst case, not just the two-model
figure originally measured).
One model for both task categories, not two specialized ones — considered and deliberately declined for v1 despite both fitting easily on cost alone (a code-tuned model + a prose-tuned model together would still total well under 4GB): keeps the catalog/config/eval footprint to one row instead of two for the first pass, revisit only if real eval data shows one task category suffering for want of a more specialized model.
Rationale for doing this at all: there is real spare VRAM, and the
(model_id, provider) catalog key was explicitly kept around for "a second
provider... without a migration." This is that extension, scoped to local
only — full multi-cloud-provider support is a separate, larger effort the
user explicitly deferred, not in scope here.
Hard constraint carried over from the local-energy-ledger work: this must not repeat the "local is free" mistake CLAUDE.md already had to unlearn once (the classifier/verifier comparison). A local dispatch model's real cost is electricity, and it must compete in ranking on a real number, not a None/zero that makes it look artificially free.
Sequencing dependency, not just a nice-to-have: this plan reuses the
local-energy metering machinery (local_energy.py, LocalEnergyConfig,
_log_local_energy, cfg.local_energy_call_sites) that Atlas is currently
building from plans/local-energy-ledger-and-session-attribution.md.
Verified live in the current (in-progress) src/dispatcher.py and
src/config.py: classify() and _run_local_vision() already call
local_energy.measure(...) / _log_local_energy(...), and
RouterConfig.local_energy_call_sites is a cached dict computed once at
config load (classify/verify/local_vision keys, each gated on the
corresponding call site's base_url being loopback). Execution of this
plan should not start until that work has landed — both to avoid two
concurrent streams editing dispatcher.py/config.py at once, and because
this plan extends that exact machinery rather than re-implementing it.
Verified against the current repo before writing this:
poller.upsert()(src/poller.py:216) isINSERT ... ON CONFLICT(model_id, provider) DO UPDATE ...— already keyed exactly right for a second provider value (e.g.ollama-local) to coexist with NeuralWatt rows with zero schema change to the key itself.routing.rejection_reason()/is_eligible()(src/routing.py:62-155) is a straight chain of independent checks againstdictrows, each returning early with a reason string — a new "is this category allowed for this row" check is a one-function addition in the same shape as the existing tool-proficiency gate.dispatcher.pyalready has a non-NeuralWatt provider value in production use: the vision fallback setsselected_provider="local"(confirmed atdispatcher.py:2516) and the main dispatch path already readsdecision.selected_providerto decide behavior. A new local-dispatch branch is the same shape, not new territory.scoring.normalize_inverted()gives a candidate withcost=Nonea neutral 0.5, not a winning 1.0 — so leaving a local row's cost unset would not make it falsely look cheapest, but it also means no real evidence is competing, which is the wrong default given the point of this feature. Seed it with a real number (below).proficiency.categories(config/config.yaml) flows directly into the classifier's system prompt (classify(),src/dispatcher.py) — adding new categories there is sufficient; no separate classifier-prompt wiring needed.evals/tasks.yamlcategories use four scoring kinds (code/exact/tool/judge), and the project's own stated preference is "objective wherever the category allows" — worth pushing forexactwhere possible ondiff_checkingrather than defaulting straight tojudge.
Scope for this pass
- Add local Ollama models as a new, distinct provider in the
modelscatalog, restricted to specific task categories. - Add two new proficiency categories:
file_summarization,diff_checking. - Add eval tasks for both, scored as objectively as each admits.
- Wire real dispatch: when routing selects a local-provider row, call
local Ollama directly instead of NeuralWatt, reusing the local-energy
metering pattern already established for
local_vision. - Seed real (not fabricated, not free) cost figures for the local row via measurement, so it competes fairly in ranking.
Explicitly not in scope: tool-calling support for the local model, vision, multi-cloud-provider work, or opening the local model up to any category beyond the two above (start narrow; broaden only once real eval/ outcome data justifies it — same discipline this project already applies to leaderboard priors).
1. Catalog: local models as a new provider
New config section, config/config.yaml (empty list by default — no
behavior change until populated, same opt-in posture as local_vision
before its model is pulled):
local_dispatch_models:
# Each entry becomes one `models` row with provider='ollama-local'.
# eligible_categories is a hard filter: this row is never a candidate
# for any other category, however good its cost/proficiency might look.
- model_id: "nemotron-mini-router:4b" # Modelfile-tagged from nemotron-mini:4b — see setup below
context_window: 16384 # match the tag's PARAMETER num_ctx; confirm against nemotron-mini's real base context at execution time
max_output_tokens: 2048
tier: 1
eligible_categories: [file_summarization, diff_checking]
Setup mirrors the classifier/vision Modelfile-tagging convention already established in this repo:
ollama pull nemotron-mini:4b
printf 'FROM nemotron-mini:4b\nPARAMETER num_ctx 16384\n' > Modelfile.local-dispatch
ollama create nemotron-mini-router:4b -f Modelfile.local-dispatch
num_ctx here is a starting estimate, not a measured value — confirm
against real file/diff sizes and nemotron-mini's actual base context window
during implementation, same as the classifier/vision tags were tuned from
measurement rather than guessed.
Schema: config/schema.sql — add eligible_categories TEXT (nullable,
comma-joined category list) to models. NULL for every existing NeuralWatt
row (no restriction, unchanged behavior). Additive migration, same pattern
as proficiency_store.ensure_columns / _ensure_route_decisions_table: a
live router.db predating this column needs a code-side
ALTER TABLE ... ADD COLUMN guarded by IF NOT EXISTS-equivalent handling
(SQLite has no ADD COLUMN IF NOT EXISTS; catch the "duplicate column"
OperationalError the way the rest of this codebase already does for
additive migrations).
Seeding the rows: extend poller.py with a small
upsert_local_dispatch_models(conn, cfg) run right after the NeuralWatt
upsert() in main() — same INSERT ... ON CONFLICT(model_id, provider)
shape, provider='ollama-local', base_model_id=model_id (no
-fast/-flex/-short suffixes to parse for a local tag),
access_level='public', availability='active',
cost_per_1m_prompt/cost_per_1m_completion filled from the seeding step
below (not left NULL — see §5), supports_tools=0, supports_vision=0,
supports_json_mode=0 unless a specific entry says otherwise,
eligible_categories joined from config. Runs from static config each poll,
so it never goes stale the way a NeuralWatt row would from a missed catalog
fetch — nothing to reconcile against an upstream source of truth.
2 & 3. New categories + eval tasks
config/config.yaml: add file_summarization and diff_checking to
proficiency.categories.
evals/tasks.yaml:
file_summarization(kind:judge— inherently prose, same treatment asdocs_writing/summarization). A handful of small real-shaped files (a config file with a non-obvious gotcha, a function with a subtle edge case) with a rubric that checks the summary states the gotcha, not just the happy path — same lessondocs_writing's rubric already learned about one-line-decides-the-score risk. Don't let this be 2-sample thin; budget for enough passes to crossself_eval_min_samples(10), the same trapdocs_writingfell into at n=2 before six more passes spread it out.diff_checking— push forexactortoolscoring where the task can be framed that way: "does this diff introduce bug X, and on which line" has a discrete right answer and shouldn't default tojudgejust because "diff review" sounds like a prose task. Model a few tasks on the same shape as thedebuggingcategory's bowling/dominoes pairs (one diff that's correct, one with a deliberately planted regression, model must tell them apart) — reuse that pattern rather than inventing a new one.
tests/test_task_set.py already asserts every code task's reference
solution passes and every exact answer is independently recomputed —
extend it to cover the new tasks the same way, not a separate mechanism.
4. Dispatch: routing to a local-provider row
src/routing.py: add eligible_categories: frozenset[str] | None
handling to rejection_reason()/is_eligible()/select_candidates() —
same shape as the existing tool-proficiency gate: a row with a non-null
eligible_categories that doesn't contain the request's task_category is
rejected with reason category_ineligible. A row with eligible_categories
unset (every NeuralWatt row) is never affected — this only ever restricts,
never expands, a candidate set.
src/dispatcher.py: when route()'s winning candidate has
provider == 'ollama-local', the main dispatch branch in chat_completions()
needs a local-call path instead of the NeuralWatt upstream client. Model
this directly on _run_local_vision() (dispatcher.py:2083): POST to the
local model's OpenAI-compatible /chat/completions, wrap the call in
local_energy.measure(...) / _log_local_energy(...) exactly as that
function already does, and reuse whatever response-shaping helper
_local_vision_response uses to return a normal OpenAI-shaped completion
(including the streamed variant) — don't reinvent that shape. Log a
route_decisions row the same as any other dispatch (selected_provider
already has precedent for a non-NeuralWatt value here).
Add a new call-site key to RouterConfig.local_energy_call_sites
(model_post_init, src/config.py:621) — e.g. "local_dispatch" — checked
against the local model's own base_url for loopback, same as the other
three sites. Local dispatch energy should land in the same
local_energy_observations table Atlas is building (not energy_observations
— that separation is the whole point of the other plan), with call_type
either "local_dispatch" or the actual task category
("file_summarization" / "diff_checking") — prefer the actual category:
these are real answered tasks, worth breaking out individually in the
ledger, unlike the fixed internal roles classify/verify play. Update
that table's call_type doc comment (config/schema.sql) to note it's not
a closed enum.
5. Real cost, not free and not fabricated
Mirror seed_energy.py's approach, scaled down: a small
seed_local_dispatch_energy.py (or a mode of the existing script — decide
at implementation time which reads cleaner) that runs a handful of
reference file-summarization/diff-checking prompts against the local model,
metering real watts via local_energy.py's sampler, and derives an
effective $/1M prompt / $/1M completion rate from
tariff_usd_per_kwh × measured energy / tokens. Write those into the
models row's cost_per_1m_prompt/cost_per_1m_completion columns.
This is the key design choice of this plan: it means routing.estimated_cost()
and rank_candidates() need zero changes to handle a local row
correctly — they already work purely off catalog price columns scaled to
request shape. A local model with a real, measured, tariff-priced cost
competes on the same axis as every cloud row, honestly.
Verification
tests/test_task_set.pyextended and passing for the two new categories.python leaderboard.py --checkstill reports correctly (local models have no leaderboard family — they start on self-eval alone, same as any unlisted family; confirm this doesn't error on a provider it's never seen).PYTHONPATH=src python -m poller(or the new local-seed step) inserts theollama-localrow(s); confirm viasqlite3 router.db "SELECT model_id, provider, eligible_categories, cost_per_1m_prompt FROM models WHERE provider='ollama-local'"./routewithtask_category: file_summarizationreturns the local model as a candidate;/routewithtask_category: coding_general(or any category outside the eligible set) does not, even if cost/proficiency would otherwise favor it — confirms the hard filter, not just a soft preference.- A real
/dispatchcall againstfile_summarizationproduces a row inlocal_energy_observations(notenergy_observations) with a sanecost_usd— confirms the metering/logging path and the separation constraint both hold. - Full
pyteststays green.