plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
16 KiB
Local energy ledger and session-directory attribution
Status: done -- local_energy.py
Spec covering the five "not built yet" items in CLAUDE.md. Real
implementation work here is scoped to two items (session-directory
attribution, local energy ledger); the other three are a research pass and
two already-decided non-builds, included for completeness so nothing on the
open-items list gets lost.
Context
CLAUDE.md's "What's NOT built yet" section lists five open items. They are not equally sized or equally actionable, and treating them as one uniform backlog would be wrong:
- #5 (local energy not on the ledger) matters more than its "low priority" position suggests — the user's own Ollama classifier/verifier runs on hardware they personally pay the electric bill for, so this is a real number they currently can't see, not an abstract accounting gap.
- #4 (session-directory attribution) is a concrete, scoped bug in one pure function with existing test coverage to extend.
- #1 (leaderboard priors) is explicitly not something to code around —
leaderboards.yamlships empty on purpose because fabricating scores would put fake data straight into routing. The tooling (leaderboard.py) is already done; what's missing is real, sourced numbers, which is a research task with a "only write what's verifiably cited" constraint, not an implementation task. - #2 (sampling depth for 3 models) no longer touches cost (only
eco, which isn't an objective) and the tooling to fix it (seed_energy.py --models ..., the now-fixedllm-router-seed.timer) already exists. This is "let it accumulate and check back," not new code. - #3 (retry doesn't reach streaming) CLAUDE.md already concludes
buffering to fix it "would cost streaming itself, a worse trade for
interactive work," and that
POST /outcomeis the accepted answer. This is a decision already made, not a gap waiting on code.
So this plan calls for real code on #4 and #5, treats #1 as a bounded research pass with a hard anti-fabrication rule, treats #2 as pure monitoring (no code), and treats #3 as a closed decision (optionally reworded in CLAUDE.md so it stops reading like an open TODO).
Verified against the current repo before writing this:
src/leaderboard.py/config/leaderboards.yaml— loader, validator,--checkcoverage report all exist and work; only the data is missing.src/dispatcher.py:1682session_directory()scans onlym["content"]text viaCWD_RE, with no use ofm["role"]ortool_calls— that's exactly why a read-heavy dependency directory can outweigh the directory actually being edited.src/metrics.py:quota_burn()sumsenergy_kwhover all ofenergy_observationswith noproviderfilter, againstcfg.objective.plan_kwh_per_period(the NeuralWatt plan allowance) — confirmed by reading it directly. Dropping local rows into that table under a sentinelprovidervalue would silently inflate the reported NeuralWatt quota burn with the user's own electricity draw, and every other unaudited aggregate over that table (per-model usage, admin history charts, CSV exports) would carry the same risk going forward — exactly the class of silent-conflation bug this project has been bitten by more than once. Local energy gets its own table, not a sentinel value in the existing one — see Phase 2.dispatcher.gross_energy_kwh(avg_power_watts, duration_seconds)already implementswatts * seconds / 3_600_000— reusable as-is for local energy math.- The reference deployment has
nvidia-smi(Quadro RTX 6000), and the liveconfig/config.yamlpointsclassifier.base_url,verification.base_url, and the vision fallback all atlocalhost:11434— so in-process GPU sampling is valid for this deployment. The VPN/remote-Ollama case CLAUDE.md documents as normal is out of scope for v1 and must fail visibly, not silently mismeasure.
Phase 1 — Fix session-directory attribution (item #4)
File: src/dispatcher.py (session_directory, ~line 1682), plus
CWD_RE if tool-call arguments need scanning.
Tests: tests/test_session_identity.py.
Today session_directory() builds an unweighted histogram of every absolute
path mentioned anywhere in the last 40 messages' content text, regardless
of role. The real observed failure: a session that read a dependency's
source 24 times (inside a venv) and wrote to the actual project 22 times
resolved to the dependency directory, because reads and writes count the
same.
Fix: weight path mentions by what kind of message they came from, using
role (already present on every message, currently unused by this
function):
role == "tool"/ tool-result content (reads: file dumps, search/listing output) → weight 1.role == "assistant"mentions that come from atool_callsentry whosefunction.namematches a write-like pattern (edit,write,patch,apply_patch,str_replace,create) → weight 3. This requires also scanningm.get("tool_calls", [])[*].function.argumentsfor paths, which the function does not currently look at at all.- Everything else (plain
assistant/usertext) → weight 1, unchanged.
Keep the existing "deepest directory that explains most of the weighted
mass" selection logic and the threshold = max(2, best * 0.6) tie-breaking
— only the counts feeding it change from +1 to +weight.
Add a regression case to tests/test_session_identity.py shaped like the
real incident: many read-mentions of /…/.venv/lib/…/somepkg alongside
fewer write-mentions (via a tool_calls edit call) of the actual project
dir, asserting the project dir wins. Keep all existing tests passing
unchanged — this is a weighting change, not a behavior rewrite, so the
already-covered "shared root wins", "single mention insufficient",
"filename vs directory" cases must still hold.
Phase 2 — Local energy ledger (item #5)
New file: src/local_energy.py
Changed: config/schema.sql (new local_energy_observations table),
src/config.py (new LocalEnergyConfig), config/config.yaml (new
local_energy: section, disabled by default), src/dispatcher.py (wrap
the three local-Ollama call sites), src/metrics.py (a fully separate
aggregate), admin/frontend/index.html (a distinct dashboard section).
New tests: tests/test_local_energy.py; a small addition to dispatcher
tests confirming the three call sites invoke the meter only when enabled.
Ground rule for this whole phase, per explicit user instruction: local
energy must be recorded, aggregated, and displayed SEPARATELY from
NeuralWatt usage everywhere — DB, /metrics, and the admin portal. Not a
provider tag inside the existing cloud table; a structurally separate
table and a structurally separate UI section, so nothing can silently sum
the two together.
Metering
src/local_energy.py, mirroring the existing circuit_breaker.py /
session_cache.py shape (injectable dependencies for testability, no
import of dispatcher/config types beyond what's passed in):
- A sampler function that shells out to
nvidia-smi --query-gpu=power.draw --format=csv,noheader,nounits(no new pinned dependency —pynvmlis not inrequirements.txtand this is a one-line CSV parse). Injectable so tests can stub a fixed wattage sequence instead of touching real hardware. - A context manager / helper,
measure(sample_interval_seconds), that starts a background thread polling the sampler everysample_interval_secondswhile the wrapped call runs, then returns(avg_power_watts, duration_seconds)on exit. Averaging over the call (not a single start/end sample) matters because classify calls are short (~1-2s) and power ramps. energy_kwhcomes from the existingdispatcher.gross_energy_kwh— no reimplementation.cost_usd = energy_kwh * cfg.local_energy.tariff_usd_per_kwh.carbon_g_co2eq = energy_kwh * 1000 * cfg.local_energy.grid_intensity_g_per_kwhwhen that config value is set, elseNone— absent tariff/grid data should degrade to "not reported," matching this project's existing rule that measurements admit on absent evidence rather than fail closed (unlike the capability flags inrouting.py).
Config
New LocalEnergyConfig(StrictModel) in src/config.py:
enabled: bool = False # off by default — depends on nvidia-smi being present
meter: Literal["nvidia_smi"] = "nvidia_smi" # room for rapl/smart_plug later, only nvidia_smi built now
sample_interval_seconds: float = 0.25
tariff_usd_per_kwh: Optional[float] = None # required when enabled
grid_intensity_g_per_kwh: Optional[float] = None # optional; carbon omitted when unset
Config load refuses enabled: true with tariff_usd_per_kwh: null —
same pattern as verification.model refusing null once classifier/verifier
hostnames differ: a silently-unpriced meter is worse than an explicit
startup error. Add the section to config/config.yaml with the same
heavily-commented style as the rest of the file, defaulted OFF, with a
comment telling the user to set their real per-kWh rate before turning it
on (no invented number — same anti-fabrication rule as leaderboards).
Remote-Ollama safety valve: at startup, if local_energy.enabled and
any of classifier.base_url / verification.base_url / the vision
fallback's base URL resolve to a non-loopback hostname, log a clear warning
and skip metering for that call site (nvidia-smi run from the dispatcher
process would measure the wrong machine). Don't silently report a bogus
number.
Storage — a dedicated table, not a tag in the cloud one
config/schema.sql:
CREATE TABLE IF NOT EXISTS local_energy_observations (
id INTEGER PRIMARY KEY AUTOINCREMENT,
model_id TEXT NOT NULL, -- the local Ollama tag, e.g. mistral-nemo-router:12b
call_type TEXT NOT NULL, -- 'classify' | 'verify' | 'local_vision'
avg_power_watts REAL,
duration_seconds REAL,
energy_kwh REAL,
cost_usd REAL, -- energy_kwh * tariff_usd_per_kwh; NULL if tariff unset
carbon_g_co2eq REAL, -- NULL unless grid_intensity_g_per_kwh is configured
meter TEXT, -- 'nvidia_smi' (room for rapl/smart_plug later)
observed_at TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_local_energy_model ON local_energy_observations (model_id);
Same code-side migration pattern already used for route_decisions
(_ensure_route_decisions_table, created at module load and again inside
the write path so a live router.db predating this table gets it without
a manual sqlite3 router.db < schema.sql): a
local_energy.ensure_local_energy_table(conn) called from both module
load and log_local_energy().
Nothing about local energy ever touches energy_observations. That's
deliberate, not an oversight to be filtered around later: quota_burn()
sums that table unconditionally against the NeuralWatt plan allowance, and
a dedicated table makes it structurally impossible for local electricity to
leak into that number, instead of relying on every current and future
query remembering a provider != 'ollama-local' filter.
Wiring
Wrap the three local-Ollama call sites with the meter, each logging a row
via log_local_energy(model_id, call_type, avg_power_watts, duration_seconds, energy_kwh, cost_usd, carbon_g_co2eq, meter)
in dispatcher.py (or local_energy.py — small enough either way, but
keep it out of log_observation(), whose parameter shape — request_id,
prompt_tokens, cloud Telemetry — doesn't map onto a classify/verify/
vision call and shouldn't be stretched to cover both):
classify()(~dispatcher.py:457) →call_type='classify'- the local LLM verification check (~
dispatcher.py:1208) →call_type='verify' _run_local_vision()(~dispatcher.py:1969) →call_type='local_vision'
Surfacing the number — DB, /metrics, and admin portal all separate
src/metrics.py: a new, standalone local_energy_summary(conn, cfg) that
queries only local_energy_observations — total kWh, total $, and a
breakdown by call_type, over the same 30-day window quota_burn() uses.
It must not join or union with energy_observations in any form. /metrics
gets a new top-level "local_energy" key, structurally parallel to but
independent from "quota" and "per_model" — never merged into either.
admin/frontend/index.html: a distinct "Local compute" card, placed
separately from the existing per-model NeuralWatt usage table and the quota
chip — not a row inside either of them. Shows total local kWh/$ and the
classify/verify/vision breakdown, sourced from the new /metrics key. This
is a required part of this phase per the user's instruction, not a
follow-up.
Phase 3 — Leaderboard priors (item #1) — research pass, not a build
No code changes; leaderboard.py and its --check coverage report already
work. The task is: for each active model family currently missing a prior
(kimi-k3, kimi-k2.7-code, kimi-k3-fast/-flex inherit from kimi-k3,
qwen3.6-35b, gemma-4-31b, deepseek-v4-flash, glm-5.2), look for real,
citable published benchmark numbers (LMArena, LiveBench, Aider polyglot,
BigCodeBench, or the vendor's own model card), and only add an entry to
config/leaderboards.yaml when a real source is found — with the source:
field filled in, matching the file's own documented shape.
Hard rule carried over from the file's own header comment: a family with no locatable real benchmark stays out of the file entirely (as it is today) rather than getting an invented plausible-sounding number. Given how recent these model names are, it's likely some families simply have no public benchmark yet — that's an acceptable outcome to report, not a gap to paper over.
After edits: python leaderboard.py --check to confirm coverage improved,
then python leaderboard.py to apply and re-blend.
Phase 4 — Sampling depth (item #2) — monitoring only, no code
kimi-k2.7-code-fast, kimi-k3, and glm-5.2-flex are still unstable on
split-half agreement, but this only affects eco (not an objective) and
the fix is coverage across time, which llm-router-seed.timer already
provides now that its log_observation keyword-arg bug (6e729ad) is fixed.
No new code. Action: let the timer accumulate for a couple of weeks, then
re-run the same split-half comparison CLAUDE.md already describes (median
of the first half of each model's seed_reference rows vs. the second
half) against the three unstable models. If it settles, done. If it
doesn't, that's the signal it's serving-condition noise rather than a
sampling-depth problem, which is a different (and lower-priority) question.
Phase 5 — Streaming retry (item #3) — confirm as closed, not build
CLAUDE.md already reasons through this and lands on "don't build it" —
buffering to enable retry on streamed responses would cost streaming
itself, and POST /outcome already covers the same traffic after the fact
with no such trade-off. No code change proposed. Optional: reword the
CLAUDE.md bullet from "not built yet" framing to state plainly that this is
a deliberate non-goal, so it stops reading like an open TODO next to items
that actually are.
Verification
- Phase 1:
pytest tests/test_session_identity.py -v— new regression case plus all existing cases green. - Phase 2:
pytest tests/test_local_energy.py -vwith the sampler stubbed; then a manual live check — enablelocal_energyin a localconfig.yaml, hit/routea few times, confirm rows land in the newlocal_energy_observationstable (and confirmenergy_observationsandquota_burn()'s reported figure are completely unchanged by this — that's the property the separate table exists to guarantee). Confirm/metricsreports a nonzerolocal_energytotal and that it's a separate key, not folded intoquotaorper_model. Load/adminand confirm the "Local compute" card renders distinctly from the NeuralWatt usage table. Confirm the remote-host safety valve by pointingclassifier.base_urlat a non-loopback host and checking metering is skipped with a log line, not a bogus number. - Phase 3:
python leaderboard.py --checkbefore/after to show coverage moved, and cite the source used for each new entry in the commit message. - Full suite:
pytest(should stay at 100% pass, count grows by the new test files).