Files
6krrt/plans/local-energy-ledger-and-session-attribution.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

16 KiB

Local energy ledger and session-directory attribution

Status: done -- local_energy.py

Spec covering the five "not built yet" items in CLAUDE.md. Real implementation work here is scoped to two items (session-directory attribution, local energy ledger); the other three are a research pass and two already-decided non-builds, included for completeness so nothing on the open-items list gets lost.

Context

CLAUDE.md's "What's NOT built yet" section lists five open items. They are not equally sized or equally actionable, and treating them as one uniform backlog would be wrong:

  • #5 (local energy not on the ledger) matters more than its "low priority" position suggests — the user's own Ollama classifier/verifier runs on hardware they personally pay the electric bill for, so this is a real number they currently can't see, not an abstract accounting gap.
  • #4 (session-directory attribution) is a concrete, scoped bug in one pure function with existing test coverage to extend.
  • #1 (leaderboard priors) is explicitly not something to code around — leaderboards.yaml ships empty on purpose because fabricating scores would put fake data straight into routing. The tooling (leaderboard.py) is already done; what's missing is real, sourced numbers, which is a research task with a "only write what's verifiably cited" constraint, not an implementation task.
  • #2 (sampling depth for 3 models) no longer touches cost (only eco, which isn't an objective) and the tooling to fix it (seed_energy.py --models ..., the now-fixed llm-router-seed.timer) already exists. This is "let it accumulate and check back," not new code.
  • #3 (retry doesn't reach streaming) CLAUDE.md already concludes buffering to fix it "would cost streaming itself, a worse trade for interactive work," and that POST /outcome is the accepted answer. This is a decision already made, not a gap waiting on code.

So this plan calls for real code on #4 and #5, treats #1 as a bounded research pass with a hard anti-fabrication rule, treats #2 as pure monitoring (no code), and treats #3 as a closed decision (optionally reworded in CLAUDE.md so it stops reading like an open TODO).

Verified against the current repo before writing this:

  • src/leaderboard.py / config/leaderboards.yaml — loader, validator, --check coverage report all exist and work; only the data is missing.
  • src/dispatcher.py:1682 session_directory() scans only m["content"] text via CWD_RE, with no use of m["role"] or tool_calls — that's exactly why a read-heavy dependency directory can outweigh the directory actually being edited.
  • src/metrics.py:quota_burn() sums energy_kwh over all of energy_observations with no provider filter, against cfg.objective.plan_kwh_per_period (the NeuralWatt plan allowance) — confirmed by reading it directly. Dropping local rows into that table under a sentinel provider value would silently inflate the reported NeuralWatt quota burn with the user's own electricity draw, and every other unaudited aggregate over that table (per-model usage, admin history charts, CSV exports) would carry the same risk going forward — exactly the class of silent-conflation bug this project has been bitten by more than once. Local energy gets its own table, not a sentinel value in the existing one — see Phase 2.
  • dispatcher.gross_energy_kwh(avg_power_watts, duration_seconds) already implements watts * seconds / 3_600_000 — reusable as-is for local energy math.
  • The reference deployment has nvidia-smi (Quadro RTX 6000), and the live config/config.yaml points classifier.base_url, verification.base_url, and the vision fallback all at localhost:11434 — so in-process GPU sampling is valid for this deployment. The VPN/remote-Ollama case CLAUDE.md documents as normal is out of scope for v1 and must fail visibly, not silently mismeasure.

Phase 1 — Fix session-directory attribution (item #4)

File: src/dispatcher.py (session_directory, ~line 1682), plus CWD_RE if tool-call arguments need scanning. Tests: tests/test_session_identity.py.

Today session_directory() builds an unweighted histogram of every absolute path mentioned anywhere in the last 40 messages' content text, regardless of role. The real observed failure: a session that read a dependency's source 24 times (inside a venv) and wrote to the actual project 22 times resolved to the dependency directory, because reads and writes count the same.

Fix: weight path mentions by what kind of message they came from, using role (already present on every message, currently unused by this function):

  • role == "tool" / tool-result content (reads: file dumps, search/listing output) → weight 1.
  • role == "assistant" mentions that come from a tool_calls entry whose function.name matches a write-like pattern (edit, write, patch, apply_patch, str_replace, create) → weight 3. This requires also scanning m.get("tool_calls", [])[*].function.arguments for paths, which the function does not currently look at at all.
  • Everything else (plain assistant/user text) → weight 1, unchanged.

Keep the existing "deepest directory that explains most of the weighted mass" selection logic and the threshold = max(2, best * 0.6) tie-breaking — only the counts feeding it change from +1 to +weight.

Add a regression case to tests/test_session_identity.py shaped like the real incident: many read-mentions of /…/.venv/lib/…/somepkg alongside fewer write-mentions (via a tool_calls edit call) of the actual project dir, asserting the project dir wins. Keep all existing tests passing unchanged — this is a weighting change, not a behavior rewrite, so the already-covered "shared root wins", "single mention insufficient", "filename vs directory" cases must still hold.


Phase 2 — Local energy ledger (item #5)

New file: src/local_energy.py Changed: config/schema.sql (new local_energy_observations table), src/config.py (new LocalEnergyConfig), config/config.yaml (new local_energy: section, disabled by default), src/dispatcher.py (wrap the three local-Ollama call sites), src/metrics.py (a fully separate aggregate), admin/frontend/index.html (a distinct dashboard section). New tests: tests/test_local_energy.py; a small addition to dispatcher tests confirming the three call sites invoke the meter only when enabled.

Ground rule for this whole phase, per explicit user instruction: local energy must be recorded, aggregated, and displayed SEPARATELY from NeuralWatt usage everywhere — DB, /metrics, and the admin portal. Not a provider tag inside the existing cloud table; a structurally separate table and a structurally separate UI section, so nothing can silently sum the two together.

Metering

src/local_energy.py, mirroring the existing circuit_breaker.py / session_cache.py shape (injectable dependencies for testability, no import of dispatcher/config types beyond what's passed in):

  • A sampler function that shells out to nvidia-smi --query-gpu=power.draw --format=csv,noheader,nounits (no new pinned dependency — pynvml is not in requirements.txt and this is a one-line CSV parse). Injectable so tests can stub a fixed wattage sequence instead of touching real hardware.
  • A context manager / helper, measure(sample_interval_seconds), that starts a background thread polling the sampler every sample_interval_seconds while the wrapped call runs, then returns (avg_power_watts, duration_seconds) on exit. Averaging over the call (not a single start/end sample) matters because classify calls are short (~1-2s) and power ramps.
  • energy_kwh comes from the existing dispatcher.gross_energy_kwh — no reimplementation.
  • cost_usd = energy_kwh * cfg.local_energy.tariff_usd_per_kwh.
  • carbon_g_co2eq = energy_kwh * 1000 * cfg.local_energy.grid_intensity_g_per_kwh when that config value is set, else None — absent tariff/grid data should degrade to "not reported," matching this project's existing rule that measurements admit on absent evidence rather than fail closed (unlike the capability flags in routing.py).

Config

New LocalEnergyConfig(StrictModel) in src/config.py:

enabled: bool = False                       # off by default — depends on nvidia-smi being present
meter: Literal["nvidia_smi"] = "nvidia_smi"  # room for rapl/smart_plug later, only nvidia_smi built now
sample_interval_seconds: float = 0.25
tariff_usd_per_kwh: Optional[float] = None   # required when enabled
grid_intensity_g_per_kwh: Optional[float] = None  # optional; carbon omitted when unset

Config load refuses enabled: true with tariff_usd_per_kwh: null — same pattern as verification.model refusing null once classifier/verifier hostnames differ: a silently-unpriced meter is worse than an explicit startup error. Add the section to config/config.yaml with the same heavily-commented style as the rest of the file, defaulted OFF, with a comment telling the user to set their real per-kWh rate before turning it on (no invented number — same anti-fabrication rule as leaderboards).

Remote-Ollama safety valve: at startup, if local_energy.enabled and any of classifier.base_url / verification.base_url / the vision fallback's base URL resolve to a non-loopback hostname, log a clear warning and skip metering for that call site (nvidia-smi run from the dispatcher process would measure the wrong machine). Don't silently report a bogus number.

Storage — a dedicated table, not a tag in the cloud one

config/schema.sql:

CREATE TABLE IF NOT EXISTS local_energy_observations (
    id                INTEGER PRIMARY KEY AUTOINCREMENT,
    model_id          TEXT NOT NULL,   -- the local Ollama tag, e.g. mistral-nemo-router:12b
    call_type         TEXT NOT NULL,   -- 'classify' | 'verify' | 'local_vision'
    avg_power_watts   REAL,
    duration_seconds  REAL,
    energy_kwh        REAL,
    cost_usd          REAL,            -- energy_kwh * tariff_usd_per_kwh; NULL if tariff unset
    carbon_g_co2eq    REAL,            -- NULL unless grid_intensity_g_per_kwh is configured
    meter             TEXT,            -- 'nvidia_smi' (room for rapl/smart_plug later)
    observed_at       TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_local_energy_model ON local_energy_observations (model_id);

Same code-side migration pattern already used for route_decisions (_ensure_route_decisions_table, created at module load and again inside the write path so a live router.db predating this table gets it without a manual sqlite3 router.db < schema.sql): a local_energy.ensure_local_energy_table(conn) called from both module load and log_local_energy().

Nothing about local energy ever touches energy_observations. That's deliberate, not an oversight to be filtered around later: quota_burn() sums that table unconditionally against the NeuralWatt plan allowance, and a dedicated table makes it structurally impossible for local electricity to leak into that number, instead of relying on every current and future query remembering a provider != 'ollama-local' filter.

Wiring

Wrap the three local-Ollama call sites with the meter, each logging a row via log_local_energy(model_id, call_type, avg_power_watts, duration_seconds, energy_kwh, cost_usd, carbon_g_co2eq, meter) in dispatcher.py (or local_energy.py — small enough either way, but keep it out of log_observation(), whose parameter shape — request_id, prompt_tokens, cloud Telemetry — doesn't map onto a classify/verify/ vision call and shouldn't be stretched to cover both):

  1. classify() (~dispatcher.py:457) → call_type='classify'
  2. the local LLM verification check (~dispatcher.py:1208) → call_type='verify'
  3. _run_local_vision() (~dispatcher.py:1969) → call_type='local_vision'

Surfacing the number — DB, /metrics, and admin portal all separate

src/metrics.py: a new, standalone local_energy_summary(conn, cfg) that queries only local_energy_observations — total kWh, total $, and a breakdown by call_type, over the same 30-day window quota_burn() uses. It must not join or union with energy_observations in any form. /metrics gets a new top-level "local_energy" key, structurally parallel to but independent from "quota" and "per_model" — never merged into either.

admin/frontend/index.html: a distinct "Local compute" card, placed separately from the existing per-model NeuralWatt usage table and the quota chip — not a row inside either of them. Shows total local kWh/$ and the classify/verify/vision breakdown, sourced from the new /metrics key. This is a required part of this phase per the user's instruction, not a follow-up.


Phase 3 — Leaderboard priors (item #1) — research pass, not a build

No code changes; leaderboard.py and its --check coverage report already work. The task is: for each active model family currently missing a prior (kimi-k3, kimi-k2.7-code, kimi-k3-fast/-flex inherit from kimi-k3, qwen3.6-35b, gemma-4-31b, deepseek-v4-flash, glm-5.2), look for real, citable published benchmark numbers (LMArena, LiveBench, Aider polyglot, BigCodeBench, or the vendor's own model card), and only add an entry to config/leaderboards.yaml when a real source is found — with the source: field filled in, matching the file's own documented shape.

Hard rule carried over from the file's own header comment: a family with no locatable real benchmark stays out of the file entirely (as it is today) rather than getting an invented plausible-sounding number. Given how recent these model names are, it's likely some families simply have no public benchmark yet — that's an acceptable outcome to report, not a gap to paper over.

After edits: python leaderboard.py --check to confirm coverage improved, then python leaderboard.py to apply and re-blend.


Phase 4 — Sampling depth (item #2) — monitoring only, no code

kimi-k2.7-code-fast, kimi-k3, and glm-5.2-flex are still unstable on split-half agreement, but this only affects eco (not an objective) and the fix is coverage across time, which llm-router-seed.timer already provides now that its log_observation keyword-arg bug (6e729ad) is fixed.

No new code. Action: let the timer accumulate for a couple of weeks, then re-run the same split-half comparison CLAUDE.md already describes (median of the first half of each model's seed_reference rows vs. the second half) against the three unstable models. If it settles, done. If it doesn't, that's the signal it's serving-condition noise rather than a sampling-depth problem, which is a different (and lower-priority) question.


Phase 5 — Streaming retry (item #3) — confirm as closed, not build

CLAUDE.md already reasons through this and lands on "don't build it" — buffering to enable retry on streamed responses would cost streaming itself, and POST /outcome already covers the same traffic after the fact with no such trade-off. No code change proposed. Optional: reword the CLAUDE.md bullet from "not built yet" framing to state plainly that this is a deliberate non-goal, so it stops reading like an open TODO next to items that actually are.


Verification

  • Phase 1: pytest tests/test_session_identity.py -v — new regression case plus all existing cases green.
  • Phase 2: pytest tests/test_local_energy.py -v with the sampler stubbed; then a manual live check — enable local_energy in a local config.yaml, hit /route a few times, confirm rows land in the new local_energy_observations table (and confirm energy_observations and quota_burn()'s reported figure are completely unchanged by this — that's the property the separate table exists to guarantee). Confirm /metrics reports a nonzero local_energy total and that it's a separate key, not folded into quota or per_model. Load /admin and confirm the "Local compute" card renders distinctly from the NeuralWatt usage table. Confirm the remote-host safety valve by pointing classifier.base_url at a non-loopback host and checking metering is skipped with a log line, not a bogus number.
  • Phase 3: python leaderboard.py --check before/after to show coverage moved, and cite the source used for each new entry in the commit message.
  • Full suite: pytest (should stay at 100% pass, count grows by the new test files).