# Local energy ledger and session-directory attribution Status: done -- local_energy.py Spec covering the five "not built yet" items in `CLAUDE.md`. Real implementation work here is scoped to two items (session-directory attribution, local energy ledger); the other three are a research pass and two already-decided non-builds, included for completeness so nothing on the open-items list gets lost. ## Context CLAUDE.md's "What's NOT built yet" section lists five open items. They are not equally sized or equally actionable, and treating them as one uniform backlog would be wrong: - **#5 (local energy not on the ledger)** matters more than its "low priority" position suggests — the user's own Ollama classifier/verifier runs on hardware they personally pay the electric bill for, so this is a real number they currently can't see, not an abstract accounting gap. - **#4 (session-directory attribution)** is a concrete, scoped bug in one pure function with existing test coverage to extend. - **#1 (leaderboard priors)** is explicitly *not* something to code around — `leaderboards.yaml` ships empty on purpose because fabricating scores would put fake data straight into routing. The tooling (`leaderboard.py`) is already done; what's missing is real, sourced numbers, which is a research task with a "only write what's verifiably cited" constraint, not an implementation task. - **#2 (sampling depth for 3 models)** no longer touches cost (only `eco`, which isn't an objective) and the tooling to fix it (`seed_energy.py --models ...`, the now-fixed `llm-router-seed.timer`) already exists. This is "let it accumulate and check back," not new code. - **#3 (retry doesn't reach streaming)** CLAUDE.md already concludes buffering to fix it "would cost streaming itself, a worse trade for interactive work," and that `POST /outcome` is the accepted answer. This is a decision already made, not a gap waiting on code. So this plan calls for real code on #4 and #5, treats #1 as a bounded research pass with a hard anti-fabrication rule, treats #2 as pure monitoring (no code), and treats #3 as a closed decision (optionally reworded in CLAUDE.md so it stops reading like an open TODO). Verified against the current repo before writing this: - `src/leaderboard.py` / `config/leaderboards.yaml` — loader, validator, `--check` coverage report all exist and work; only the data is missing. - `src/dispatcher.py:1682` `session_directory()` scans only `m["content"]` text via `CWD_RE`, with no use of `m["role"]` or `tool_calls` — that's exactly why a read-heavy dependency directory can outweigh the directory actually being edited. - `src/metrics.py:quota_burn()` sums `energy_kwh` over **all** of `energy_observations` with no `provider` filter, against `cfg.objective.plan_kwh_per_period` (the NeuralWatt plan allowance) — confirmed by reading it directly. Dropping local rows into that table under a sentinel `provider` value would silently inflate the reported NeuralWatt quota burn with the user's own electricity draw, and every other unaudited aggregate over that table (per-model usage, admin history charts, CSV exports) would carry the same risk going forward — exactly the class of silent-conflation bug this project has been bitten by more than once. **Local energy gets its own table**, not a sentinel value in the existing one — see Phase 2. - `dispatcher.gross_energy_kwh(avg_power_watts, duration_seconds)` already implements `watts * seconds / 3_600_000` — reusable as-is for local energy math. - The reference deployment has `nvidia-smi` (`Quadro RTX 6000`), and the live `config/config.yaml` points `classifier.base_url`, `verification.base_url`, and the vision fallback all at `localhost:11434` — so in-process GPU sampling is valid for this deployment. The VPN/remote-Ollama case CLAUDE.md documents as normal is out of scope for v1 and must fail visibly, not silently mismeasure. --- ## Phase 1 — Fix session-directory attribution (item #4) **File:** `src/dispatcher.py` (`session_directory`, ~line 1682), plus `CWD_RE` if tool-call arguments need scanning. **Tests:** `tests/test_session_identity.py`. Today `session_directory()` builds an unweighted histogram of every absolute path mentioned anywhere in the last 40 messages' `content` text, regardless of `role`. The real observed failure: a session that read a dependency's source 24 times (inside a venv) and wrote to the actual project 22 times resolved to the dependency directory, because reads and writes count the same. Fix: weight path mentions by what kind of message they came from, using `role` (already present on every message, currently unused by this function): - `role == "tool"` / tool-result content (reads: file dumps, search/listing output) → weight 1. - `role == "assistant"` mentions that come from a `tool_calls` entry whose `function.name` matches a write-like pattern (`edit`, `write`, `patch`, `apply_patch`, `str_replace`, `create`) → weight 3. This requires also scanning `m.get("tool_calls", [])[*].function.arguments` for paths, which the function does not currently look at at all. - Everything else (plain `assistant`/`user` text) → weight 1, unchanged. Keep the existing "deepest directory that explains most of the weighted mass" selection logic and the `threshold = max(2, best * 0.6)` tie-breaking — only the counts feeding it change from `+1` to `+weight`. Add a regression case to `tests/test_session_identity.py` shaped like the real incident: many read-mentions of `/…/.venv/lib/…/somepkg` alongside fewer write-mentions (via a `tool_calls` edit call) of the actual project dir, asserting the project dir wins. Keep all existing tests passing unchanged — this is a weighting change, not a behavior rewrite, so the already-covered "shared root wins", "single mention insufficient", "filename vs directory" cases must still hold. --- ## Phase 2 — Local energy ledger (item #5) **New file:** `src/local_energy.py` **Changed:** `config/schema.sql` (new `local_energy_observations` table), `src/config.py` (new `LocalEnergyConfig`), `config/config.yaml` (new `local_energy:` section, disabled by default), `src/dispatcher.py` (wrap the three local-Ollama call sites), `src/metrics.py` (a fully separate aggregate), `admin/frontend/index.html` (a distinct dashboard section). **New tests:** `tests/test_local_energy.py`; a small addition to dispatcher tests confirming the three call sites invoke the meter only when enabled. **Ground rule for this whole phase, per explicit user instruction: local energy must be recorded, aggregated, and displayed SEPARATELY from NeuralWatt usage everywhere — DB, `/metrics`, and the admin portal. Not a `provider` tag inside the existing cloud table; a structurally separate table and a structurally separate UI section, so nothing can silently sum the two together.** ### Metering `src/local_energy.py`, mirroring the existing `circuit_breaker.py` / `session_cache.py` shape (injectable dependencies for testability, no import of `dispatcher`/`config` types beyond what's passed in): - A sampler function that shells out to `nvidia-smi --query-gpu=power.draw --format=csv,noheader,nounits` (no new pinned dependency — `pynvml` is not in `requirements.txt` and this is a one-line CSV parse). Injectable so tests can stub a fixed wattage sequence instead of touching real hardware. - A context manager / helper, `measure(sample_interval_seconds)`, that starts a background thread polling the sampler every `sample_interval_seconds` while the wrapped call runs, then returns `(avg_power_watts, duration_seconds)` on exit. Averaging over the call (not a single start/end sample) matters because classify calls are short (~1-2s) and power ramps. - `energy_kwh` comes from the **existing** `dispatcher.gross_energy_kwh` — no reimplementation. - `cost_usd = energy_kwh * cfg.local_energy.tariff_usd_per_kwh`. - `carbon_g_co2eq = energy_kwh * 1000 * cfg.local_energy.grid_intensity_g_per_kwh` when that config value is set, else `None` — absent tariff/grid data should degrade to "not reported," matching this project's existing rule that measurements admit on absent evidence rather than fail closed (unlike the capability flags in `routing.py`). ### Config New `LocalEnergyConfig(StrictModel)` in `src/config.py`: ``` enabled: bool = False # off by default — depends on nvidia-smi being present meter: Literal["nvidia_smi"] = "nvidia_smi" # room for rapl/smart_plug later, only nvidia_smi built now sample_interval_seconds: float = 0.25 tariff_usd_per_kwh: Optional[float] = None # required when enabled grid_intensity_g_per_kwh: Optional[float] = None # optional; carbon omitted when unset ``` Config load **refuses** `enabled: true` with `tariff_usd_per_kwh: null` — same pattern as `verification.model` refusing null once classifier/verifier hostnames differ: a silently-unpriced meter is worse than an explicit startup error. Add the section to `config/config.yaml` with the same heavily-commented style as the rest of the file, defaulted OFF, with a comment telling the user to set their real per-kWh rate before turning it on (no invented number — same anti-fabrication rule as leaderboards). **Remote-Ollama safety valve:** at startup, if `local_energy.enabled` and any of `classifier.base_url` / `verification.base_url` / the vision fallback's base URL resolve to a non-loopback hostname, log a clear warning and skip metering for that call site (nvidia-smi run from the dispatcher process would measure the wrong machine). Don't silently report a bogus number. ### Storage — a dedicated table, not a tag in the cloud one `config/schema.sql`: ```sql CREATE TABLE IF NOT EXISTS local_energy_observations ( id INTEGER PRIMARY KEY AUTOINCREMENT, model_id TEXT NOT NULL, -- the local Ollama tag, e.g. mistral-nemo-router:12b call_type TEXT NOT NULL, -- 'classify' | 'verify' | 'local_vision' avg_power_watts REAL, duration_seconds REAL, energy_kwh REAL, cost_usd REAL, -- energy_kwh * tariff_usd_per_kwh; NULL if tariff unset carbon_g_co2eq REAL, -- NULL unless grid_intensity_g_per_kwh is configured meter TEXT, -- 'nvidia_smi' (room for rapl/smart_plug later) observed_at TEXT NOT NULL ); CREATE INDEX IF NOT EXISTS idx_local_energy_model ON local_energy_observations (model_id); ``` Same code-side migration pattern already used for `route_decisions` (`_ensure_route_decisions_table`, created at module load and again inside the write path so a live `router.db` predating this table gets it without a manual `sqlite3 router.db < schema.sql`): a `local_energy.ensure_local_energy_table(conn)` called from both module load and `log_local_energy()`. **Nothing about local energy ever touches `energy_observations`.** That's deliberate, not an oversight to be filtered around later: `quota_burn()` sums that table unconditionally against the NeuralWatt plan allowance, and a dedicated table makes it structurally impossible for local electricity to leak into that number, instead of relying on every current and future query remembering a `provider != 'ollama-local'` filter. ### Wiring Wrap the three local-Ollama call sites with the meter, each logging a row via `log_local_energy(model_id, call_type, avg_power_watts, duration_seconds, energy_kwh, cost_usd, carbon_g_co2eq, meter)` in `dispatcher.py` (or `local_energy.py` — small enough either way, but keep it out of `log_observation()`, whose parameter shape — `request_id`, `prompt_tokens`, cloud `Telemetry` — doesn't map onto a classify/verify/ vision call and shouldn't be stretched to cover both): 1. `classify()` (~`dispatcher.py:457`) → `call_type='classify'` 2. the local LLM verification check (~`dispatcher.py:1208`) → `call_type='verify'` 3. `_run_local_vision()` (~`dispatcher.py:1969`) → `call_type='local_vision'` ### Surfacing the number — DB, `/metrics`, and admin portal all separate `src/metrics.py`: a new, standalone `local_energy_summary(conn, cfg)` that queries **only** `local_energy_observations` — total kWh, total \$, and a breakdown by `call_type`, over the same 30-day window `quota_burn()` uses. It must not join or union with `energy_observations` in any form. `/metrics` gets a new top-level `"local_energy"` key, structurally parallel to but independent from `"quota"` and `"per_model"` — never merged into either. `admin/frontend/index.html`: a distinct **"Local compute"** card, placed separately from the existing per-model NeuralWatt usage table and the quota chip — not a row inside either of them. Shows total local kWh/\$ and the classify/verify/vision breakdown, sourced from the new `/metrics` key. This is a required part of this phase per the user's instruction, not a follow-up. --- ## Phase 3 — Leaderboard priors (item #1) — research pass, not a build No code changes; `leaderboard.py` and its `--check` coverage report already work. The task is: for each active model family currently missing a prior (`kimi-k3`, `kimi-k2.7-code`, `kimi-k3-fast`/`-flex` inherit from `kimi-k3`, `qwen3.6-35b`, `gemma-4-31b`, `deepseek-v4-flash`, `glm-5.2`), look for real, citable published benchmark numbers (LMArena, LiveBench, Aider polyglot, BigCodeBench, or the vendor's own model card), and only add an entry to `config/leaderboards.yaml` when a real source is found — with the `source:` field filled in, matching the file's own documented shape. **Hard rule carried over from the file's own header comment:** a family with no locatable real benchmark stays out of the file entirely (as it is today) rather than getting an invented plausible-sounding number. Given how recent these model names are, it's likely some families simply have no public benchmark yet — that's an acceptable outcome to report, not a gap to paper over. After edits: `python leaderboard.py --check` to confirm coverage improved, then `python leaderboard.py` to apply and re-blend. --- ## Phase 4 — Sampling depth (item #2) — monitoring only, no code `kimi-k2.7-code-fast`, `kimi-k3`, and `glm-5.2-flex` are still unstable on split-half agreement, but this only affects `eco` (not an objective) and the fix is coverage across time, which `llm-router-seed.timer` already provides now that its `log_observation` keyword-arg bug (6e729ad) is fixed. No new code. Action: let the timer accumulate for a couple of weeks, then re-run the same split-half comparison CLAUDE.md already describes (median of the first half of each model's `seed_reference` rows vs. the second half) against the three unstable models. If it settles, done. If it doesn't, that's the signal it's serving-condition noise rather than a sampling-depth problem, which is a different (and lower-priority) question. --- ## Phase 5 — Streaming retry (item #3) — confirm as closed, not build CLAUDE.md already reasons through this and lands on "don't build it" — buffering to enable retry on streamed responses would cost streaming itself, and `POST /outcome` already covers the same traffic after the fact with no such trade-off. No code change proposed. Optional: reword the CLAUDE.md bullet from "not built yet" framing to state plainly that this is a deliberate non-goal, so it stops reading like an open TODO next to items that actually are. --- ## Verification - Phase 1: `pytest tests/test_session_identity.py -v` — new regression case plus all existing cases green. - Phase 2: `pytest tests/test_local_energy.py -v` with the sampler stubbed; then a manual live check — enable `local_energy` in a local `config.yaml`, hit `/route` a few times, confirm rows land in the new `local_energy_observations` table (and confirm `energy_observations` and `quota_burn()`'s reported figure are completely unchanged by this — that's the property the separate table exists to guarantee). Confirm `/metrics` reports a nonzero `local_energy` total and that it's a separate key, not folded into `quota` or `per_model`. Load `/admin` and confirm the "Local compute" card renders distinctly from the NeuralWatt usage table. Confirm the remote-host safety valve by pointing `classifier.base_url` at a non-loopback host and checking metering is skipped with a log line, not a bogus number. - Phase 3: `python leaderboard.py --check` before/after to show coverage moved, and cite the source used for each new entry in the commit message. - Full suite: `pytest` (should stay at 100% pass, count grows by the new test files).