Files
6krrt/plans/local-energy-ledger-and-session-attribution.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

321 lines
16 KiB
Markdown

# Local energy ledger and session-directory attribution
Status: done -- local_energy.py
Spec covering the five "not built yet" items in `CLAUDE.md`. Real
implementation work here is scoped to two items (session-directory
attribution, local energy ledger); the other three are a research pass and
two already-decided non-builds, included for completeness so nothing on the
open-items list gets lost.
## Context
CLAUDE.md's "What's NOT built yet" section lists five open items. They are
not equally sized or equally actionable, and treating them as one uniform
backlog would be wrong:
- **#5 (local energy not on the ledger)** matters more than its "low
priority" position suggests — the user's own Ollama classifier/verifier
runs on hardware they personally pay the electric bill for, so this is a
real number they currently can't see, not an abstract accounting gap.
- **#4 (session-directory attribution)** is a concrete, scoped bug in one
pure function with existing test coverage to extend.
- **#1 (leaderboard priors)** is explicitly *not* something to code around —
`leaderboards.yaml` ships empty on purpose because fabricating scores would
put fake data straight into routing. The tooling (`leaderboard.py`) is
already done; what's missing is real, sourced numbers, which is a research
task with a "only write what's verifiably cited" constraint, not an
implementation task.
- **#2 (sampling depth for 3 models)** no longer touches cost (only `eco`,
which isn't an objective) and the tooling to fix it (`seed_energy.py
--models ...`, the now-fixed `llm-router-seed.timer`) already exists. This
is "let it accumulate and check back," not new code.
- **#3 (retry doesn't reach streaming)** CLAUDE.md already concludes
buffering to fix it "would cost streaming itself, a worse trade for
interactive work," and that `POST /outcome` is the accepted answer. This
is a decision already made, not a gap waiting on code.
So this plan calls for real code on #4 and #5, treats #1 as a bounded
research pass with a hard anti-fabrication rule, treats #2 as pure
monitoring (no code), and treats #3 as a closed decision (optionally
reworded in CLAUDE.md so it stops reading like an open TODO).
Verified against the current repo before writing this:
- `src/leaderboard.py` / `config/leaderboards.yaml` — loader, validator,
`--check` coverage report all exist and work; only the data is missing.
- `src/dispatcher.py:1682` `session_directory()` scans only `m["content"]`
text via `CWD_RE`, with no use of `m["role"]` or `tool_calls` — that's
exactly why a read-heavy dependency directory can outweigh the directory
actually being edited.
- `src/metrics.py:quota_burn()` sums `energy_kwh` over **all** of
`energy_observations` with no `provider` filter, against
`cfg.objective.plan_kwh_per_period` (the NeuralWatt plan allowance) —
confirmed by reading it directly. Dropping local rows into that table
under a sentinel `provider` value would silently inflate the reported
NeuralWatt quota burn with the user's own electricity draw, and every
other unaudited aggregate over that table (per-model usage, admin history
charts, CSV exports) would carry the same risk going forward — exactly
the class of silent-conflation bug this project has been bitten by more
than once. **Local energy gets its own table**, not a sentinel value in
the existing one — see Phase 2.
- `dispatcher.gross_energy_kwh(avg_power_watts, duration_seconds)` already
implements `watts * seconds / 3_600_000` — reusable as-is for local energy
math.
- The reference deployment has `nvidia-smi` (`Quadro RTX 6000`), and the
live `config/config.yaml` points `classifier.base_url`,
`verification.base_url`, and the vision fallback all at
`localhost:11434` — so in-process GPU sampling is valid for this
deployment. The VPN/remote-Ollama case CLAUDE.md documents as normal is
out of scope for v1 and must fail visibly, not silently mismeasure.
---
## Phase 1 — Fix session-directory attribution (item #4)
**File:** `src/dispatcher.py` (`session_directory`, ~line 1682), plus
`CWD_RE` if tool-call arguments need scanning.
**Tests:** `tests/test_session_identity.py`.
Today `session_directory()` builds an unweighted histogram of every absolute
path mentioned anywhere in the last 40 messages' `content` text, regardless
of `role`. The real observed failure: a session that read a dependency's
source 24 times (inside a venv) and wrote to the actual project 22 times
resolved to the dependency directory, because reads and writes count the
same.
Fix: weight path mentions by what kind of message they came from, using
`role` (already present on every message, currently unused by this
function):
- `role == "tool"` / tool-result content (reads: file dumps, search/listing
output) → weight 1.
- `role == "assistant"` mentions that come from a `tool_calls` entry whose
`function.name` matches a write-like pattern (`edit`, `write`, `patch`,
`apply_patch`, `str_replace`, `create`) → weight 3. This requires also
scanning `m.get("tool_calls", [])[*].function.arguments` for paths, which
the function does not currently look at at all.
- Everything else (plain `assistant`/`user` text) → weight 1, unchanged.
Keep the existing "deepest directory that explains most of the weighted
mass" selection logic and the `threshold = max(2, best * 0.6)` tie-breaking
— only the counts feeding it change from `+1` to `+weight`.
Add a regression case to `tests/test_session_identity.py` shaped like the
real incident: many read-mentions of `/…/.venv/lib/…/somepkg` alongside
fewer write-mentions (via a `tool_calls` edit call) of the actual project
dir, asserting the project dir wins. Keep all existing tests passing
unchanged — this is a weighting change, not a behavior rewrite, so the
already-covered "shared root wins", "single mention insufficient",
"filename vs directory" cases must still hold.
---
## Phase 2 — Local energy ledger (item #5)
**New file:** `src/local_energy.py`
**Changed:** `config/schema.sql` (new `local_energy_observations` table),
`src/config.py` (new `LocalEnergyConfig`), `config/config.yaml` (new
`local_energy:` section, disabled by default), `src/dispatcher.py` (wrap
the three local-Ollama call sites), `src/metrics.py` (a fully separate
aggregate), `admin/frontend/index.html` (a distinct dashboard section).
**New tests:** `tests/test_local_energy.py`; a small addition to dispatcher
tests confirming the three call sites invoke the meter only when enabled.
**Ground rule for this whole phase, per explicit user instruction: local
energy must be recorded, aggregated, and displayed SEPARATELY from
NeuralWatt usage everywhere — DB, `/metrics`, and the admin portal. Not a
`provider` tag inside the existing cloud table; a structurally separate
table and a structurally separate UI section, so nothing can silently sum
the two together.**
### Metering
`src/local_energy.py`, mirroring the existing `circuit_breaker.py` /
`session_cache.py` shape (injectable dependencies for testability, no
import of `dispatcher`/`config` types beyond what's passed in):
- A sampler function that shells out to
`nvidia-smi --query-gpu=power.draw --format=csv,noheader,nounits` (no new
pinned dependency — `pynvml` is not in `requirements.txt` and this is a
one-line CSV parse). Injectable so tests can stub a fixed wattage
sequence instead of touching real hardware.
- A context manager / helper, `measure(sample_interval_seconds)`, that
starts a background thread polling the sampler every
`sample_interval_seconds` while the wrapped call runs, then returns
`(avg_power_watts, duration_seconds)` on exit. Averaging over the call
(not a single start/end sample) matters because classify calls are short
(~1-2s) and power ramps.
- `energy_kwh` comes from the **existing** `dispatcher.gross_energy_kwh` —
no reimplementation.
- `cost_usd = energy_kwh * cfg.local_energy.tariff_usd_per_kwh`.
- `carbon_g_co2eq = energy_kwh * 1000 * cfg.local_energy.grid_intensity_g_per_kwh`
when that config value is set, else `None` — absent tariff/grid data
should degrade to "not reported," matching this project's existing rule
that measurements admit on absent evidence rather than fail closed
(unlike the capability flags in `routing.py`).
### Config
New `LocalEnergyConfig(StrictModel)` in `src/config.py`:
```
enabled: bool = False # off by default — depends on nvidia-smi being present
meter: Literal["nvidia_smi"] = "nvidia_smi" # room for rapl/smart_plug later, only nvidia_smi built now
sample_interval_seconds: float = 0.25
tariff_usd_per_kwh: Optional[float] = None # required when enabled
grid_intensity_g_per_kwh: Optional[float] = None # optional; carbon omitted when unset
```
Config load **refuses** `enabled: true` with `tariff_usd_per_kwh: null` —
same pattern as `verification.model` refusing null once classifier/verifier
hostnames differ: a silently-unpriced meter is worse than an explicit
startup error. Add the section to `config/config.yaml` with the same
heavily-commented style as the rest of the file, defaulted OFF, with a
comment telling the user to set their real per-kWh rate before turning it
on (no invented number — same anti-fabrication rule as leaderboards).
**Remote-Ollama safety valve:** at startup, if `local_energy.enabled` and
any of `classifier.base_url` / `verification.base_url` / the vision
fallback's base URL resolve to a non-loopback hostname, log a clear warning
and skip metering for that call site (nvidia-smi run from the dispatcher
process would measure the wrong machine). Don't silently report a bogus
number.
### Storage — a dedicated table, not a tag in the cloud one
`config/schema.sql`:
```sql
CREATE TABLE IF NOT EXISTS local_energy_observations (
id INTEGER PRIMARY KEY AUTOINCREMENT,
model_id TEXT NOT NULL, -- the local Ollama tag, e.g. mistral-nemo-router:12b
call_type TEXT NOT NULL, -- 'classify' | 'verify' | 'local_vision'
avg_power_watts REAL,
duration_seconds REAL,
energy_kwh REAL,
cost_usd REAL, -- energy_kwh * tariff_usd_per_kwh; NULL if tariff unset
carbon_g_co2eq REAL, -- NULL unless grid_intensity_g_per_kwh is configured
meter TEXT, -- 'nvidia_smi' (room for rapl/smart_plug later)
observed_at TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_local_energy_model ON local_energy_observations (model_id);
```
Same code-side migration pattern already used for `route_decisions`
(`_ensure_route_decisions_table`, created at module load and again inside
the write path so a live `router.db` predating this table gets it without
a manual `sqlite3 router.db < schema.sql`): a
`local_energy.ensure_local_energy_table(conn)` called from both module
load and `log_local_energy()`.
**Nothing about local energy ever touches `energy_observations`.** That's
deliberate, not an oversight to be filtered around later: `quota_burn()`
sums that table unconditionally against the NeuralWatt plan allowance, and
a dedicated table makes it structurally impossible for local electricity to
leak into that number, instead of relying on every current and future
query remembering a `provider != 'ollama-local'` filter.
### Wiring
Wrap the three local-Ollama call sites with the meter, each logging a row
via `log_local_energy(model_id, call_type, avg_power_watts, duration_seconds, energy_kwh, cost_usd, carbon_g_co2eq, meter)`
in `dispatcher.py` (or `local_energy.py` — small enough either way, but
keep it out of `log_observation()`, whose parameter shape — `request_id`,
`prompt_tokens`, cloud `Telemetry` — doesn't map onto a classify/verify/
vision call and shouldn't be stretched to cover both):
1. `classify()` (~`dispatcher.py:457`) → `call_type='classify'`
2. the local LLM verification check (~`dispatcher.py:1208`) →
`call_type='verify'`
3. `_run_local_vision()` (~`dispatcher.py:1969`) → `call_type='local_vision'`
### Surfacing the number — DB, `/metrics`, and admin portal all separate
`src/metrics.py`: a new, standalone `local_energy_summary(conn, cfg)` that
queries **only** `local_energy_observations` — total kWh, total \$, and a
breakdown by `call_type`, over the same 30-day window `quota_burn()` uses.
It must not join or union with `energy_observations` in any form. `/metrics`
gets a new top-level `"local_energy"` key, structurally parallel to but
independent from `"quota"` and `"per_model"` — never merged into either.
`admin/frontend/index.html`: a distinct **"Local compute"** card, placed
separately from the existing per-model NeuralWatt usage table and the quota
chip — not a row inside either of them. Shows total local kWh/\$ and the
classify/verify/vision breakdown, sourced from the new `/metrics` key. This
is a required part of this phase per the user's instruction, not a
follow-up.
---
## Phase 3 — Leaderboard priors (item #1) — research pass, not a build
No code changes; `leaderboard.py` and its `--check` coverage report already
work. The task is: for each active model family currently missing a prior
(`kimi-k3`, `kimi-k2.7-code`, `kimi-k3-fast`/`-flex` inherit from `kimi-k3`,
`qwen3.6-35b`, `gemma-4-31b`, `deepseek-v4-flash`, `glm-5.2`), look for real,
citable published benchmark numbers (LMArena, LiveBench, Aider polyglot,
BigCodeBench, or the vendor's own model card), and only add an entry to
`config/leaderboards.yaml` when a real source is found — with the `source:`
field filled in, matching the file's own documented shape.
**Hard rule carried over from the file's own header comment:** a family
with no locatable real benchmark stays out of the file entirely (as it is
today) rather than getting an invented plausible-sounding number. Given how
recent these model names are, it's likely some families simply have no
public benchmark yet — that's an acceptable outcome to report, not a gap to
paper over.
After edits: `python leaderboard.py --check` to confirm coverage improved,
then `python leaderboard.py` to apply and re-blend.
---
## Phase 4 — Sampling depth (item #2) — monitoring only, no code
`kimi-k2.7-code-fast`, `kimi-k3`, and `glm-5.2-flex` are still unstable on
split-half agreement, but this only affects `eco` (not an objective) and
the fix is coverage across time, which `llm-router-seed.timer` already
provides now that its `log_observation` keyword-arg bug (6e729ad) is fixed.
No new code. Action: let the timer accumulate for a couple of weeks, then
re-run the same split-half comparison CLAUDE.md already describes (median
of the first half of each model's `seed_reference` rows vs. the second
half) against the three unstable models. If it settles, done. If it
doesn't, that's the signal it's serving-condition noise rather than a
sampling-depth problem, which is a different (and lower-priority) question.
---
## Phase 5 — Streaming retry (item #3) — confirm as closed, not build
CLAUDE.md already reasons through this and lands on "don't build it" —
buffering to enable retry on streamed responses would cost streaming
itself, and `POST /outcome` already covers the same traffic after the fact
with no such trade-off. No code change proposed. Optional: reword the
CLAUDE.md bullet from "not built yet" framing to state plainly that this is
a deliberate non-goal, so it stops reading like an open TODO next to items
that actually are.
---
## Verification
- Phase 1: `pytest tests/test_session_identity.py -v` — new regression case
plus all existing cases green.
- Phase 2: `pytest tests/test_local_energy.py -v` with the sampler stubbed;
then a manual live check — enable `local_energy` in a local
`config.yaml`, hit `/route` a few times, confirm rows land in the new
`local_energy_observations` table (and confirm `energy_observations` and
`quota_burn()`'s reported figure are completely unchanged by this —
that's the property the separate table exists to guarantee). Confirm
`/metrics` reports a nonzero `local_energy` total and that it's a
separate key, not folded into `quota` or `per_model`. Load `/admin` and
confirm the "Local compute" card renders distinctly from the NeuralWatt
usage table. Confirm the remote-host safety valve by pointing
`classifier.base_url` at a non-loopback host and checking metering is
skipped with a log line, not a bogus number.
- Phase 3: `python leaderboard.py --check` before/after to show coverage
moved, and cite the source used for each new entry in the commit message.
- Full suite: `pytest` (should stay at 100% pass, count grows by the new
test files).