plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
321 lines
16 KiB
Markdown
321 lines
16 KiB
Markdown
# Local energy ledger and session-directory attribution
|
|
|
|
Status: done -- local_energy.py
|
|
|
|
Spec covering the five "not built yet" items in `CLAUDE.md`. Real
|
|
implementation work here is scoped to two items (session-directory
|
|
attribution, local energy ledger); the other three are a research pass and
|
|
two already-decided non-builds, included for completeness so nothing on the
|
|
open-items list gets lost.
|
|
|
|
## Context
|
|
|
|
CLAUDE.md's "What's NOT built yet" section lists five open items. They are
|
|
not equally sized or equally actionable, and treating them as one uniform
|
|
backlog would be wrong:
|
|
|
|
- **#5 (local energy not on the ledger)** matters more than its "low
|
|
priority" position suggests — the user's own Ollama classifier/verifier
|
|
runs on hardware they personally pay the electric bill for, so this is a
|
|
real number they currently can't see, not an abstract accounting gap.
|
|
- **#4 (session-directory attribution)** is a concrete, scoped bug in one
|
|
pure function with existing test coverage to extend.
|
|
- **#1 (leaderboard priors)** is explicitly *not* something to code around —
|
|
`leaderboards.yaml` ships empty on purpose because fabricating scores would
|
|
put fake data straight into routing. The tooling (`leaderboard.py`) is
|
|
already done; what's missing is real, sourced numbers, which is a research
|
|
task with a "only write what's verifiably cited" constraint, not an
|
|
implementation task.
|
|
- **#2 (sampling depth for 3 models)** no longer touches cost (only `eco`,
|
|
which isn't an objective) and the tooling to fix it (`seed_energy.py
|
|
--models ...`, the now-fixed `llm-router-seed.timer`) already exists. This
|
|
is "let it accumulate and check back," not new code.
|
|
- **#3 (retry doesn't reach streaming)** CLAUDE.md already concludes
|
|
buffering to fix it "would cost streaming itself, a worse trade for
|
|
interactive work," and that `POST /outcome` is the accepted answer. This
|
|
is a decision already made, not a gap waiting on code.
|
|
|
|
So this plan calls for real code on #4 and #5, treats #1 as a bounded
|
|
research pass with a hard anti-fabrication rule, treats #2 as pure
|
|
monitoring (no code), and treats #3 as a closed decision (optionally
|
|
reworded in CLAUDE.md so it stops reading like an open TODO).
|
|
|
|
Verified against the current repo before writing this:
|
|
- `src/leaderboard.py` / `config/leaderboards.yaml` — loader, validator,
|
|
`--check` coverage report all exist and work; only the data is missing.
|
|
- `src/dispatcher.py:1682` `session_directory()` scans only `m["content"]`
|
|
text via `CWD_RE`, with no use of `m["role"]` or `tool_calls` — that's
|
|
exactly why a read-heavy dependency directory can outweigh the directory
|
|
actually being edited.
|
|
- `src/metrics.py:quota_burn()` sums `energy_kwh` over **all** of
|
|
`energy_observations` with no `provider` filter, against
|
|
`cfg.objective.plan_kwh_per_period` (the NeuralWatt plan allowance) —
|
|
confirmed by reading it directly. Dropping local rows into that table
|
|
under a sentinel `provider` value would silently inflate the reported
|
|
NeuralWatt quota burn with the user's own electricity draw, and every
|
|
other unaudited aggregate over that table (per-model usage, admin history
|
|
charts, CSV exports) would carry the same risk going forward — exactly
|
|
the class of silent-conflation bug this project has been bitten by more
|
|
than once. **Local energy gets its own table**, not a sentinel value in
|
|
the existing one — see Phase 2.
|
|
- `dispatcher.gross_energy_kwh(avg_power_watts, duration_seconds)` already
|
|
implements `watts * seconds / 3_600_000` — reusable as-is for local energy
|
|
math.
|
|
- The reference deployment has `nvidia-smi` (`Quadro RTX 6000`), and the
|
|
live `config/config.yaml` points `classifier.base_url`,
|
|
`verification.base_url`, and the vision fallback all at
|
|
`localhost:11434` — so in-process GPU sampling is valid for this
|
|
deployment. The VPN/remote-Ollama case CLAUDE.md documents as normal is
|
|
out of scope for v1 and must fail visibly, not silently mismeasure.
|
|
|
|
---
|
|
|
|
## Phase 1 — Fix session-directory attribution (item #4)
|
|
|
|
**File:** `src/dispatcher.py` (`session_directory`, ~line 1682), plus
|
|
`CWD_RE` if tool-call arguments need scanning.
|
|
**Tests:** `tests/test_session_identity.py`.
|
|
|
|
Today `session_directory()` builds an unweighted histogram of every absolute
|
|
path mentioned anywhere in the last 40 messages' `content` text, regardless
|
|
of `role`. The real observed failure: a session that read a dependency's
|
|
source 24 times (inside a venv) and wrote to the actual project 22 times
|
|
resolved to the dependency directory, because reads and writes count the
|
|
same.
|
|
|
|
Fix: weight path mentions by what kind of message they came from, using
|
|
`role` (already present on every message, currently unused by this
|
|
function):
|
|
|
|
- `role == "tool"` / tool-result content (reads: file dumps, search/listing
|
|
output) → weight 1.
|
|
- `role == "assistant"` mentions that come from a `tool_calls` entry whose
|
|
`function.name` matches a write-like pattern (`edit`, `write`, `patch`,
|
|
`apply_patch`, `str_replace`, `create`) → weight 3. This requires also
|
|
scanning `m.get("tool_calls", [])[*].function.arguments` for paths, which
|
|
the function does not currently look at at all.
|
|
- Everything else (plain `assistant`/`user` text) → weight 1, unchanged.
|
|
|
|
Keep the existing "deepest directory that explains most of the weighted
|
|
mass" selection logic and the `threshold = max(2, best * 0.6)` tie-breaking
|
|
— only the counts feeding it change from `+1` to `+weight`.
|
|
|
|
Add a regression case to `tests/test_session_identity.py` shaped like the
|
|
real incident: many read-mentions of `/…/.venv/lib/…/somepkg` alongside
|
|
fewer write-mentions (via a `tool_calls` edit call) of the actual project
|
|
dir, asserting the project dir wins. Keep all existing tests passing
|
|
unchanged — this is a weighting change, not a behavior rewrite, so the
|
|
already-covered "shared root wins", "single mention insufficient",
|
|
"filename vs directory" cases must still hold.
|
|
|
|
---
|
|
|
|
## Phase 2 — Local energy ledger (item #5)
|
|
|
|
**New file:** `src/local_energy.py`
|
|
**Changed:** `config/schema.sql` (new `local_energy_observations` table),
|
|
`src/config.py` (new `LocalEnergyConfig`), `config/config.yaml` (new
|
|
`local_energy:` section, disabled by default), `src/dispatcher.py` (wrap
|
|
the three local-Ollama call sites), `src/metrics.py` (a fully separate
|
|
aggregate), `admin/frontend/index.html` (a distinct dashboard section).
|
|
**New tests:** `tests/test_local_energy.py`; a small addition to dispatcher
|
|
tests confirming the three call sites invoke the meter only when enabled.
|
|
|
|
**Ground rule for this whole phase, per explicit user instruction: local
|
|
energy must be recorded, aggregated, and displayed SEPARATELY from
|
|
NeuralWatt usage everywhere — DB, `/metrics`, and the admin portal. Not a
|
|
`provider` tag inside the existing cloud table; a structurally separate
|
|
table and a structurally separate UI section, so nothing can silently sum
|
|
the two together.**
|
|
|
|
### Metering
|
|
|
|
`src/local_energy.py`, mirroring the existing `circuit_breaker.py` /
|
|
`session_cache.py` shape (injectable dependencies for testability, no
|
|
import of `dispatcher`/`config` types beyond what's passed in):
|
|
|
|
- A sampler function that shells out to
|
|
`nvidia-smi --query-gpu=power.draw --format=csv,noheader,nounits` (no new
|
|
pinned dependency — `pynvml` is not in `requirements.txt` and this is a
|
|
one-line CSV parse). Injectable so tests can stub a fixed wattage
|
|
sequence instead of touching real hardware.
|
|
- A context manager / helper, `measure(sample_interval_seconds)`, that
|
|
starts a background thread polling the sampler every
|
|
`sample_interval_seconds` while the wrapped call runs, then returns
|
|
`(avg_power_watts, duration_seconds)` on exit. Averaging over the call
|
|
(not a single start/end sample) matters because classify calls are short
|
|
(~1-2s) and power ramps.
|
|
- `energy_kwh` comes from the **existing** `dispatcher.gross_energy_kwh` —
|
|
no reimplementation.
|
|
- `cost_usd = energy_kwh * cfg.local_energy.tariff_usd_per_kwh`.
|
|
- `carbon_g_co2eq = energy_kwh * 1000 * cfg.local_energy.grid_intensity_g_per_kwh`
|
|
when that config value is set, else `None` — absent tariff/grid data
|
|
should degrade to "not reported," matching this project's existing rule
|
|
that measurements admit on absent evidence rather than fail closed
|
|
(unlike the capability flags in `routing.py`).
|
|
|
|
### Config
|
|
|
|
New `LocalEnergyConfig(StrictModel)` in `src/config.py`:
|
|
|
|
```
|
|
enabled: bool = False # off by default — depends on nvidia-smi being present
|
|
meter: Literal["nvidia_smi"] = "nvidia_smi" # room for rapl/smart_plug later, only nvidia_smi built now
|
|
sample_interval_seconds: float = 0.25
|
|
tariff_usd_per_kwh: Optional[float] = None # required when enabled
|
|
grid_intensity_g_per_kwh: Optional[float] = None # optional; carbon omitted when unset
|
|
```
|
|
|
|
Config load **refuses** `enabled: true` with `tariff_usd_per_kwh: null` —
|
|
same pattern as `verification.model` refusing null once classifier/verifier
|
|
hostnames differ: a silently-unpriced meter is worse than an explicit
|
|
startup error. Add the section to `config/config.yaml` with the same
|
|
heavily-commented style as the rest of the file, defaulted OFF, with a
|
|
comment telling the user to set their real per-kWh rate before turning it
|
|
on (no invented number — same anti-fabrication rule as leaderboards).
|
|
|
|
**Remote-Ollama safety valve:** at startup, if `local_energy.enabled` and
|
|
any of `classifier.base_url` / `verification.base_url` / the vision
|
|
fallback's base URL resolve to a non-loopback hostname, log a clear warning
|
|
and skip metering for that call site (nvidia-smi run from the dispatcher
|
|
process would measure the wrong machine). Don't silently report a bogus
|
|
number.
|
|
|
|
### Storage — a dedicated table, not a tag in the cloud one
|
|
|
|
`config/schema.sql`:
|
|
|
|
```sql
|
|
CREATE TABLE IF NOT EXISTS local_energy_observations (
|
|
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
|
model_id TEXT NOT NULL, -- the local Ollama tag, e.g. mistral-nemo-router:12b
|
|
call_type TEXT NOT NULL, -- 'classify' | 'verify' | 'local_vision'
|
|
avg_power_watts REAL,
|
|
duration_seconds REAL,
|
|
energy_kwh REAL,
|
|
cost_usd REAL, -- energy_kwh * tariff_usd_per_kwh; NULL if tariff unset
|
|
carbon_g_co2eq REAL, -- NULL unless grid_intensity_g_per_kwh is configured
|
|
meter TEXT, -- 'nvidia_smi' (room for rapl/smart_plug later)
|
|
observed_at TEXT NOT NULL
|
|
);
|
|
CREATE INDEX IF NOT EXISTS idx_local_energy_model ON local_energy_observations (model_id);
|
|
```
|
|
|
|
Same code-side migration pattern already used for `route_decisions`
|
|
(`_ensure_route_decisions_table`, created at module load and again inside
|
|
the write path so a live `router.db` predating this table gets it without
|
|
a manual `sqlite3 router.db < schema.sql`): a
|
|
`local_energy.ensure_local_energy_table(conn)` called from both module
|
|
load and `log_local_energy()`.
|
|
|
|
**Nothing about local energy ever touches `energy_observations`.** That's
|
|
deliberate, not an oversight to be filtered around later: `quota_burn()`
|
|
sums that table unconditionally against the NeuralWatt plan allowance, and
|
|
a dedicated table makes it structurally impossible for local electricity to
|
|
leak into that number, instead of relying on every current and future
|
|
query remembering a `provider != 'ollama-local'` filter.
|
|
|
|
### Wiring
|
|
|
|
Wrap the three local-Ollama call sites with the meter, each logging a row
|
|
via `log_local_energy(model_id, call_type, avg_power_watts, duration_seconds, energy_kwh, cost_usd, carbon_g_co2eq, meter)`
|
|
in `dispatcher.py` (or `local_energy.py` — small enough either way, but
|
|
keep it out of `log_observation()`, whose parameter shape — `request_id`,
|
|
`prompt_tokens`, cloud `Telemetry` — doesn't map onto a classify/verify/
|
|
vision call and shouldn't be stretched to cover both):
|
|
|
|
1. `classify()` (~`dispatcher.py:457`) → `call_type='classify'`
|
|
2. the local LLM verification check (~`dispatcher.py:1208`) →
|
|
`call_type='verify'`
|
|
3. `_run_local_vision()` (~`dispatcher.py:1969`) → `call_type='local_vision'`
|
|
|
|
### Surfacing the number — DB, `/metrics`, and admin portal all separate
|
|
|
|
`src/metrics.py`: a new, standalone `local_energy_summary(conn, cfg)` that
|
|
queries **only** `local_energy_observations` — total kWh, total \$, and a
|
|
breakdown by `call_type`, over the same 30-day window `quota_burn()` uses.
|
|
It must not join or union with `energy_observations` in any form. `/metrics`
|
|
gets a new top-level `"local_energy"` key, structurally parallel to but
|
|
independent from `"quota"` and `"per_model"` — never merged into either.
|
|
|
|
`admin/frontend/index.html`: a distinct **"Local compute"** card, placed
|
|
separately from the existing per-model NeuralWatt usage table and the quota
|
|
chip — not a row inside either of them. Shows total local kWh/\$ and the
|
|
classify/verify/vision breakdown, sourced from the new `/metrics` key. This
|
|
is a required part of this phase per the user's instruction, not a
|
|
follow-up.
|
|
|
|
---
|
|
|
|
## Phase 3 — Leaderboard priors (item #1) — research pass, not a build
|
|
|
|
No code changes; `leaderboard.py` and its `--check` coverage report already
|
|
work. The task is: for each active model family currently missing a prior
|
|
(`kimi-k3`, `kimi-k2.7-code`, `kimi-k3-fast`/`-flex` inherit from `kimi-k3`,
|
|
`qwen3.6-35b`, `gemma-4-31b`, `deepseek-v4-flash`, `glm-5.2`), look for real,
|
|
citable published benchmark numbers (LMArena, LiveBench, Aider polyglot,
|
|
BigCodeBench, or the vendor's own model card), and only add an entry to
|
|
`config/leaderboards.yaml` when a real source is found — with the `source:`
|
|
field filled in, matching the file's own documented shape.
|
|
|
|
**Hard rule carried over from the file's own header comment:** a family
|
|
with no locatable real benchmark stays out of the file entirely (as it is
|
|
today) rather than getting an invented plausible-sounding number. Given how
|
|
recent these model names are, it's likely some families simply have no
|
|
public benchmark yet — that's an acceptable outcome to report, not a gap to
|
|
paper over.
|
|
|
|
After edits: `python leaderboard.py --check` to confirm coverage improved,
|
|
then `python leaderboard.py` to apply and re-blend.
|
|
|
|
---
|
|
|
|
## Phase 4 — Sampling depth (item #2) — monitoring only, no code
|
|
|
|
`kimi-k2.7-code-fast`, `kimi-k3`, and `glm-5.2-flex` are still unstable on
|
|
split-half agreement, but this only affects `eco` (not an objective) and
|
|
the fix is coverage across time, which `llm-router-seed.timer` already
|
|
provides now that its `log_observation` keyword-arg bug (6e729ad) is fixed.
|
|
|
|
No new code. Action: let the timer accumulate for a couple of weeks, then
|
|
re-run the same split-half comparison CLAUDE.md already describes (median
|
|
of the first half of each model's `seed_reference` rows vs. the second
|
|
half) against the three unstable models. If it settles, done. If it
|
|
doesn't, that's the signal it's serving-condition noise rather than a
|
|
sampling-depth problem, which is a different (and lower-priority) question.
|
|
|
|
---
|
|
|
|
## Phase 5 — Streaming retry (item #3) — confirm as closed, not build
|
|
|
|
CLAUDE.md already reasons through this and lands on "don't build it" —
|
|
buffering to enable retry on streamed responses would cost streaming
|
|
itself, and `POST /outcome` already covers the same traffic after the fact
|
|
with no such trade-off. No code change proposed. Optional: reword the
|
|
CLAUDE.md bullet from "not built yet" framing to state plainly that this is
|
|
a deliberate non-goal, so it stops reading like an open TODO next to items
|
|
that actually are.
|
|
|
|
---
|
|
|
|
## Verification
|
|
|
|
- Phase 1: `pytest tests/test_session_identity.py -v` — new regression case
|
|
plus all existing cases green.
|
|
- Phase 2: `pytest tests/test_local_energy.py -v` with the sampler stubbed;
|
|
then a manual live check — enable `local_energy` in a local
|
|
`config.yaml`, hit `/route` a few times, confirm rows land in the new
|
|
`local_energy_observations` table (and confirm `energy_observations` and
|
|
`quota_burn()`'s reported figure are completely unchanged by this —
|
|
that's the property the separate table exists to guarantee). Confirm
|
|
`/metrics` reports a nonzero `local_energy` total and that it's a
|
|
separate key, not folded into `quota` or `per_model`. Load `/admin` and
|
|
confirm the "Local compute" card renders distinctly from the NeuralWatt
|
|
usage table. Confirm the remote-host safety valve by pointing
|
|
`classifier.base_url` at a non-loopback host and checking metering is
|
|
skipped with a log line, not a bogus number.
|
|
- Phase 3: `python leaderboard.py --check` before/after to show coverage
|
|
moved, and cite the source used for each new entry in the commit message.
|
|
- Full suite: `pytest` (should stay at 100% pass, count grows by the new
|
|
test files).
|