plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
7.2 KiB
Quota: track balance and burn rate, not percentage of plan
Status: done -- balance/burn/runway per provider
Status: FINAL — decision-complete. Written 2026-09-04 from live data.
The modelling error
metrics.quota_burn reports usage as a fraction of plan_kwh_per_period, and
coverage.warnings renders it as:
metered usage is 146% of the 6.25 kWh plan allowance. A quota is a wall, not a bill — requests fail rather than costing more.
That claim is false for this account. Usage sailed past 100% and nothing failed, because the provider bills overage against a credit balance. The warning is simultaneously alarming and unactionable: it says a wall was hit when no wall exists, and it offers no number an operator can act on.
Meanwhile the authoritative signal — the provider reports
allowance_remaining_usd on every single response — is captured
(config/schema.sql:156, read at src/dispatcher.py:1416), persisted to
energy_observations, and then used by exactly one thing: seed_energy.py,
for sweep accounting. Nothing watches it.
What the recorded data actually shows
Queried 2026-09-04 from energy_observations:
| date | calls | balance range (USD) |
|---|---|---|
| 2026-09-04 | 1,853 | 19.70 → 12.19 |
| 2026-09-03 | 174 | 19.95 → 19.70 |
| 2026-09-02 | 3,573 | 20.01 → 0.0071 |
| 2026-09-01 | 3,116 | 14.34 → 4.34 |
The 2026-09-02 provider outage is in the table. The balance walked down through
$0.4991 → $0.4967 → $0.4924 at 05:54:26–05:54:32 and bottomed at $0.0071.
The router observed the balance approaching zero, request by request, and
said nothing. Hours of warning were available and discarded.
Current state: $12.19 remaining, burning $7.51 over 1,853 calls in 14.3 hours ≈ $0.52/hour ≈ ~23 hours of runway.
That is the sentence the dashboard should be showing. "146% of plan" is not.
Three defects, in order of value
1. The wrong signal
Warn on balance and projected time-to-zero, sourced from
allowance_remaining_usd, not on percentage of a kWh plan.
- Headline: current balance, recent burn rate, projected hours remaining.
- Warn when projected runway drops below a configurable threshold (hours, not percent). Default it to something an operator can act within — a few hours.
- Keep the kWh plan figure as secondary context. It is still the right unit for the subscription; it is the wrong thing to alarm on.
The implementation trap, and it is the whole difficulty: a top-up makes
the balance JUMP UP. On 2026-09-02 it went 0.0071 → 20.0071. A naive
MAX - MIN over a window reports a burn of ~$20 when the real burn was ~$20
down plus a $20 credit. Compute burn from consecutive decreasing deltas
only, or segment the series at every increase and use the most recent
segment. A test must cover a window containing a top-up; getting this wrong
produces a confidently wrong runway estimate, which is worse than none.
Handle allowance_remaining_usd IS NULL (older rows, and any response that
omits it) by excluding those rows, not by treating them as zero.
2. The wrong window
quota_burn sums energy_kwh over a 30-day rolling window while the
subscription resets monthly on objective.billing_reset_day (currently 6).
Those are different periods, so the reported fraction does not correspond to
the billing period it appears to describe. It also returns
reset_date = today − 30 days, which is the rolling-window start named as
if it were a billing reset — the same confusion billing_reset_day was added
to fix in the admin modal.
Compute the kWh figure over the current billing period (since the most
recent billing_reset_day), and rename the rolling-window field so it cannot
be mistaken for a reset date. Keep the rolling figure if it is useful, but
label it honestly.
3. The wrong words
"A quota is a wall, not a bill — requests fail rather than costing more"
appears in the warning text and the same framing is in config/config.yaml
(max_energy_per_request, plan_kwh_per_period comments) and in CLAUDE.md.
It is demonstrably false for this account and it changes what an operator
does: a wall means "stop", overage means "you are being billed, decide if that
is fine."
Correct all three places. State plainly that overage is billed against a
credit balance, and that plan_kwh_per_period gates nothing — verified:
it appears only in metrics.py reporting and the admin allowlist.
Surfacing
- Admin dashboard quota chip: show balance and runway as the headline, percentage-of-plan demoted to detail.
/metricsquotablock gains the balance/burn fields alongside the existing ones. Do not remove existing keys — the TUI and admin snapshot read them.
Non-goals
- Do not gate or refuse requests on quota.
plan_kwh_per_periodgates nothing today and this plan does not change that. Refusing traffic because a local estimate says the balance is low would turn a billing question into an outage — and the estimate can be wrong (see the top-up trap). - Do not add a second source of truth.
allowance_remaining_usdis the provider's own number; do not reconstruct a balance from summedcost_usd. - Do not touch
max_energy_per_requestbehaviour. - Do not change how
energy_observationsrows are written.
Success criteria
/metricsquotareports current balance, burn rate over a stated window, and projected hours to zero, all derived fromallowance_remaining_usd.- A test with a top-up inside the window (balance increasing) produces a correct burn rate, not one inflated by the credit.
- A test with all-NULL
allowance_remaining_usddegrades gracefully rather than reporting a zero balance. - The kWh figure is computed over the current billing period, and no field
named
reset_datereturns a rolling-window start. - The "wall, not a bill" framing is corrected in
metrics.py,config/config.yamlandCLAUDE.md, and states thatplan_kwh_per_periodgates nothing. - The admin quota chip leads with balance and runway.
- Existing
/metricsquotakeys still present. - Full suite green with
local_energy.enabledboth true and false. - The user's
config/config.local.yamlis byte-identical after the run (local_energy.enabled: true,tariff_usd_per_kwh: 0.159). Corrected 2026-09-04: this criterion previously said the tariff lines inconfig/config.yamlmust stay uncommitted. That is stale — PR #26 moved deployment values into the gitignored overlay, soconfig/config.yamlis now clean and tracked. Note the consequence for this plan specifically: §3 edits comments inconfig/config.yaml, which is now an ordinary tracked edit to commit deliberately, not a file to tiptoe around. The file to protect is the overlay, and it is not in git history at all.
The pattern worth naming
This is the third plan in a row whose finding is "the signal was already
recorded and nothing surfaced it" — after the capability-gated ceiling
collapse and the unread route_decisions rejections. The router is good at
capturing evidence and poor at putting it where an operator looks. Worth
considering a general review of what is persisted versus what is displayed,
rather than a fourth one-off detector.