plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
143 lines
7.4 KiB
Markdown
143 lines
7.4 KiB
Markdown
# Fix admin portal quota: undercounted usage + wrong "Resets" date
|
|
|
|
Status: done -- quota.by_provider in /metrics
|
|
|
|
## Context
|
|
|
|
User-reported, from the live admin portal's quota modal (screenshot,
|
|
2026-08-31): displays 93.0% metered against the plan, but the user is
|
|
already on NeuralWatt overage credits — the real figure should be at or
|
|
above 100%. Separately, the modal's "Resets" field reads `2026-08-02`, but
|
|
the user's actual NeuralWatt subscription billing period resets
|
|
2026-09-06. Both traced to root cause in the code, not guessed at.
|
|
|
|
### Bug 1: undercounting — `eval_proficiency.py` never logs its calls
|
|
|
|
`metrics.py:quota_burn()` computes `metered_fraction_of_plan` from
|
|
`SUM(energy_kwh) FROM energy_observations` over the trailing 30 days. That
|
|
table is populated by `dispatcher.log_observation()`, called from every
|
|
routed/dispatched request **and** from `seed_energy.py`'s sweeps (confirmed:
|
|
`seed_energy.py` imports and calls `log_observation` directly).
|
|
|
|
`eval_proficiency.py` does not. Confirmed by grep: it calls NeuralWatt via
|
|
`requests.post` (twice, for the primary and judge calls) and never once
|
|
calls `log_observation` or touches `energy_observations`. This is
|
|
intentional for a different reason — CLAUDE.md documents that
|
|
`eval_proficiency.py` deliberately bypasses the *dispatcher* so a transient
|
|
eval-harness failure can't trip the circuit breaker against production
|
|
traffic — but logging energy for quota purposes is an unrelated concern
|
|
that fix never considered, and got dropped as a side effect.
|
|
|
|
The practical impact: every real, billed NeuralWatt call the eval harness
|
|
makes — and CLAUDE.md's own history describes many: the original 43-task
|
|
benchmark sweep, six additional `docs_writing` passes, benchmark-sourced
|
|
hardening passes, judge-model calls (a second real call per sample) — is
|
|
invisible to `quota_burn()`. Given how much eval traffic this project has
|
|
actually run, this plausibly accounts for most or all of the gap between
|
|
the router's self-reported 93.0% and the real account being in overage.
|
|
|
|
**This is very likely the primary fix**, and it's a small one: make
|
|
`eval_proficiency.py` log the same way `seed_energy.py` already does.
|
|
|
|
Residual honesty check for this fix: even after it, `metered_fraction_of_plan`
|
|
remains a **router + eval-harness** tally, not literal ground truth from
|
|
NeuralWatt's own account — anything hitting the NeuralWatt API through some
|
|
other path entirely (a manual curl, a different tool) still wouldn't count,
|
|
and the payload's own `note` field already discloses this ("router-metered
|
|
only; traffic bypassing the router is not counted"). Worth a quick check
|
|
during implementation: does NeuralWatt's API expose any account-level
|
|
usage/quota endpoint the poller could read as ground truth instead of (or
|
|
alongside) the self-tallied figure? Don't assume one exists — check, and if
|
|
not, closing the eval-harness gap is the real fix available.
|
|
|
|
### Bug 2: `reset_date` is the window start, not a reset date
|
|
|
|
`metrics.py:quota_burn()`, confirmed directly:
|
|
|
|
```python
|
|
"reset_date": (datetime.now(timezone.utc).date() - timedelta(days=30)).isoformat(),
|
|
```
|
|
|
|
This computes *today minus 30 days* — the start of the rolling 30-day
|
|
window the SUM above covers — not any actual billing reset date. The
|
|
function's own docstring already says as much ("the ISO date of today
|
|
minus 30 days, the rolling-window start"). The admin frontend then labels
|
|
it plainly wrong: `admin/frontend/index.html:685` —
|
|
`<span class="quota-label">Resets</span><span class="quota-value">${quota.reset_date}</span>`.
|
|
Nothing in this codebase knows the user's real NeuralWatt billing-cycle
|
|
anchor date at all — there's no config field for it anywhere.
|
|
|
|
Fix: add a real anchor the user actually has (they stated it: the 6th of
|
|
each month) as config, and compute a genuine next-reset date from it.
|
|
Don't just relabel the window-start field to something less misleading —
|
|
that would be honest but still useless; the user has the real date, use it.
|
|
|
|
---
|
|
|
|
## Fix 1: log eval_proficiency.py's calls
|
|
|
|
**File**: `src/eval_proficiency.py`.
|
|
|
|
Add a sentinel `task_category` for these rows, the same pattern
|
|
`seed_energy.py` uses (`SEED_CATEGORY = "seed_reference"`) — e.g.
|
|
`EVAL_CATEGORY = "eval_proficiency"`. After each real completion (both the
|
|
primary model call and the judge call, since both are real billed
|
|
NeuralWatt requests), call `dispatcher.log_observation(...)` exactly as
|
|
`seed_energy.py` does at its call site — same `extract_telemetry`/
|
|
`gross_energy_kwh` plumbing, no new machinery needed, just wiring the
|
|
existing one into a file that never used it.
|
|
|
|
Keep it tagged distinctly (`EVAL_CATEGORY`, not the real task category being
|
|
evaluated) so it stays excluded from anything that treats `task_category` as
|
|
"a real routing decision" — mirror however `SEED_CATEGORY` rows are already
|
|
excluded elsewhere (e.g. `metrics.scoring_coverage()` filters
|
|
`task_category = SEED_CATEGORY` deliberately; check for other such filters
|
|
and extend the same way rather than inventing a second pattern).
|
|
|
|
## Fix 2: a real billing-reset anchor
|
|
|
|
**Config**: `src/config.py`'s `Objective` (currently
|
|
`quality_tolerance`, `assumed_cache_rate`, `assumed_completion_tokens`,
|
|
`max_energy_per_request`, `plan_kwh_per_period`) — add
|
|
`billing_reset_day: Optional[int] = None`, validated to `1..28` (skip the
|
|
29-31 edge cases of shorter months rather than handle them — simplest
|
|
correct thing; a user on the 31st can round down). `config/config.yaml`
|
|
gets the section with a comment telling the operator to set their real
|
|
NeuralWatt billing date, same anti-fabrication posture as every other "set
|
|
your own real number" field in this file (the local-energy tariff being the
|
|
most recent precedent).
|
|
|
|
**Backend**: `metrics.py:quota_burn()` — when `billing_reset_day` is set,
|
|
compute the real next occurrence (this month's `billing_reset_day` if
|
|
today's day-of-month hasn't reached it yet, else next month's) and return
|
|
it as a distinctly-named field — e.g. `next_reset_date` — rather than
|
|
overloading `reset_date`. When `billing_reset_day` is unset, omit
|
|
`next_reset_date` (or `None`) rather than falling back to the window-start
|
|
value under a misleading label — absent evidence should not masquerade as
|
|
an answer, same rule this project already applies everywhere else. Decide
|
|
at implementation time whether the window-start value is worth keeping
|
|
under an honest label (e.g. `window_start_date`) for diagnostic purposes,
|
|
or dropping — it's not what a user asks "when does this reset" to mean.
|
|
|
|
**Frontend**: `admin/frontend/index.html`'s `renderQuotaModal` — show
|
|
`next_reset_date` under "Resets" when present; when absent, show something
|
|
honest ("not configured") rather than silently keeping the old wrong value
|
|
around under the same label.
|
|
|
|
---
|
|
|
|
## Verification
|
|
|
|
- `tests/test_eval_proficiency.py` (or wherever its tests live — check) gets
|
|
a case confirming a real call logs an `energy_observations` row tagged
|
|
`EVAL_CATEGORY`, and that `metrics.scoring_coverage()` (or whatever else
|
|
filters `SEED_CATEGORY`) correctly also excludes `EVAL_CATEGORY` from
|
|
wherever that exclusion matters.
|
|
- `metrics.quota_burn()` gets test cases for: `billing_reset_day` unset
|
|
(no `next_reset_date` in the response), set with today before the
|
|
reset day this month, and set with today after it (rolls to next month).
|
|
- Manual check against the live DB after landing: re-open the admin quota
|
|
modal and confirm the percentage moved up (closer to the real account
|
|
state) and "Resets" shows a real upcoming date, not a past one.
|
|
- Full `pytest` stays green.
|