Files
6krrt/plans/admin-quota-undercounting-and-reset-date.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

7.4 KiB

Fix admin portal quota: undercounted usage + wrong "Resets" date

Status: done -- quota.by_provider in /metrics

Context

User-reported, from the live admin portal's quota modal (screenshot, 2026-08-31): displays 93.0% metered against the plan, but the user is already on NeuralWatt overage credits — the real figure should be at or above 100%. Separately, the modal's "Resets" field reads 2026-08-02, but the user's actual NeuralWatt subscription billing period resets 2026-09-06. Both traced to root cause in the code, not guessed at.

Bug 1: undercounting — eval_proficiency.py never logs its calls

metrics.py:quota_burn() computes metered_fraction_of_plan from SUM(energy_kwh) FROM energy_observations over the trailing 30 days. That table is populated by dispatcher.log_observation(), called from every routed/dispatched request and from seed_energy.py's sweeps (confirmed: seed_energy.py imports and calls log_observation directly).

eval_proficiency.py does not. Confirmed by grep: it calls NeuralWatt via requests.post (twice, for the primary and judge calls) and never once calls log_observation or touches energy_observations. This is intentional for a different reason — CLAUDE.md documents that eval_proficiency.py deliberately bypasses the dispatcher so a transient eval-harness failure can't trip the circuit breaker against production traffic — but logging energy for quota purposes is an unrelated concern that fix never considered, and got dropped as a side effect.

The practical impact: every real, billed NeuralWatt call the eval harness makes — and CLAUDE.md's own history describes many: the original 43-task benchmark sweep, six additional docs_writing passes, benchmark-sourced hardening passes, judge-model calls (a second real call per sample) — is invisible to quota_burn(). Given how much eval traffic this project has actually run, this plausibly accounts for most or all of the gap between the router's self-reported 93.0% and the real account being in overage.

This is very likely the primary fix, and it's a small one: make eval_proficiency.py log the same way seed_energy.py already does.

Residual honesty check for this fix: even after it, metered_fraction_of_plan remains a router + eval-harness tally, not literal ground truth from NeuralWatt's own account — anything hitting the NeuralWatt API through some other path entirely (a manual curl, a different tool) still wouldn't count, and the payload's own note field already discloses this ("router-metered only; traffic bypassing the router is not counted"). Worth a quick check during implementation: does NeuralWatt's API expose any account-level usage/quota endpoint the poller could read as ground truth instead of (or alongside) the self-tallied figure? Don't assume one exists — check, and if not, closing the eval-harness gap is the real fix available.

Bug 2: reset_date is the window start, not a reset date

metrics.py:quota_burn(), confirmed directly:

"reset_date": (datetime.now(timezone.utc).date() - timedelta(days=30)).isoformat(),

This computes today minus 30 days — the start of the rolling 30-day window the SUM above covers — not any actual billing reset date. The function's own docstring already says as much ("the ISO date of today minus 30 days, the rolling-window start"). The admin frontend then labels it plainly wrong: admin/frontend/index.html:685 — <span class="quota-label">Resets</span><span class="quota-value">${quota.reset_date}</span>. Nothing in this codebase knows the user's real NeuralWatt billing-cycle anchor date at all — there's no config field for it anywhere.

Fix: add a real anchor the user actually has (they stated it: the 6th of each month) as config, and compute a genuine next-reset date from it. Don't just relabel the window-start field to something less misleading — that would be honest but still useless; the user has the real date, use it.


Fix 1: log eval_proficiency.py's calls

File: src/eval_proficiency.py.

Add a sentinel task_category for these rows, the same pattern seed_energy.py uses (SEED_CATEGORY = "seed_reference") — e.g. EVAL_CATEGORY = "eval_proficiency". After each real completion (both the primary model call and the judge call, since both are real billed NeuralWatt requests), call dispatcher.log_observation(...) exactly as seed_energy.py does at its call site — same extract_telemetry/ gross_energy_kwh plumbing, no new machinery needed, just wiring the existing one into a file that never used it.

Keep it tagged distinctly (EVAL_CATEGORY, not the real task category being evaluated) so it stays excluded from anything that treats task_category as "a real routing decision" — mirror however SEED_CATEGORY rows are already excluded elsewhere (e.g. metrics.scoring_coverage() filters task_category = SEED_CATEGORY deliberately; check for other such filters and extend the same way rather than inventing a second pattern).

Fix 2: a real billing-reset anchor

Config: src/config.py's Objective (currently quality_tolerance, assumed_cache_rate, assumed_completion_tokens, max_energy_per_request, plan_kwh_per_period) — add billing_reset_day: Optional[int] = None, validated to 1..28 (skip the 29-31 edge cases of shorter months rather than handle them — simplest correct thing; a user on the 31st can round down). config/config.yaml gets the section with a comment telling the operator to set their real NeuralWatt billing date, same anti-fabrication posture as every other "set your own real number" field in this file (the local-energy tariff being the most recent precedent).

Backend: metrics.py:quota_burn() — when billing_reset_day is set, compute the real next occurrence (this month's billing_reset_day if today's day-of-month hasn't reached it yet, else next month's) and return it as a distinctly-named field — e.g. next_reset_date — rather than overloading reset_date. When billing_reset_day is unset, omit next_reset_date (or None) rather than falling back to the window-start value under a misleading label — absent evidence should not masquerade as an answer, same rule this project already applies everywhere else. Decide at implementation time whether the window-start value is worth keeping under an honest label (e.g. window_start_date) for diagnostic purposes, or dropping — it's not what a user asks "when does this reset" to mean.

Frontend: admin/frontend/index.html's renderQuotaModal — show next_reset_date under "Resets" when present; when absent, show something honest ("not configured") rather than silently keeping the old wrong value around under the same label.


Verification

  • tests/test_eval_proficiency.py (or wherever its tests live — check) gets a case confirming a real call logs an energy_observations row tagged EVAL_CATEGORY, and that metrics.scoring_coverage() (or whatever else filters SEED_CATEGORY) correctly also excludes EVAL_CATEGORY from wherever that exclusion matters.
  • metrics.quota_burn() gets test cases for: billing_reset_day unset (no next_reset_date in the response), set with today before the reset day this month, and set with today after it (rolls to next month).
  • Manual check against the live DB after landing: re-open the admin quota modal and confirm the percentage moved up (closer to the real account state) and "Resets" shows a real upcoming date, not a past one.
  • Full pytest stays green.