The dashboard was eight cards restating what the other seven pages own, so it was the page you passed through rather than the one you started from. It is now the landing page: no nav entry of its own, reached by the logo, with two views over the one /api/snapshot payload so switching costs no request. Board is one tile per destination, identical anatomy every time: what it is, one number, one qualifier, a link. Under them, a full-width Activity band plots routed decisions against billed requests on one axis, so the gap between them -- local dispatch, cache hits, requests that never reached a provider -- is readable, which it is in neither series alone. Its range (24h/7d/30d) is separate from the tile sparklines, because it is the chart you come to the page to read. Beside it, the eight busiest models ranked by CALLS, not cost: cost coverage is partial, and an OpenRouter row with two priced calls out of hundreds otherwise outranks the model doing the work. Live loads the row itself -- category, model, provider, required context, tier, classify time, exploratory flag, cost -- so the usual question does not need the Decisions page, and groups rows by minute so a burst reads as one. The dot is the classification source: green for a real classification, amber for a degraded one, which still routes but never moves proficiency, and until now was visible only in a /metrics warning. The rail beside it answers what Board cannot: throughput against the last hour, classifier p50 and p95 (the latency floor every routed request pays before an upstream token), degraded share, and the top five models by share of the live window keyed on model AND provider, since the same model on two providers is two different routing outcomes. Spend is not restated here; it belongs to Quota. Warnings move behind a navbar bell. They are recomputed from live state on every poll, so the only honest dismissal is "hide until the condition changes": a dismissal is keyed on the warning's TEXT, so when the numbers move the key stops matching and the warning returns on its own. Nothing has to expire it. Per browser, in localStorage, wrapped in try/catch -- a reading state, not a fact about the router. Two tests changed shape rather than being deleted. The verdict-bar tests pinned a card this page no longer has, but the lesson they recorded is the expensive part and is preserved verbatim in the replacement's docstring: `unverifiable` is not a finding, it is the default for prose and for any turn ending in a tool call, and any bar scaled against it renders every meaningful verdict at the 2px minimum. Exclude it from the scale, never from the denominator. The Live rail states the same fact as one ratio, which needs no scale. The nav test's rule is inverted for index.html rather than skipped: the landing page must have no nav entry AND must claim no active one, so neither can drift back in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
35 KiB
Quota, redesigned around billing shape (subscription vs paygo vs self-hosted)
Status: done -- shipped as the quota page redesign; decision-complete. Written 2026-09-10 from the live database,
the live /metrics, and two live probes against the OpenRouter API.
Intended executor: opencode (Prometheus → Atlas). Every open question in this document is closed; where a choice existed, the choice and its reason are recorded rather than left to the implementer.
Read plans/admin-design-standards.md and plans/admin-work-framework.md
before touching admin/frontend/*. This document does not restate the visual
language; it specifies what the pages must say.
On the file:line anchors below. Every anchor in the first draft of this
document was taken against a working tree five commits behind origin/main,
and PR #76 (feat/mangled-output-breaker, merged 2026-09-10T23:48Z) added ~77
lines to src/metrics.py in between. Four anchors were wrong as a result. They
are corrected here and re-verified against f89cefc, but the lesson stands:
treat every anchor as a hint and confirm it by grepping the landmark named
beside it. Each anchor below is paired with the symbol or string it points
at, so the grep is always available.
Why this exists
The word "Quota" currently covers three unrelated quantities, and the portal
sums two of them into a figure CLAUDE.md itself disclaims as "not a single
spendable figure". Meanwhile the one number an operator asks for first — money
spent this month — appears nowhere in the portal at all.
All evidence below is from the live system on 2026-09-10/11, billing period 2026-09-06 → 2026-10-06.
Defect 1 — one word, three quantities
metrics.quota_burn returns a kWh-plan report with a per-provider balance map
bolted onto it (src/metrics.py:334, merging quota_balance_and_burn at
src/metrics.py:203). The live payload carries, under one key:
| field | what it actually is | magnitude |
|---|---|---|
metered_fraction_of_plan |
NeuralWatt kWh subscription consumption | 0.4607 |
by_provider.neuralwatt.balance_usd |
overage-billing accounting artifact | −$0.004198 |
by_provider.openrouter.balance_usd |
a real prepaid credit pool | $40.45 |
total_balance_usd |
the sum of the previous two | $40.4495 |
Those are a gauge, a rounding residue, and a wallet. The sum is not a quantity.
Defect 2 — money spent is not in the portal
Actual billed spend this period, from the live DB:
| provider | period spend | source |
|---|---|---|
| neuralwatt | $22.66 | SUM(energy_observations.cost_usd), 6,079 rows |
| openrouter | $9.57 | account poll: total_credits 50 − total_usage 9.5714 |
| ollama-local | $0.0027 | local_energy_observations.cost_usd, 202 rows |
≈ $32.23 this period, and the dashboard renders none of it. The NeuralWatt figure has been captured per request since 2026-08-17 and is read by exactly nothing.
Defect 3 — OpenRouter has no per-request cost, so it renders as free
extract_telemetry (src/dispatcher.py:2004) reads only NeuralWatt's
top-level energy / cost blocks. Every OpenRouter row therefore lands with
cost_usd IS NULL: 3,106 rows this period. metrics.per_model
(src/metrics.py:1386, def per_model) COALESCEs those NULLs to 0, so the
live dashboard shows:
xiaomi/mimo-v2.5 openrouter calls=655 cost=$0.0000 kwh=0.00000
That model is #8 by volume and cost real money — the pool moved $1.23 on 2026-09-10 alone, the day it served 598 of those calls at ~175k prompt tokens each.
This is fixable, and it was verified live rather than assumed. Sending
usage: {"include": true} makes OpenRouter return billed cost per request.
Probed 2026-09-10 against xiaomi/mimo-v2.5:
{"prompt_tokens": 254, "completion_tokens": 8, "cost": 1.14576e-05,
"prompt_tokens_details": {"cached_tokens": 192},
"cost_details": {"upstream_inference_cost": 1.14576e-05,
"upstream_inference_prompt_cost": 9.2176e-06,
"upstream_inference_completions_cost": 2.24e-06}}
A :free model returns the same shape with cost: 0, correctly. Note
cached_tokens: 192/254 — a measured cache rate on a field
routing.assumed_cache_rate currently guesses at.
Defect 4 — the plan gauge hides the only actionable number
The chip renders 46.1% · 6.25 kWh plan. At the time of writing, 5.01 of the
period's 30 days had elapsed — 16.7%. Burning 46.1% of the allowance in 16.7%
of the period is 2.76x sustainable pace, projecting ~17.2 kWh against a
6.25 kWh plan.
That ratio is the number the operator needed on 2026-09-06 and computed by
hand against metered_kwh_period and elapsed time. Nothing in the portal
computes it. A percentage without a pace reads as reassuring at exactly the
moment it should not.
Defect 5 — the electricity the operator personally pays for is not in Quota
local_energy_summary (src/metrics.py:1547) exists and is rendered in its
own dashboard card, but it is absent from Quota — while ollama-local rows
are present in energy_observations (60 rows) and therefore in the "Cloud
energy" pane, which double-counts them and files self-hosted draw under cloud.
Also found, and deliberately NOT fixed here
Joined on request_id over this period, route_decisions.est_cost_usd
overprices NeuralWatt 4.14x against its own billed cost (n=4,958), while
OpenRouter's estimate lands within ~13% of the observed pool delta:
| model | n | est | billed | ratio |
|---|---|---|---|---|
| kimi-k2.7-code | 2,853 | $62.78 | $11.21 | 5.60 |
| qwen3.6-35b | 1,024 | $4.01 | $0.48 | 8.40 |
| glm-5.3 | 288 | $14.75 | $6.76 | 2.18 |
| deepseek-v4-flash | 714 | $2.72 | $1.15 | 2.36 |
This is expected in direction — NeuralWatt bills
min($8/kWh × kWh, 3 × list) and models that batch well land far under list,
which is the whole subject of CLAUDE.md's cost section — but the size means
the cost tiebreak systematically overprices NeuralWatt relative to OpenRouter
by roughly 4x. That is a routing-correctness question, not a dashboard one.
Scope decision: this plan surfaces the ratio and changes no routing. Work item 3 puts est-vs-billed on the page so the bias is visible and measurable; acting on it belongs in a separate plan, because the fix is either a per-provider calibration factor or a switch to billed-cost feedback, and neither should be decided inside a dashboard change.
The model: three billing shapes, one comparable unit
The frankenstein comes from trying to express three billing shapes in one schema. The redesign names the shapes and gives each its native unit, then uses dollars — the only unit all three share — for the cross-provider view.
| shape | provider | allowance | remaining | spend truth | granularity |
|---|---|---|---|---|---|
metered_plan |
neuralwatt | 6.25 kWh/period | plan − metered kWh | cost_usd per request |
per request |
prepaid_credit |
openrouter | $50 pool | total_credits − total_usage |
account poll; per-request after work item 1 | 2 h (account), per request (attribution) |
self_hosted |
ollama-local | none — no cap exists | n/a | GPU draw × tariff_usd_per_kwh |
per request |
Two rules fall out of that table, and both are load-bearing:
Rule 1 — account truth for totals, request truth for attribution, never add
them. An account poll is authoritative for "what did this provider charge
me"; summed per-request costs are authoritative for "which model spent it".
They are two measurements of overlapping quantities at different fidelities.
Adding them double-counts; preferring the request sum for a total silently
undercounts whatever the router did not observe (plan_kwh_per_period's own
note already says as much: "router-metered only"). So: totals come from the
account, breakdowns come from requests, and the page states when the
breakdown covers less than the total.
Rule 2 — a missing number is not zero. COALESCE(SUM(cost_usd), 0) is how
OpenRouter came to read $0.0000. Unknown renders as — with a reason,
everywhere, in every card and every row.
Work item 1 — capture per-request cost where the provider reports it
Config (src/config.py, config/config.yaml). Add to DispatchProvider:
# dispatch_providers.openrouter
# OpenRouter returns billed cost per request when the request asks for usage
# accounting. Verified live 2026-09-10: usage.cost, usage.cost_details and
# usage.prompt_tokens_details.cached_tokens all come back populated.
reports_cost_in_usage: true
Default false. It must appear in config/config.yaml explicitly for both
providers, not only as a Pydantic default — CLAUDE.md's "every knob belongs
in config.yaml" rule, and the reason outcome_attribution_window_seconds was
deleted.
Gate the request-body change on this flag rather than sending
usage: {"include": true} unconditionally: NeuralWatt is OpenAI-compatible and
an unknown top-level key is a plausible 400 on a provider we have no reason to
probe.
Request (src/dispatcher.py:4212, where upstream_body is built):
if getattr(settings, "reports_cost_in_usage", False):
upstream_body.setdefault("usage", {"include": True})
setdefault, so a client that sent its own usage block wins.
Extraction (src/dispatcher.py:2004). Extend extract_telemetry to fall
back to the usage block when NeuralWatt's cost block is absent:
cost_usd←cost.request_cost_usd, elseusage.cost- new
Telemetry.cached_prompt_tokens←usage.prompt_tokens_details.cached_tokens
Keep the precedence in that order. A provider that reports both is reporting
the same number twice, and request_cost_usd is the field this project has
already validated against billing.
THE TRAP — the streamed path does not pass usage to extract_telemetry,
and every agent client streams. On the buffered path
(src/dispatcher.py:4385) extract_telemetry(payload) sees the whole response
including usage. On the streamed path (src/dispatcher.py:4637) it is called
as extract_telemetry(collected), where collected holds only the sniffed
NeuralWatt SSE comment lines; the streamed usage dict lives in a separate
local (src/dispatcher.py:4603, populated from chunk["usage"]) and is passed
to log_observation for token counts only.
So a change that only touches extract_telemetry fixes cost for buffered
requests and leaves it NULL for essentially all real traffic.
stream_options.include_usage is already set (src/dispatcher.py:4217), so
the final chunk does carry it; it simply never reaches the extractor.
The fix, decided — not implementer's choice. Signature becomes:
def extract_telemetry(payload: dict, *, usage: Optional[dict] = None) -> Telemetry:
- When
usageis not passed, the function falls back topayload.get("usage"). That alone fixes every buffered call site with no change at the call sites at all, becausepayloadthere is the whole response. - The streamed call site passes it explicitly —
extract_telemetry(collected, usage=usage)— becausecollectedholds only sniffed SSE comment blocks and structurally cannot contain a usage dict. - Cost precedence inside the function, in this order:
payload["cost"]["request_cost_usd"], thenusage["cost"]. First non-None wins.
The rejected alternative was extract_telemetry({**collected, "usage": usage}).
It needs no signature change, which is why it is tempting, and it is wrong:
collected is by contract the output of _sniff_telemetry_line
(src/dispatcher.py:3735), whose keys are provider SSE comment names. Injecting
a synthetic usage key into it makes that dict a half-response/half-sniff
hybrid, and the next person to read either function has to hold both meanings at
once. The keyword argument keeps the two sources named.
Note the precedence order is what makes the existing NeuralWatt behaviour
unchanged: NW streaming populates collected["cost"], so it never reaches the
usage branch, and its cost_usd stays the figure this project has already
validated against billing. The streamed path must be covered by a test
asserting a non-null cost_usd from usage.cost alone.
Audit the other log_observation call sites while in here
(src/dispatcher.py:4871 and the local-dispatch path) so no branch silently
drops the new field.
Schema. energy_observations gains one nullable column:
ALTER TABLE energy_observations ADD COLUMN cached_prompt_tokens INTEGER;
Add it to config/schema.sql with a comment recording why it exists (a
measured cache rate against routing.assumed_cache_rate's guess) and that
NULL means "provider did not report", not zero.
What this does NOT do: it does not backfill. OpenRouter cost exists from the deploy forward only, which is precisely why Rule 1 keeps account polls as the source of account totals.
Work item 2 — store the credit pool, not just the remainder
_BALANCE_PARSERS["openrouter"] (src/poller.py:66) computes
total_credits − total_usage and discards both inputs. The pool size is what
turns a bare "$40.45" into "used $9.57 of $50", and it is already in the
response.
provider_balance_observationsgains two nullable columns:total_credits_usd,total_usage_usd. Extend_ensure_provider_balance_table(src/poller.py:422) with idempotentALTER TABLE ... ADD COLUMNguarded the way the rest of this project guards live-DB migrations, and mirror them intoconfig/schema.sql.- The parser registry changes shape: a parser returns
(balance_usd, total_credits_usd | None, total_usage_usd | None).record_balance(src/poller.py:449) persists all three. A provider whose endpoint reports only a remainder stores NULL for the other two, and the UI falls back to a bare remaining figure. tests/test_multi_provider_poller.py:728-731pins_BALANCE_PARSERS.keys() == config.PROVIDERS_WITH_BALANCE_PARSERS(src/config.py:1056); keep that pin working. It is in that file rather than a poller-parsing one, which is easy to miss when scanning by name.
Nullable columns because a NULL pool size must render as "balance $40.45" and never as "$40.45 of $0".
Work item 3 — the /metrics quota block, rewritten
Replace quota_burn and quota_balance_and_burn with one function,
metrics.quota_accounts(conn, cfg). No deprecated aliases — CLAUDE.md
records that the last quota reshape updated every consumer in the same change,
and that is the standard here too.
Shape
"quota": {
"period": {
"start": "2026-09-06",
"next_reset": "2026-10-06",
"elapsed_fraction": 0.167,
"source": "objective.billing_reset_day"
},
"accounts": [
{
"provider": "neuralwatt",
"shape": "metered_plan",
"plan": {
"unit": "kwh",
"allowance": 6.25,
"used": 2.87933,
"used_fraction": 0.4607,
"pace_ratio": 2.76,
"projected_period_total": 17.24,
"pace_note": null
},
"spend_usd": {
"period": 22.66,
"window_30d": 61.2,
"source": "per_request_billed",
"attribution_coverage": 1.0
},
"credit": {
"balance_usd": -0.004198,
"meaning": "overage_allowance",
"observed_at": "2026-09-10T18:49:43Z",
"age_seconds": 19980
},
"fidelity": "per_request"
},
{
"provider": "openrouter",
"shape": "prepaid_credit",
"pool": {
"total_credits_usd": 50.0,
"used_usd": 9.5714,
"remaining_usd": 40.4286,
"used_fraction": 0.1914,
"observed_at": "2026-09-10T23:17:32Z",
"age_seconds": 7200
},
"burn": {
"window_hours": 24,
"usd_per_hour": 0.0337,
"projected_hours_remaining": 1200.9,
"runway_low_warning": false,
"note": null
},
"spend_usd": {
"period": 9.5714,
"window_30d": 9.5714,
"source": "account_poll_delta",
"attribution_coverage": 0.0
},
"fidelity": "account_poll"
},
{
"provider": "ollama-local",
"shape": "self_hosted",
"energy": { "unit": "kwh", "period": 0.01725, "window_30d": 0.0402 },
"spend_usd": {
"period": 0.00274,
"window_30d": 0.0064,
"source": "derived_from_draw",
"attribution_coverage": 1.0
},
"tariff_usd_per_kwh": 0.15856,
"fidelity": "derived"
}
],
"spend": {
"by_provider_usd": { "neuralwatt": 22.66, "openrouter": 9.5714, "ollama-local": 0.00274 },
"total_usd": 32.23,
"estimated_usd": { "neuralwatt": 86.05, "openrouter": 10.07 },
"estimate_ratio": { "neuralwatt": 4.14, "openrouter": 1.13 }
},
"alarm": {
"kind": "plan_pace",
"provider": "neuralwatt",
"severity": "warn",
"headline": "neuralwatt: 2.8x sustainable pace, projecting 17.2 kWh against a 6.25 kWh plan"
},
"note": "router-metered only; traffic bypassing the router is not counted"
}
Decisions embedded in that shape
-
shapeis derived from config, never hardcoded per provider name. The rule, in precedence order, evaluated percfg.dispatch_providersentry plus the local-dispatch provider:# test shape 1 provider is the local-dispatch provider ( ollama-local)self_hosted2 has_energy_telemetryandobjective.plan_kwh_per_periodis setmetered_plan3 balance_urlis setprepaid_credit4 none of the above unmeteredFirst match wins, and the order matters: a provider could satisfy both 2 and 3, and the plan is the thing with a period and a ceiling, so it leads.
unmeteredis a real state, not an error — a configured provider with no telemetry, no plan and no balance endpoint genuinely has nothing to meter, and it renders as a name plus its per-request spend if any, with no gauge. Ametered_planprovider whenplan_kwh_per_periodis null degrades to case 3 or 4 rather than rendering a gauge against a null allowance. -
spend.by_provider_usd[p]is exactlyaccounts[p].spend_usd.period, andspend.total_usdis the sum of those values — nothing else. It is never computed from a second query, and a balance figure never enters it. This is Rule 1 as an implementation constraint: each account has already decided whether its own period spend comes from summed per-request cost, an account poll delta, or derived draw, and the ledger just adds up those decisions.estimated_usd[p]is a separateSUM(route_decisions.est_cost_usd)over the same period and provider, andestimate_ratio[p]isestimated_usd[p] / by_provider_usd[p], null when the denominator is 0. -
accountsis a list, not a map keyed by provider, ordered by alarm severity then provider name. The oldby_providermap forced every consumer to re-derive which provider mattered; the list makes the answer positional and makes the TUI's row order free. -
total_balance_usdis deleted. It summed a prepaid pool with an overage residue.spend.total_usdreplaces it as the one honest cross-provider number, because dollars spent are additive across billing shapes. -
shapedrives rendering,fidelityqualifies it. A consumer switches layout onshapeand prints a caveat fromfidelity; neither is inferred from provider name or from which config keys happen to be set. -
spend_usd.sourceis one ofper_request_billed,account_poll_delta,derived_from_draw.attribution_coverageis the fraction of that provider's period spend that per-request rows can account for — 1.0 for NeuralWatt, 0.0 for OpenRouter until work item 1 lands, and a partial value for the period that straddles the deploy. This is Rule 1 made machine readable: it is how a card knows to say "breakdown covers $4.10 of $9.57". -
credit.meaningreplacesbalance_source.telemetry/polled/unconfigureddescribed where the number came from;overage_allowance/prepaid_pool/not_configureddescribe what it means, which is what decides whether it deserves a gauge or a footnote.
pace_ratio — the computation, and its trap
pace_ratio = used_fraction / elapsed_fraction, with
elapsed_fraction = (now − period_start) / (next_reset − period_start).
Trap: elapsed_fraction → 0 at period start, and the ratio explodes. In
the first hours of a period a handful of calls reads as a 40x overrun. Guard
it: when elapsed_fraction < 0.02 (roughly the first 14 hours of a 30-day
period), return pace_ratio: null and projected_period_total: null with
pace_note: "too early in the period to project". A null pace must render as
"—", never as 0 or 1.0.
Second trap, and this project has been burned by it once already. The
projection must not read as a wall. objective.plan_kwh_per_period gates
nothing; overage is billed. plans/quota-balance-and-burn-rate.md records the
last warning that got this wrong ("a quota is a wall, not a bill — requests
fail rather than costing more") and why it was both alarming and false. Project
in kWh and say what it costs, never "requests will fail".
alarm — what the chip shows
One alarm, the worst across accounts, or null:
plan_pace:pace_ratio >= objective.plan_pace_warn_ratio(new knob, default1.25, which must be written intoconfig/config.yaml, not left as a Pydantic default).severity: "critical"at>= 2.0.credit_runway: existingprojected_hours_remaining < objective.quota_runway_warning_hours(6).criticalbelow half that.stale_reading: aprepaid_creditaccount whoseage_secondsexceeds 3x the poll interval. This one is new and it matters: the live pool sat flat at $40.453746437 for six consecutive polls across 18 hours while 705 completions ran, and the current UI renders that as a healthy account with no burn. It is poll granularity, not calm. A flat reading and a stale reading must be distinguishable on the page.
coverage.warnings (src/metrics.py:493) currently emits the >80%-of-plan
string. Replace it with a pace-based warning sourced from alarm, so the
navbar bell and the chip cannot disagree.
Consumers to update in the same change
Measured inventory — grep -rln "total_balance_usd\|by_provider" over the
repo. There is no deprecation path, so all of these move together:
| file | what breaks |
|---|---|
src/metrics.py:203,334 |
the two functions being replaced |
src/dispatcher.py |
/metrics assembly |
src/admin.py:1849 |
/api/snapshot's quota key, plus the new GET /admin/api/quota |
src/tui_model.py:41-60 |
builds {label, value} quota rows from the flat keys |
src/tui.py:78,105,566,601,615 |
per-provider row formatting, the kWh lead line, and the total_balance_usd line CLAUDE.md documents |
tests/test_metrics.py |
49 references — the bulk of the work |
tests/test_metrics_endpoint.py |
6 |
tests/test_tui.py |
6 |
tests/test_credit_attenuation_routing.py |
6 — asserts attenuation reads balances, so it must keep passing against the new shape without weakening what it pins |
CLAUDE.md, docs/api.md, docs/admin-portal.md |
describe the old contract by name |
test_credit_attenuation_routing.py deserves a second look rather than a
mechanical rename: credit_attenuation derives a routing multiplier from
polled balances, so its fixtures encode the old assumption that a balance is
one flat number per provider. Its behaviour must not change here — only where
it reads the balance from.
Work item 4 — /admin/quota, a page
The chip-plus-modal cannot hold three shaped cards and a ledger. Follow the
portal's page-per-domain pattern (providers, proficiency, profiles).
Mechanics (src/admin.py:823-870): add _admin_quota →
admin/frontend/quota.html, a @router.get("/quota") returning it with
_NO_CACHE_HEADERS, and a nav item after Providers in the navbar of all
admin pages (admin/frontend/*.html:281-287) — the navbar is duplicated per
page, so a nav item added to one page only is a half-shipped nav.
Data: GET /admin/api/quota → metrics.quota_accounts(conn, cfg) plus the
spend ledger's per-model rows. A dedicated endpoint rather than reusing
/api/snapshot, because the page needs per-model spend the snapshot does not
carry and none of the snapshot's proficiency/verdict payload.
Cards, one per account, laid out by shape
Three peer cards on one row is .col-xl-4 — the Standards doc is explicit that
.col-xl-6 strands a third card and opens two dead columns.
metered_plan (NeuralWatt). The kWh gauge is the hero, and the pace marker
is what makes it a redesign rather than a reskin: draw the gauge with a tick at
elapsed_fraction so "burned" and "elapsed" are visually comparable in one
glance. Below it: pace ratio, projected period total, $ billed this period,
and the overage allowance as a de-emphasized footnote labeled
overage allowance — never as a balance or a runway, because at −$0.004 it is
neither.
prepaid_credit (OpenRouter). A depletion gauge: used_usd of
total_credits_usd, remaining in the hero position. Then burn $/hr, runway in
days (hours only under ~48h), and the reading age rendered as text
("polled 2h ago"), promoted to a warning tint past the staleness threshold.
When total_credits_usd is NULL, degrade to a bare remaining figure with no
gauge.
self_hosted (local). kWh and $ this period, the tariff it was priced at,
and a one-line statement that this is billed to the operator's utility rather
than a provider. No gauge — there is no allowance to fill. Hidden entirely
when local_energy.enabled is false.
Spend ledger
Below the cards, full width:
spend.total_usdas the page's single headline number, with a stacked bar by provider. This is the answer to "what did I spend this month" and it is the reason the page exists.- Ranked per-model spend — model, provider, calls, $, share of total.
—in the $ column where cost is unknown, with the row still present and ranked by calls. Whereattribution_coverage < 1, a single line above the table states the gap in dollars, e.g. "breakdown accounts for $4.10 of OpenRouter's $9.57; the remainder predates per-request cost capture". - Estimate calibration, a small two-column strip: est vs billed per provider with the ratio. Label it as what it is — the cost tiebreak's input measured against the bill — and do not editorialize a fix; that is the follow-up plan.
Chip
renderQuotaChip (admin/frontend/index.html:779) stops hardcoding the kWh
story. It renders alarm: headline text, severity tint, and a link to
/admin/quota. With alarm: null it shows the period's spend to date, which
is calm and still useful. Its comment at line 790 ("The chip always tells the
kWh-subscription story") is the assumption being deleted — remove it, do not
leave it stale.
The chip stops being a modal trigger and becomes a link. Four edits, all in
index.html, and leaving any one of them behind leaves a dead control:
- Line 305 — the chip
<div>carriesdata-bs-toggle="modal"data-bs-target="#quota-modal"role="button"tabindex="0". Replace the element with an<a href="/admin/quota">carrying the same.quota-chipclass, and drop all four attributes: an anchor is focusable and activatable on its own, sorole/tabindexbecome wrong rather than merely redundant. The navbar'sz-index: 1030note in the Standards doc exists because of this chip'sbackdrop-filter— keep the stacking context intact. - Lines 462-478 — the
#quota-modalblock is deleted outright. renderQuotaModal(:659) and its#quota-modal-bodywrites are deleted, and therenderQuotaModal(data.quota)call at:558with it.grep -n 'quota-modal' admin/frontend/*.htmlmust come back empty when done.formatProviderBalance/formatBalanceStaleness(:708,:731) move to the new page rather than being deleted — the staleness helper is exactly what theprepaid_creditcard's reading-age line needs, and its tiny-negative-balance handling is hard-won (it exists so NeuralWatt's −$0.004 never renders as "-$0.00").
Work item 5 — the per-model usage pane
Two panes read data.per_model today: renderCloudEnergy
(admin/frontend/index.html:906) and renderBars (:935). Three defects,
each falsifiable:
- Bar length and bar label disagree.
renderBarsranks byBARS_SORT_KEYS[_barsSortKey](:840) — calls, cost or energy — but the value label is hardcoded to$cost · kWh(:957). Sorting by calls produces bars ordered by call count, labeled in dollars. - Unknown renders as zero, giving
xiaomi/mimo-v2.5a$0.0000label on 655 real calls (Defect 3 above). ollama-localappears in a card called "Cloud energy", double-counting against the Local Compute card that owns it.
And a fourth, which is the operator's own framing: calls, cost and energy compete for one row of space, and energy is the least actionable of the three.
The rework
-
One ranked list. Sort keys become calls | spend | tokens. Energy leaves the sort keys: it is not a thing an operator ranks models by, it is a thing they inspect.
tokensreplaces it because it is the honest volume metric for a provider that reports no energy at all. -
The label follows the sort key, plus share of total: sorting by calls shows
655 calls · 8.2%; by spend,$1.23 · 3.8%. One metric, the one being ranked. -
—for unknown, never$0.0000, with a footnote naming the reason ("provider does not report per-request cost"). -
Energy, carbon, attribution and grid move to a row-click detail popup — kWh, gCO2eq,
grid_id, meanattribution_ratio, mean completion tokens, cached-token share once work item 1 lands. Render only the fields that have data; a popup of six—s is worse than a popup that admits it has nothing.Copy the pattern that already exists rather than inventing one.
admin/frontend/proficiency.htmldoes exactly this for its matrix cells: a<div class="modal modal-blur fade" id="cell-modal">with a-titleand-body(:264-276), a lazily-constructednew tabler.Modal(...)held in a module-level_cellModal(:575), and a click handler that setstextContenton the title andinnerHTMLon the body before.show()(:582-587). Mirror it as#model-modalon the dashboard.No new endpoint. The popup's fields all come from the
per_modelrow the pane already has in hand — the click handler passes the row object it rendered, so there is no fetch, no loading state, and no second source of truth to drift. That is also why the new aggregate columns in "Metrics support" below are added toper_modelitself rather than to a detail route. This is the operator's own proposal and it is the right one: attribution ratio in particular is a diagnostic on a 750x-between-models, 1.8x-within-model quantity, which is an inspection surface, not a ranking. -
Cloud pane excludes
ollama-local. Filter on provider, not on whether energy is present. -
renderCloudEnergy's 5-decimal summary trio is cut. Period spend now has a first-class home on/admin/quota; a second, differently-windowed copy of it on the dashboard is how the two drift apart.
Metrics support
metrics.per_model (src/metrics.py:1386, def per_model) adds
sum_prompt_tokens, sum_completion_tokens, sum_cached_prompt_tokens, and
cost_rows / total_rows so the UI can tell "no cost reported" from "cost was
zero". Stop COALESCEing cost and energy to 0 — return NULL and let the
consumer render —.
That coalesce has a second consumer, and removing it crashes the TUI rather
than merely uglifying it. src/tui_model.py:108-116 maps sum_cost_usd /
sum_energy_kwh / sum_carbon_g_co2eq straight through to a row dict, and
src/tui.py then renders three cells from it — but they are not equally safe:
| line | code | on None |
|---|---|---|
src/tui.py:667 |
_fmt_usd(r["cost_usd"]) |
safe — _fmt_usd returns "n/a" (src/tui.py:55-58) |
src/tui.py:668 |
f"{r['energy_kwh']:.6g}" |
TypeError |
src/tui.py:669 |
f"{r['carbon_g_co2eq']:.4g}" |
TypeError |
The COALESCE is the only thing standing between those two format strings and
an unhandled exception, so the fix for 668-669 is not "render —" as a styling
choice — it is required for the TUI to start. Give both the same None-aware
treatment _fmt_usd already has (a _fmt_num(v, spec) helper returning "n/a"
on None is the smallest change that covers both), and add a per_model fixture
row with NULL energy to tests/test_tui.py so nothing re-introduces a bare
format on a nullable column.
Tests
New or extended, offline, matching this repo's existing fixtures:
tests/test_metrics.py— the existing quota tests (49 references) are rewritten against the new shape, not deleted. Add: one account per configured provider;shape/fidelity/credit.meaningvalues;pace_ratiois null inside the first 2% of a period and correct outside it;spend.total_usdsums only per-provider spend and never a balance; nototal_balance_usdkey survives anywhere in the payload; a NULLtotal_credits_usdyields a pool block with noused_fraction.- A
stale_readingcase — a balance series that is flat and old raises the alarm; flat but fresh does not. This is the 18-hour flat-pool state observed on the live DB, and it is the one alarm nothing covers today. - Dispatcher cost capture — four cases: flag off → no
usagekey inupstream_body; flag on, buffered →cost_usdfromusage.cost; flag on, streamed →cost_usdnon-null (the trap in work item 1); NeuralWatt'scost.request_cost_usdstill wins when both blocks are present. - Poller — three-tuple parser, the two new columns persisted, idempotent migration against a table that predates them, and a parser returning NULLs for pool size.
- Admin —
/admin/quotaserves the file;/admin/api/quotareturns the documented keys. - TUI — update the quota panel's fixtures; keep
tests/test_tui_schema_drift.pyandtests/test_tui_warnings.pygreen.test_tui_warnings.py's fixture is a coupled system —CLAUDE.mdrecords that adding a seed there can silence an unrelated warning class — so changing the plan-usage warning means re-checking both directions of that test, not just the failing one.
Docs to update in the same change
CLAUDE.md— thequota.by_provider/total_balance_usdparagraph in "Billing is per-kWh", and thecredit_attenuationsection's two-semantics note, which is built ontotal_balance_usdexisting.docs/api.md— the/metricsquotablock.docs/admin-portal.md— the new page, the chip's new meaning, the reworked per-model pane.docs/data-model.md— three new columns.plans/admin-design-standards.md— only if the pace-marker gauge or the depletion gauge introduces a pattern the next page should reuse. It probably does; record the grammar (gauge + reference tick), not the CSS.
Out of scope, recorded so it is not rediscovered
- The 4.14x est-vs-billed bias. Surfaced by work item 3, fixed elsewhere.
- Replacing
routing.assumed_cache_ratewith measuredcached_tokens. Work item 1 starts capturing the data; consuming it changes pricing and therefore routing, which needs its own plan and its own before/after. - Energy-budget attenuation (a
credit_attenuationanalogue keyed on plan pace). The pace signal this plan computes is its natural input, but it is a routing feature. - Backfilling OpenRouter cost. Not possible per request; Rule 1 is the reason it is not needed for account totals.