Files
6krrt/plans/quota-multi-provider-redesign.md
adlee-was-taken fd038125f7 feat(admin): the dashboard becomes a landing page, with Board and Live views
The dashboard was eight cards restating what the other seven pages own, so
it was the page you passed through rather than the one you started from. It
is now the landing page: no nav entry of its own, reached by the logo, with
two views over the one /api/snapshot payload so switching costs no request.

Board is one tile per destination, identical anatomy every time: what it is,
one number, one qualifier, a link. Under them, a full-width Activity band
plots routed decisions against billed requests on one axis, so the gap
between them -- local dispatch, cache hits, requests that never reached a
provider -- is readable, which it is in neither series alone. Its range
(24h/7d/30d) is separate from the tile sparklines, because it is the chart
you come to the page to read. Beside it, the eight busiest models ranked by
CALLS, not cost: cost coverage is partial, and an OpenRouter row with two
priced calls out of hundreds otherwise outranks the model doing the work.

Live loads the row itself -- category, model, provider, required context,
tier, classify time, exploratory flag, cost -- so the usual question does
not need the Decisions page, and groups rows by minute so a burst reads as
one. The dot is the classification source: green for a real classification,
amber for a degraded one, which still routes but never moves proficiency,
and until now was visible only in a /metrics warning. The rail beside it
answers what Board cannot: throughput against the last hour, classifier p50
and p95 (the latency floor every routed request pays before an upstream
token), degraded share, and the top five models by share of the live window
keyed on model AND provider, since the same model on two providers is two
different routing outcomes. Spend is not restated here; it belongs to Quota.

Warnings move behind a navbar bell. They are recomputed from live state on
every poll, so the only honest dismissal is "hide until the condition
changes": a dismissal is keyed on the warning's TEXT, so when the numbers
move the key stops matching and the warning returns on its own. Nothing has
to expire it. Per browser, in localStorage, wrapped in try/catch -- a
reading state, not a fact about the router.

Two tests changed shape rather than being deleted. The verdict-bar tests
pinned a card this page no longer has, but the lesson they recorded is the
expensive part and is preserved verbatim in the replacement's docstring:
`unverifiable` is not a finding, it is the default for prose and for any
turn ending in a tool call, and any bar scaled against it renders every
meaningful verdict at the 2px minimum. Exclude it from the scale, never from
the denominator. The Live rail states the same fact as one ratio, which
needs no scale. The nav test's rule is inverted for index.html rather than
skipped: the landing page must have no nav entry AND must claim no active
one, so neither can drift back in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-12 00:43:55 -04:00

35 KiB
Raw Permalink Blame History

Quota, redesigned around billing shape (subscription vs paygo vs self-hosted)

Status: done -- shipped as the quota page redesign; decision-complete. Written 2026-09-10 from the live database, the live /metrics, and two live probes against the OpenRouter API.

Intended executor: opencode (Prometheus → Atlas). Every open question in this document is closed; where a choice existed, the choice and its reason are recorded rather than left to the implementer.

Read plans/admin-design-standards.md and plans/admin-work-framework.md before touching admin/frontend/*. This document does not restate the visual language; it specifies what the pages must say.

On the file:line anchors below. Every anchor in the first draft of this document was taken against a working tree five commits behind origin/main, and PR #76 (feat/mangled-output-breaker, merged 2026-09-10T23:48Z) added ~77 lines to src/metrics.py in between. Four anchors were wrong as a result. They are corrected here and re-verified against f89cefc, but the lesson stands: treat every anchor as a hint and confirm it by grepping the landmark named beside it. Each anchor below is paired with the symbol or string it points at, so the grep is always available.

Why this exists

The word "Quota" currently covers three unrelated quantities, and the portal sums two of them into a figure CLAUDE.md itself disclaims as "not a single spendable figure". Meanwhile the one number an operator asks for first — money spent this month — appears nowhere in the portal at all.

All evidence below is from the live system on 2026-09-10/11, billing period 2026-09-06 → 2026-10-06.

Defect 1 — one word, three quantities

metrics.quota_burn returns a kWh-plan report with a per-provider balance map bolted onto it (src/metrics.py:334, merging quota_balance_and_burn at src/metrics.py:203). The live payload carries, under one key:

field what it actually is magnitude
metered_fraction_of_plan NeuralWatt kWh subscription consumption 0.4607
by_provider.neuralwatt.balance_usd overage-billing accounting artifact −$0.004198
by_provider.openrouter.balance_usd a real prepaid credit pool $40.45
total_balance_usd the sum of the previous two $40.4495

Those are a gauge, a rounding residue, and a wallet. The sum is not a quantity.

Defect 2 — money spent is not in the portal

Actual billed spend this period, from the live DB:

provider period spend source
neuralwatt $22.66 SUM(energy_observations.cost_usd), 6,079 rows
openrouter $9.57 account poll: total_credits 50 − total_usage 9.5714
ollama-local $0.0027 local_energy_observations.cost_usd, 202 rows

≈ $32.23 this period, and the dashboard renders none of it. The NeuralWatt figure has been captured per request since 2026-08-17 and is read by exactly nothing.

Defect 3 — OpenRouter has no per-request cost, so it renders as free

extract_telemetry (src/dispatcher.py:2004) reads only NeuralWatt's top-level energy / cost blocks. Every OpenRouter row therefore lands with cost_usd IS NULL: 3,106 rows this period. metrics.per_model (src/metrics.py:1386, def per_model) COALESCEs those NULLs to 0, so the live dashboard shows:

xiaomi/mimo-v2.5   openrouter   calls=655   cost=$0.0000   kwh=0.00000

That model is #8 by volume and cost real money — the pool moved $1.23 on 2026-09-10 alone, the day it served 598 of those calls at ~175k prompt tokens each.

This is fixable, and it was verified live rather than assumed. Sending usage: {"include": true} makes OpenRouter return billed cost per request. Probed 2026-09-10 against xiaomi/mimo-v2.5:

{"prompt_tokens": 254, "completion_tokens": 8, "cost": 1.14576e-05,
 "prompt_tokens_details": {"cached_tokens": 192},
 "cost_details": {"upstream_inference_cost": 1.14576e-05,
                  "upstream_inference_prompt_cost": 9.2176e-06,
                  "upstream_inference_completions_cost": 2.24e-06}}

A :free model returns the same shape with cost: 0, correctly. Note cached_tokens: 192/254 — a measured cache rate on a field routing.assumed_cache_rate currently guesses at.

Defect 4 — the plan gauge hides the only actionable number

The chip renders 46.1% · 6.25 kWh plan. At the time of writing, 5.01 of the period's 30 days had elapsed — 16.7%. Burning 46.1% of the allowance in 16.7% of the period is 2.76x sustainable pace, projecting ~17.2 kWh against a 6.25 kWh plan.

That ratio is the number the operator needed on 2026-09-06 and computed by hand against metered_kwh_period and elapsed time. Nothing in the portal computes it. A percentage without a pace reads as reassuring at exactly the moment it should not.

Defect 5 — the electricity the operator personally pays for is not in Quota

local_energy_summary (src/metrics.py:1547) exists and is rendered in its own dashboard card, but it is absent from Quota — while ollama-local rows are present in energy_observations (60 rows) and therefore in the "Cloud energy" pane, which double-counts them and files self-hosted draw under cloud.

Also found, and deliberately NOT fixed here

Joined on request_id over this period, route_decisions.est_cost_usd overprices NeuralWatt 4.14x against its own billed cost (n=4,958), while OpenRouter's estimate lands within ~13% of the observed pool delta:

model n est billed ratio
kimi-k2.7-code 2,853 $62.78 $11.21 5.60
qwen3.6-35b 1,024 $4.01 $0.48 8.40
glm-5.3 288 $14.75 $6.76 2.18
deepseek-v4-flash 714 $2.72 $1.15 2.36

This is expected in direction — NeuralWatt bills min($8/kWh × kWh, 3 × list) and models that batch well land far under list, which is the whole subject of CLAUDE.md's cost section — but the size means the cost tiebreak systematically overprices NeuralWatt relative to OpenRouter by roughly 4x. That is a routing-correctness question, not a dashboard one.

Scope decision: this plan surfaces the ratio and changes no routing. Work item 3 puts est-vs-billed on the page so the bias is visible and measurable; acting on it belongs in a separate plan, because the fix is either a per-provider calibration factor or a switch to billed-cost feedback, and neither should be decided inside a dashboard change.

The model: three billing shapes, one comparable unit

The frankenstein comes from trying to express three billing shapes in one schema. The redesign names the shapes and gives each its native unit, then uses dollars — the only unit all three share — for the cross-provider view.

shape provider allowance remaining spend truth granularity
metered_plan neuralwatt 6.25 kWh/period plan − metered kWh cost_usd per request per request
prepaid_credit openrouter $50 pool total_credits − total_usage account poll; per-request after work item 1 2 h (account), per request (attribution)
self_hosted ollama-local none — no cap exists n/a GPU draw × tariff_usd_per_kwh per request

Two rules fall out of that table, and both are load-bearing:

Rule 1 — account truth for totals, request truth for attribution, never add them. An account poll is authoritative for "what did this provider charge me"; summed per-request costs are authoritative for "which model spent it". They are two measurements of overlapping quantities at different fidelities. Adding them double-counts; preferring the request sum for a total silently undercounts whatever the router did not observe (plan_kwh_per_period's own note already says as much: "router-metered only"). So: totals come from the account, breakdowns come from requests, and the page states when the breakdown covers less than the total.

Rule 2 — a missing number is not zero. COALESCE(SUM(cost_usd), 0) is how OpenRouter came to read $0.0000. Unknown renders as — with a reason, everywhere, in every card and every row.

Work item 1 — capture per-request cost where the provider reports it

Config (src/config.py, config/config.yaml). Add to DispatchProvider:

# dispatch_providers.openrouter
  # OpenRouter returns billed cost per request when the request asks for usage
  # accounting. Verified live 2026-09-10: usage.cost, usage.cost_details and
  # usage.prompt_tokens_details.cached_tokens all come back populated.
  reports_cost_in_usage: true

Default false. It must appear in config/config.yaml explicitly for both providers, not only as a Pydantic default — CLAUDE.md's "every knob belongs in config.yaml" rule, and the reason outcome_attribution_window_seconds was deleted.

Gate the request-body change on this flag rather than sending usage: {"include": true} unconditionally: NeuralWatt is OpenAI-compatible and an unknown top-level key is a plausible 400 on a provider we have no reason to probe.

Request (src/dispatcher.py:4212, where upstream_body is built):

if getattr(settings, "reports_cost_in_usage", False):
    upstream_body.setdefault("usage", {"include": True})

setdefault, so a client that sent its own usage block wins.

Extraction (src/dispatcher.py:2004). Extend extract_telemetry to fall back to the usage block when NeuralWatt's cost block is absent:

  • cost_usd ← cost.request_cost_usd, else usage.cost
  • new Telemetry.cached_prompt_tokens ← usage.prompt_tokens_details.cached_tokens

Keep the precedence in that order. A provider that reports both is reporting the same number twice, and request_cost_usd is the field this project has already validated against billing.

THE TRAP — the streamed path does not pass usage to extract_telemetry, and every agent client streams. On the buffered path (src/dispatcher.py:4385) extract_telemetry(payload) sees the whole response including usage. On the streamed path (src/dispatcher.py:4637) it is called as extract_telemetry(collected), where collected holds only the sniffed NeuralWatt SSE comment lines; the streamed usage dict lives in a separate local (src/dispatcher.py:4603, populated from chunk["usage"]) and is passed to log_observation for token counts only.

So a change that only touches extract_telemetry fixes cost for buffered requests and leaves it NULL for essentially all real traffic. stream_options.include_usage is already set (src/dispatcher.py:4217), so the final chunk does carry it; it simply never reaches the extractor.

The fix, decided — not implementer's choice. Signature becomes:

def extract_telemetry(payload: dict, *, usage: Optional[dict] = None) -> Telemetry:
  • When usage is not passed, the function falls back to payload.get("usage"). That alone fixes every buffered call site with no change at the call sites at all, because payload there is the whole response.
  • The streamed call site passes it explicitly — extract_telemetry(collected, usage=usage) — because collected holds only sniffed SSE comment blocks and structurally cannot contain a usage dict.
  • Cost precedence inside the function, in this order: payload["cost"]["request_cost_usd"], then usage["cost"]. First non-None wins.

The rejected alternative was extract_telemetry({**collected, "usage": usage}). It needs no signature change, which is why it is tempting, and it is wrong: collected is by contract the output of _sniff_telemetry_line (src/dispatcher.py:3735), whose keys are provider SSE comment names. Injecting a synthetic usage key into it makes that dict a half-response/half-sniff hybrid, and the next person to read either function has to hold both meanings at once. The keyword argument keeps the two sources named.

Note the precedence order is what makes the existing NeuralWatt behaviour unchanged: NW streaming populates collected["cost"], so it never reaches the usage branch, and its cost_usd stays the figure this project has already validated against billing. The streamed path must be covered by a test asserting a non-null cost_usd from usage.cost alone.

Audit the other log_observation call sites while in here (src/dispatcher.py:4871 and the local-dispatch path) so no branch silently drops the new field.

Schema. energy_observations gains one nullable column:

ALTER TABLE energy_observations ADD COLUMN cached_prompt_tokens INTEGER;

Add it to config/schema.sql with a comment recording why it exists (a measured cache rate against routing.assumed_cache_rate's guess) and that NULL means "provider did not report", not zero.

What this does NOT do: it does not backfill. OpenRouter cost exists from the deploy forward only, which is precisely why Rule 1 keeps account polls as the source of account totals.

Work item 2 — store the credit pool, not just the remainder

_BALANCE_PARSERS["openrouter"] (src/poller.py:66) computes total_credits − total_usage and discards both inputs. The pool size is what turns a bare "$40.45" into "used $9.57 of $50", and it is already in the response.

  • provider_balance_observations gains two nullable columns: total_credits_usd, total_usage_usd. Extend _ensure_provider_balance_table (src/poller.py:422) with idempotent ALTER TABLE ... ADD COLUMN guarded the way the rest of this project guards live-DB migrations, and mirror them into config/schema.sql.
  • The parser registry changes shape: a parser returns (balance_usd, total_credits_usd | None, total_usage_usd | None). record_balance (src/poller.py:449) persists all three. A provider whose endpoint reports only a remainder stores NULL for the other two, and the UI falls back to a bare remaining figure.
  • tests/test_multi_provider_poller.py:728-731 pins _BALANCE_PARSERS.keys() == config.PROVIDERS_WITH_BALANCE_PARSERS (src/config.py:1056); keep that pin working. It is in that file rather than a poller-parsing one, which is easy to miss when scanning by name.

Nullable columns because a NULL pool size must render as "balance $40.45" and never as "$40.45 of $0".

Work item 3 — the /metrics quota block, rewritten

Replace quota_burn and quota_balance_and_burn with one function, metrics.quota_accounts(conn, cfg). No deprecated aliases — CLAUDE.md records that the last quota reshape updated every consumer in the same change, and that is the standard here too.

Shape

"quota": {
  "period": {
    "start": "2026-09-06",
    "next_reset": "2026-10-06",
    "elapsed_fraction": 0.167,
    "source": "objective.billing_reset_day"
  },
  "accounts": [
    {
      "provider": "neuralwatt",
      "shape": "metered_plan",
      "plan": {
        "unit": "kwh",
        "allowance": 6.25,
        "used": 2.87933,
        "used_fraction": 0.4607,
        "pace_ratio": 2.76,
        "projected_period_total": 17.24,
        "pace_note": null
      },
      "spend_usd": {
        "period": 22.66,
        "window_30d": 61.2,
        "source": "per_request_billed",
        "attribution_coverage": 1.0
      },
      "credit": {
        "balance_usd": -0.004198,
        "meaning": "overage_allowance",
        "observed_at": "2026-09-10T18:49:43Z",
        "age_seconds": 19980
      },
      "fidelity": "per_request"
    },
    {
      "provider": "openrouter",
      "shape": "prepaid_credit",
      "pool": {
        "total_credits_usd": 50.0,
        "used_usd": 9.5714,
        "remaining_usd": 40.4286,
        "used_fraction": 0.1914,
        "observed_at": "2026-09-10T23:17:32Z",
        "age_seconds": 7200
      },
      "burn": {
        "window_hours": 24,
        "usd_per_hour": 0.0337,
        "projected_hours_remaining": 1200.9,
        "runway_low_warning": false,
        "note": null
      },
      "spend_usd": {
        "period": 9.5714,
        "window_30d": 9.5714,
        "source": "account_poll_delta",
        "attribution_coverage": 0.0
      },
      "fidelity": "account_poll"
    },
    {
      "provider": "ollama-local",
      "shape": "self_hosted",
      "energy": { "unit": "kwh", "period": 0.01725, "window_30d": 0.0402 },
      "spend_usd": {
        "period": 0.00274,
        "window_30d": 0.0064,
        "source": "derived_from_draw",
        "attribution_coverage": 1.0
      },
      "tariff_usd_per_kwh": 0.15856,
      "fidelity": "derived"
    }
  ],
  "spend": {
    "by_provider_usd": { "neuralwatt": 22.66, "openrouter": 9.5714, "ollama-local": 0.00274 },
    "total_usd": 32.23,
    "estimated_usd": { "neuralwatt": 86.05, "openrouter": 10.07 },
    "estimate_ratio": { "neuralwatt": 4.14, "openrouter": 1.13 }
  },
  "alarm": {
    "kind": "plan_pace",
    "provider": "neuralwatt",
    "severity": "warn",
    "headline": "neuralwatt: 2.8x sustainable pace, projecting 17.2 kWh against a 6.25 kWh plan"
  },
  "note": "router-metered only; traffic bypassing the router is not counted"
}

Decisions embedded in that shape

  • shape is derived from config, never hardcoded per provider name. The rule, in precedence order, evaluated per cfg.dispatch_providers entry plus the local-dispatch provider:

    # test shape
    1 provider is the local-dispatch provider (ollama-local) self_hosted
    2 has_energy_telemetry and objective.plan_kwh_per_period is set metered_plan
    3 balance_url is set prepaid_credit
    4 none of the above unmetered

    First match wins, and the order matters: a provider could satisfy both 2 and 3, and the plan is the thing with a period and a ceiling, so it leads. unmetered is a real state, not an error — a configured provider with no telemetry, no plan and no balance endpoint genuinely has nothing to meter, and it renders as a name plus its per-request spend if any, with no gauge. A metered_plan provider when plan_kwh_per_period is null degrades to case 3 or 4 rather than rendering a gauge against a null allowance.

  • spend.by_provider_usd[p] is exactly accounts[p].spend_usd.period, and spend.total_usd is the sum of those values — nothing else. It is never computed from a second query, and a balance figure never enters it. This is Rule 1 as an implementation constraint: each account has already decided whether its own period spend comes from summed per-request cost, an account poll delta, or derived draw, and the ledger just adds up those decisions. estimated_usd[p] is a separate SUM(route_decisions.est_cost_usd) over the same period and provider, and estimate_ratio[p] is estimated_usd[p] / by_provider_usd[p], null when the denominator is 0.

  • accounts is a list, not a map keyed by provider, ordered by alarm severity then provider name. The old by_provider map forced every consumer to re-derive which provider mattered; the list makes the answer positional and makes the TUI's row order free.

  • total_balance_usd is deleted. It summed a prepaid pool with an overage residue. spend.total_usd replaces it as the one honest cross-provider number, because dollars spent are additive across billing shapes.

  • shape drives rendering, fidelity qualifies it. A consumer switches layout on shape and prints a caveat from fidelity; neither is inferred from provider name or from which config keys happen to be set.

  • spend_usd.source is one of per_request_billed, account_poll_delta, derived_from_draw. attribution_coverage is the fraction of that provider's period spend that per-request rows can account for — 1.0 for NeuralWatt, 0.0 for OpenRouter until work item 1 lands, and a partial value for the period that straddles the deploy. This is Rule 1 made machine readable: it is how a card knows to say "breakdown covers $4.10 of $9.57".

  • credit.meaning replaces balance_source. telemetry / polled / unconfigured described where the number came from; overage_allowance / prepaid_pool / not_configured describe what it means, which is what decides whether it deserves a gauge or a footnote.

pace_ratio — the computation, and its trap

pace_ratio = used_fraction / elapsed_fraction, with elapsed_fraction = (now − period_start) / (next_reset − period_start).

Trap: elapsed_fraction → 0 at period start, and the ratio explodes. In the first hours of a period a handful of calls reads as a 40x overrun. Guard it: when elapsed_fraction < 0.02 (roughly the first 14 hours of a 30-day period), return pace_ratio: null and projected_period_total: null with pace_note: "too early in the period to project". A null pace must render as "—", never as 0 or 1.0.

Second trap, and this project has been burned by it once already. The projection must not read as a wall. objective.plan_kwh_per_period gates nothing; overage is billed. plans/quota-balance-and-burn-rate.md records the last warning that got this wrong ("a quota is a wall, not a bill — requests fail rather than costing more") and why it was both alarming and false. Project in kWh and say what it costs, never "requests will fail".

alarm — what the chip shows

One alarm, the worst across accounts, or null:

  • plan_pace: pace_ratio >= objective.plan_pace_warn_ratio (new knob, default 1.25, which must be written into config/config.yaml, not left as a Pydantic default). severity: "critical" at >= 2.0.
  • credit_runway: existing projected_hours_remaining < objective.quota_runway_warning_hours (6). critical below half that.
  • stale_reading: a prepaid_credit account whose age_seconds exceeds 3x the poll interval. This one is new and it matters: the live pool sat flat at $40.453746437 for six consecutive polls across 18 hours while 705 completions ran, and the current UI renders that as a healthy account with no burn. It is poll granularity, not calm. A flat reading and a stale reading must be distinguishable on the page.

coverage.warnings (src/metrics.py:493) currently emits the >80%-of-plan string. Replace it with a pace-based warning sourced from alarm, so the navbar bell and the chip cannot disagree.

Consumers to update in the same change

Measured inventory — grep -rln "total_balance_usd\|by_provider" over the repo. There is no deprecation path, so all of these move together:

file what breaks
src/metrics.py:203,334 the two functions being replaced
src/dispatcher.py /metrics assembly
src/admin.py:1849 /api/snapshot's quota key, plus the new GET /admin/api/quota
src/tui_model.py:41-60 builds {label, value} quota rows from the flat keys
src/tui.py:78,105,566,601,615 per-provider row formatting, the kWh lead line, and the total_balance_usd line CLAUDE.md documents
tests/test_metrics.py 49 references — the bulk of the work
tests/test_metrics_endpoint.py 6
tests/test_tui.py 6
tests/test_credit_attenuation_routing.py 6 — asserts attenuation reads balances, so it must keep passing against the new shape without weakening what it pins
CLAUDE.md, docs/api.md, docs/admin-portal.md describe the old contract by name

test_credit_attenuation_routing.py deserves a second look rather than a mechanical rename: credit_attenuation derives a routing multiplier from polled balances, so its fixtures encode the old assumption that a balance is one flat number per provider. Its behaviour must not change here — only where it reads the balance from.

Work item 4 — /admin/quota, a page

The chip-plus-modal cannot hold three shaped cards and a ledger. Follow the portal's page-per-domain pattern (providers, proficiency, profiles).

Mechanics (src/admin.py:823-870): add _admin_quota → admin/frontend/quota.html, a @router.get("/quota") returning it with _NO_CACHE_HEADERS, and a nav item after Providers in the navbar of all admin pages (admin/frontend/*.html:281-287) — the navbar is duplicated per page, so a nav item added to one page only is a half-shipped nav.

Data: GET /admin/api/quota → metrics.quota_accounts(conn, cfg) plus the spend ledger's per-model rows. A dedicated endpoint rather than reusing /api/snapshot, because the page needs per-model spend the snapshot does not carry and none of the snapshot's proficiency/verdict payload.

Cards, one per account, laid out by shape

Three peer cards on one row is .col-xl-4 — the Standards doc is explicit that .col-xl-6 strands a third card and opens two dead columns.

metered_plan (NeuralWatt). The kWh gauge is the hero, and the pace marker is what makes it a redesign rather than a reskin: draw the gauge with a tick at elapsed_fraction so "burned" and "elapsed" are visually comparable in one glance. Below it: pace ratio, projected period total, $ billed this period, and the overage allowance as a de-emphasized footnote labeled overage allowance — never as a balance or a runway, because at −$0.004 it is neither.

prepaid_credit (OpenRouter). A depletion gauge: used_usd of total_credits_usd, remaining in the hero position. Then burn $/hr, runway in days (hours only under ~48h), and the reading age rendered as text ("polled 2h ago"), promoted to a warning tint past the staleness threshold. When total_credits_usd is NULL, degrade to a bare remaining figure with no gauge.

self_hosted (local). kWh and $ this period, the tariff it was priced at, and a one-line statement that this is billed to the operator's utility rather than a provider. No gauge — there is no allowance to fill. Hidden entirely when local_energy.enabled is false.

Spend ledger

Below the cards, full width:

  1. spend.total_usd as the page's single headline number, with a stacked bar by provider. This is the answer to "what did I spend this month" and it is the reason the page exists.
  2. Ranked per-model spend — model, provider, calls, $, share of total. — in the $ column where cost is unknown, with the row still present and ranked by calls. Where attribution_coverage < 1, a single line above the table states the gap in dollars, e.g. "breakdown accounts for $4.10 of OpenRouter's $9.57; the remainder predates per-request cost capture".
  3. Estimate calibration, a small two-column strip: est vs billed per provider with the ratio. Label it as what it is — the cost tiebreak's input measured against the bill — and do not editorialize a fix; that is the follow-up plan.

Chip

renderQuotaChip (admin/frontend/index.html:779) stops hardcoding the kWh story. It renders alarm: headline text, severity tint, and a link to /admin/quota. With alarm: null it shows the period's spend to date, which is calm and still useful. Its comment at line 790 ("The chip always tells the kWh-subscription story") is the assumption being deleted — remove it, do not leave it stale.

The chip stops being a modal trigger and becomes a link. Four edits, all in index.html, and leaving any one of them behind leaves a dead control:

  1. Line 305 — the chip <div> carries data-bs-toggle="modal" data-bs-target="#quota-modal" role="button" tabindex="0". Replace the element with an <a href="/admin/quota"> carrying the same .quota-chip class, and drop all four attributes: an anchor is focusable and activatable on its own, so role/tabindex become wrong rather than merely redundant. The navbar's z-index: 1030 note in the Standards doc exists because of this chip's backdrop-filter — keep the stacking context intact.
  2. Lines 462-478 — the #quota-modal block is deleted outright.
  3. renderQuotaModal (:659) and its #quota-modal-body writes are deleted, and the renderQuotaModal(data.quota) call at :558 with it.
  4. grep -n 'quota-modal' admin/frontend/*.html must come back empty when done. formatProviderBalance / formatBalanceStaleness (:708, :731) move to the new page rather than being deleted — the staleness helper is exactly what the prepaid_credit card's reading-age line needs, and its tiny-negative-balance handling is hard-won (it exists so NeuralWatt's −$0.004 never renders as "-$0.00").

Work item 5 — the per-model usage pane

Two panes read data.per_model today: renderCloudEnergy (admin/frontend/index.html:906) and renderBars (:935). Three defects, each falsifiable:

  1. Bar length and bar label disagree. renderBars ranks by BARS_SORT_KEYS[_barsSortKey] (:840) — calls, cost or energy — but the value label is hardcoded to $cost · kWh (:957). Sorting by calls produces bars ordered by call count, labeled in dollars.
  2. Unknown renders as zero, giving xiaomi/mimo-v2.5 a $0.0000 label on 655 real calls (Defect 3 above).
  3. ollama-local appears in a card called "Cloud energy", double-counting against the Local Compute card that owns it.

And a fourth, which is the operator's own framing: calls, cost and energy compete for one row of space, and energy is the least actionable of the three.

The rework

  • One ranked list. Sort keys become calls | spend | tokens. Energy leaves the sort keys: it is not a thing an operator ranks models by, it is a thing they inspect. tokens replaces it because it is the honest volume metric for a provider that reports no energy at all.

  • The label follows the sort key, plus share of total: sorting by calls shows 655 calls · 8.2%; by spend, $1.23 · 3.8%. One metric, the one being ranked.

  • — for unknown, never $0.0000, with a footnote naming the reason ("provider does not report per-request cost").

  • Energy, carbon, attribution and grid move to a row-click detail popup — kWh, gCO2eq, grid_id, mean attribution_ratio, mean completion tokens, cached-token share once work item 1 lands. Render only the fields that have data; a popup of six —s is worse than a popup that admits it has nothing.

    Copy the pattern that already exists rather than inventing one. admin/frontend/proficiency.html does exactly this for its matrix cells: a <div class="modal modal-blur fade" id="cell-modal"> with a -title and -body (:264-276), a lazily-constructed new tabler.Modal(...) held in a module-level _cellModal (:575), and a click handler that sets textContent on the title and innerHTML on the body before .show() (:582-587). Mirror it as #model-modal on the dashboard.

    No new endpoint. The popup's fields all come from the per_model row the pane already has in hand — the click handler passes the row object it rendered, so there is no fetch, no loading state, and no second source of truth to drift. That is also why the new aggregate columns in "Metrics support" below are added to per_model itself rather than to a detail route. This is the operator's own proposal and it is the right one: attribution ratio in particular is a diagnostic on a 750x-between-models, 1.8x-within-model quantity, which is an inspection surface, not a ranking.

  • Cloud pane excludes ollama-local. Filter on provider, not on whether energy is present.

  • renderCloudEnergy's 5-decimal summary trio is cut. Period spend now has a first-class home on /admin/quota; a second, differently-windowed copy of it on the dashboard is how the two drift apart.

Metrics support

metrics.per_model (src/metrics.py:1386, def per_model) adds sum_prompt_tokens, sum_completion_tokens, sum_cached_prompt_tokens, and cost_rows / total_rows so the UI can tell "no cost reported" from "cost was zero". Stop COALESCEing cost and energy to 0 — return NULL and let the consumer render —.

That coalesce has a second consumer, and removing it crashes the TUI rather than merely uglifying it. src/tui_model.py:108-116 maps sum_cost_usd / sum_energy_kwh / sum_carbon_g_co2eq straight through to a row dict, and src/tui.py then renders three cells from it — but they are not equally safe:

line code on None
src/tui.py:667 _fmt_usd(r["cost_usd"]) safe — _fmt_usd returns "n/a" (src/tui.py:55-58)
src/tui.py:668 f"{r['energy_kwh']:.6g}" TypeError
src/tui.py:669 f"{r['carbon_g_co2eq']:.4g}" TypeError

The COALESCE is the only thing standing between those two format strings and an unhandled exception, so the fix for 668-669 is not "render —" as a styling choice — it is required for the TUI to start. Give both the same None-aware treatment _fmt_usd already has (a _fmt_num(v, spec) helper returning "n/a" on None is the smallest change that covers both), and add a per_model fixture row with NULL energy to tests/test_tui.py so nothing re-introduces a bare format on a nullable column.

Tests

New or extended, offline, matching this repo's existing fixtures:

  • tests/test_metrics.py — the existing quota tests (49 references) are rewritten against the new shape, not deleted. Add: one account per configured provider; shape/fidelity/credit.meaning values; pace_ratio is null inside the first 2% of a period and correct outside it; spend.total_usd sums only per-provider spend and never a balance; no total_balance_usd key survives anywhere in the payload; a NULL total_credits_usd yields a pool block with no used_fraction.
  • A stale_reading case — a balance series that is flat and old raises the alarm; flat but fresh does not. This is the 18-hour flat-pool state observed on the live DB, and it is the one alarm nothing covers today.
  • Dispatcher cost capture — four cases: flag off → no usage key in upstream_body; flag on, buffered → cost_usd from usage.cost; flag on, streamed → cost_usd non-null (the trap in work item 1); NeuralWatt's cost.request_cost_usd still wins when both blocks are present.
  • Poller — three-tuple parser, the two new columns persisted, idempotent migration against a table that predates them, and a parser returning NULLs for pool size.
  • Admin — /admin/quota serves the file; /admin/api/quota returns the documented keys.
  • TUI — update the quota panel's fixtures; keep tests/test_tui_schema_drift.py and tests/test_tui_warnings.py green. test_tui_warnings.py's fixture is a coupled system — CLAUDE.md records that adding a seed there can silence an unrelated warning class — so changing the plan-usage warning means re-checking both directions of that test, not just the failing one.

Docs to update in the same change

  • CLAUDE.md — the quota.by_provider / total_balance_usd paragraph in "Billing is per-kWh", and the credit_attenuation section's two-semantics note, which is built on total_balance_usd existing.
  • docs/api.md — the /metrics quota block.
  • docs/admin-portal.md — the new page, the chip's new meaning, the reworked per-model pane.
  • docs/data-model.md — three new columns.
  • plans/admin-design-standards.md — only if the pace-marker gauge or the depletion gauge introduces a pattern the next page should reuse. It probably does; record the grammar (gauge + reference tick), not the CSS.

Out of scope, recorded so it is not rediscovered

  • The 4.14x est-vs-billed bias. Surfaced by work item 3, fixed elsewhere.
  • Replacing routing.assumed_cache_rate with measured cached_tokens. Work item 1 starts capturing the data; consuming it changes pricing and therefore routing, which needs its own plan and its own before/after.
  • Energy-budget attenuation (a credit_attenuation analogue keyed on plan pace). The pace signal this plan computes is its natural input, but it is a routing feature.
  • Backfilling OpenRouter cost. Not possible per request; Rule 1 is the reason it is not needed for account totals.