Files
6krrt/plans/quota-multi-provider-redesign.md
adlee-was-taken fd038125f7 feat(admin): the dashboard becomes a landing page, with Board and Live views
The dashboard was eight cards restating what the other seven pages own, so
it was the page you passed through rather than the one you started from. It
is now the landing page: no nav entry of its own, reached by the logo, with
two views over the one /api/snapshot payload so switching costs no request.

Board is one tile per destination, identical anatomy every time: what it is,
one number, one qualifier, a link. Under them, a full-width Activity band
plots routed decisions against billed requests on one axis, so the gap
between them -- local dispatch, cache hits, requests that never reached a
provider -- is readable, which it is in neither series alone. Its range
(24h/7d/30d) is separate from the tile sparklines, because it is the chart
you come to the page to read. Beside it, the eight busiest models ranked by
CALLS, not cost: cost coverage is partial, and an OpenRouter row with two
priced calls out of hundreds otherwise outranks the model doing the work.

Live loads the row itself -- category, model, provider, required context,
tier, classify time, exploratory flag, cost -- so the usual question does
not need the Decisions page, and groups rows by minute so a burst reads as
one. The dot is the classification source: green for a real classification,
amber for a degraded one, which still routes but never moves proficiency,
and until now was visible only in a /metrics warning. The rail beside it
answers what Board cannot: throughput against the last hour, classifier p50
and p95 (the latency floor every routed request pays before an upstream
token), degraded share, and the top five models by share of the live window
keyed on model AND provider, since the same model on two providers is two
different routing outcomes. Spend is not restated here; it belongs to Quota.

Warnings move behind a navbar bell. They are recomputed from live state on
every poll, so the only honest dismissal is "hide until the condition
changes": a dismissal is keyed on the warning's TEXT, so when the numbers
move the key stops matching and the warning returns on its own. Nothing has
to expire it. Per browser, in localStorage, wrapped in try/catch -- a
reading state, not a fact about the router.

Two tests changed shape rather than being deleted. The verdict-bar tests
pinned a card this page no longer has, but the lesson they recorded is the
expensive part and is preserved verbatim in the replacement's docstring:
`unverifiable` is not a finding, it is the default for prose and for any
turn ending in a tool call, and any bar scaled against it renders every
meaningful verdict at the 2px minimum. Exclude it from the scale, never from
the denominator. The Live rail states the same fact as one ratio, which
needs no scale. The nav test's rule is inverted for index.html rather than
skipped: the landing page must have no nav entry AND must claim no active
one, so neither can drift back in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-12 00:43:55 -04:00

733 lines
35 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Quota, redesigned around billing shape (subscription vs paygo vs self-hosted)
Status: done -- shipped as the quota page redesign; decision-complete. Written 2026-09-10 from the live database,
the live `/metrics`, and two live probes against the OpenRouter API.
Intended executor: opencode (Prometheus → Atlas). Every open question in this
document is closed; where a choice existed, the choice and its reason are
recorded rather than left to the implementer.
Read `plans/admin-design-standards.md` and `plans/admin-work-framework.md`
before touching `admin/frontend/*`. This document does not restate the visual
language; it specifies what the pages must say.
**On the file:line anchors below.** Every anchor in the first draft of this
document was taken against a working tree five commits behind `origin/main`,
and PR #76 (`feat/mangled-output-breaker`, merged 2026-09-10T23:48Z) added ~77
lines to `src/metrics.py` in between. Four anchors were wrong as a result. They
are corrected here and re-verified against `f89cefc`, but the lesson stands:
**treat every anchor as a hint and confirm it by grepping the landmark named
beside it.** Each anchor below is paired with the symbol or string it points
at, so the grep is always available.
## Why this exists
The word "Quota" currently covers three unrelated quantities, and the portal
sums two of them into a figure `CLAUDE.md` itself disclaims as "not a single
spendable figure". Meanwhile the one number an operator asks for first — money
spent this month — appears nowhere in the portal at all.
All evidence below is from the live system on 2026-09-10/11, billing period
2026-09-06 → 2026-10-06.
### Defect 1 — one word, three quantities
`metrics.quota_burn` returns a kWh-plan report with a per-provider balance map
bolted onto it (`src/metrics.py:334`, merging `quota_balance_and_burn` at
`src/metrics.py:203`). The live payload carries, under one key:
| field | what it actually is | magnitude |
|---|---|---|
| `metered_fraction_of_plan` | NeuralWatt kWh subscription consumption | 0.4607 |
| `by_provider.neuralwatt.balance_usd` | overage-billing accounting artifact | −$0.004198 |
| `by_provider.openrouter.balance_usd` | a real prepaid credit pool | $40.45 |
| `total_balance_usd` | the sum of the previous two | $40.4495 |
Those are a gauge, a rounding residue, and a wallet. The sum is not a quantity.
### Defect 2 — money spent is not in the portal
Actual billed spend this period, from the live DB:
| provider | period spend | source |
|---|---|---|
| neuralwatt | **$22.66** | `SUM(energy_observations.cost_usd)`, 6,079 rows |
| openrouter | **$9.57** | account poll: `total_credits 50 − total_usage 9.5714` |
| ollama-local | $0.0027 | `local_energy_observations.cost_usd`, 202 rows |
≈ **$32.23 this period**, and the dashboard renders none of it. The NeuralWatt
figure has been captured per request since 2026-08-17 and is read by exactly
nothing.
### Defect 3 — OpenRouter has no per-request cost, so it renders as free
`extract_telemetry` (`src/dispatcher.py:2004`) reads only NeuralWatt's
top-level `energy` / `cost` blocks. Every OpenRouter row therefore lands with
`cost_usd IS NULL`: 3,106 rows this period. `metrics.per_model`
(`src/metrics.py:1386`, `def per_model`) `COALESCE`s those NULLs to 0, so the
live dashboard shows:
```
xiaomi/mimo-v2.5 openrouter calls=655 cost=$0.0000 kwh=0.00000
```
That model is #8 by volume and cost real money — the pool moved $1.23 on
2026-09-10 alone, the day it served 598 of those calls at ~175k prompt tokens
each.
**This is fixable, and it was verified live rather than assumed.** Sending
`usage: {"include": true}` makes OpenRouter return billed cost per request.
Probed 2026-09-10 against `xiaomi/mimo-v2.5`:
```json
{"prompt_tokens": 254, "completion_tokens": 8, "cost": 1.14576e-05,
"prompt_tokens_details": {"cached_tokens": 192},
"cost_details": {"upstream_inference_cost": 1.14576e-05,
"upstream_inference_prompt_cost": 9.2176e-06,
"upstream_inference_completions_cost": 2.24e-06}}
```
A `:free` model returns the same shape with `cost: 0`, correctly. Note
`cached_tokens: 192/254` — a **measured** cache rate on a field
`routing.assumed_cache_rate` currently guesses at.
### Defect 4 — the plan gauge hides the only actionable number
The chip renders `46.1% · 6.25 kWh plan`. At the time of writing, 5.01 of the
period's 30 days had elapsed — 16.7%. Burning 46.1% of the allowance in 16.7%
of the period is **2.76x sustainable pace**, projecting ~17.2 kWh against a
6.25 kWh plan.
That ratio is the number the operator needed on 2026-09-06 and computed by
hand against `metered_kwh_period` and elapsed time. Nothing in the portal
computes it. A percentage without a pace reads as reassuring at exactly the
moment it should not.
### Defect 5 — the electricity the operator personally pays for is not in Quota
`local_energy_summary` (`src/metrics.py:1547`) exists and is rendered in its
own dashboard card, but it is absent from Quota — while `ollama-local` rows
*are* present in `energy_observations` (60 rows) and therefore in the "Cloud
energy" pane, which double-counts them and files self-hosted draw under cloud.
### Also found, and deliberately NOT fixed here
Joined on `request_id` over this period, `route_decisions.est_cost_usd`
overprices NeuralWatt **4.14x** against its own billed cost (n=4,958), while
OpenRouter's estimate lands within ~13% of the observed pool delta:
| model | n | est | billed | ratio |
|---|---|---|---|---|
| kimi-k2.7-code | 2,853 | $62.78 | $11.21 | 5.60 |
| qwen3.6-35b | 1,024 | $4.01 | $0.48 | 8.40 |
| glm-5.3 | 288 | $14.75 | $6.76 | 2.18 |
| deepseek-v4-flash | 714 | $2.72 | $1.15 | 2.36 |
This is expected in direction — NeuralWatt bills
`min($8/kWh × kWh, 3 × list)` and models that batch well land far under list,
which is the whole subject of `CLAUDE.md`'s cost section — but the *size* means
the cost tiebreak systematically overprices NeuralWatt relative to OpenRouter
by roughly 4x. That is a routing-correctness question, not a dashboard one.
**Scope decision: this plan surfaces the ratio and changes no routing.** Work
item 3 puts est-vs-billed on the page so the bias is visible and measurable;
acting on it belongs in a separate plan, because the fix is either a
per-provider calibration factor or a switch to billed-cost feedback, and
neither should be decided inside a dashboard change.
## The model: three billing shapes, one comparable unit
The frankenstein comes from trying to express three billing shapes in one
schema. The redesign names the shapes and gives each its native unit, then
uses dollars — the only unit all three share — for the cross-provider view.
| shape | provider | allowance | remaining | spend truth | granularity |
|---|---|---|---|---|---|
| `metered_plan` | neuralwatt | 6.25 kWh/period | plan − metered kWh | `cost_usd` per request | per request |
| `prepaid_credit` | openrouter | $50 pool | `total_credits − total_usage` | account poll; per-request after work item 1 | 2 h (account), per request (attribution) |
| `self_hosted` | ollama-local | none — no cap exists | n/a | GPU draw × `tariff_usd_per_kwh` | per request |
Two rules fall out of that table, and both are load-bearing:
**Rule 1 — account truth for totals, request truth for attribution, never add
them.** An account poll is authoritative for "what did this provider charge
me"; summed per-request costs are authoritative for "which model spent it".
They are two measurements of overlapping quantities at different fidelities.
Adding them double-counts; preferring the request sum for a total silently
undercounts whatever the router did not observe (`plan_kwh_per_period`'s own
note already says as much: "router-metered only"). So: **totals come from the
account, breakdowns come from requests**, and the page states when the
breakdown covers less than the total.
**Rule 2 — a missing number is not zero.** `COALESCE(SUM(cost_usd), 0)` is how
OpenRouter came to read `$0.0000`. Unknown renders as `—` with a reason,
everywhere, in every card and every row.
## Work item 1 — capture per-request cost where the provider reports it
**Config** (`src/config.py`, `config/config.yaml`). Add to `DispatchProvider`:
```yaml
# dispatch_providers.openrouter
# OpenRouter returns billed cost per request when the request asks for usage
# accounting. Verified live 2026-09-10: usage.cost, usage.cost_details and
# usage.prompt_tokens_details.cached_tokens all come back populated.
reports_cost_in_usage: true
```
Default `false`. It must appear in `config/config.yaml` explicitly for both
providers, not only as a Pydantic default — `CLAUDE.md`'s "every knob belongs
in config.yaml" rule, and the reason `outcome_attribution_window_seconds` was
deleted.
Gate the request-body change on this flag rather than sending
`usage: {"include": true}` unconditionally: NeuralWatt is OpenAI-compatible and
an unknown top-level key is a plausible 400 on a provider we have no reason to
probe.
**Request** (`src/dispatcher.py:4212`, where `upstream_body` is built):
```python
if getattr(settings, "reports_cost_in_usage", False):
upstream_body.setdefault("usage", {"include": True})
```
`setdefault`, so a client that sent its own `usage` block wins.
**Extraction** (`src/dispatcher.py:2004`). Extend `extract_telemetry` to fall
back to the `usage` block when NeuralWatt's `cost` block is absent:
- `cost_usd` ← `cost.request_cost_usd`, else `usage.cost`
- new `Telemetry.cached_prompt_tokens` ← `usage.prompt_tokens_details.cached_tokens`
Keep the precedence in that order. A provider that reports both is reporting
the same number twice, and `request_cost_usd` is the field this project has
already validated against billing.
**THE TRAP — the streamed path does not pass `usage` to `extract_telemetry`,
and every agent client streams.** On the buffered path
(`src/dispatcher.py:4385`) `extract_telemetry(payload)` sees the whole response
including `usage`. On the streamed path (`src/dispatcher.py:4637`) it is called
as `extract_telemetry(collected)`, where `collected` holds only the sniffed
NeuralWatt SSE *comment* lines; the streamed `usage` dict lives in a separate
local (`src/dispatcher.py:4603`, populated from `chunk["usage"]`) and is passed
to `log_observation` for token counts only.
So a change that only touches `extract_telemetry` fixes cost for buffered
requests and leaves it NULL for essentially all real traffic.
`stream_options.include_usage` is already set (`src/dispatcher.py:4217`), so
the final chunk does carry it; it simply never reaches the extractor.
**The fix, decided — not implementer's choice.** Signature becomes:
```python
def extract_telemetry(payload: dict, *, usage: Optional[dict] = None) -> Telemetry:
```
- When `usage` is not passed, the function falls back to `payload.get("usage")`.
That alone fixes every **buffered** call site with no change at the call
sites at all, because `payload` there is the whole response.
- The **streamed** call site passes it explicitly —
`extract_telemetry(collected, usage=usage)` — because `collected` holds only
sniffed SSE comment blocks and structurally cannot contain a usage dict.
- Cost precedence inside the function, in this order:
`payload["cost"]["request_cost_usd"]`, then `usage["cost"]`. First non-None
wins.
The rejected alternative was `extract_telemetry({**collected, "usage": usage})`.
It needs no signature change, which is why it is tempting, and it is wrong:
`collected` is by contract the output of `_sniff_telemetry_line`
(`src/dispatcher.py:3735`), whose keys are provider SSE comment names. Injecting
a synthetic `usage` key into it makes that dict a half-response/half-sniff
hybrid, and the next person to read either function has to hold both meanings at
once. The keyword argument keeps the two sources named.
Note the precedence order is what makes the existing NeuralWatt behaviour
unchanged: NW streaming populates `collected["cost"]`, so it never reaches the
`usage` branch, and its `cost_usd` stays the figure this project has already
validated against billing. The streamed path must be covered by a test
asserting a non-null `cost_usd` from `usage.cost` alone.
Audit the other `log_observation` call sites while in here
(`src/dispatcher.py:4871` and the local-dispatch path) so no branch silently
drops the new field.
**Schema.** `energy_observations` gains one nullable column:
```sql
ALTER TABLE energy_observations ADD COLUMN cached_prompt_tokens INTEGER;
```
Add it to `config/schema.sql` with a comment recording why it exists (a
measured cache rate against `routing.assumed_cache_rate`'s guess) and that
NULL means "provider did not report", not zero.
**What this does NOT do:** it does not backfill. OpenRouter cost exists from
the deploy forward only, which is precisely why Rule 1 keeps account polls as
the source of account totals.
## Work item 2 — store the credit pool, not just the remainder
`_BALANCE_PARSERS["openrouter"]` (`src/poller.py:66`) computes
`total_credits − total_usage` and discards both inputs. The pool size is what
turns a bare "$40.45" into "used $9.57 of $50", and it is already in the
response.
- `provider_balance_observations` gains two nullable columns:
`total_credits_usd`, `total_usage_usd`. Extend
`_ensure_provider_balance_table` (`src/poller.py:422`) with idempotent
`ALTER TABLE ... ADD COLUMN` guarded the way the rest of this project guards
live-DB migrations, and mirror them into `config/schema.sql`.
- The parser registry changes shape: a parser returns
`(balance_usd, total_credits_usd | None, total_usage_usd | None)`.
`record_balance` (`src/poller.py:449`) persists all three. A provider whose
endpoint reports only a remainder stores NULL for the other two, and the UI
falls back to a bare remaining figure.
- `tests/test_multi_provider_poller.py:728-731` pins
`_BALANCE_PARSERS.keys() == config.PROVIDERS_WITH_BALANCE_PARSERS`
(`src/config.py:1056`); keep that pin working. It is in that file rather than
a poller-parsing one, which is easy to miss when scanning by name.
Nullable columns because a NULL pool size must render as "balance $40.45" and
never as "$40.45 of $0".
## Work item 3 — the `/metrics` `quota` block, rewritten
Replace `quota_burn` and `quota_balance_and_burn` with one function,
`metrics.quota_accounts(conn, cfg)`. **No deprecated aliases** — `CLAUDE.md`
records that the last quota reshape updated every consumer in the same change,
and that is the standard here too.
### Shape
```json
"quota": {
"period": {
"start": "2026-09-06",
"next_reset": "2026-10-06",
"elapsed_fraction": 0.167,
"source": "objective.billing_reset_day"
},
"accounts": [
{
"provider": "neuralwatt",
"shape": "metered_plan",
"plan": {
"unit": "kwh",
"allowance": 6.25,
"used": 2.87933,
"used_fraction": 0.4607,
"pace_ratio": 2.76,
"projected_period_total": 17.24,
"pace_note": null
},
"spend_usd": {
"period": 22.66,
"window_30d": 61.2,
"source": "per_request_billed",
"attribution_coverage": 1.0
},
"credit": {
"balance_usd": -0.004198,
"meaning": "overage_allowance",
"observed_at": "2026-09-10T18:49:43Z",
"age_seconds": 19980
},
"fidelity": "per_request"
},
{
"provider": "openrouter",
"shape": "prepaid_credit",
"pool": {
"total_credits_usd": 50.0,
"used_usd": 9.5714,
"remaining_usd": 40.4286,
"used_fraction": 0.1914,
"observed_at": "2026-09-10T23:17:32Z",
"age_seconds": 7200
},
"burn": {
"window_hours": 24,
"usd_per_hour": 0.0337,
"projected_hours_remaining": 1200.9,
"runway_low_warning": false,
"note": null
},
"spend_usd": {
"period": 9.5714,
"window_30d": 9.5714,
"source": "account_poll_delta",
"attribution_coverage": 0.0
},
"fidelity": "account_poll"
},
{
"provider": "ollama-local",
"shape": "self_hosted",
"energy": { "unit": "kwh", "period": 0.01725, "window_30d": 0.0402 },
"spend_usd": {
"period": 0.00274,
"window_30d": 0.0064,
"source": "derived_from_draw",
"attribution_coverage": 1.0
},
"tariff_usd_per_kwh": 0.15856,
"fidelity": "derived"
}
],
"spend": {
"by_provider_usd": { "neuralwatt": 22.66, "openrouter": 9.5714, "ollama-local": 0.00274 },
"total_usd": 32.23,
"estimated_usd": { "neuralwatt": 86.05, "openrouter": 10.07 },
"estimate_ratio": { "neuralwatt": 4.14, "openrouter": 1.13 }
},
"alarm": {
"kind": "plan_pace",
"provider": "neuralwatt",
"severity": "warn",
"headline": "neuralwatt: 2.8x sustainable pace, projecting 17.2 kWh against a 6.25 kWh plan"
},
"note": "router-metered only; traffic bypassing the router is not counted"
}
```
### Decisions embedded in that shape
- **`shape` is derived from config, never hardcoded per provider name.** The
rule, in precedence order, evaluated per `cfg.dispatch_providers` entry plus
the local-dispatch provider:
| # | test | shape |
|---|---|---|
| 1 | provider is the local-dispatch provider (`ollama-local`) | `self_hosted` |
| 2 | `has_energy_telemetry` and `objective.plan_kwh_per_period` is set | `metered_plan` |
| 3 | `balance_url` is set | `prepaid_credit` |
| 4 | none of the above | `unmetered` |
First match wins, and the order matters: a provider could satisfy both 2 and
3, and the plan is the thing with a period and a ceiling, so it leads.
`unmetered` is a real state, not an error — a configured provider with no
telemetry, no plan and no balance endpoint genuinely has nothing to meter,
and it renders as a name plus its per-request spend if any, with no gauge.
A `metered_plan` provider when `plan_kwh_per_period` is null degrades to
case 3 or 4 rather than rendering a gauge against a null allowance.
- **`spend.by_provider_usd[p]` is exactly `accounts[p].spend_usd.period`**, and
`spend.total_usd` is the sum of those values — nothing else. It is never
computed from a second query, and a balance figure never enters it. This is
Rule 1 as an implementation constraint: each account has already decided
whether its own period spend comes from summed per-request cost, an account
poll delta, or derived draw, and the ledger just adds up those decisions.
`estimated_usd[p]` is a separate `SUM(route_decisions.est_cost_usd)` over the
same period and provider, and `estimate_ratio[p]` is
`estimated_usd[p] / by_provider_usd[p]`, null when the denominator is 0.
- **`accounts` is a list, not a map keyed by provider**, ordered by alarm
severity then provider name. The old `by_provider` map forced every consumer
to re-derive which provider mattered; the list makes the answer positional
and makes the TUI's row order free.
- **`total_balance_usd` is deleted.** It summed a prepaid pool with an overage
residue. `spend.total_usd` replaces it as the one honest cross-provider
number, because dollars spent *are* additive across billing shapes.
- **`shape` drives rendering, `fidelity` qualifies it.** A consumer switches
layout on `shape` and prints a caveat from `fidelity`; neither is inferred
from provider name or from which config keys happen to be set.
- **`spend_usd.source`** is one of `per_request_billed`, `account_poll_delta`,
`derived_from_draw`. **`attribution_coverage`** is the fraction of that
provider's period spend that per-request rows can account for — 1.0 for
NeuralWatt, 0.0 for OpenRouter until work item 1 lands, and a partial value
for the period that straddles the deploy. This is Rule 1 made machine
readable: it is how a card knows to say "breakdown covers $4.10 of $9.57".
- **`credit.meaning`** replaces `balance_source`. `telemetry` / `polled` /
`unconfigured` described *where the number came from*; `overage_allowance` /
`prepaid_pool` / `not_configured` describe *what it means*, which is what
decides whether it deserves a gauge or a footnote.
### `pace_ratio` — the computation, and its trap
`pace_ratio = used_fraction / elapsed_fraction`, with
`elapsed_fraction = (now − period_start) / (next_reset − period_start)`.
**Trap: `elapsed_fraction` → 0 at period start, and the ratio explodes.** In
the first hours of a period a handful of calls reads as a 40x overrun. Guard
it: when `elapsed_fraction < 0.02` (roughly the first 14 hours of a 30-day
period), return `pace_ratio: null` and `projected_period_total: null` with
`pace_note: "too early in the period to project"`. A null pace must render as
"—", never as 0 or 1.0.
**Second trap, and this project has been burned by it once already.** The
projection must not read as a wall. `objective.plan_kwh_per_period` gates
nothing; overage is billed. `plans/quota-balance-and-burn-rate.md` records the
last warning that got this wrong ("a quota is a wall, not a bill — requests
fail rather than costing more") and why it was both alarming and false. Project
in kWh and say what it costs, never "requests will fail".
### `alarm` — what the chip shows
One alarm, the worst across accounts, or `null`:
- `plan_pace`: `pace_ratio >= objective.plan_pace_warn_ratio` (new knob,
default `1.25`, **which must be written into `config/config.yaml`**, not
left as a Pydantic default). `severity: "critical"` at `>= 2.0`.
- `credit_runway`: existing `projected_hours_remaining <
objective.quota_runway_warning_hours` (6). `critical` below half that.
- `stale_reading`: a `prepaid_credit` account whose `age_seconds` exceeds 3x
the poll interval. **This one is new and it matters**: the live pool sat flat
at $40.453746437 for six consecutive polls across 18 hours while 705
completions ran, and the current UI renders that as a healthy account with no
burn. It is poll granularity, not calm. A flat reading and a stale reading
must be distinguishable on the page.
`coverage.warnings` (`src/metrics.py:493`) currently emits the >80%-of-plan
string. Replace it with a pace-based warning sourced from `alarm`, so the
navbar bell and the chip cannot disagree.
### Consumers to update in the same change
Measured inventory — `grep -rln "total_balance_usd\|by_provider"` over the
repo. There is no deprecation path, so all of these move together:
| file | what breaks |
|---|---|
| `src/metrics.py:203,334` | the two functions being replaced |
| `src/dispatcher.py` | `/metrics` assembly |
| `src/admin.py:1849` | `/api/snapshot`'s `quota` key, plus the new `GET /admin/api/quota` |
| `src/tui_model.py:41-60` | builds `{label, value}` quota rows from the flat keys |
| `src/tui.py:78,105,566,601,615` | per-provider row formatting, the kWh lead line, and the `total_balance_usd` line `CLAUDE.md` documents |
| `tests/test_metrics.py` | **49 references** — the bulk of the work |
| `tests/test_metrics_endpoint.py` | 6 |
| `tests/test_tui.py` | 6 |
| `tests/test_credit_attenuation_routing.py` | 6 — asserts attenuation reads balances, so it must keep passing against the new shape without weakening what it pins |
| `CLAUDE.md`, `docs/api.md`, `docs/admin-portal.md` | describe the old contract by name |
`test_credit_attenuation_routing.py` deserves a second look rather than a
mechanical rename: `credit_attenuation` derives a routing multiplier from
polled balances, so its fixtures encode the *old* assumption that a balance is
one flat number per provider. Its behaviour must not change here — only where
it reads the balance from.
## Work item 4 — `/admin/quota`, a page
The chip-plus-modal cannot hold three shaped cards and a ledger. Follow the
portal's page-per-domain pattern (`providers`, `proficiency`, `profiles`).
**Mechanics** (`src/admin.py:823-870`): add `_admin_quota` →
`admin/frontend/quota.html`, a `@router.get("/quota")` returning it with
`_NO_CACHE_HEADERS`, and a nav item after `Providers` in the navbar of **all**
admin pages (`admin/frontend/*.html:281-287`) — the navbar is duplicated per
page, so a nav item added to one page only is a half-shipped nav.
**Data**: `GET /admin/api/quota` → `metrics.quota_accounts(conn, cfg)` plus the
spend ledger's per-model rows. A dedicated endpoint rather than reusing
`/api/snapshot`, because the page needs per-model spend the snapshot does not
carry and none of the snapshot's proficiency/verdict payload.
### Cards, one per account, laid out by `shape`
Three peer cards on one row is `.col-xl-4` — the Standards doc is explicit that
`.col-xl-6` strands a third card and opens two dead columns.
**`metered_plan` (NeuralWatt).** The kWh gauge is the hero, and the pace marker
is what makes it a redesign rather than a reskin: draw the gauge with a tick at
`elapsed_fraction` so "burned" and "elapsed" are visually comparable in one
glance. Below it: pace ratio, projected period total, **$ billed this period**,
and the overage allowance as a de-emphasized footnote labeled
`overage allowance` — never as a balance or a runway, because at −$0.004 it is
neither.
**`prepaid_credit` (OpenRouter).** A depletion gauge: `used_usd` of
`total_credits_usd`, remaining in the hero position. Then burn $/hr, runway in
days (hours only under ~48h), and the **reading age** rendered as text
("polled 2h ago"), promoted to a warning tint past the staleness threshold.
When `total_credits_usd` is NULL, degrade to a bare remaining figure with no
gauge.
**`self_hosted` (local).** kWh and $ this period, the tariff it was priced at,
and a one-line statement that this is billed to the operator's utility rather
than a provider. No gauge — there is no allowance to fill. Hidden entirely
when `local_energy.enabled` is false.
### Spend ledger
Below the cards, full width:
1. **`spend.total_usd` as the page's single headline number**, with a stacked
bar by provider. This is the answer to "what did I spend this month" and it
is the reason the page exists.
2. **Ranked per-model spend** — model, provider, calls, $, share of total.
`—` in the $ column where cost is unknown, with the row still present and
ranked by calls. Where `attribution_coverage < 1`, a single line above the
table states the gap in dollars, e.g. "breakdown accounts for $4.10 of
OpenRouter's $9.57; the remainder predates per-request cost capture".
3. **Estimate calibration**, a small two-column strip: est vs billed per
provider with the ratio. Label it as what it is — the cost tiebreak's input
measured against the bill — and do not editorialize a fix; that is the
follow-up plan.
### Chip
`renderQuotaChip` (`admin/frontend/index.html:779`) stops hardcoding the kWh
story. It renders `alarm`: headline text, severity tint, and a link to
`/admin/quota`. With `alarm: null` it shows the period's spend to date, which
is calm and still useful. Its comment at line 790 ("The chip always tells the
kWh-subscription story") is the assumption being deleted — remove it, do not
leave it stale.
**The chip stops being a modal trigger and becomes a link.** Four edits, all in
`index.html`, and leaving any one of them behind leaves a dead control:
1. Line 305 — the chip `<div>` carries `data-bs-toggle="modal"`
`data-bs-target="#quota-modal"` `role="button"` `tabindex="0"`. Replace the
element with an `<a href="/admin/quota">` carrying the same `.quota-chip`
class, and drop all four attributes: an anchor is focusable and activatable
on its own, so `role`/`tabindex` become wrong rather than merely redundant.
The navbar's `z-index: 1030` note in the Standards doc exists because of
this chip's `backdrop-filter` — keep the stacking context intact.
2. Lines 462-478 — the `#quota-modal` block is deleted outright.
3. `renderQuotaModal` (`:659`) and its `#quota-modal-body` writes are deleted,
and the `renderQuotaModal(data.quota)` call at `:558` with it.
4. `grep -n 'quota-modal' admin/frontend/*.html` must come back empty when
done. `formatProviderBalance` / `formatBalanceStaleness` (`:708`, `:731`)
move to the new page rather than being deleted — the staleness helper is
exactly what the `prepaid_credit` card's reading-age line needs, and its
tiny-negative-balance handling is hard-won (it exists so NeuralWatt's
−$0.004 never renders as "-$0.00").
## Work item 5 — the per-model usage pane
Two panes read `data.per_model` today: `renderCloudEnergy`
(`admin/frontend/index.html:906`) and `renderBars` (`:935`). Three defects,
each falsifiable:
1. **Bar length and bar label disagree.** `renderBars` ranks by
`BARS_SORT_KEYS[_barsSortKey]` (`:840`) — calls, cost or energy — but the
value label is hardcoded to `$cost · kWh` (`:957`). Sorting by calls
produces bars ordered by call count, labeled in dollars.
2. **Unknown renders as zero**, giving `xiaomi/mimo-v2.5` a `$0.0000` label on
655 real calls (Defect 3 above).
3. **`ollama-local` appears in a card called "Cloud energy"**, double-counting
against the Local Compute card that owns it.
And a fourth, which is the operator's own framing: calls, cost and energy
compete for one row of space, and energy is the least actionable of the three.
### The rework
- **One ranked list.** Sort keys become **calls | spend | tokens**. Energy
leaves the sort keys: it is not a thing an operator ranks models by, it is a
thing they inspect. `tokens` replaces it because it is the honest volume
metric for a provider that reports no energy at all.
- **The label follows the sort key**, plus share of total: sorting by calls
shows `655 calls · 8.2%`; by spend, `$1.23 · 3.8%`. One metric, the one
being ranked.
- **`—` for unknown**, never `$0.0000`, with a footnote naming the reason
("provider does not report per-request cost").
- **Energy, carbon, attribution and grid move to a row-click detail popup** —
kWh, gCO2eq, `grid_id`, mean `attribution_ratio`, mean completion tokens,
cached-token share once work item 1 lands. Render only the fields that have
data; a popup of six `—`s is worse than a popup that admits it has nothing.
**Copy the pattern that already exists rather than inventing one.**
`admin/frontend/proficiency.html` does exactly this for its matrix cells: a
`<div class="modal modal-blur fade" id="cell-modal">` with a `-title` and
`-body` (`:264-276`), a lazily-constructed `new tabler.Modal(...)` held in a
module-level `_cellModal` (`:575`), and a click handler that sets
`textContent` on the title and `innerHTML` on the body before `.show()`
(`:582-587`). Mirror it as `#model-modal` on the dashboard.
**No new endpoint.** The popup's fields all come from the `per_model` row the
pane already has in hand — the click handler passes the row object it
rendered, so there is no fetch, no loading state, and no second source of
truth to drift. That is also why the new aggregate columns in "Metrics
support" below are added to `per_model` itself rather than to a detail route.
This is the operator's own proposal and it is the right one: attribution
ratio in particular is a diagnostic on a 750x-between-models, 1.8x-within-model
quantity, which is an inspection surface, not a ranking.
- **Cloud pane excludes `ollama-local`.** Filter on provider, not on whether
energy is present.
- **`renderCloudEnergy`'s 5-decimal summary trio is cut.** Period spend now has
a first-class home on `/admin/quota`; a second, differently-windowed copy of
it on the dashboard is how the two drift apart.
### Metrics support
`metrics.per_model` (`src/metrics.py:1386`, `def per_model`) adds
`sum_prompt_tokens`, `sum_completion_tokens`, `sum_cached_prompt_tokens`, and
`cost_rows` / `total_rows` so the UI can tell "no cost reported" from "cost was
zero". Stop `COALESCE`ing cost and energy to 0 — return NULL and let the
consumer render `—`.
**That coalesce has a second consumer, and removing it crashes the TUI rather
than merely uglifying it.** `src/tui_model.py:108-116` maps `sum_cost_usd` /
`sum_energy_kwh` / `sum_carbon_g_co2eq` straight through to a row dict, and
`src/tui.py` then renders three cells from it — but they are not equally safe:
| line | code | on `None` |
|---|---|---|
| `src/tui.py:667` | `_fmt_usd(r["cost_usd"])` | safe — `_fmt_usd` returns `"n/a"` (`src/tui.py:55-58`) |
| `src/tui.py:668` | `f"{r['energy_kwh']:.6g}"` | **TypeError** |
| `src/tui.py:669` | `f"{r['carbon_g_co2eq']:.4g}"` | **TypeError** |
The `COALESCE` is the only thing standing between those two format strings and
an unhandled exception, so the fix for 668-669 is not "render `—`" as a styling
choice — it is required for the TUI to start. Give both the same None-aware
treatment `_fmt_usd` already has (a `_fmt_num(v, spec)` helper returning `"n/a"`
on None is the smallest change that covers both), and add a `per_model` fixture
row with NULL energy to `tests/test_tui.py` so nothing re-introduces a bare
format on a nullable column.
## Tests
New or extended, offline, matching this repo's existing fixtures:
- `tests/test_metrics.py` — the existing quota tests (49 references) are
rewritten against the new shape, not deleted. Add: one account per configured
provider; `shape`/`fidelity`/`credit.meaning` values; **`pace_ratio` is null
inside the first 2% of a period** and correct outside it; `spend.total_usd`
sums only per-provider spend and never a balance; no `total_balance_usd` key
survives anywhere in the payload; a NULL `total_credits_usd` yields a pool
block with no `used_fraction`.
- A `stale_reading` case — a balance series that is flat *and* old raises the
alarm; flat but fresh does not. This is the 18-hour flat-pool state observed
on the live DB, and it is the one alarm nothing covers today.
- Dispatcher cost capture — four cases: flag off → no `usage` key in
`upstream_body`; flag on, buffered → `cost_usd` from `usage.cost`; flag on,
**streamed** → `cost_usd` non-null (the trap in work item 1); NeuralWatt's
`cost.request_cost_usd` still wins when both blocks are present.
- Poller — three-tuple parser, the two new columns persisted, idempotent
migration against a table that predates them, and a parser returning NULLs
for pool size.
- Admin — `/admin/quota` serves the file; `/admin/api/quota` returns the
documented keys.
- TUI — update the quota panel's fixtures; keep
`tests/test_tui_schema_drift.py` and `tests/test_tui_warnings.py` green.
**`test_tui_warnings.py`'s fixture is a coupled system** — `CLAUDE.md`
records that adding a seed there can silence an unrelated warning class — so
changing the plan-usage warning means re-checking both directions of that
test, not just the failing one.
## Docs to update in the same change
- `CLAUDE.md` — the `quota.by_provider` / `total_balance_usd` paragraph in
"Billing is per-kWh", and the `credit_attenuation` section's two-semantics
note, which is built on `total_balance_usd` existing.
- `docs/api.md` — the `/metrics` `quota` block.
- `docs/admin-portal.md` — the new page, the chip's new meaning, the reworked
per-model pane.
- `docs/data-model.md` — three new columns.
- `plans/admin-design-standards.md` — only if the pace-marker gauge or the
depletion gauge introduces a pattern the next page should reuse. It probably
does; record the grammar (gauge + reference tick), not the CSS.
## Out of scope, recorded so it is not rediscovered
- **The 4.14x est-vs-billed bias.** Surfaced by work item 3, fixed elsewhere.
- **Replacing `routing.assumed_cache_rate` with measured `cached_tokens`.**
Work item 1 starts capturing the data; consuming it changes pricing and
therefore routing, which needs its own plan and its own before/after.
- **Energy-budget attenuation** (a `credit_attenuation` analogue keyed on plan
pace). The pace signal this plan computes is its natural input, but it is a
routing feature.
- **Backfilling OpenRouter cost.** Not possible per request; Rule 1 is the
reason it is not needed for account totals.