The dashboard was eight cards restating what the other seven pages own, so it was the page you passed through rather than the one you started from. It is now the landing page: no nav entry of its own, reached by the logo, with two views over the one /api/snapshot payload so switching costs no request. Board is one tile per destination, identical anatomy every time: what it is, one number, one qualifier, a link. Under them, a full-width Activity band plots routed decisions against billed requests on one axis, so the gap between them -- local dispatch, cache hits, requests that never reached a provider -- is readable, which it is in neither series alone. Its range (24h/7d/30d) is separate from the tile sparklines, because it is the chart you come to the page to read. Beside it, the eight busiest models ranked by CALLS, not cost: cost coverage is partial, and an OpenRouter row with two priced calls out of hundreds otherwise outranks the model doing the work. Live loads the row itself -- category, model, provider, required context, tier, classify time, exploratory flag, cost -- so the usual question does not need the Decisions page, and groups rows by minute so a burst reads as one. The dot is the classification source: green for a real classification, amber for a degraded one, which still routes but never moves proficiency, and until now was visible only in a /metrics warning. The rail beside it answers what Board cannot: throughput against the last hour, classifier p50 and p95 (the latency floor every routed request pays before an upstream token), degraded share, and the top five models by share of the live window keyed on model AND provider, since the same model on two providers is two different routing outcomes. Spend is not restated here; it belongs to Quota. Warnings move behind a navbar bell. They are recomputed from live state on every poll, so the only honest dismissal is "hide until the condition changes": a dismissal is keyed on the warning's TEXT, so when the numbers move the key stops matching and the warning returns on its own. Nothing has to expire it. Per browser, in localStorage, wrapped in try/catch -- a reading state, not a fact about the router. Two tests changed shape rather than being deleted. The verdict-bar tests pinned a card this page no longer has, but the lesson they recorded is the expensive part and is preserved verbatim in the replacement's docstring: `unverifiable` is not a finding, it is the default for prose and for any turn ending in a tool call, and any bar scaled against it renders every meaningful verdict at the 2px minimum. Exclude it from the scale, never from the denominator. The Live rail states the same fact as one ratio, which needs no scale. The nav test's rule is inverted for index.html rather than skipped: the landing page must have no nav entry AND must claim no active one, so neither can drift back in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
733 lines
35 KiB
Markdown
733 lines
35 KiB
Markdown
# Quota, redesigned around billing shape (subscription vs paygo vs self-hosted)
|
||
|
||
Status: done -- shipped as the quota page redesign; decision-complete. Written 2026-09-10 from the live database,
|
||
the live `/metrics`, and two live probes against the OpenRouter API.
|
||
|
||
Intended executor: opencode (Prometheus → Atlas). Every open question in this
|
||
document is closed; where a choice existed, the choice and its reason are
|
||
recorded rather than left to the implementer.
|
||
|
||
Read `plans/admin-design-standards.md` and `plans/admin-work-framework.md`
|
||
before touching `admin/frontend/*`. This document does not restate the visual
|
||
language; it specifies what the pages must say.
|
||
|
||
**On the file:line anchors below.** Every anchor in the first draft of this
|
||
document was taken against a working tree five commits behind `origin/main`,
|
||
and PR #76 (`feat/mangled-output-breaker`, merged 2026-09-10T23:48Z) added ~77
|
||
lines to `src/metrics.py` in between. Four anchors were wrong as a result. They
|
||
are corrected here and re-verified against `f89cefc`, but the lesson stands:
|
||
**treat every anchor as a hint and confirm it by grepping the landmark named
|
||
beside it.** Each anchor below is paired with the symbol or string it points
|
||
at, so the grep is always available.
|
||
|
||
## Why this exists
|
||
|
||
The word "Quota" currently covers three unrelated quantities, and the portal
|
||
sums two of them into a figure `CLAUDE.md` itself disclaims as "not a single
|
||
spendable figure". Meanwhile the one number an operator asks for first — money
|
||
spent this month — appears nowhere in the portal at all.
|
||
|
||
All evidence below is from the live system on 2026-09-10/11, billing period
|
||
2026-09-06 → 2026-10-06.
|
||
|
||
### Defect 1 — one word, three quantities
|
||
|
||
`metrics.quota_burn` returns a kWh-plan report with a per-provider balance map
|
||
bolted onto it (`src/metrics.py:334`, merging `quota_balance_and_burn` at
|
||
`src/metrics.py:203`). The live payload carries, under one key:
|
||
|
||
| field | what it actually is | magnitude |
|
||
|---|---|---|
|
||
| `metered_fraction_of_plan` | NeuralWatt kWh subscription consumption | 0.4607 |
|
||
| `by_provider.neuralwatt.balance_usd` | overage-billing accounting artifact | −$0.004198 |
|
||
| `by_provider.openrouter.balance_usd` | a real prepaid credit pool | $40.45 |
|
||
| `total_balance_usd` | the sum of the previous two | $40.4495 |
|
||
|
||
Those are a gauge, a rounding residue, and a wallet. The sum is not a quantity.
|
||
|
||
### Defect 2 — money spent is not in the portal
|
||
|
||
Actual billed spend this period, from the live DB:
|
||
|
||
| provider | period spend | source |
|
||
|---|---|---|
|
||
| neuralwatt | **$22.66** | `SUM(energy_observations.cost_usd)`, 6,079 rows |
|
||
| openrouter | **$9.57** | account poll: `total_credits 50 − total_usage 9.5714` |
|
||
| ollama-local | $0.0027 | `local_energy_observations.cost_usd`, 202 rows |
|
||
|
||
≈ **$32.23 this period**, and the dashboard renders none of it. The NeuralWatt
|
||
figure has been captured per request since 2026-08-17 and is read by exactly
|
||
nothing.
|
||
|
||
### Defect 3 — OpenRouter has no per-request cost, so it renders as free
|
||
|
||
`extract_telemetry` (`src/dispatcher.py:2004`) reads only NeuralWatt's
|
||
top-level `energy` / `cost` blocks. Every OpenRouter row therefore lands with
|
||
`cost_usd IS NULL`: 3,106 rows this period. `metrics.per_model`
|
||
(`src/metrics.py:1386`, `def per_model`) `COALESCE`s those NULLs to 0, so the
|
||
live dashboard shows:
|
||
|
||
```
|
||
xiaomi/mimo-v2.5 openrouter calls=655 cost=$0.0000 kwh=0.00000
|
||
```
|
||
|
||
That model is #8 by volume and cost real money — the pool moved $1.23 on
|
||
2026-09-10 alone, the day it served 598 of those calls at ~175k prompt tokens
|
||
each.
|
||
|
||
**This is fixable, and it was verified live rather than assumed.** Sending
|
||
`usage: {"include": true}` makes OpenRouter return billed cost per request.
|
||
Probed 2026-09-10 against `xiaomi/mimo-v2.5`:
|
||
|
||
```json
|
||
{"prompt_tokens": 254, "completion_tokens": 8, "cost": 1.14576e-05,
|
||
"prompt_tokens_details": {"cached_tokens": 192},
|
||
"cost_details": {"upstream_inference_cost": 1.14576e-05,
|
||
"upstream_inference_prompt_cost": 9.2176e-06,
|
||
"upstream_inference_completions_cost": 2.24e-06}}
|
||
```
|
||
|
||
A `:free` model returns the same shape with `cost: 0`, correctly. Note
|
||
`cached_tokens: 192/254` — a **measured** cache rate on a field
|
||
`routing.assumed_cache_rate` currently guesses at.
|
||
|
||
### Defect 4 — the plan gauge hides the only actionable number
|
||
|
||
The chip renders `46.1% · 6.25 kWh plan`. At the time of writing, 5.01 of the
|
||
period's 30 days had elapsed — 16.7%. Burning 46.1% of the allowance in 16.7%
|
||
of the period is **2.76x sustainable pace**, projecting ~17.2 kWh against a
|
||
6.25 kWh plan.
|
||
|
||
That ratio is the number the operator needed on 2026-09-06 and computed by
|
||
hand against `metered_kwh_period` and elapsed time. Nothing in the portal
|
||
computes it. A percentage without a pace reads as reassuring at exactly the
|
||
moment it should not.
|
||
|
||
### Defect 5 — the electricity the operator personally pays for is not in Quota
|
||
|
||
`local_energy_summary` (`src/metrics.py:1547`) exists and is rendered in its
|
||
own dashboard card, but it is absent from Quota — while `ollama-local` rows
|
||
*are* present in `energy_observations` (60 rows) and therefore in the "Cloud
|
||
energy" pane, which double-counts them and files self-hosted draw under cloud.
|
||
|
||
### Also found, and deliberately NOT fixed here
|
||
|
||
Joined on `request_id` over this period, `route_decisions.est_cost_usd`
|
||
overprices NeuralWatt **4.14x** against its own billed cost (n=4,958), while
|
||
OpenRouter's estimate lands within ~13% of the observed pool delta:
|
||
|
||
| model | n | est | billed | ratio |
|
||
|---|---|---|---|---|
|
||
| kimi-k2.7-code | 2,853 | $62.78 | $11.21 | 5.60 |
|
||
| qwen3.6-35b | 1,024 | $4.01 | $0.48 | 8.40 |
|
||
| glm-5.3 | 288 | $14.75 | $6.76 | 2.18 |
|
||
| deepseek-v4-flash | 714 | $2.72 | $1.15 | 2.36 |
|
||
|
||
This is expected in direction — NeuralWatt bills
|
||
`min($8/kWh × kWh, 3 × list)` and models that batch well land far under list,
|
||
which is the whole subject of `CLAUDE.md`'s cost section — but the *size* means
|
||
the cost tiebreak systematically overprices NeuralWatt relative to OpenRouter
|
||
by roughly 4x. That is a routing-correctness question, not a dashboard one.
|
||
|
||
**Scope decision: this plan surfaces the ratio and changes no routing.** Work
|
||
item 3 puts est-vs-billed on the page so the bias is visible and measurable;
|
||
acting on it belongs in a separate plan, because the fix is either a
|
||
per-provider calibration factor or a switch to billed-cost feedback, and
|
||
neither should be decided inside a dashboard change.
|
||
|
||
## The model: three billing shapes, one comparable unit
|
||
|
||
The frankenstein comes from trying to express three billing shapes in one
|
||
schema. The redesign names the shapes and gives each its native unit, then
|
||
uses dollars — the only unit all three share — for the cross-provider view.
|
||
|
||
| shape | provider | allowance | remaining | spend truth | granularity |
|
||
|---|---|---|---|---|---|
|
||
| `metered_plan` | neuralwatt | 6.25 kWh/period | plan − metered kWh | `cost_usd` per request | per request |
|
||
| `prepaid_credit` | openrouter | $50 pool | `total_credits − total_usage` | account poll; per-request after work item 1 | 2 h (account), per request (attribution) |
|
||
| `self_hosted` | ollama-local | none — no cap exists | n/a | GPU draw × `tariff_usd_per_kwh` | per request |
|
||
|
||
Two rules fall out of that table, and both are load-bearing:
|
||
|
||
**Rule 1 — account truth for totals, request truth for attribution, never add
|
||
them.** An account poll is authoritative for "what did this provider charge
|
||
me"; summed per-request costs are authoritative for "which model spent it".
|
||
They are two measurements of overlapping quantities at different fidelities.
|
||
Adding them double-counts; preferring the request sum for a total silently
|
||
undercounts whatever the router did not observe (`plan_kwh_per_period`'s own
|
||
note already says as much: "router-metered only"). So: **totals come from the
|
||
account, breakdowns come from requests**, and the page states when the
|
||
breakdown covers less than the total.
|
||
|
||
**Rule 2 — a missing number is not zero.** `COALESCE(SUM(cost_usd), 0)` is how
|
||
OpenRouter came to read `$0.0000`. Unknown renders as `—` with a reason,
|
||
everywhere, in every card and every row.
|
||
|
||
## Work item 1 — capture per-request cost where the provider reports it
|
||
|
||
**Config** (`src/config.py`, `config/config.yaml`). Add to `DispatchProvider`:
|
||
|
||
```yaml
|
||
# dispatch_providers.openrouter
|
||
# OpenRouter returns billed cost per request when the request asks for usage
|
||
# accounting. Verified live 2026-09-10: usage.cost, usage.cost_details and
|
||
# usage.prompt_tokens_details.cached_tokens all come back populated.
|
||
reports_cost_in_usage: true
|
||
```
|
||
|
||
Default `false`. It must appear in `config/config.yaml` explicitly for both
|
||
providers, not only as a Pydantic default — `CLAUDE.md`'s "every knob belongs
|
||
in config.yaml" rule, and the reason `outcome_attribution_window_seconds` was
|
||
deleted.
|
||
|
||
Gate the request-body change on this flag rather than sending
|
||
`usage: {"include": true}` unconditionally: NeuralWatt is OpenAI-compatible and
|
||
an unknown top-level key is a plausible 400 on a provider we have no reason to
|
||
probe.
|
||
|
||
**Request** (`src/dispatcher.py:4212`, where `upstream_body` is built):
|
||
|
||
```python
|
||
if getattr(settings, "reports_cost_in_usage", False):
|
||
upstream_body.setdefault("usage", {"include": True})
|
||
```
|
||
|
||
`setdefault`, so a client that sent its own `usage` block wins.
|
||
|
||
**Extraction** (`src/dispatcher.py:2004`). Extend `extract_telemetry` to fall
|
||
back to the `usage` block when NeuralWatt's `cost` block is absent:
|
||
|
||
- `cost_usd` ← `cost.request_cost_usd`, else `usage.cost`
|
||
- new `Telemetry.cached_prompt_tokens` ← `usage.prompt_tokens_details.cached_tokens`
|
||
|
||
Keep the precedence in that order. A provider that reports both is reporting
|
||
the same number twice, and `request_cost_usd` is the field this project has
|
||
already validated against billing.
|
||
|
||
**THE TRAP — the streamed path does not pass `usage` to `extract_telemetry`,
|
||
and every agent client streams.** On the buffered path
|
||
(`src/dispatcher.py:4385`) `extract_telemetry(payload)` sees the whole response
|
||
including `usage`. On the streamed path (`src/dispatcher.py:4637`) it is called
|
||
as `extract_telemetry(collected)`, where `collected` holds only the sniffed
|
||
NeuralWatt SSE *comment* lines; the streamed `usage` dict lives in a separate
|
||
local (`src/dispatcher.py:4603`, populated from `chunk["usage"]`) and is passed
|
||
to `log_observation` for token counts only.
|
||
|
||
So a change that only touches `extract_telemetry` fixes cost for buffered
|
||
requests and leaves it NULL for essentially all real traffic.
|
||
`stream_options.include_usage` is already set (`src/dispatcher.py:4217`), so
|
||
the final chunk does carry it; it simply never reaches the extractor.
|
||
|
||
**The fix, decided — not implementer's choice.** Signature becomes:
|
||
|
||
```python
|
||
def extract_telemetry(payload: dict, *, usage: Optional[dict] = None) -> Telemetry:
|
||
```
|
||
|
||
- When `usage` is not passed, the function falls back to `payload.get("usage")`.
|
||
That alone fixes every **buffered** call site with no change at the call
|
||
sites at all, because `payload` there is the whole response.
|
||
- The **streamed** call site passes it explicitly —
|
||
`extract_telemetry(collected, usage=usage)` — because `collected` holds only
|
||
sniffed SSE comment blocks and structurally cannot contain a usage dict.
|
||
- Cost precedence inside the function, in this order:
|
||
`payload["cost"]["request_cost_usd"]`, then `usage["cost"]`. First non-None
|
||
wins.
|
||
|
||
The rejected alternative was `extract_telemetry({**collected, "usage": usage})`.
|
||
It needs no signature change, which is why it is tempting, and it is wrong:
|
||
`collected` is by contract the output of `_sniff_telemetry_line`
|
||
(`src/dispatcher.py:3735`), whose keys are provider SSE comment names. Injecting
|
||
a synthetic `usage` key into it makes that dict a half-response/half-sniff
|
||
hybrid, and the next person to read either function has to hold both meanings at
|
||
once. The keyword argument keeps the two sources named.
|
||
|
||
Note the precedence order is what makes the existing NeuralWatt behaviour
|
||
unchanged: NW streaming populates `collected["cost"]`, so it never reaches the
|
||
`usage` branch, and its `cost_usd` stays the figure this project has already
|
||
validated against billing. The streamed path must be covered by a test
|
||
asserting a non-null `cost_usd` from `usage.cost` alone.
|
||
|
||
Audit the other `log_observation` call sites while in here
|
||
(`src/dispatcher.py:4871` and the local-dispatch path) so no branch silently
|
||
drops the new field.
|
||
|
||
**Schema.** `energy_observations` gains one nullable column:
|
||
|
||
```sql
|
||
ALTER TABLE energy_observations ADD COLUMN cached_prompt_tokens INTEGER;
|
||
```
|
||
|
||
Add it to `config/schema.sql` with a comment recording why it exists (a
|
||
measured cache rate against `routing.assumed_cache_rate`'s guess) and that
|
||
NULL means "provider did not report", not zero.
|
||
|
||
**What this does NOT do:** it does not backfill. OpenRouter cost exists from
|
||
the deploy forward only, which is precisely why Rule 1 keeps account polls as
|
||
the source of account totals.
|
||
|
||
## Work item 2 — store the credit pool, not just the remainder
|
||
|
||
`_BALANCE_PARSERS["openrouter"]` (`src/poller.py:66`) computes
|
||
`total_credits − total_usage` and discards both inputs. The pool size is what
|
||
turns a bare "$40.45" into "used $9.57 of $50", and it is already in the
|
||
response.
|
||
|
||
- `provider_balance_observations` gains two nullable columns:
|
||
`total_credits_usd`, `total_usage_usd`. Extend
|
||
`_ensure_provider_balance_table` (`src/poller.py:422`) with idempotent
|
||
`ALTER TABLE ... ADD COLUMN` guarded the way the rest of this project guards
|
||
live-DB migrations, and mirror them into `config/schema.sql`.
|
||
- The parser registry changes shape: a parser returns
|
||
`(balance_usd, total_credits_usd | None, total_usage_usd | None)`.
|
||
`record_balance` (`src/poller.py:449`) persists all three. A provider whose
|
||
endpoint reports only a remainder stores NULL for the other two, and the UI
|
||
falls back to a bare remaining figure.
|
||
- `tests/test_multi_provider_poller.py:728-731` pins
|
||
`_BALANCE_PARSERS.keys() == config.PROVIDERS_WITH_BALANCE_PARSERS`
|
||
(`src/config.py:1056`); keep that pin working. It is in that file rather than
|
||
a poller-parsing one, which is easy to miss when scanning by name.
|
||
|
||
Nullable columns because a NULL pool size must render as "balance $40.45" and
|
||
never as "$40.45 of $0".
|
||
|
||
## Work item 3 — the `/metrics` `quota` block, rewritten
|
||
|
||
Replace `quota_burn` and `quota_balance_and_burn` with one function,
|
||
`metrics.quota_accounts(conn, cfg)`. **No deprecated aliases** — `CLAUDE.md`
|
||
records that the last quota reshape updated every consumer in the same change,
|
||
and that is the standard here too.
|
||
|
||
### Shape
|
||
|
||
```json
|
||
"quota": {
|
||
"period": {
|
||
"start": "2026-09-06",
|
||
"next_reset": "2026-10-06",
|
||
"elapsed_fraction": 0.167,
|
||
"source": "objective.billing_reset_day"
|
||
},
|
||
"accounts": [
|
||
{
|
||
"provider": "neuralwatt",
|
||
"shape": "metered_plan",
|
||
"plan": {
|
||
"unit": "kwh",
|
||
"allowance": 6.25,
|
||
"used": 2.87933,
|
||
"used_fraction": 0.4607,
|
||
"pace_ratio": 2.76,
|
||
"projected_period_total": 17.24,
|
||
"pace_note": null
|
||
},
|
||
"spend_usd": {
|
||
"period": 22.66,
|
||
"window_30d": 61.2,
|
||
"source": "per_request_billed",
|
||
"attribution_coverage": 1.0
|
||
},
|
||
"credit": {
|
||
"balance_usd": -0.004198,
|
||
"meaning": "overage_allowance",
|
||
"observed_at": "2026-09-10T18:49:43Z",
|
||
"age_seconds": 19980
|
||
},
|
||
"fidelity": "per_request"
|
||
},
|
||
{
|
||
"provider": "openrouter",
|
||
"shape": "prepaid_credit",
|
||
"pool": {
|
||
"total_credits_usd": 50.0,
|
||
"used_usd": 9.5714,
|
||
"remaining_usd": 40.4286,
|
||
"used_fraction": 0.1914,
|
||
"observed_at": "2026-09-10T23:17:32Z",
|
||
"age_seconds": 7200
|
||
},
|
||
"burn": {
|
||
"window_hours": 24,
|
||
"usd_per_hour": 0.0337,
|
||
"projected_hours_remaining": 1200.9,
|
||
"runway_low_warning": false,
|
||
"note": null
|
||
},
|
||
"spend_usd": {
|
||
"period": 9.5714,
|
||
"window_30d": 9.5714,
|
||
"source": "account_poll_delta",
|
||
"attribution_coverage": 0.0
|
||
},
|
||
"fidelity": "account_poll"
|
||
},
|
||
{
|
||
"provider": "ollama-local",
|
||
"shape": "self_hosted",
|
||
"energy": { "unit": "kwh", "period": 0.01725, "window_30d": 0.0402 },
|
||
"spend_usd": {
|
||
"period": 0.00274,
|
||
"window_30d": 0.0064,
|
||
"source": "derived_from_draw",
|
||
"attribution_coverage": 1.0
|
||
},
|
||
"tariff_usd_per_kwh": 0.15856,
|
||
"fidelity": "derived"
|
||
}
|
||
],
|
||
"spend": {
|
||
"by_provider_usd": { "neuralwatt": 22.66, "openrouter": 9.5714, "ollama-local": 0.00274 },
|
||
"total_usd": 32.23,
|
||
"estimated_usd": { "neuralwatt": 86.05, "openrouter": 10.07 },
|
||
"estimate_ratio": { "neuralwatt": 4.14, "openrouter": 1.13 }
|
||
},
|
||
"alarm": {
|
||
"kind": "plan_pace",
|
||
"provider": "neuralwatt",
|
||
"severity": "warn",
|
||
"headline": "neuralwatt: 2.8x sustainable pace, projecting 17.2 kWh against a 6.25 kWh plan"
|
||
},
|
||
"note": "router-metered only; traffic bypassing the router is not counted"
|
||
}
|
||
```
|
||
|
||
### Decisions embedded in that shape
|
||
|
||
- **`shape` is derived from config, never hardcoded per provider name.** The
|
||
rule, in precedence order, evaluated per `cfg.dispatch_providers` entry plus
|
||
the local-dispatch provider:
|
||
|
||
| # | test | shape |
|
||
|---|---|---|
|
||
| 1 | provider is the local-dispatch provider (`ollama-local`) | `self_hosted` |
|
||
| 2 | `has_energy_telemetry` and `objective.plan_kwh_per_period` is set | `metered_plan` |
|
||
| 3 | `balance_url` is set | `prepaid_credit` |
|
||
| 4 | none of the above | `unmetered` |
|
||
|
||
First match wins, and the order matters: a provider could satisfy both 2 and
|
||
3, and the plan is the thing with a period and a ceiling, so it leads.
|
||
`unmetered` is a real state, not an error — a configured provider with no
|
||
telemetry, no plan and no balance endpoint genuinely has nothing to meter,
|
||
and it renders as a name plus its per-request spend if any, with no gauge.
|
||
A `metered_plan` provider when `plan_kwh_per_period` is null degrades to
|
||
case 3 or 4 rather than rendering a gauge against a null allowance.
|
||
|
||
- **`spend.by_provider_usd[p]` is exactly `accounts[p].spend_usd.period`**, and
|
||
`spend.total_usd` is the sum of those values — nothing else. It is never
|
||
computed from a second query, and a balance figure never enters it. This is
|
||
Rule 1 as an implementation constraint: each account has already decided
|
||
whether its own period spend comes from summed per-request cost, an account
|
||
poll delta, or derived draw, and the ledger just adds up those decisions.
|
||
`estimated_usd[p]` is a separate `SUM(route_decisions.est_cost_usd)` over the
|
||
same period and provider, and `estimate_ratio[p]` is
|
||
`estimated_usd[p] / by_provider_usd[p]`, null when the denominator is 0.
|
||
|
||
- **`accounts` is a list, not a map keyed by provider**, ordered by alarm
|
||
severity then provider name. The old `by_provider` map forced every consumer
|
||
to re-derive which provider mattered; the list makes the answer positional
|
||
and makes the TUI's row order free.
|
||
- **`total_balance_usd` is deleted.** It summed a prepaid pool with an overage
|
||
residue. `spend.total_usd` replaces it as the one honest cross-provider
|
||
number, because dollars spent *are* additive across billing shapes.
|
||
- **`shape` drives rendering, `fidelity` qualifies it.** A consumer switches
|
||
layout on `shape` and prints a caveat from `fidelity`; neither is inferred
|
||
from provider name or from which config keys happen to be set.
|
||
- **`spend_usd.source`** is one of `per_request_billed`, `account_poll_delta`,
|
||
`derived_from_draw`. **`attribution_coverage`** is the fraction of that
|
||
provider's period spend that per-request rows can account for — 1.0 for
|
||
NeuralWatt, 0.0 for OpenRouter until work item 1 lands, and a partial value
|
||
for the period that straddles the deploy. This is Rule 1 made machine
|
||
readable: it is how a card knows to say "breakdown covers $4.10 of $9.57".
|
||
- **`credit.meaning`** replaces `balance_source`. `telemetry` / `polled` /
|
||
`unconfigured` described *where the number came from*; `overage_allowance` /
|
||
`prepaid_pool` / `not_configured` describe *what it means*, which is what
|
||
decides whether it deserves a gauge or a footnote.
|
||
|
||
### `pace_ratio` — the computation, and its trap
|
||
|
||
`pace_ratio = used_fraction / elapsed_fraction`, with
|
||
`elapsed_fraction = (now − period_start) / (next_reset − period_start)`.
|
||
|
||
**Trap: `elapsed_fraction` → 0 at period start, and the ratio explodes.** In
|
||
the first hours of a period a handful of calls reads as a 40x overrun. Guard
|
||
it: when `elapsed_fraction < 0.02` (roughly the first 14 hours of a 30-day
|
||
period), return `pace_ratio: null` and `projected_period_total: null` with
|
||
`pace_note: "too early in the period to project"`. A null pace must render as
|
||
"—", never as 0 or 1.0.
|
||
|
||
**Second trap, and this project has been burned by it once already.** The
|
||
projection must not read as a wall. `objective.plan_kwh_per_period` gates
|
||
nothing; overage is billed. `plans/quota-balance-and-burn-rate.md` records the
|
||
last warning that got this wrong ("a quota is a wall, not a bill — requests
|
||
fail rather than costing more") and why it was both alarming and false. Project
|
||
in kWh and say what it costs, never "requests will fail".
|
||
|
||
### `alarm` — what the chip shows
|
||
|
||
One alarm, the worst across accounts, or `null`:
|
||
|
||
- `plan_pace`: `pace_ratio >= objective.plan_pace_warn_ratio` (new knob,
|
||
default `1.25`, **which must be written into `config/config.yaml`**, not
|
||
left as a Pydantic default). `severity: "critical"` at `>= 2.0`.
|
||
- `credit_runway`: existing `projected_hours_remaining <
|
||
objective.quota_runway_warning_hours` (6). `critical` below half that.
|
||
- `stale_reading`: a `prepaid_credit` account whose `age_seconds` exceeds 3x
|
||
the poll interval. **This one is new and it matters**: the live pool sat flat
|
||
at $40.453746437 for six consecutive polls across 18 hours while 705
|
||
completions ran, and the current UI renders that as a healthy account with no
|
||
burn. It is poll granularity, not calm. A flat reading and a stale reading
|
||
must be distinguishable on the page.
|
||
|
||
`coverage.warnings` (`src/metrics.py:493`) currently emits the >80%-of-plan
|
||
string. Replace it with a pace-based warning sourced from `alarm`, so the
|
||
navbar bell and the chip cannot disagree.
|
||
|
||
### Consumers to update in the same change
|
||
|
||
Measured inventory — `grep -rln "total_balance_usd\|by_provider"` over the
|
||
repo. There is no deprecation path, so all of these move together:
|
||
|
||
| file | what breaks |
|
||
|---|---|
|
||
| `src/metrics.py:203,334` | the two functions being replaced |
|
||
| `src/dispatcher.py` | `/metrics` assembly |
|
||
| `src/admin.py:1849` | `/api/snapshot`'s `quota` key, plus the new `GET /admin/api/quota` |
|
||
| `src/tui_model.py:41-60` | builds `{label, value}` quota rows from the flat keys |
|
||
| `src/tui.py:78,105,566,601,615` | per-provider row formatting, the kWh lead line, and the `total_balance_usd` line `CLAUDE.md` documents |
|
||
| `tests/test_metrics.py` | **49 references** — the bulk of the work |
|
||
| `tests/test_metrics_endpoint.py` | 6 |
|
||
| `tests/test_tui.py` | 6 |
|
||
| `tests/test_credit_attenuation_routing.py` | 6 — asserts attenuation reads balances, so it must keep passing against the new shape without weakening what it pins |
|
||
| `CLAUDE.md`, `docs/api.md`, `docs/admin-portal.md` | describe the old contract by name |
|
||
|
||
`test_credit_attenuation_routing.py` deserves a second look rather than a
|
||
mechanical rename: `credit_attenuation` derives a routing multiplier from
|
||
polled balances, so its fixtures encode the *old* assumption that a balance is
|
||
one flat number per provider. Its behaviour must not change here — only where
|
||
it reads the balance from.
|
||
|
||
## Work item 4 — `/admin/quota`, a page
|
||
|
||
The chip-plus-modal cannot hold three shaped cards and a ledger. Follow the
|
||
portal's page-per-domain pattern (`providers`, `proficiency`, `profiles`).
|
||
|
||
**Mechanics** (`src/admin.py:823-870`): add `_admin_quota` →
|
||
`admin/frontend/quota.html`, a `@router.get("/quota")` returning it with
|
||
`_NO_CACHE_HEADERS`, and a nav item after `Providers` in the navbar of **all**
|
||
admin pages (`admin/frontend/*.html:281-287`) — the navbar is duplicated per
|
||
page, so a nav item added to one page only is a half-shipped nav.
|
||
|
||
**Data**: `GET /admin/api/quota` → `metrics.quota_accounts(conn, cfg)` plus the
|
||
spend ledger's per-model rows. A dedicated endpoint rather than reusing
|
||
`/api/snapshot`, because the page needs per-model spend the snapshot does not
|
||
carry and none of the snapshot's proficiency/verdict payload.
|
||
|
||
### Cards, one per account, laid out by `shape`
|
||
|
||
Three peer cards on one row is `.col-xl-4` — the Standards doc is explicit that
|
||
`.col-xl-6` strands a third card and opens two dead columns.
|
||
|
||
**`metered_plan` (NeuralWatt).** The kWh gauge is the hero, and the pace marker
|
||
is what makes it a redesign rather than a reskin: draw the gauge with a tick at
|
||
`elapsed_fraction` so "burned" and "elapsed" are visually comparable in one
|
||
glance. Below it: pace ratio, projected period total, **$ billed this period**,
|
||
and the overage allowance as a de-emphasized footnote labeled
|
||
`overage allowance` — never as a balance or a runway, because at −$0.004 it is
|
||
neither.
|
||
|
||
**`prepaid_credit` (OpenRouter).** A depletion gauge: `used_usd` of
|
||
`total_credits_usd`, remaining in the hero position. Then burn $/hr, runway in
|
||
days (hours only under ~48h), and the **reading age** rendered as text
|
||
("polled 2h ago"), promoted to a warning tint past the staleness threshold.
|
||
When `total_credits_usd` is NULL, degrade to a bare remaining figure with no
|
||
gauge.
|
||
|
||
**`self_hosted` (local).** kWh and $ this period, the tariff it was priced at,
|
||
and a one-line statement that this is billed to the operator's utility rather
|
||
than a provider. No gauge — there is no allowance to fill. Hidden entirely
|
||
when `local_energy.enabled` is false.
|
||
|
||
### Spend ledger
|
||
|
||
Below the cards, full width:
|
||
|
||
1. **`spend.total_usd` as the page's single headline number**, with a stacked
|
||
bar by provider. This is the answer to "what did I spend this month" and it
|
||
is the reason the page exists.
|
||
2. **Ranked per-model spend** — model, provider, calls, $, share of total.
|
||
`—` in the $ column where cost is unknown, with the row still present and
|
||
ranked by calls. Where `attribution_coverage < 1`, a single line above the
|
||
table states the gap in dollars, e.g. "breakdown accounts for $4.10 of
|
||
OpenRouter's $9.57; the remainder predates per-request cost capture".
|
||
3. **Estimate calibration**, a small two-column strip: est vs billed per
|
||
provider with the ratio. Label it as what it is — the cost tiebreak's input
|
||
measured against the bill — and do not editorialize a fix; that is the
|
||
follow-up plan.
|
||
|
||
### Chip
|
||
|
||
`renderQuotaChip` (`admin/frontend/index.html:779`) stops hardcoding the kWh
|
||
story. It renders `alarm`: headline text, severity tint, and a link to
|
||
`/admin/quota`. With `alarm: null` it shows the period's spend to date, which
|
||
is calm and still useful. Its comment at line 790 ("The chip always tells the
|
||
kWh-subscription story") is the assumption being deleted — remove it, do not
|
||
leave it stale.
|
||
|
||
**The chip stops being a modal trigger and becomes a link.** Four edits, all in
|
||
`index.html`, and leaving any one of them behind leaves a dead control:
|
||
|
||
1. Line 305 — the chip `<div>` carries `data-bs-toggle="modal"`
|
||
`data-bs-target="#quota-modal"` `role="button"` `tabindex="0"`. Replace the
|
||
element with an `<a href="/admin/quota">` carrying the same `.quota-chip`
|
||
class, and drop all four attributes: an anchor is focusable and activatable
|
||
on its own, so `role`/`tabindex` become wrong rather than merely redundant.
|
||
The navbar's `z-index: 1030` note in the Standards doc exists because of
|
||
this chip's `backdrop-filter` — keep the stacking context intact.
|
||
2. Lines 462-478 — the `#quota-modal` block is deleted outright.
|
||
3. `renderQuotaModal` (`:659`) and its `#quota-modal-body` writes are deleted,
|
||
and the `renderQuotaModal(data.quota)` call at `:558` with it.
|
||
4. `grep -n 'quota-modal' admin/frontend/*.html` must come back empty when
|
||
done. `formatProviderBalance` / `formatBalanceStaleness` (`:708`, `:731`)
|
||
move to the new page rather than being deleted — the staleness helper is
|
||
exactly what the `prepaid_credit` card's reading-age line needs, and its
|
||
tiny-negative-balance handling is hard-won (it exists so NeuralWatt's
|
||
−$0.004 never renders as "-$0.00").
|
||
|
||
## Work item 5 — the per-model usage pane
|
||
|
||
Two panes read `data.per_model` today: `renderCloudEnergy`
|
||
(`admin/frontend/index.html:906`) and `renderBars` (`:935`). Three defects,
|
||
each falsifiable:
|
||
|
||
1. **Bar length and bar label disagree.** `renderBars` ranks by
|
||
`BARS_SORT_KEYS[_barsSortKey]` (`:840`) — calls, cost or energy — but the
|
||
value label is hardcoded to `$cost · kWh` (`:957`). Sorting by calls
|
||
produces bars ordered by call count, labeled in dollars.
|
||
2. **Unknown renders as zero**, giving `xiaomi/mimo-v2.5` a `$0.0000` label on
|
||
655 real calls (Defect 3 above).
|
||
3. **`ollama-local` appears in a card called "Cloud energy"**, double-counting
|
||
against the Local Compute card that owns it.
|
||
|
||
And a fourth, which is the operator's own framing: calls, cost and energy
|
||
compete for one row of space, and energy is the least actionable of the three.
|
||
|
||
### The rework
|
||
|
||
- **One ranked list.** Sort keys become **calls | spend | tokens**. Energy
|
||
leaves the sort keys: it is not a thing an operator ranks models by, it is a
|
||
thing they inspect. `tokens` replaces it because it is the honest volume
|
||
metric for a provider that reports no energy at all.
|
||
- **The label follows the sort key**, plus share of total: sorting by calls
|
||
shows `655 calls · 8.2%`; by spend, `$1.23 · 3.8%`. One metric, the one
|
||
being ranked.
|
||
- **`—` for unknown**, never `$0.0000`, with a footnote naming the reason
|
||
("provider does not report per-request cost").
|
||
- **Energy, carbon, attribution and grid move to a row-click detail popup** —
|
||
kWh, gCO2eq, `grid_id`, mean `attribution_ratio`, mean completion tokens,
|
||
cached-token share once work item 1 lands. Render only the fields that have
|
||
data; a popup of six `—`s is worse than a popup that admits it has nothing.
|
||
|
||
**Copy the pattern that already exists rather than inventing one.**
|
||
`admin/frontend/proficiency.html` does exactly this for its matrix cells: a
|
||
`<div class="modal modal-blur fade" id="cell-modal">` with a `-title` and
|
||
`-body` (`:264-276`), a lazily-constructed `new tabler.Modal(...)` held in a
|
||
module-level `_cellModal` (`:575`), and a click handler that sets
|
||
`textContent` on the title and `innerHTML` on the body before `.show()`
|
||
(`:582-587`). Mirror it as `#model-modal` on the dashboard.
|
||
|
||
**No new endpoint.** The popup's fields all come from the `per_model` row the
|
||
pane already has in hand — the click handler passes the row object it
|
||
rendered, so there is no fetch, no loading state, and no second source of
|
||
truth to drift. That is also why the new aggregate columns in "Metrics
|
||
support" below are added to `per_model` itself rather than to a detail route.
|
||
This is the operator's own proposal and it is the right one: attribution
|
||
ratio in particular is a diagnostic on a 750x-between-models, 1.8x-within-model
|
||
quantity, which is an inspection surface, not a ranking.
|
||
- **Cloud pane excludes `ollama-local`.** Filter on provider, not on whether
|
||
energy is present.
|
||
- **`renderCloudEnergy`'s 5-decimal summary trio is cut.** Period spend now has
|
||
a first-class home on `/admin/quota`; a second, differently-windowed copy of
|
||
it on the dashboard is how the two drift apart.
|
||
|
||
### Metrics support
|
||
|
||
`metrics.per_model` (`src/metrics.py:1386`, `def per_model`) adds
|
||
`sum_prompt_tokens`, `sum_completion_tokens`, `sum_cached_prompt_tokens`, and
|
||
`cost_rows` / `total_rows` so the UI can tell "no cost reported" from "cost was
|
||
zero". Stop `COALESCE`ing cost and energy to 0 — return NULL and let the
|
||
consumer render `—`.
|
||
|
||
**That coalesce has a second consumer, and removing it crashes the TUI rather
|
||
than merely uglifying it.** `src/tui_model.py:108-116` maps `sum_cost_usd` /
|
||
`sum_energy_kwh` / `sum_carbon_g_co2eq` straight through to a row dict, and
|
||
`src/tui.py` then renders three cells from it — but they are not equally safe:
|
||
|
||
| line | code | on `None` |
|
||
|---|---|---|
|
||
| `src/tui.py:667` | `_fmt_usd(r["cost_usd"])` | safe — `_fmt_usd` returns `"n/a"` (`src/tui.py:55-58`) |
|
||
| `src/tui.py:668` | `f"{r['energy_kwh']:.6g}"` | **TypeError** |
|
||
| `src/tui.py:669` | `f"{r['carbon_g_co2eq']:.4g}"` | **TypeError** |
|
||
|
||
The `COALESCE` is the only thing standing between those two format strings and
|
||
an unhandled exception, so the fix for 668-669 is not "render `—`" as a styling
|
||
choice — it is required for the TUI to start. Give both the same None-aware
|
||
treatment `_fmt_usd` already has (a `_fmt_num(v, spec)` helper returning `"n/a"`
|
||
on None is the smallest change that covers both), and add a `per_model` fixture
|
||
row with NULL energy to `tests/test_tui.py` so nothing re-introduces a bare
|
||
format on a nullable column.
|
||
|
||
## Tests
|
||
|
||
New or extended, offline, matching this repo's existing fixtures:
|
||
|
||
- `tests/test_metrics.py` — the existing quota tests (49 references) are
|
||
rewritten against the new shape, not deleted. Add: one account per configured
|
||
provider; `shape`/`fidelity`/`credit.meaning` values; **`pace_ratio` is null
|
||
inside the first 2% of a period** and correct outside it; `spend.total_usd`
|
||
sums only per-provider spend and never a balance; no `total_balance_usd` key
|
||
survives anywhere in the payload; a NULL `total_credits_usd` yields a pool
|
||
block with no `used_fraction`.
|
||
- A `stale_reading` case — a balance series that is flat *and* old raises the
|
||
alarm; flat but fresh does not. This is the 18-hour flat-pool state observed
|
||
on the live DB, and it is the one alarm nothing covers today.
|
||
- Dispatcher cost capture — four cases: flag off → no `usage` key in
|
||
`upstream_body`; flag on, buffered → `cost_usd` from `usage.cost`; flag on,
|
||
**streamed** → `cost_usd` non-null (the trap in work item 1); NeuralWatt's
|
||
`cost.request_cost_usd` still wins when both blocks are present.
|
||
- Poller — three-tuple parser, the two new columns persisted, idempotent
|
||
migration against a table that predates them, and a parser returning NULLs
|
||
for pool size.
|
||
- Admin — `/admin/quota` serves the file; `/admin/api/quota` returns the
|
||
documented keys.
|
||
- TUI — update the quota panel's fixtures; keep
|
||
`tests/test_tui_schema_drift.py` and `tests/test_tui_warnings.py` green.
|
||
**`test_tui_warnings.py`'s fixture is a coupled system** — `CLAUDE.md`
|
||
records that adding a seed there can silence an unrelated warning class — so
|
||
changing the plan-usage warning means re-checking both directions of that
|
||
test, not just the failing one.
|
||
|
||
## Docs to update in the same change
|
||
|
||
- `CLAUDE.md` — the `quota.by_provider` / `total_balance_usd` paragraph in
|
||
"Billing is per-kWh", and the `credit_attenuation` section's two-semantics
|
||
note, which is built on `total_balance_usd` existing.
|
||
- `docs/api.md` — the `/metrics` `quota` block.
|
||
- `docs/admin-portal.md` — the new page, the chip's new meaning, the reworked
|
||
per-model pane.
|
||
- `docs/data-model.md` — three new columns.
|
||
- `plans/admin-design-standards.md` — only if the pace-marker gauge or the
|
||
depletion gauge introduces a pattern the next page should reuse. It probably
|
||
does; record the grammar (gauge + reference tick), not the CSS.
|
||
|
||
## Out of scope, recorded so it is not rediscovered
|
||
|
||
- **The 4.14x est-vs-billed bias.** Surfaced by work item 3, fixed elsewhere.
|
||
- **Replacing `routing.assumed_cache_rate` with measured `cached_tokens`.**
|
||
Work item 1 starts capturing the data; consuming it changes pricing and
|
||
therefore routing, which needs its own plan and its own before/after.
|
||
- **Energy-budget attenuation** (a `credit_attenuation` analogue keyed on plan
|
||
pace). The pace signal this plan computes is its natural input, but it is a
|
||
routing feature.
|
||
- **Backfilling OpenRouter cost.** Not possible per request; Rule 1 is the
|
||
reason it is not needed for account totals.
|