feat: per-provider quota measurement and credit-aware routing attenuation #40

Merged
alee merged 13 commits from feat/per-provider-balance-attenuation into main 2026-09-06 19:08:45 +00:00
Owner

What this PR does

Per-provider quota measurement (fixes NeuralWatt-only global balance/burn)

  • Adds provider_balance_observations account-balance table and polls any provider with balance_url inside the existing 2h catalog poller.
  • OpenRouter configured with balance_url: https://openrouter.ai/api/v1/credits; its real credit balance now surfaces in /metrics.
  • Rewrites metrics.quota_balance_and_burn to return per-provider balance/burn/runway under quota.by_provider plus a total_balance_usd sum.
  • Removes the old flat quota keys (balance_usd, burn_rate_usd_per_hour, runway_*) from /metrics — a clean break, all in-repo consumers updated.
  • quota_burn meters kWh only against telemetry-capable providers (avoiding empty IN ()).

Opt-in credit-aware routing attenuation

  • New objective.credit_attenuation config block (enabled: false by default).
  • credit_attenuation_multiplier and a new rank_candidates effective_cost tiebreak: healthy providers stay at 1.0; low-balance providers get a multiplier up to max_multiplier.
  • Attenuation applies ONLY to providers with balance_url set (prepaid-account path). Telemetry-allowance-billed providers (NeuralWatt) are pinned to multiplier 1.0, so normal ~-0.004 USD overage readings never bias routing toward OpenRouter.
  • est_cost_usd and logged cost remain honest catalog estimates; effective_cost is routing-only.
  • Dispatcher resolver caches via module globals (refresh_seconds) and emits a structured log line when multipliers != 1.0.

Reviewer note: F2 mutation-testing catch-and-fix

The initial end-to-end attenuation test priced NeuralWatt cheaper than OpenRouter, so the test "flipped" to the healthy provider even without attenuation (tautological). A final-wave reviewer temporarily removed the provider_cost_multipliers passthrough from route() and the whole suite stayed green — the wiring had zero regression coverage. This was fixed by repricing the fixture so OpenRouter is raw-cheaper and only a working passthrough produces the flip; the mutation check now fails exactly the two flip tests (assert 'openrouter' == 'neuralwatt'). Full suite remains green at 1463 tests.

Verification

  • python -m pytest -q → 1463 passed, 0 failures.
  • Live throwaway-port QA: /metrics shows independent per-provider balances; attenuation probe flips selection from OpenRouter → NeuralWatt when OpenRouter credit is low.
  • No config/config.local.yaml changes; port 8080/production untouched.

Post-merge action

After merge, restart the production dispatcher: systemctl --user restart llm-router.service.

## What this PR does ### Per-provider quota measurement (fixes NeuralWatt-only global balance/burn) - Adds `provider_balance_observations` account-balance table and polls any provider with `balance_url` inside the existing 2h catalog poller. - OpenRouter configured with `balance_url: https://openrouter.ai/api/v1/credits`; its real credit balance now surfaces in `/metrics`. - Rewrites `metrics.quota_balance_and_burn` to return per-provider balance/burn/runway under `quota.by_provider` plus a `total_balance_usd` sum. - Removes the old flat quota keys (`balance_usd`, `burn_rate_usd_per_hour`, `runway_*`) from `/metrics` — a clean break, all in-repo consumers updated. - `quota_burn` meters kWh only against telemetry-capable providers (avoiding empty `IN ()`). ### Opt-in credit-aware routing attenuation - New `objective.credit_attenuation` config block (`enabled: false` by default). - `credit_attenuation_multiplier` and a new `rank_candidates` `effective_cost` tiebreak: healthy providers stay at 1.0; low-balance providers get a multiplier up to `max_multiplier`. - Attenuation applies ONLY to providers with `balance_url` set (prepaid-account path). Telemetry-`allowance`-billed providers (NeuralWatt) are pinned to multiplier 1.0, so normal ~-0.004 USD overage readings never bias routing toward OpenRouter. - `est_cost_usd` and logged `cost` remain honest catalog estimates; `effective_cost` is routing-only. - Dispatcher resolver caches via module globals (`refresh_seconds`) and emits a structured log line when multipliers != 1.0. ### Reviewer note: F2 mutation-testing catch-and-fix The initial end-to-end attenuation test priced NeuralWatt cheaper than OpenRouter, so the test "flipped" to the healthy provider even without attenuation (tautological). A final-wave reviewer temporarily removed the `provider_cost_multipliers` passthrough from `route()` and the whole suite stayed green — the wiring had zero regression coverage. This was fixed by repricing the fixture so OpenRouter is raw-cheaper and only a working passthrough produces the flip; the mutation check now fails exactly the two flip tests (`assert 'openrouter' == 'neuralwatt'`). Full suite remains green at 1463 tests. ## Verification - `python -m pytest -q` → 1463 passed, 0 failures. - Live throwaway-port QA: `/metrics` shows independent per-provider balances; attenuation probe flips selection from OpenRouter → NeuralWatt when OpenRouter credit is low. - No `config/config.local.yaml` changes; port 8080/production untouched. ## Post-merge action After merge, restart the production dispatcher: `systemctl --user restart llm-router.service`.
alee added 13 commits 2026-09-06 19:01:29 +00:00
alee merged commit 7811a069d9 into main 2026-09-06 19:08:45 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#40