Files
6krrt/plans/multi-provider-support.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

261 lines
15 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Generalizing to multiple providers
Status: parked -- parked on provider selection; the coupling surface in it is still valid
**Status: PARKED 2026-09-04 — DRAFT, not decision-complete, and now blocked on
provider selection rather than on design.** Written 2026-09-04 against `main`
at `340453e`; coupling surface measured and three open questions closed by code
inspection on the same commit.
Parked by the user: **Z.ai is no longer a settled first target.** The GLM
overlap that made it attractive is a real argument, but fit is unconfirmed and
the candidate search is still open. Nothing below should be read as a
commitment to a specific provider.
Remaining gaps before this can go to the opencode/Prometheus pipeline:
1. **Which provider goes first** (was treated as answered; it is not).
2. The proficiency-sharing decision — a judgement call, not a lookup.
3. The eco-without-telemetry decision.
4. Success criteria / test plan / phasing, still empty.
**What does NOT need to wait.** The measured coupling surface below is
provider-agnostic: the 20 hardcoded `neuralwatt` references are literals that
are wrong regardless of which provider lands second, and three of them are
latent defects today. Re-verified on 2026-09-04 against
`feat/config-local-overlay` — still 20 references, same six-file distribution.
Those can be parameterised independently of this plan and should not be held
hostage to the provider question.
## Why this, why now
Everything downstream of the catalog — `scoring.py`, `tiering.py`,
`routing.py`, `proficiency.py`, `verification.py`, `feedback.py`, the admin
portal, the TUI — already operates on normalized DB rows with no idea which
provider produced them. `dispatch_providers: dict[str, DispatchProvider]` in
config and the `(model_id, provider)` composite key were deliberately kept
when OpenRouter was dropped (`c3484f0`) for exactly this. The coupling that's
actually left is narrow:
- **`poller.py`** — `fetch_neuralwatt()` is a bespoke function: hardcoded
URL, NeuralWatt's specific catalog JSON shape, `provider="neuralwatt"`
written as a literal.
- **`dispatcher.py`** — three things tangled together that need separating:
1. Usage/telemetry parsing (NeuralWatt reports energy via SSE **comment**
lines — a protocol quirk, not an OpenAI standard).
2. Cost semantics (the $8/kWh-billed-capped-at-3x-list model is
NeuralWatt-specific billing behavior). What `routing.estimated_cost`
actually scores on — catalog price × request shape — is already
provider-agnostic; the kWh math is validation for one provider's
billing quirk, not the load-bearing input.
3. Account-refusal handling (`degrade to local` assumes one cloud account
— see dispatcher.py ~2537). With N providers this mostly *disappears*:
circuit-break the failing provider's rows and let ranking fail over to
the next-best candidate on a different provider.
**On re-adding OpenRouter specifically:** it was dropped for a real reason,
not a bad one — the project pivoted to scoring on *measured* billing and
carbon, and OpenRouter (an aggregator) doesn't expose per-request
energy/carbon telemetry the way NeuralWatt does. That's not a reason to
avoid it now — it's the first real test case for a `has_energy_telemetry`
capability flag, since eco/energy needs to degrade gracefully per-provider
rather than assuming every row has NeuralWatt's shape. Worth remembering
before re-proposing it as if it were untried.
## Constraint: cheap, no-commitment testing
Budget is small and this is for testing the plumbing, not production spend —
prepaid credits in ~$10 increments, no subscriptions. Candidates, **pricing
and minimums need re-verification before committing to one** (this kind of
thing moves):
| provider | fit | notes |
|---|---|---|
| Z.ai | best first target | Pay-as-you-go API, OpenAI-compatible (`https://api.z.ai/api/paas/v4` or `/api/openai/v1`), no stated minimum top-up. Three free-tier models (GLM-4.7-Flash, GLM-4.5-Flash, GLM-4.6V-Flash) for zero-cost plumbing tests. **Actually the publisher of the GLM family this catalog already serves via NeuralWatt** (`glm-5.2-fast`, `glm-5.2-flex`) — a same-weights cross-provider comparison, which is a stronger generalization test than an unrelated model set, and a way to sanity-check NeuralWatt's `static_fallback` carbon figure for GLM. That is not hypothetical: `CLAUDE.md` records the GLM rows reporting `grid_id: FI` at 475 gCO2/kWh as `carbon_source: static_fallback` — a substituted constant, not a measurement — and they are currently EXCLUDED from eco scoring for that reason. A second provider serving the same weights is a free experiment on one of this project's standing open questions. No response-level energy/carbon telemetry — same capability-flag case as OpenRouter, not a differentiator there. Don't confuse this with the separate "GLM Coding Plan" subscription ($18-168/mo) — that's a different product and not what fits the budget constraint here. |
| OpenRouter | good second | prepaid, no minimum, OpenAI-compatible, many free models for zero-cost plumbing tests, broad catalog overlap (kimi/deepseek/qwen/gemma too, not just GLM), and this repo already has git history (`c3484f0` and its parent) to mine for catalog-normalization shape. No per-request energy/carbon. It's an aggregator of aggregators, so if it ever *does* report grid/energy data, treat it as less trustworthy than a direct provider's own figure. |
| DeepInfra | plausible | prepaid, no subscription, OpenAI-compatible, cheap open-weight catalog. No energy telemetry. |
| Together AI | plausible | prepaid credits, OpenAI-compatible, per-token billing (straightforward vs. NeuralWatt's kWh math). |
| Fireworks AI | plausible | same shape as Together. |
| Groq | maybe later | OpenAI-compatible, free tier + pay-as-you-go, but a much smaller catalog — more useful for a latency-tolerance test than a routing-breadth test. |
**Recommendation, downgraded 2026-09-04.** The *shape* still holds: **one
provider first**, all the way through poller → dispatch → a handful of real
routed requests, before touching a second. That is what confirms the
abstraction generalizes instead of just looking like it does on paper.
Which provider is **reopened**. Z.ai was the pick on the strength of the GLM
overlap — a same-weights cross-provider comparison, and a free experiment on
the `static_fallback` carbon question. That argument is still good and should
be re-used to score whatever candidate comes next; it was never a claim that
Z.ai's API, pricing or terms actually fit. Treat the table above as a
shortlist to re-verify, not a ranking to execute.
Worth noting for the search: the criterion that made Z.ai attractive —
**catalog overlap with what NeuralWatt already serves** — is separable from
Z.ai itself. OpenRouter carries kimi, deepseek, qwen and gemma alongside GLM,
so it satisfies the same-weights test more broadly, and this repo already has
git history (`c3484f0` and its parent) with a working catalog normalizer to
mine. It was dropped for a reason that no longer disqualifies it, now that
`has_energy_telemetry` is the planned answer to missing energy data.
## Architecture sketch
A `Provider` protocol/interface — not fleshed out yet, but the shape:
- `fetch_catalog() -> list[ModelRow]`
- `parse_usage(response) -> UsageInfo` (cost, tokens, energy if present)
- `detect_refusal(error) -> bool`
- capability flags: `has_energy_telemetry`, `has_regional_carbon`
`fetch_neuralwatt` and the SSE-comment parser move behind it as the first
concrete implementation. `dispatch_providers` entries get a `type:`
discriminator so config knows which implementation to instantiate.
## Measured coupling surface
Claims about how narrow the coupling is should be counted, not asserted.
`grep -rn neuralwatt src/*.py` on `main` at `340453e` returns **20 hardcoded
references across 6 files**:
| file | refs | what they are |
|---|---|---|
| `poller.py` | 8 | the bespoke fetch: URL, `fetch_neuralwatt()`, literal `provider="neuralwatt"`, the sanity-floor count query, log prefixes |
| `eval_proficiency.py` | 5 | `provider: str = "neuralwatt"` defaults, the judge's `dispatch_providers["neuralwatt"]` lookup, and an `identity["provider"] != "neuralwatt"` skip |
| `dispatcher.py` | 4 | passthrough default, and a capability lookup (see below) |
| `metrics.py` | 1 | catalog-staleness query |
| `leaderboard.py` | 1 | `set_leaderboard(..., "neuralwatt", ...)` |
| `seed_energy.py` | 1 | `dispatch_providers["neuralwatt"]` |
That is the actual work list. Two entries deserve calling out because they
**contradict the "needs ~zero change" section below.**
### `metrics.py:175` — catalog staleness would ignore a second provider
```sql
SELECT MAX(last_updated) AS last_updated FROM models WHERE provider='neuralwatt'
```
The staleness warning is computed over NeuralWatt rows only. Add a provider
whose poller silently stops and the warning stays green while its catalog
freezes — the exact silent-and-open failure `CLAUDE.md`'s "Run as a service"
section describes, now with a second way in. `metrics.py` is listed under
"needs ~zero change"; it does not.
### `leaderboard.py:125` — priors are written to a hardcoded provider
`set_leaderboard(conn, cfg, model_id, "neuralwatt", category, score)` pins the
provider literal, so a curated prior for a shared model (GLM, say) attaches to
the NeuralWatt row and never to the Z.ai one. `leaderboards.yaml` ships empty
today so nothing is broken yet, but this decides the answer to the
proficiency-sharing question above by accident rather than on purpose.
### `dispatcher.py:2879` — `_model_exists` silently strips a real model id
**Corrected 2026-09-04.** An earlier revision of this plan attributed this line
to `_check_pinned_capabilities`. That was wrong: `_check_pinned_capabilities`
(`dispatcher.py:2887`) already takes `provider` as a parameter and passes it to
its query. It is not a coupling site. Line 2879 belongs to `_model_exists`,
and the defect there is worse.
```sql
SELECT 1 FROM models WHERE model_id = ? AND provider = 'neuralwatt'
```
`_model_exists` decides whether a `provider/model` string a client sent is an
opencode-style alias to **strip** (`llm-router/...`) or a real id that merely
contains a slash. Scoped to NeuralWatt, a real second-provider id returns
False and gets treated as an alias.
That matters more than it looks, because **`vendor/model` is the native id
format for the two strongest candidates**: OpenRouter ids are always
`deepseek/deepseek-chat`-shaped, and several other aggregators follow suit. So
the failure is not exotic — it is the default case for a whole class of
provider.
And it fails **silently in the wrong direction**. A provider-scoped 422 would
at least be loud. This one strips the prefix and proceeds with the remainder,
so a client pinning `deepseek/deepseek-chat` gets whatever `deepseek-chat`
resolves to — a different row, on a different provider, at a different price —
with nothing in the response indicating a substitution occurred. Fail-closed
was the design intent everywhere else in this capability path; here it fails
open.
Requirement: `_model_exists` must resolve across all configured providers, and
alias-stripping must be decided by something other than "no NeuralWatt row has
this id".
### Risk not in the draft: per-provider fetch isolation
`poller.mark_stale` runs only inside `poller.main()`, and `main()` returns
early on `RequestException` — **before** `upsert` and **before** `mark_stale`.
With two providers in one run, a failure fetching provider B can skip
`mark_stale` for provider A entirely, or a partial/empty `data` array from B
can mark rows stale that were never B's. `docs/incidents.md` records the
empty-`data` path as the one that can empty the candidate set.
Requirement: **each provider's fetch, upsert and staleness marking must be
isolated.** One provider's outage must not affect another's rows in either
direction, and the sanity floor (`poller.py:420`) must be per-provider.
## What should need ~zero change
Scoring, tiering, routing, proficiency, verification, feedback, admin, TUI.
If any of these turn out to need provider-aware branching, that's a sign the
interface boundary is in the wrong place — worth treating as a red flag
during implementation, not a shrug.
**Two of them already fail that test**, per the inventory above: `metrics.py`
hardcodes the provider in its staleness query and `leaderboard.py` hardcodes it
when writing priors. Neither is a deep coupling — both are literals, not
branching — but the list should be read as "should need ~zero change *after*
those two literals are parameterised", not as a claim that they are already
clean. The red flag to watch for during implementation is provider-aware
*logic* appearing in these modules; a hardcoded string is a different and much
cheaper problem.
## Open questions
- **Eco/carbon for a provider with no telemetry.** Exclude those rows from
eco ranking entirely, or treat eco as unweighted/missing for them without
disqualifying them? Affects whether a non-NeuralWatt row can ever win on
the eco axis, or only ever competes on cost/quality.
- **Same model, two providers — per-model or per-(model, provider)?**
PARTLY RESOLVED, and the remaining half is a design decision rather than a
verification. The schema already supports divergence: `proficiency`'s
primary key is `(model_id, provider, category)`. But
`propagate_to_variants` copies scores across variants keyed on
`base_model_id`, on the stated principle that proficiency is "a property of
the weights, not the queue". Same-weights-different-serving-stack is exactly
the case that principle does not decide: quantization and serving
differences could justify separate scores, while the existing logic would
share one. **Decide this explicitly before implementation** — it is the one
question here that code inspection cannot answer.
- **RESOLVED — circuit breaker is genuinely provider-generic.** Verified in
code, not prose: `is_down`, `record_failure` and `record_success` all take
`(model_id, provider)` and `_store` is keyed on that tuple
(`src/circuit_breaker.py:29,37,57`). Per-provider breaking works today with
no change.
- **Does the account-refusal-degrades-to-local path actually get simpler or
just get an `if` added?** Stated above as an assumption; worth confirming
against the real code before it's a success criterion.
- **RESOLVED — "no serving class" is already a real state.**
`poller.parse_serving_class` returns the schema defaults
(`'standard'`, `'default'`, `'full'`) for a base id with no suffix, so a
Z.ai or OpenRouter row lands `standard` and is interactive-eligible rather
than null. No change needed.
## Non-goals (this pass)
- Don't wire up more than one new provider before the first one round-trips
end to end.
- Don't try to make eco/carbon methodology uniform across providers that
don't expose the same telemetry — graceful absence beats a fabricated
number (same principle as the empty `leaderboards.yaml`).
- Don't rebuild the cost model — catalog-price × shape already generalizes;
confirm that rather than redesigning it.
## Not yet defined
Success criteria, test plan, and phasing/milestones — fill in once the open
questions above are resolved and this graduates to FINAL.