# Generalizing to multiple providers Status: parked -- parked on provider selection; the coupling surface in it is still valid **Status: PARKED 2026-09-04 — DRAFT, not decision-complete, and now blocked on provider selection rather than on design.** Written 2026-09-04 against `main` at `340453e`; coupling surface measured and three open questions closed by code inspection on the same commit. Parked by the user: **Z.ai is no longer a settled first target.** The GLM overlap that made it attractive is a real argument, but fit is unconfirmed and the candidate search is still open. Nothing below should be read as a commitment to a specific provider. Remaining gaps before this can go to the opencode/Prometheus pipeline: 1. **Which provider goes first** (was treated as answered; it is not). 2. The proficiency-sharing decision — a judgement call, not a lookup. 3. The eco-without-telemetry decision. 4. Success criteria / test plan / phasing, still empty. **What does NOT need to wait.** The measured coupling surface below is provider-agnostic: the 20 hardcoded `neuralwatt` references are literals that are wrong regardless of which provider lands second, and three of them are latent defects today. Re-verified on 2026-09-04 against `feat/config-local-overlay` — still 20 references, same six-file distribution. Those can be parameterised independently of this plan and should not be held hostage to the provider question. ## Why this, why now Everything downstream of the catalog — `scoring.py`, `tiering.py`, `routing.py`, `proficiency.py`, `verification.py`, `feedback.py`, the admin portal, the TUI — already operates on normalized DB rows with no idea which provider produced them. `dispatch_providers: dict[str, DispatchProvider]` in config and the `(model_id, provider)` composite key were deliberately kept when OpenRouter was dropped (`c3484f0`) for exactly this. The coupling that's actually left is narrow: - **`poller.py`** — `fetch_neuralwatt()` is a bespoke function: hardcoded URL, NeuralWatt's specific catalog JSON shape, `provider="neuralwatt"` written as a literal. - **`dispatcher.py`** — three things tangled together that need separating: 1. Usage/telemetry parsing (NeuralWatt reports energy via SSE **comment** lines — a protocol quirk, not an OpenAI standard). 2. Cost semantics (the $8/kWh-billed-capped-at-3x-list model is NeuralWatt-specific billing behavior). What `routing.estimated_cost` actually scores on — catalog price × request shape — is already provider-agnostic; the kWh math is validation for one provider's billing quirk, not the load-bearing input. 3. Account-refusal handling (`degrade to local` assumes one cloud account — see dispatcher.py ~2537). With N providers this mostly *disappears*: circuit-break the failing provider's rows and let ranking fail over to the next-best candidate on a different provider. **On re-adding OpenRouter specifically:** it was dropped for a real reason, not a bad one — the project pivoted to scoring on *measured* billing and carbon, and OpenRouter (an aggregator) doesn't expose per-request energy/carbon telemetry the way NeuralWatt does. That's not a reason to avoid it now — it's the first real test case for a `has_energy_telemetry` capability flag, since eco/energy needs to degrade gracefully per-provider rather than assuming every row has NeuralWatt's shape. Worth remembering before re-proposing it as if it were untried. ## Constraint: cheap, no-commitment testing Budget is small and this is for testing the plumbing, not production spend — prepaid credits in ~$10 increments, no subscriptions. Candidates, **pricing and minimums need re-verification before committing to one** (this kind of thing moves): | provider | fit | notes | |---|---|---| | Z.ai | best first target | Pay-as-you-go API, OpenAI-compatible (`https://api.z.ai/api/paas/v4` or `/api/openai/v1`), no stated minimum top-up. Three free-tier models (GLM-4.7-Flash, GLM-4.5-Flash, GLM-4.6V-Flash) for zero-cost plumbing tests. **Actually the publisher of the GLM family this catalog already serves via NeuralWatt** (`glm-5.2-fast`, `glm-5.2-flex`) — a same-weights cross-provider comparison, which is a stronger generalization test than an unrelated model set, and a way to sanity-check NeuralWatt's `static_fallback` carbon figure for GLM. That is not hypothetical: `CLAUDE.md` records the GLM rows reporting `grid_id: FI` at 475 gCO2/kWh as `carbon_source: static_fallback` — a substituted constant, not a measurement — and they are currently EXCLUDED from eco scoring for that reason. A second provider serving the same weights is a free experiment on one of this project's standing open questions. No response-level energy/carbon telemetry — same capability-flag case as OpenRouter, not a differentiator there. Don't confuse this with the separate "GLM Coding Plan" subscription ($18-168/mo) — that's a different product and not what fits the budget constraint here. | | OpenRouter | good second | prepaid, no minimum, OpenAI-compatible, many free models for zero-cost plumbing tests, broad catalog overlap (kimi/deepseek/qwen/gemma too, not just GLM), and this repo already has git history (`c3484f0` and its parent) to mine for catalog-normalization shape. No per-request energy/carbon. It's an aggregator of aggregators, so if it ever *does* report grid/energy data, treat it as less trustworthy than a direct provider's own figure. | | DeepInfra | plausible | prepaid, no subscription, OpenAI-compatible, cheap open-weight catalog. No energy telemetry. | | Together AI | plausible | prepaid credits, OpenAI-compatible, per-token billing (straightforward vs. NeuralWatt's kWh math). | | Fireworks AI | plausible | same shape as Together. | | Groq | maybe later | OpenAI-compatible, free tier + pay-as-you-go, but a much smaller catalog — more useful for a latency-tolerance test than a routing-breadth test. | **Recommendation, downgraded 2026-09-04.** The *shape* still holds: **one provider first**, all the way through poller → dispatch → a handful of real routed requests, before touching a second. That is what confirms the abstraction generalizes instead of just looking like it does on paper. Which provider is **reopened**. Z.ai was the pick on the strength of the GLM overlap — a same-weights cross-provider comparison, and a free experiment on the `static_fallback` carbon question. That argument is still good and should be re-used to score whatever candidate comes next; it was never a claim that Z.ai's API, pricing or terms actually fit. Treat the table above as a shortlist to re-verify, not a ranking to execute. Worth noting for the search: the criterion that made Z.ai attractive — **catalog overlap with what NeuralWatt already serves** — is separable from Z.ai itself. OpenRouter carries kimi, deepseek, qwen and gemma alongside GLM, so it satisfies the same-weights test more broadly, and this repo already has git history (`c3484f0` and its parent) with a working catalog normalizer to mine. It was dropped for a reason that no longer disqualifies it, now that `has_energy_telemetry` is the planned answer to missing energy data. ## Architecture sketch A `Provider` protocol/interface — not fleshed out yet, but the shape: - `fetch_catalog() -> list[ModelRow]` - `parse_usage(response) -> UsageInfo` (cost, tokens, energy if present) - `detect_refusal(error) -> bool` - capability flags: `has_energy_telemetry`, `has_regional_carbon` `fetch_neuralwatt` and the SSE-comment parser move behind it as the first concrete implementation. `dispatch_providers` entries get a `type:` discriminator so config knows which implementation to instantiate. ## Measured coupling surface Claims about how narrow the coupling is should be counted, not asserted. `grep -rn neuralwatt src/*.py` on `main` at `340453e` returns **20 hardcoded references across 6 files**: | file | refs | what they are | |---|---|---| | `poller.py` | 8 | the bespoke fetch: URL, `fetch_neuralwatt()`, literal `provider="neuralwatt"`, the sanity-floor count query, log prefixes | | `eval_proficiency.py` | 5 | `provider: str = "neuralwatt"` defaults, the judge's `dispatch_providers["neuralwatt"]` lookup, and an `identity["provider"] != "neuralwatt"` skip | | `dispatcher.py` | 4 | passthrough default, and a capability lookup (see below) | | `metrics.py` | 1 | catalog-staleness query | | `leaderboard.py` | 1 | `set_leaderboard(..., "neuralwatt", ...)` | | `seed_energy.py` | 1 | `dispatch_providers["neuralwatt"]` | That is the actual work list. Two entries deserve calling out because they **contradict the "needs ~zero change" section below.** ### `metrics.py:175` — catalog staleness would ignore a second provider ```sql SELECT MAX(last_updated) AS last_updated FROM models WHERE provider='neuralwatt' ``` The staleness warning is computed over NeuralWatt rows only. Add a provider whose poller silently stops and the warning stays green while its catalog freezes — the exact silent-and-open failure `CLAUDE.md`'s "Run as a service" section describes, now with a second way in. `metrics.py` is listed under "needs ~zero change"; it does not. ### `leaderboard.py:125` — priors are written to a hardcoded provider `set_leaderboard(conn, cfg, model_id, "neuralwatt", category, score)` pins the provider literal, so a curated prior for a shared model (GLM, say) attaches to the NeuralWatt row and never to the Z.ai one. `leaderboards.yaml` ships empty today so nothing is broken yet, but this decides the answer to the proficiency-sharing question above by accident rather than on purpose. ### `dispatcher.py:2879` — `_model_exists` silently strips a real model id **Corrected 2026-09-04.** An earlier revision of this plan attributed this line to `_check_pinned_capabilities`. That was wrong: `_check_pinned_capabilities` (`dispatcher.py:2887`) already takes `provider` as a parameter and passes it to its query. It is not a coupling site. Line 2879 belongs to `_model_exists`, and the defect there is worse. ```sql SELECT 1 FROM models WHERE model_id = ? AND provider = 'neuralwatt' ``` `_model_exists` decides whether a `provider/model` string a client sent is an opencode-style alias to **strip** (`llm-router/...`) or a real id that merely contains a slash. Scoped to NeuralWatt, a real second-provider id returns False and gets treated as an alias. That matters more than it looks, because **`vendor/model` is the native id format for the two strongest candidates**: OpenRouter ids are always `deepseek/deepseek-chat`-shaped, and several other aggregators follow suit. So the failure is not exotic — it is the default case for a whole class of provider. And it fails **silently in the wrong direction**. A provider-scoped 422 would at least be loud. This one strips the prefix and proceeds with the remainder, so a client pinning `deepseek/deepseek-chat` gets whatever `deepseek-chat` resolves to — a different row, on a different provider, at a different price — with nothing in the response indicating a substitution occurred. Fail-closed was the design intent everywhere else in this capability path; here it fails open. Requirement: `_model_exists` must resolve across all configured providers, and alias-stripping must be decided by something other than "no NeuralWatt row has this id". ### Risk not in the draft: per-provider fetch isolation `poller.mark_stale` runs only inside `poller.main()`, and `main()` returns early on `RequestException` — **before** `upsert` and **before** `mark_stale`. With two providers in one run, a failure fetching provider B can skip `mark_stale` for provider A entirely, or a partial/empty `data` array from B can mark rows stale that were never B's. `docs/incidents.md` records the empty-`data` path as the one that can empty the candidate set. Requirement: **each provider's fetch, upsert and staleness marking must be isolated.** One provider's outage must not affect another's rows in either direction, and the sanity floor (`poller.py:420`) must be per-provider. ## What should need ~zero change Scoring, tiering, routing, proficiency, verification, feedback, admin, TUI. If any of these turn out to need provider-aware branching, that's a sign the interface boundary is in the wrong place — worth treating as a red flag during implementation, not a shrug. **Two of them already fail that test**, per the inventory above: `metrics.py` hardcodes the provider in its staleness query and `leaderboard.py` hardcodes it when writing priors. Neither is a deep coupling — both are literals, not branching — but the list should be read as "should need ~zero change *after* those two literals are parameterised", not as a claim that they are already clean. The red flag to watch for during implementation is provider-aware *logic* appearing in these modules; a hardcoded string is a different and much cheaper problem. ## Open questions - **Eco/carbon for a provider with no telemetry.** Exclude those rows from eco ranking entirely, or treat eco as unweighted/missing for them without disqualifying them? Affects whether a non-NeuralWatt row can ever win on the eco axis, or only ever competes on cost/quality. - **Same model, two providers — per-model or per-(model, provider)?** PARTLY RESOLVED, and the remaining half is a design decision rather than a verification. The schema already supports divergence: `proficiency`'s primary key is `(model_id, provider, category)`. But `propagate_to_variants` copies scores across variants keyed on `base_model_id`, on the stated principle that proficiency is "a property of the weights, not the queue". Same-weights-different-serving-stack is exactly the case that principle does not decide: quantization and serving differences could justify separate scores, while the existing logic would share one. **Decide this explicitly before implementation** — it is the one question here that code inspection cannot answer. - **RESOLVED — circuit breaker is genuinely provider-generic.** Verified in code, not prose: `is_down`, `record_failure` and `record_success` all take `(model_id, provider)` and `_store` is keyed on that tuple (`src/circuit_breaker.py:29,37,57`). Per-provider breaking works today with no change. - **Does the account-refusal-degrades-to-local path actually get simpler or just get an `if` added?** Stated above as an assumption; worth confirming against the real code before it's a success criterion. - **RESOLVED — "no serving class" is already a real state.** `poller.parse_serving_class` returns the schema defaults (`'standard'`, `'default'`, `'full'`) for a base id with no suffix, so a Z.ai or OpenRouter row lands `standard` and is interactive-eligible rather than null. No change needed. ## Non-goals (this pass) - Don't wire up more than one new provider before the first one round-trips end to end. - Don't try to make eco/carbon methodology uniform across providers that don't expose the same telemetry — graceful absence beats a fabricated number (same principle as the empty `leaderboards.yaml`). - Don't rebuild the cost model — catalog-price × shape already generalizes; confirm that rather than redesigning it. ## Not yet defined Success criteria, test plan, and phasing/milestones — fill in once the open questions above are resolved and this graduates to FINAL.