plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
261 lines
15 KiB
Markdown
261 lines
15 KiB
Markdown
# Generalizing to multiple providers
|
||
|
||
Status: parked -- parked on provider selection; the coupling surface in it is still valid
|
||
|
||
**Status: PARKED 2026-09-04 — DRAFT, not decision-complete, and now blocked on
|
||
provider selection rather than on design.** Written 2026-09-04 against `main`
|
||
at `340453e`; coupling surface measured and three open questions closed by code
|
||
inspection on the same commit.
|
||
|
||
Parked by the user: **Z.ai is no longer a settled first target.** The GLM
|
||
overlap that made it attractive is a real argument, but fit is unconfirmed and
|
||
the candidate search is still open. Nothing below should be read as a
|
||
commitment to a specific provider.
|
||
|
||
Remaining gaps before this can go to the opencode/Prometheus pipeline:
|
||
|
||
1. **Which provider goes first** (was treated as answered; it is not).
|
||
2. The proficiency-sharing decision — a judgement call, not a lookup.
|
||
3. The eco-without-telemetry decision.
|
||
4. Success criteria / test plan / phasing, still empty.
|
||
|
||
**What does NOT need to wait.** The measured coupling surface below is
|
||
provider-agnostic: the 20 hardcoded `neuralwatt` references are literals that
|
||
are wrong regardless of which provider lands second, and three of them are
|
||
latent defects today. Re-verified on 2026-09-04 against
|
||
`feat/config-local-overlay` — still 20 references, same six-file distribution.
|
||
Those can be parameterised independently of this plan and should not be held
|
||
hostage to the provider question.
|
||
|
||
## Why this, why now
|
||
|
||
Everything downstream of the catalog — `scoring.py`, `tiering.py`,
|
||
`routing.py`, `proficiency.py`, `verification.py`, `feedback.py`, the admin
|
||
portal, the TUI — already operates on normalized DB rows with no idea which
|
||
provider produced them. `dispatch_providers: dict[str, DispatchProvider]` in
|
||
config and the `(model_id, provider)` composite key were deliberately kept
|
||
when OpenRouter was dropped (`c3484f0`) for exactly this. The coupling that's
|
||
actually left is narrow:
|
||
|
||
- **`poller.py`** — `fetch_neuralwatt()` is a bespoke function: hardcoded
|
||
URL, NeuralWatt's specific catalog JSON shape, `provider="neuralwatt"`
|
||
written as a literal.
|
||
- **`dispatcher.py`** — three things tangled together that need separating:
|
||
1. Usage/telemetry parsing (NeuralWatt reports energy via SSE **comment**
|
||
lines — a protocol quirk, not an OpenAI standard).
|
||
2. Cost semantics (the $8/kWh-billed-capped-at-3x-list model is
|
||
NeuralWatt-specific billing behavior). What `routing.estimated_cost`
|
||
actually scores on — catalog price × request shape — is already
|
||
provider-agnostic; the kWh math is validation for one provider's
|
||
billing quirk, not the load-bearing input.
|
||
3. Account-refusal handling (`degrade to local` assumes one cloud account
|
||
— see dispatcher.py ~2537). With N providers this mostly *disappears*:
|
||
circuit-break the failing provider's rows and let ranking fail over to
|
||
the next-best candidate on a different provider.
|
||
|
||
**On re-adding OpenRouter specifically:** it was dropped for a real reason,
|
||
not a bad one — the project pivoted to scoring on *measured* billing and
|
||
carbon, and OpenRouter (an aggregator) doesn't expose per-request
|
||
energy/carbon telemetry the way NeuralWatt does. That's not a reason to
|
||
avoid it now — it's the first real test case for a `has_energy_telemetry`
|
||
capability flag, since eco/energy needs to degrade gracefully per-provider
|
||
rather than assuming every row has NeuralWatt's shape. Worth remembering
|
||
before re-proposing it as if it were untried.
|
||
|
||
## Constraint: cheap, no-commitment testing
|
||
|
||
Budget is small and this is for testing the plumbing, not production spend —
|
||
prepaid credits in ~$10 increments, no subscriptions. Candidates, **pricing
|
||
and minimums need re-verification before committing to one** (this kind of
|
||
thing moves):
|
||
|
||
| provider | fit | notes |
|
||
|---|---|---|
|
||
| Z.ai | best first target | Pay-as-you-go API, OpenAI-compatible (`https://api.z.ai/api/paas/v4` or `/api/openai/v1`), no stated minimum top-up. Three free-tier models (GLM-4.7-Flash, GLM-4.5-Flash, GLM-4.6V-Flash) for zero-cost plumbing tests. **Actually the publisher of the GLM family this catalog already serves via NeuralWatt** (`glm-5.2-fast`, `glm-5.2-flex`) — a same-weights cross-provider comparison, which is a stronger generalization test than an unrelated model set, and a way to sanity-check NeuralWatt's `static_fallback` carbon figure for GLM. That is not hypothetical: `CLAUDE.md` records the GLM rows reporting `grid_id: FI` at 475 gCO2/kWh as `carbon_source: static_fallback` — a substituted constant, not a measurement — and they are currently EXCLUDED from eco scoring for that reason. A second provider serving the same weights is a free experiment on one of this project's standing open questions. No response-level energy/carbon telemetry — same capability-flag case as OpenRouter, not a differentiator there. Don't confuse this with the separate "GLM Coding Plan" subscription ($18-168/mo) — that's a different product and not what fits the budget constraint here. |
|
||
| OpenRouter | good second | prepaid, no minimum, OpenAI-compatible, many free models for zero-cost plumbing tests, broad catalog overlap (kimi/deepseek/qwen/gemma too, not just GLM), and this repo already has git history (`c3484f0` and its parent) to mine for catalog-normalization shape. No per-request energy/carbon. It's an aggregator of aggregators, so if it ever *does* report grid/energy data, treat it as less trustworthy than a direct provider's own figure. |
|
||
| DeepInfra | plausible | prepaid, no subscription, OpenAI-compatible, cheap open-weight catalog. No energy telemetry. |
|
||
| Together AI | plausible | prepaid credits, OpenAI-compatible, per-token billing (straightforward vs. NeuralWatt's kWh math). |
|
||
| Fireworks AI | plausible | same shape as Together. |
|
||
| Groq | maybe later | OpenAI-compatible, free tier + pay-as-you-go, but a much smaller catalog — more useful for a latency-tolerance test than a routing-breadth test. |
|
||
|
||
**Recommendation, downgraded 2026-09-04.** The *shape* still holds: **one
|
||
provider first**, all the way through poller → dispatch → a handful of real
|
||
routed requests, before touching a second. That is what confirms the
|
||
abstraction generalizes instead of just looking like it does on paper.
|
||
|
||
Which provider is **reopened**. Z.ai was the pick on the strength of the GLM
|
||
overlap — a same-weights cross-provider comparison, and a free experiment on
|
||
the `static_fallback` carbon question. That argument is still good and should
|
||
be re-used to score whatever candidate comes next; it was never a claim that
|
||
Z.ai's API, pricing or terms actually fit. Treat the table above as a
|
||
shortlist to re-verify, not a ranking to execute.
|
||
|
||
Worth noting for the search: the criterion that made Z.ai attractive —
|
||
**catalog overlap with what NeuralWatt already serves** — is separable from
|
||
Z.ai itself. OpenRouter carries kimi, deepseek, qwen and gemma alongside GLM,
|
||
so it satisfies the same-weights test more broadly, and this repo already has
|
||
git history (`c3484f0` and its parent) with a working catalog normalizer to
|
||
mine. It was dropped for a reason that no longer disqualifies it, now that
|
||
`has_energy_telemetry` is the planned answer to missing energy data.
|
||
|
||
## Architecture sketch
|
||
|
||
A `Provider` protocol/interface — not fleshed out yet, but the shape:
|
||
|
||
- `fetch_catalog() -> list[ModelRow]`
|
||
- `parse_usage(response) -> UsageInfo` (cost, tokens, energy if present)
|
||
- `detect_refusal(error) -> bool`
|
||
- capability flags: `has_energy_telemetry`, `has_regional_carbon`
|
||
|
||
`fetch_neuralwatt` and the SSE-comment parser move behind it as the first
|
||
concrete implementation. `dispatch_providers` entries get a `type:`
|
||
discriminator so config knows which implementation to instantiate.
|
||
|
||
## Measured coupling surface
|
||
|
||
Claims about how narrow the coupling is should be counted, not asserted.
|
||
`grep -rn neuralwatt src/*.py` on `main` at `340453e` returns **20 hardcoded
|
||
references across 6 files**:
|
||
|
||
| file | refs | what they are |
|
||
|---|---|---|
|
||
| `poller.py` | 8 | the bespoke fetch: URL, `fetch_neuralwatt()`, literal `provider="neuralwatt"`, the sanity-floor count query, log prefixes |
|
||
| `eval_proficiency.py` | 5 | `provider: str = "neuralwatt"` defaults, the judge's `dispatch_providers["neuralwatt"]` lookup, and an `identity["provider"] != "neuralwatt"` skip |
|
||
| `dispatcher.py` | 4 | passthrough default, and a capability lookup (see below) |
|
||
| `metrics.py` | 1 | catalog-staleness query |
|
||
| `leaderboard.py` | 1 | `set_leaderboard(..., "neuralwatt", ...)` |
|
||
| `seed_energy.py` | 1 | `dispatch_providers["neuralwatt"]` |
|
||
|
||
That is the actual work list. Two entries deserve calling out because they
|
||
**contradict the "needs ~zero change" section below.**
|
||
|
||
### `metrics.py:175` — catalog staleness would ignore a second provider
|
||
|
||
```sql
|
||
SELECT MAX(last_updated) AS last_updated FROM models WHERE provider='neuralwatt'
|
||
```
|
||
|
||
The staleness warning is computed over NeuralWatt rows only. Add a provider
|
||
whose poller silently stops and the warning stays green while its catalog
|
||
freezes — the exact silent-and-open failure `CLAUDE.md`'s "Run as a service"
|
||
section describes, now with a second way in. `metrics.py` is listed under
|
||
"needs ~zero change"; it does not.
|
||
|
||
### `leaderboard.py:125` — priors are written to a hardcoded provider
|
||
|
||
`set_leaderboard(conn, cfg, model_id, "neuralwatt", category, score)` pins the
|
||
provider literal, so a curated prior for a shared model (GLM, say) attaches to
|
||
the NeuralWatt row and never to the Z.ai one. `leaderboards.yaml` ships empty
|
||
today so nothing is broken yet, but this decides the answer to the
|
||
proficiency-sharing question above by accident rather than on purpose.
|
||
|
||
### `dispatcher.py:2879` — `_model_exists` silently strips a real model id
|
||
|
||
**Corrected 2026-09-04.** An earlier revision of this plan attributed this line
|
||
to `_check_pinned_capabilities`. That was wrong: `_check_pinned_capabilities`
|
||
(`dispatcher.py:2887`) already takes `provider` as a parameter and passes it to
|
||
its query. It is not a coupling site. Line 2879 belongs to `_model_exists`,
|
||
and the defect there is worse.
|
||
|
||
```sql
|
||
SELECT 1 FROM models WHERE model_id = ? AND provider = 'neuralwatt'
|
||
```
|
||
|
||
`_model_exists` decides whether a `provider/model` string a client sent is an
|
||
opencode-style alias to **strip** (`llm-router/...`) or a real id that merely
|
||
contains a slash. Scoped to NeuralWatt, a real second-provider id returns
|
||
False and gets treated as an alias.
|
||
|
||
That matters more than it looks, because **`vendor/model` is the native id
|
||
format for the two strongest candidates**: OpenRouter ids are always
|
||
`deepseek/deepseek-chat`-shaped, and several other aggregators follow suit. So
|
||
the failure is not exotic — it is the default case for a whole class of
|
||
provider.
|
||
|
||
And it fails **silently in the wrong direction**. A provider-scoped 422 would
|
||
at least be loud. This one strips the prefix and proceeds with the remainder,
|
||
so a client pinning `deepseek/deepseek-chat` gets whatever `deepseek-chat`
|
||
resolves to — a different row, on a different provider, at a different price —
|
||
with nothing in the response indicating a substitution occurred. Fail-closed
|
||
was the design intent everywhere else in this capability path; here it fails
|
||
open.
|
||
|
||
Requirement: `_model_exists` must resolve across all configured providers, and
|
||
alias-stripping must be decided by something other than "no NeuralWatt row has
|
||
this id".
|
||
|
||
### Risk not in the draft: per-provider fetch isolation
|
||
|
||
`poller.mark_stale` runs only inside `poller.main()`, and `main()` returns
|
||
early on `RequestException` — **before** `upsert` and **before** `mark_stale`.
|
||
With two providers in one run, a failure fetching provider B can skip
|
||
`mark_stale` for provider A entirely, or a partial/empty `data` array from B
|
||
can mark rows stale that were never B's. `docs/incidents.md` records the
|
||
empty-`data` path as the one that can empty the candidate set.
|
||
|
||
Requirement: **each provider's fetch, upsert and staleness marking must be
|
||
isolated.** One provider's outage must not affect another's rows in either
|
||
direction, and the sanity floor (`poller.py:420`) must be per-provider.
|
||
|
||
## What should need ~zero change
|
||
|
||
Scoring, tiering, routing, proficiency, verification, feedback, admin, TUI.
|
||
If any of these turn out to need provider-aware branching, that's a sign the
|
||
interface boundary is in the wrong place — worth treating as a red flag
|
||
during implementation, not a shrug.
|
||
|
||
**Two of them already fail that test**, per the inventory above: `metrics.py`
|
||
hardcodes the provider in its staleness query and `leaderboard.py` hardcodes it
|
||
when writing priors. Neither is a deep coupling — both are literals, not
|
||
branching — but the list should be read as "should need ~zero change *after*
|
||
those two literals are parameterised", not as a claim that they are already
|
||
clean. The red flag to watch for during implementation is provider-aware
|
||
*logic* appearing in these modules; a hardcoded string is a different and much
|
||
cheaper problem.
|
||
|
||
## Open questions
|
||
|
||
- **Eco/carbon for a provider with no telemetry.** Exclude those rows from
|
||
eco ranking entirely, or treat eco as unweighted/missing for them without
|
||
disqualifying them? Affects whether a non-NeuralWatt row can ever win on
|
||
the eco axis, or only ever competes on cost/quality.
|
||
- **Same model, two providers — per-model or per-(model, provider)?**
|
||
PARTLY RESOLVED, and the remaining half is a design decision rather than a
|
||
verification. The schema already supports divergence: `proficiency`'s
|
||
primary key is `(model_id, provider, category)`. But
|
||
`propagate_to_variants` copies scores across variants keyed on
|
||
`base_model_id`, on the stated principle that proficiency is "a property of
|
||
the weights, not the queue". Same-weights-different-serving-stack is exactly
|
||
the case that principle does not decide: quantization and serving
|
||
differences could justify separate scores, while the existing logic would
|
||
share one. **Decide this explicitly before implementation** — it is the one
|
||
question here that code inspection cannot answer.
|
||
- **RESOLVED — circuit breaker is genuinely provider-generic.** Verified in
|
||
code, not prose: `is_down`, `record_failure` and `record_success` all take
|
||
`(model_id, provider)` and `_store` is keyed on that tuple
|
||
(`src/circuit_breaker.py:29,37,57`). Per-provider breaking works today with
|
||
no change.
|
||
- **Does the account-refusal-degrades-to-local path actually get simpler or
|
||
just get an `if` added?** Stated above as an assumption; worth confirming
|
||
against the real code before it's a success criterion.
|
||
- **RESOLVED — "no serving class" is already a real state.**
|
||
`poller.parse_serving_class` returns the schema defaults
|
||
(`'standard'`, `'default'`, `'full'`) for a base id with no suffix, so a
|
||
Z.ai or OpenRouter row lands `standard` and is interactive-eligible rather
|
||
than null. No change needed.
|
||
|
||
## Non-goals (this pass)
|
||
|
||
- Don't wire up more than one new provider before the first one round-trips
|
||
end to end.
|
||
- Don't try to make eco/carbon methodology uniform across providers that
|
||
don't expose the same telemetry — graceful absence beats a fabricated
|
||
number (same principle as the empty `leaderboards.yaml`).
|
||
- Don't rebuild the cost model — catalog-price × shape already generalizes;
|
||
confirm that rather than redesigning it.
|
||
|
||
## Not yet defined
|
||
|
||
Success criteria, test plan, and phasing/milestones — fill in once the open
|
||
questions above are resolved and this graduates to FINAL.
|