plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
15 KiB
Generalizing to multiple providers
Status: parked -- parked on provider selection; the coupling surface in it is still valid
Status: PARKED 2026-09-04 — DRAFT, not decision-complete, and now blocked on
provider selection rather than on design. Written 2026-09-04 against main
at 340453e; coupling surface measured and three open questions closed by code
inspection on the same commit.
Parked by the user: Z.ai is no longer a settled first target. The GLM overlap that made it attractive is a real argument, but fit is unconfirmed and the candidate search is still open. Nothing below should be read as a commitment to a specific provider.
Remaining gaps before this can go to the opencode/Prometheus pipeline:
- Which provider goes first (was treated as answered; it is not).
- The proficiency-sharing decision — a judgement call, not a lookup.
- The eco-without-telemetry decision.
- Success criteria / test plan / phasing, still empty.
What does NOT need to wait. The measured coupling surface below is
provider-agnostic: the 20 hardcoded neuralwatt references are literals that
are wrong regardless of which provider lands second, and three of them are
latent defects today. Re-verified on 2026-09-04 against
feat/config-local-overlay — still 20 references, same six-file distribution.
Those can be parameterised independently of this plan and should not be held
hostage to the provider question.
Why this, why now
Everything downstream of the catalog — scoring.py, tiering.py,
routing.py, proficiency.py, verification.py, feedback.py, the admin
portal, the TUI — already operates on normalized DB rows with no idea which
provider produced them. dispatch_providers: dict[str, DispatchProvider] in
config and the (model_id, provider) composite key were deliberately kept
when OpenRouter was dropped (c3484f0) for exactly this. The coupling that's
actually left is narrow:
poller.py—fetch_neuralwatt()is a bespoke function: hardcoded URL, NeuralWatt's specific catalog JSON shape,provider="neuralwatt"written as a literal.dispatcher.py— three things tangled together that need separating:- Usage/telemetry parsing (NeuralWatt reports energy via SSE comment lines — a protocol quirk, not an OpenAI standard).
- Cost semantics (the $8/kWh-billed-capped-at-3x-list model is
NeuralWatt-specific billing behavior). What
routing.estimated_costactually scores on — catalog price × request shape — is already provider-agnostic; the kWh math is validation for one provider's billing quirk, not the load-bearing input. - Account-refusal handling (
degrade to localassumes one cloud account — see dispatcher.py ~2537). With N providers this mostly disappears: circuit-break the failing provider's rows and let ranking fail over to the next-best candidate on a different provider.
On re-adding OpenRouter specifically: it was dropped for a real reason,
not a bad one — the project pivoted to scoring on measured billing and
carbon, and OpenRouter (an aggregator) doesn't expose per-request
energy/carbon telemetry the way NeuralWatt does. That's not a reason to
avoid it now — it's the first real test case for a has_energy_telemetry
capability flag, since eco/energy needs to degrade gracefully per-provider
rather than assuming every row has NeuralWatt's shape. Worth remembering
before re-proposing it as if it were untried.
Constraint: cheap, no-commitment testing
Budget is small and this is for testing the plumbing, not production spend — prepaid credits in ~$10 increments, no subscriptions. Candidates, pricing and minimums need re-verification before committing to one (this kind of thing moves):
| provider | fit | notes |
|---|---|---|
| Z.ai | best first target | Pay-as-you-go API, OpenAI-compatible (https://api.z.ai/api/paas/v4 or /api/openai/v1), no stated minimum top-up. Three free-tier models (GLM-4.7-Flash, GLM-4.5-Flash, GLM-4.6V-Flash) for zero-cost plumbing tests. Actually the publisher of the GLM family this catalog already serves via NeuralWatt (glm-5.2-fast, glm-5.2-flex) — a same-weights cross-provider comparison, which is a stronger generalization test than an unrelated model set, and a way to sanity-check NeuralWatt's static_fallback carbon figure for GLM. That is not hypothetical: CLAUDE.md records the GLM rows reporting grid_id: FI at 475 gCO2/kWh as carbon_source: static_fallback — a substituted constant, not a measurement — and they are currently EXCLUDED from eco scoring for that reason. A second provider serving the same weights is a free experiment on one of this project's standing open questions. No response-level energy/carbon telemetry — same capability-flag case as OpenRouter, not a differentiator there. Don't confuse this with the separate "GLM Coding Plan" subscription ($18-168/mo) — that's a different product and not what fits the budget constraint here. |
| OpenRouter | good second | prepaid, no minimum, OpenAI-compatible, many free models for zero-cost plumbing tests, broad catalog overlap (kimi/deepseek/qwen/gemma too, not just GLM), and this repo already has git history (c3484f0 and its parent) to mine for catalog-normalization shape. No per-request energy/carbon. It's an aggregator of aggregators, so if it ever does report grid/energy data, treat it as less trustworthy than a direct provider's own figure. |
| DeepInfra | plausible | prepaid, no subscription, OpenAI-compatible, cheap open-weight catalog. No energy telemetry. |
| Together AI | plausible | prepaid credits, OpenAI-compatible, per-token billing (straightforward vs. NeuralWatt's kWh math). |
| Fireworks AI | plausible | same shape as Together. |
| Groq | maybe later | OpenAI-compatible, free tier + pay-as-you-go, but a much smaller catalog — more useful for a latency-tolerance test than a routing-breadth test. |
Recommendation, downgraded 2026-09-04. The shape still holds: one provider first, all the way through poller → dispatch → a handful of real routed requests, before touching a second. That is what confirms the abstraction generalizes instead of just looking like it does on paper.
Which provider is reopened. Z.ai was the pick on the strength of the GLM
overlap — a same-weights cross-provider comparison, and a free experiment on
the static_fallback carbon question. That argument is still good and should
be re-used to score whatever candidate comes next; it was never a claim that
Z.ai's API, pricing or terms actually fit. Treat the table above as a
shortlist to re-verify, not a ranking to execute.
Worth noting for the search: the criterion that made Z.ai attractive —
catalog overlap with what NeuralWatt already serves — is separable from
Z.ai itself. OpenRouter carries kimi, deepseek, qwen and gemma alongside GLM,
so it satisfies the same-weights test more broadly, and this repo already has
git history (c3484f0 and its parent) with a working catalog normalizer to
mine. It was dropped for a reason that no longer disqualifies it, now that
has_energy_telemetry is the planned answer to missing energy data.
Architecture sketch
A Provider protocol/interface — not fleshed out yet, but the shape:
fetch_catalog() -> list[ModelRow]parse_usage(response) -> UsageInfo(cost, tokens, energy if present)detect_refusal(error) -> bool- capability flags:
has_energy_telemetry,has_regional_carbon
fetch_neuralwatt and the SSE-comment parser move behind it as the first
concrete implementation. dispatch_providers entries get a type:
discriminator so config knows which implementation to instantiate.
Measured coupling surface
Claims about how narrow the coupling is should be counted, not asserted.
grep -rn neuralwatt src/*.py on main at 340453e returns 20 hardcoded
references across 6 files:
| file | refs | what they are |
|---|---|---|
poller.py |
8 | the bespoke fetch: URL, fetch_neuralwatt(), literal provider="neuralwatt", the sanity-floor count query, log prefixes |
eval_proficiency.py |
5 | provider: str = "neuralwatt" defaults, the judge's dispatch_providers["neuralwatt"] lookup, and an identity["provider"] != "neuralwatt" skip |
dispatcher.py |
4 | passthrough default, and a capability lookup (see below) |
metrics.py |
1 | catalog-staleness query |
leaderboard.py |
1 | set_leaderboard(..., "neuralwatt", ...) |
seed_energy.py |
1 | dispatch_providers["neuralwatt"] |
That is the actual work list. Two entries deserve calling out because they contradict the "needs ~zero change" section below.
metrics.py:175 — catalog staleness would ignore a second provider
SELECT MAX(last_updated) AS last_updated FROM models WHERE provider='neuralwatt'
The staleness warning is computed over NeuralWatt rows only. Add a provider
whose poller silently stops and the warning stays green while its catalog
freezes — the exact silent-and-open failure CLAUDE.md's "Run as a service"
section describes, now with a second way in. metrics.py is listed under
"needs ~zero change"; it does not.
leaderboard.py:125 — priors are written to a hardcoded provider
set_leaderboard(conn, cfg, model_id, "neuralwatt", category, score) pins the
provider literal, so a curated prior for a shared model (GLM, say) attaches to
the NeuralWatt row and never to the Z.ai one. leaderboards.yaml ships empty
today so nothing is broken yet, but this decides the answer to the
proficiency-sharing question above by accident rather than on purpose.
dispatcher.py:2879 — _model_exists silently strips a real model id
Corrected 2026-09-04. An earlier revision of this plan attributed this line
to _check_pinned_capabilities. That was wrong: _check_pinned_capabilities
(dispatcher.py:2887) already takes provider as a parameter and passes it to
its query. It is not a coupling site. Line 2879 belongs to _model_exists,
and the defect there is worse.
SELECT 1 FROM models WHERE model_id = ? AND provider = 'neuralwatt'
_model_exists decides whether a provider/model string a client sent is an
opencode-style alias to strip (llm-router/...) or a real id that merely
contains a slash. Scoped to NeuralWatt, a real second-provider id returns
False and gets treated as an alias.
That matters more than it looks, because vendor/model is the native id
format for the two strongest candidates: OpenRouter ids are always
deepseek/deepseek-chat-shaped, and several other aggregators follow suit. So
the failure is not exotic — it is the default case for a whole class of
provider.
And it fails silently in the wrong direction. A provider-scoped 422 would
at least be loud. This one strips the prefix and proceeds with the remainder,
so a client pinning deepseek/deepseek-chat gets whatever deepseek-chat
resolves to — a different row, on a different provider, at a different price —
with nothing in the response indicating a substitution occurred. Fail-closed
was the design intent everywhere else in this capability path; here it fails
open.
Requirement: _model_exists must resolve across all configured providers, and
alias-stripping must be decided by something other than "no NeuralWatt row has
this id".
Risk not in the draft: per-provider fetch isolation
poller.mark_stale runs only inside poller.main(), and main() returns
early on RequestException — before upsert and before mark_stale.
With two providers in one run, a failure fetching provider B can skip
mark_stale for provider A entirely, or a partial/empty data array from B
can mark rows stale that were never B's. docs/incidents.md records the
empty-data path as the one that can empty the candidate set.
Requirement: each provider's fetch, upsert and staleness marking must be
isolated. One provider's outage must not affect another's rows in either
direction, and the sanity floor (poller.py:420) must be per-provider.
What should need ~zero change
Scoring, tiering, routing, proficiency, verification, feedback, admin, TUI. If any of these turn out to need provider-aware branching, that's a sign the interface boundary is in the wrong place — worth treating as a red flag during implementation, not a shrug.
Two of them already fail that test, per the inventory above: metrics.py
hardcodes the provider in its staleness query and leaderboard.py hardcodes it
when writing priors. Neither is a deep coupling — both are literals, not
branching — but the list should be read as "should need ~zero change after
those two literals are parameterised", not as a claim that they are already
clean. The red flag to watch for during implementation is provider-aware
logic appearing in these modules; a hardcoded string is a different and much
cheaper problem.
Open questions
- Eco/carbon for a provider with no telemetry. Exclude those rows from eco ranking entirely, or treat eco as unweighted/missing for them without disqualifying them? Affects whether a non-NeuralWatt row can ever win on the eco axis, or only ever competes on cost/quality.
- Same model, two providers — per-model or per-(model, provider)?
PARTLY RESOLVED, and the remaining half is a design decision rather than a
verification. The schema already supports divergence:
proficiency's primary key is(model_id, provider, category). Butpropagate_to_variantscopies scores across variants keyed onbase_model_id, on the stated principle that proficiency is "a property of the weights, not the queue". Same-weights-different-serving-stack is exactly the case that principle does not decide: quantization and serving differences could justify separate scores, while the existing logic would share one. Decide this explicitly before implementation — it is the one question here that code inspection cannot answer. - RESOLVED — circuit breaker is genuinely provider-generic. Verified in
code, not prose:
is_down,record_failureandrecord_successall take(model_id, provider)and_storeis keyed on that tuple (src/circuit_breaker.py:29,37,57). Per-provider breaking works today with no change. - Does the account-refusal-degrades-to-local path actually get simpler or
just get an
ifadded? Stated above as an assumption; worth confirming against the real code before it's a success criterion. - RESOLVED — "no serving class" is already a real state.
poller.parse_serving_classreturns the schema defaults ('standard','default','full') for a base id with no suffix, so a Z.ai or OpenRouter row landsstandardand is interactive-eligible rather than null. No change needed.
Non-goals (this pass)
- Don't wire up more than one new provider before the first one round-trips end to end.
- Don't try to make eco/carbon methodology uniform across providers that
don't expose the same telemetry — graceful absence beats a fabricated
number (same principle as the empty
leaderboards.yaml). - Don't rebuild the cost model — catalog-price × shape already generalizes; confirm that rather than redesigning it.
Not yet defined
Success criteria, test plan, and phasing/milestones — fill in once the open questions above are resolved and this graduates to FINAL.