classify() has one endpoint and one failure behaviour: return
fallback_category/fallback_tier, currently general_chat/tier 2. That is the
worst possible guess, because proficiency_score is the ONLY category-dependent
term in ranking -- a request mislabelled general_chat loses the single signal
that makes routing category-aware.
The session cache cannot help because of WHERE it sits: session_cache.get() is
called at dispatcher.py:3077, before classifying, while classify() resolves its
own failure internally and never sees the session key. A session with 200 prior
coding_refactor turns still degrades to general_chat the moment the local model
goes away.
Replaces the static fallback with an ordered cascade, cheapest first: stale
session reuse (free, offline, covers a running session whose model vanished
mid-conversation), then durable route_decisions history (survives a router
restart, which empties the in-memory cache), then an optional cloud classifier,
then the existing static fallback. Each step records a DISTINCT
classification_source so an operator can tell "sessions are being reused
correctly" from "everything failed, we are guessing" -- today both collapse
into one bucket.
Cloud step is opt-in and must be skipped when the account is out of credits;
Plan 5's _account_level_refusal already knows. Its timeout must be short, since
the request has already spent the local attempt's budget.
Note on priority: source='fallback' appears ZERO times in 16,744 recorded
decisions, so this path is real but cold -- which is why it deserves hardening
rather than trust. It is also about to get warmer: the operator stops Ollama to
play games on the same GPU, which is a deliberate recurring whole-session
outage of the local classifier.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
A hand-corrupted config.yaml (duplicate profiles: keys, invalid or
non-mapping profile entries) crashed GET /admin/api/profiles, the
profile CRUD writes, and GET /admin/api/config with raw 500s: the
store was parsed by load_config_store with nobody catching
ruamel YAMLError, and stored profile entries fed RoutingProfile(**entry)
uncaught.
Add load_config_store_safe, which maps YAMLError to a 503 whose detail
names the parse error (ruamel includes line/column) and points at the
config.yaml.bak.* backups, and use it in _persist_config_block (the
first statement under the write lock, so an unparseable store aborts
before any mutation, backup, or byte touches disk), _persisted_profiles,
the delete guard's store read, and admin_config_get. In the GET list,
a stored profile entry that fails RoutingProfile validation now 503s
naming profiles[<name>] instead of the file; non-mapping entries hit the
same guard via TypeError.
Error paths only: success shapes, write discipline, and the existing
200/403/404/409/422 codes are unchanged. Existing 422/404 mappings on
writes are preserved.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
CHEAP wins under both interactive and batch tolerance in the router
fixture, so the X-Router-Model header alone cannot tell whether bare
auto resolved through the configured batch-local profile. The persisted
route_decisions row can: assert its last entry carries
latency_tolerance='batch' and profile='batch-local' after a bare auto
request, alongside the existing response assertions. Verified
discriminating by mutation: flipping the configured default to
'default' fails the test on the latency_tolerance assertion.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Update the auto/auto:batch equivalence test to stop monkeypatching
cfg.profiles["default"] (a builtin name since Wave A): it now sets
routing.default_profile to a config profile batch-local, and keeps the
same flex-twin assertions for bare auto and auto:batch.
Add chat-completions coverage that bare auto routes through the
configured default profile and that auto:<name> still overrides it.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Cover the profile CRUD endpoints end to end against a temp config copy:
round-trip create/update/delete with no dangling profiles: block,
read-only builtins (403), delete/rename guards naming
routing.default_profile (422), zero-admission save warning, null
allowed_model_ids round-trip (never []), 409 on create-existing,
404 on update/delete of unknown profiles, 422 on an out-of-range tier.
Pin the write path through the CRUD endpoint: comment preservation
(sentinel + local_energy block), backup-before-write with pre-write
contents, 403 for the non-allowlisted profiles.foo config path, and
_persist_config_block concurrency (two threads, no zero-byte backups).
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Generalize _persist_config_value into _persist_config_block: nested
writes create intermediate CommentedMap blocks, delete mode prunes the
leaf and an empty parent, and the whole-config RouterConfig validation,
backup-first, .tmp+os.replace atomic swap discipline is preserved
unchanged under the same _config_write_lock. _persist_config_value
remains as a thin wrapper for scalar paths.
Add POST /admin/api/profiles/ (create, 403 on builtin, 409 on existing),
POST /admin/api/profiles/{name} (full replace, 404 unknown, 403 builtin;
the current routing.default_profile may be edited), and DELETE
/admin/api/profiles/{name} (404 unknown, 422 naming routing.default_profile
when it targets the profile, no dangling empty profiles: block).
Both save endpoints probe the CANDIDATE profile through the shared
_profile_probe (extracted from GET) before writing and return
zero_admit plus a warning when the profile admits zero models under
the current catalog; every response notes a restart is required.
GET /admin/api/profiles now enumerates config profiles from the
persisted store so CRUD writes are visible without a reload.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Render routing.default_profile as a live-profile select on Controls,
flag dirty selections in the existing Save flow, and show an amber
'admits 0 models' hint when the selected profile admits nothing.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The current warning says "metered usage is 146% of the 6.25 kWh plan
allowance. A quota is a wall, not a bill -- requests fail rather than costing
more." That is false for this account: usage sailed past 100% and nothing
failed, because the provider bills overage against a credit balance. The
warning is alarming and unactionable at once.
Meanwhile the authoritative signal is already captured and ignored. The
provider reports allowance_remaining_usd on EVERY response; it is persisted to
energy_observations and read by exactly one thing -- seed_energy.py, for sweep
accounting.
The recorded data contains the 2026-09-02 outage in full: the balance walked
down through $0.4991 -> $0.4967 -> $0.4924 at 05:54 and bottomed at $0.0071.
The router watched the balance approach zero request by request and said
nothing. Current state, for scale: $12.19 left, $7.51 burned over 1,853 calls
in 14.3h, about 23 hours of runway. That is the sentence the dashboard should
show.
Three defects: the wrong signal (percent-of-plan instead of balance and
projected time-to-zero); the wrong window (30-day rolling while the
subscription resets monthly on billing_reset_day, and a `reset_date` field
that actually returns the rolling-window start); and the wrong words, since
the wall framing is in config.yaml and CLAUDE.md too and changes what an
operator does about it.
The implementation trap is called out explicitly: a top-up makes the balance
JUMP UP (0.0071 -> 20.0071 on 09-02), so a naive MAX-MIN reports a ~$20 burn
that never happened. Burn must come from consecutive decreasing deltas only,
with a test covering a window containing a top-up.
Explicit non-goal: do NOT gate requests on quota. plan_kwh_per_period gates
nothing today and must continue to gate nothing -- refusing traffic on a local
estimate would turn a billing question into an outage.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
Promoting C to in-scope left the old 'ship A and B first, defer C'
recommendation in place, which contradicted the decision two paragraphs above.
Rewritten as sequencing rationale: A's read-only panel is where a created
profile gets verified, and B's default_profile is what C's delete-guard has to
protect, so building C first would mean writing CRUD against a surface nobody
can see.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
Part C promoted from deferred to in-scope, with the edge cases nailed down so
the executor is not guessing: built-ins read-only and name collisions rejected
at load; allowed_model_ids as a catalog-populated multi-select rather than free
text; deleting the profile named by routing.default_profile refused, since that
turns Part B's refuse-at-load into an outage; zero-model profiles allowed but
warned on save.
Sequencing note kept: A and B first, C as its own wave, since C is most of the
work in the plan.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
Named routing profiles shipped in PR #24 with no portal surface: they exist
only as code defaults and an auto:<name> string a client must know to send.
Per the user's principle -- the portal need not cover every lever, but main
features should each have basic functionality and configurability there --
that fails the bar.
Part A (read-only) is the highest value per unit of work, because a profile's
DEFINITION is not the useful fact -- what it currently ADMITS is, and the two
differ invisibly. Measured live: bigboybritches (min_tier=3) matches 7 models
but only 4 are reachable interactively, since three are -flex rows the default
latency_tolerance filters out. locality matches 1, and only for two categories,
because per-model eligible_categories ANDs on top. None of that is derivable
from reading the predicate. Admitted sets must be computed by calling
routing.select_candidates, not by reimplementing it -- metrics.context_ceilings
is the precedent.
Part B adds routing.default_profile on the existing _CONFIG_ALLOWLIST, with
routing.default_flex_preference as the exact precedent. This is what makes
profiles usable without touching client config: bare `auto`, which is what
opencode.json already sends, resolves to the operator's choice.
Part C (create/edit custom profiles) is flagged as a scope decision. The
allowlist edits named SCALARS; a profile is a structured object with six
optional fields, so persisting one is most of the work in the plan.
Recommendation is to defer it and design it once operators know which fields
they actually reach for.
Also notes for the implementer that this plan writes config.yaml
programmatically, so preserving unrelated lines -- the user's tariff included
-- needs an explicit test rather than an assumption.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U