Update admin-portal.md Controls section with categories/accordion/coverage sources. Update CLIREF.md North Star 1 for completed classifier coverage. Add plans/deferred-knobs.md listing 39 outside-gate + 31 warning knobs. Ultraworked with Sisyphus Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
30 KiB
Deep dive into the admin web portal. Back: README.
Admin web portal
GET /admin/ serves a management portal from the running router on the same
loopback-only bind as the API. It has no auth layer yet, so like the other
endpoints it is reachable only from 127.0.0.1.
The portal is a six-page, glass dark-mode dashboard built on Tabler/Bootstrap
with a few lighter inline SVG icons (see plans/admin-design-standards.md
for the visual language and plans/admin-work-framework.md for the method
used to evolve it). The six pages are Dashboard, Models, Profiles, Decisions,
Controls, and Providers. It is implemented in admin.py, with
config/admin_schema.sql for its future tables and admin/frontend/ for the
browser UI.
Dashboard — GET /admin/ aggregates the router at a glance: a quota chip
(burn against objective.plan_kwh_per_period), per-model usage bars, verdict
mix and category breakdown, five history mini-charts, and a recent-decisions
table. Clicking the chip opens a detail modal with per-provider balances and
runway; each line is labeled by its billing shape, so "telemetry" balances
read as "overage allowance" and "polled" balances read as "credits". Tiny
values keep their sign: a negative balance smaller than half a cent renders as
~$0.00 (slight overage), a positive tiny balance renders as <$0.01, and an
exact zero balance renders as $0.00.
Pinch savings — the dashboard card shows pinch's effect: share pruned (as a
percent), tokens saved, median saved, and 30d dollars saved. When pinch is
disabled the card shows "Pinch is disabled". The TUI dashboard has the same
pinch rows (share_pruned, total_tokens_saved, median_tokens_saved,
dollars_saved_usd_30d) pulled from the snapshot. See docs/pinch.md
for the full model.
Loops — the dashboard carries a Loops card (admin/frontend/index.html#loops)
showing the watchdog's open alerts, one row per alert. It pulls from
GET /admin/api/watchdog/loops (watchdog_alerts rows where resolved_at is
null) via loadLoops(). Each row shows exactly three things: a severity badge
(critical / warning / else info), the raw dedup key
(opencode-loop:<root session id>), and opened_at. That is all the panel
shows today — it does not show the agent, model, top repeated target, calls,
or $ since the last landed change; those live in watchdog_verdicts and have
not shipped to this panel yet.
Models — GET /admin/models is the model-availability table. Each row shows
the serving class, tier, status, and an override dropdown (active /
blocked / deprecated / stale) that writes through to the routing hard
filters. By default the table hides deprecated and stale rows; enable the
Show deprecated / stale toggle to include them.
blocked is a fourth value in that same per-model override dropdown, distinct
from the catalog meaning of deprecated. It is an operator stop: a model an
operator has switched off, not one the catalog flagged. Routing and pinned
requests both exclude a blocked model (_admin_excluded_models in
dispatcher.py and metrics.py). Undo is setting the dropdown back to
active (clearing the override via DELETE on the availability path), not a
separate unlock — there is no per-model Block button beside evidence and no
Blocked list with reasons; those belong to the "not built yet" item below.
Profiles — GET /admin/profiles shows each routing profile and what it
admits under the live catalog.
The page uses a two-pane layout (not a card grid): a narrow master list on
the left carries one compact line per profile — name, icon, admitted/interactive
counts, default badge, and a category-coverage info icon for access-gated rows.
Clicking a line loads the detail pane on the right, which consumes the full
width. The page header has a New profile button (opens the create/edit modal
in config.local.yaml).
Built-in profiles are read-only; the detail pane shows a Duplicate button
for each — opening a create modal pre-filled from the built-in definition under a
new name, producing an ordinary config profile. Config profiles can be edited
inline in the detail pane, duplicated, and deleted (unless they are the
current routing.default_profile, in which case the delete button is disabled).
Admission is probed per task category, not once with task_category=None.
eligible_categories is a restrict-only gate, so a category-gated row is
excluded from a category-less question, and locality used to report
"admits 0 models" while its one active ollama-local row was serving the two
categories it exists for. A badge that calls a working profile broken teaches
operators to ignore the badge on the day it is real — and the day it is real
is incident #3. The card now reads "admits 1 model for 2 of
11 task categories", and the zero-admit warning is reserved for a profile
that admits nothing under any category.
Proficiency — GET /admin/proficiency is the model × category matrix of
blended_score. proficiency_score is the only category-dependent term in
the ranking, so this is the table that decides routing; before this page the
portal's entire surface for it was a top-N list for one category on the
dashboard.
The page's one load-bearing requirement is that a score traffic earned must
not look like one that was copied in. Colour carries provenance and nothing
else — value is left to the digits, because a heatmap of blended_score would
put the eye on the number and hide exactly the distinction the page is for:
| cell | meaning |
|---|---|
| green | measured here, from this row's own client outcomes |
| amber | not measured here — an inherited prior, or a copy |
| blue | benchmark, at or above proficiency.self_eval_min_samples |
| slate | benchmark, below that floor (thin) |
| corner wedge | inherited_from names the row it was copied from |
Inheritance outranks source, and that ordering matters: a -flex row
carries its family's number verbatim, source column included, so rows exist
reading source='outcome_blended' with inherited_from set. Colouring those
green would paint a copy as a measurement.
Every cell opens a detail popup with leaderboard_score, self_eval_score,
outcome_score, both sample counts, inherited_from, and last_updated.
Filter by category, source or model name; sort by model, mean score or samples.
The page is read-only and there is no write endpoint, deliberately.
Proficiency is derived from evaluation and client outcomes, so a hand-edited
score is a fabricated measurement — the same failure as the empty
leaderboards.yaml and the provider's static_fallback carbon constant this
project already excludes.
Its lower card lists recent POST /outcome reports — the only ground truth
the router gets — with each one's verdict, whether it was counted or
excluded (model_attributable), and whether feedback.py has folded it in
yet. That exclusion needed surfacing: degraded-classification outcomes are
recorded and then silently dropped from folding, and "recorded but never
surfaced" is a pattern this project has now hit four times.
Controls — GET /admin/controls holds the operational triggers, runtime
knobs, the classifier's mode config, and the persisted config editor (writes
land in config/config.local.yaml; see
config-local-overlay). Changes are marked as dirty
and written only on save.
Knob organization
The generic knob table is grouped into six categories, each with a collapsible section header:
- Routing / Quality / Cost — the core dispatch objective: logging levels, energy ceiling, plan quota, cache pricing, and default profile.
- Classification — classifier runtime scalars such as cooldown, fallback tier/category, degraded warning thresholds, and context framing.
- Local Hardware — gates over local compute paths: the local-compute master switch, local verification, and local vision fallback.
- Caching / Context — session classification cache and context pruning (pinch).
- Safety Nets — circuit breaker.
- Watchdog — response-loop detector thresholds.
Within each category, knobs marked Advanced hide behind a secondary "Advanced" accordion. These are shape parameters internal to a covered master switch (for example, pinch relevance options and watchdog detector thresholds). The accordion keeps the first view short without removing access.
Knobs reach the page through three coverage sources:
- Registry knobs — runtime toggles from
_BOOL_KNOBS,_FLOAT_KNOBS, and_INT_KNOBSinadmin.py. They take effect immediately in memory and revert on restart. - Allowlist knobs — the same scalars persisted through
_CONFIG_ALLOWLIST, written toconfig/config.local.yamland surviving a restart. - Card-backed paths — knobs with their own dedicated admin card and endpoint pair because they have cross-field structure a flat scalar input cannot safely represent.
It carries three dedicated cards, none a row in the generic runtime-knob list, because each has cross-field structure a flat scalar/boolean input can't safely represent:
-
Local Compute — "gaming mode". See Gaming mode below.
-
Classifier —
classifier.mode's config (cloud_llmneeds a pinned primary orauto;local_encoderneeds a model id;local_decisionneeds a generative model id). It shows the current mode and, forcloud_primary_auto, the live resolved cheapest candidate — computed withrouting.cheapest_classifier_candidate, the same functionclassify()calls, not a static echo of the config. An admin panel that shows a config value instead of what dispatch actually does is a real class of bug — the same lesson the profiles zero-admit badge above already taught this project — and a "live" reading that is secretly stale is worse than no reading at all.POST /admin/api/classifier-configvalidates the same way config load does: an invalid combination (e.g.cloud_llmwith neither a pinned primary norauto) is rejected before anything reachesconfig.local.yaml, and mode plus its companion block are written as one atomic change so an in-between invalid state is never even written transiently.The header badge names the mode that is running, not the one saved. A save only persists to
config.local.yaml, and the router reads this block at startup, so an amberrestart pendingbadge sits beside it until the service restarts. A blank field is left out of the overlay, so the repo default keeps applying; a typed0is saved as0.When
local_decisionis selected the panel reveals a Local Decision block with the fields fromLocalDecisionConfiginsrc/config.py:field type default meaning base_urlstring http://localhost:11434Ollama endpoint for the generative classifier modelstring qwen3.5:4bThe small generative model that picks category by choice num_ctxint 8192Context window passed to the model timeout_sint 10Seconds before the request times out confidence_minfloat 0.5Minimum logprob-derived confidence to accept the verdict (range [0.0, 1.0]); below this the classification cascades as a failure coverage_minfloat 0.3Minimum total probability mass on option letters; below this parse_logprobsraises and the classification cascadestier_enabledbool falseWhether the classifier may also decide task_tier(vs. only category) -
Watchdog — the watchdog's live state (
controls.html#watchdog-card,loadWatchdog()): the last tick time (and ano tickbadge before the first one), how many sessions the last tick saw, how many verdicts it flagged, and how many alerts are currently open. Below those, a set of notification channel toggles (each channel with an enable/disable and a minimum severity, persisted viaPOST /admin/api/watchdog/channels), a Refresh button, and a Send test alert button (POST /admin/api/watchdog/test-alert).
Providers — GET /admin/providers is the dispatch-provider management page.
The left card adds a new provider (name, base_url, api_key_env,
has_energy_telemetry, enabled); the right card lists configured providers
with their source badge (base config vs overlay) and edit/delete actions. Base
providers are read-only; overlay providers can be edited or deleted. Creating,
editing, or deleting a provider writes to config/config.local.yaml, and a
service reload is required before dispatch sees the change.
Below the provider list is the Provider allowlists card. It lists every
provider and whether require_allowlist is enabled. For providers with an
allowlist, click Manage to see the current allowed model ids, remove
entries, and browse a live upstream catalog preview to add models with an
optional note. The catalog preview calls poller.fetch_openrouter directly, so
it is only available for OpenRouter providers that require an allowlist.
Decisions — GET /admin/decisions is the full decision log with
kind/category/tier filters and free-text search.
The stale-page notice
An admin tab left open across a frontend change keeps running the JavaScript it loaded, and the way that surfaces is unhelpful: the old code posts a field the shipped code no longer sends, and the operator gets a 422 about a field that is not in the file in front of them.
This is not an HTTP caching bug, and no header fixes it. Every frontend
route in admin.py already sets Cache-Control: no-cache, and FileResponse
supplies ETag and Last-Modified, so any page load revalidates and gets
fresh bytes. Verified against the live service:
GET /admin/profiles
cache-control: no-cache
etag: "27f7f7286a7435089bd3e372f447d4a5"
last-modified: Sat, 12 Sep 2026 16:02:48 GMT
The tab never loads again, so there is no request to attach a header to. The page has to ask instead.
GET /admin/api/frontend-version returns {"digest": "..."} — a SHA-256 over
(name, st_mtime_ns, st_size) for every file the frontend routes serve. It
reads no file contents: the question is only whether the files under an open
page moved, and stat()ing them answers it. mtime is in the digest as well as
size because the case this exists for is a hand-edit during iteration, which
frequently does not change a file's length. A missing file contributes a marker
rather than raising.
admin._FRONTEND_FILES is the set that gets fingerprinted, and
tests/test_admin_frontend.py pins it against the directory listing. A page
added to the portal but not to that list would be invisible to the poll, which
reopens exactly the failure the notice exists to close.
navbar.js — one file, loaded by all eight pages — captures the digest on load
and re-checks it every 30s, plus whenever the tab becomes visible again, which
is the moment that matches the failure's shape. A change shows a fixed pill in
the bottom-right corner reading This page is out of date, with a Reload
button and a dismiss. It is position: fixed and hidden until it fires, so it
shifts nothing when it appears, and it is not a slot in the navbar cluster
because appearing there would shove the status dot and the profile switcher
sideways — and the navbar scrolls out of view while the reason to reload does
not. Dismissing hides it until the next change.
It never reloads by itself. An operator may be mid-edit in a profile modal or holding a dirty config row, and discarding that silently costs more than the staleness does. The offer is the feature.
The poll interval is a constant in navbar.js, not a config knob. The repo's
"every knob belongs in config.yaml" rule is about config the router reads;
nothing in the routing path depends on this number, and a key in config.py
plus config.yaml whose only consumer is a browser timer would be a knob
pretending to be policy.
Gaming mode
local_compute.enabled (default true) is the outer gate over every call the
router makes to local hardware. Turn it off when you stop Ollama to give the
GPU back to something else, and the router skips the local path instead of
discovering the outage one classifier.timeout_seconds at a time — 120s per
request on this deployment — at each call site independently.
Skip, not fail. The gate sits above the client construction at every site:
| call site | behaviour with local compute off |
|---|---|
| classifier | skipped without dialling; goes straight to the cascade |
/health probe |
classifier_reachable: null — not asked, not unreachable |
| local verification | both the buffered and the streamed path decline |
| local vision fallback | declines |
| local dispatch rows | dropped in load_candidates, so a local row is never selected and then 503'd at dispatch |
/v1/models |
stops listing local rows; an explicit pin gets a 503 naming the flag |
It is one flag the code reads, not a macro that writes five keys. A macro
is hard to undo cleanly, drifts the moment a sixth call site appears, and
leaves nobody able to answer "why isn't the classifier running?" from one
place. verification.local_llm_enabled, local_vision.enabled and
local_energy.enabled keep their own meanings; this ANDs over them, so
turning it back on restores exactly the state you left.
It refuses to engage without classifier.cloud_fallback — on the runtime
knob (409) and at config load (a validation error). Skipping the local
classifier does not make classification remote; without a cloud classifier it
stops classifying, and every request falls through to a static guess recorded
as general_chat, a fully scored category that is indistinguishable from a
real classification afterwards. A refusal rather than a warning, because a
warning is what nobody reads while their game is loading. Nothing auto-writes
the cloud_fallback block — uncommenting the shipped example in
config/config.yaml is the intended setup path.
With the mode on, cascade steps 1 and 2 still run ahead of the cloud call.
A stale session classification is free and was a real classification of that
same session, so paying to re-derive an answer already held would be spending
money for nothing. "Force cloud" replaces the local model, not free correct
answers. Cloud classifications carry source="classifier_cloud", which is in
the attributable set, so POST /outcome keeps training proficiency normally.
Read-only dashboards — GET /admin/api/snapshot exposes data for:
quota— per-provider billing shape fromquota_accounts():period(billing window with start/next_reset/elapsed_fraction/source),accounts[](list of per-provider dicts withprovider,shape:metered_plan | prepaid_credit | self_hosted | unmetered,spend_usd, and type-specific blocks:planfor metered_plan,poolfor prepaid_credit,burnfor burn-rate metrics,creditwhen a balance URL was polled,energyfor kWh/calls),spend(aggregate spend withby_provider_usd,total_usd,estimated_usd, andestimate_ratio), andalarm(plan-pace or stale-reading alert with kind/severity/headline). The oldby_provider/total_balance_usdshape was removed.- per-model usage from
energy_observations - live routing decisions from
route_decisions - verdict mix and scoring coverage
- history:
GET /admin/api/history?range=6h(also1h,24h,7d,30d)
GET /admin/api/proficiency serves the proficiency matrix (every row with
source, both sample counts, inherited_from, last_updated, and a
server-computed thin flag) plus the recent client outcomes. Both queries live
in metrics.py, which must never import dispatcher — that separation is what
keeps them usable from /health, the TUI and the portal without a circular
import. There is no write counterpart.
GET /admin/api/frontend-version returns the frontend fingerprint that backs
the stale-page notice above. It stats the served files and reads none of them.
Operational triggers — async, fire-and-forget maintenance jobs:
POST /admin/api/refresh-catalogruns the poller and tier passPOST /admin/api/seed-energy?samples=Nstarts a reference sweepPOST /admin/api/apply-feedback?dry_run=truerunsfeedback.pyPOST /admin/api/restart-servicerestarts the running systemd unit
Runtime toggles
GET /admin/api/runtime shows persisted-vs-runtime values. POST /admin/api/runtime/{knob}
flips in-memory settings such as log_route_decisions, local_llm_enabled
and local_compute_enabled. Changes take effect immediately but reset on
restart. One knob can be refused rather than applied: see
Gaming mode.
Most knobs are booleans (_BOOL_KNOBS in admin.py) and
default_flex_preference / active_profile are strings. Numbers get one table
per type, not one shared table:
| table | knob | range | blank |
|---|---|---|---|
_FLOAT_KNOBS |
incumbent_challenger_cache_rate |
0 - 1 | neutral |
_INT_KNOBS |
session_cache_staleness_seconds |
5 - 7200 | refused |
They are separate because the declared type has to survive the write.
_FLOAT_KNOBS coerces with float(value), and session_cache.staleness_seconds
is declared int — a float there would leave the running cfg holding a value
load_config could never produce, and a fractional one would 422 the persisted
twin at load. The float tuple also carries a fourth element, neutral_path,
which exists only because the dial has a null-means-neutral semantic; an int
knob with no neutral would need a sentinel in it.
A runtime write bypasses every Pydantic validator — StrictModel sets
extra="forbid", not validate_assignment — so the endpoint does the type and
range checking itself: a numeric knob refuses anything that is not a number
(booleans included, since True is an int in Python and would land as 1.0
or a 1-second window) and anything outside its declared bounds. _INT_KNOBS
also refuses floats rather than truncating them, and imports its bounds from
config.STALENESS_SECONDS_MIN/MAX so the runtime path and the config
validator cannot drift into disagreeing about what the service will boot with.
Persisted config edits
GET /admin/api/config lists allowlisted keys. POST /admin/api/config/{key}
writes one allowlisted key back to config/config.local.yaml with a timestamped
backup and whole-config validation of the merged base + overlay. Arbitrary keys
are rejected.
Incumbent cache pricing (Wave 2)
Both knobs appear in both panels, as two adjacent rows:
| key | runtime knob | persisted key |
|---|---|---|
| gate | incumbent_cache_pricing |
objective.incumbent_cache_pricing |
| challenger dial | incumbent_challenger_cache_rate |
objective.incumbent_challenger_cache_rate |
The runtime pair is the one that matters for tuning. The dial exists so the
feature can be walked from neutral to full penalty without a revert, with a
re-measurement between each step (see docs/routing.md § incumbency pricing);
a loop that needs a tracked-file edit and a systemctl --user restart between
readings is a loop nobody walks. dispatcher.py reads cfg.objective.* per
request, so an in-memory write is live on the next one.
Blank means neutral, never zero. They are opposite ends of the same dial:
null follows objective.assumed_cache_rate, so a challenger pays nothing for
discarding the incumbent's cache, while 0.0 prices every challenger as a
fully cold prompt — the maximum incumbent advantage. So an empty field posts
null, not 0 (Number('') is 0 in JavaScript, which is exactly how an
operator clearing the field to switch the feature off would instead have
switched it to maximum). The persisted path writes null through to the
overlay and lets Objective._resolve_challenger_cache_rate resolve it at load;
the runtime path makes the same substitution itself, so cfg only ever holds
one representation of neutral. Both directions are pinned by tests in
tests/test_admin_config.py and tests/test_admin_runtime.py.
The dial is a cache rate in [0, 1], not a percentage, and the UI does no
scaling — what is typed is what is stored. That is the
classifier.confidence_threshold lesson applied rather than relearned: that
control saved a raw 80 meaning 80% for a field the loader wanted as
0.0-1.0, and it was caught before the restart only by luck. Out-of-range
values are refused by the endpoint on the runtime path and by RouterConfig
validation of the merged config on the persisted path, before any byte reaches
disk.
Neither knob is enabled in tracked config: incumbent_cache_pricing ships
false and the dial ships blank.
Session classification cache
The same shape as the pair above — a switch and the window it gates, adjacent in both panels:
| key | runtime knob | persisted key |
|---|---|---|
| switch | session_cache_enabled |
session_cache.enabled |
| window | session_cache_staleness_seconds |
session_cache.staleness_seconds |
The window decides how long one classification keeps steering routing, so
it is the knob that sets the size of the concession CLAUDE.md's north star rule
2 describes, not an implementation detail of the switch. Measured on live
traffic: 96.6% of all classifications are cache replays, and one classification
drove 107 consecutive turns across 840s. Before this control the only
available settings were a 1200-second window or no cache at all, and the answer
is almost certainly in between — which is why the runtime half matters.
dispatcher.py reads cfg.session_cache.staleness_seconds per request, so a
narrower window is live on the next one and the replay share can be re-measured
without a restart.
Units are seconds and the UI does no scaling — what is typed is what is
stored, the classifier.confidence_threshold lesson applied rather than
relearned. The field is an int, and a float is refused rather than truncated:
20.5 quietly becoming 20 is a window nobody chose.
Blank is refused, and so is 0. They are not the same as switching the cache off, and the difference is easy to miss:
session_cache.putstill writes on every turn, so the entry exists.- The classifier-failure cascade reads it with
session_cache.stale_read, which ignores staleness entirely — so a 0-second window still replays a session's label whenever the classifier is down.
So 0 is out of range (floor 5, the pre-existing > 0 rule) and an empty field
posts null, which the runtime endpoint refuses with a message naming
session_cache_enabled as the switch the operator was reaching for. The
persisted path refuses both through whole-config validation of the merged
base + overlay, before any byte reaches disk.
The ceiling, 7200, is a judgement: an unbounded window is a cache that never
expires. It is 7200/1200 = 6x (same ratio) the shipped default and 7200/840 ≈
8.57x the longest single-classification
run measured, so it sits above every value there is a reason to try while still
guaranteeing a label cannot outlive the working session that produced it. The
bound lives in config.py (STALENESS_SECONDS_MIN/MAX) so a file edit, an
overlay write and a runtime POST all enforce the same range.
The shipped default is unchanged: enabled: true, staleness_seconds: 1200.
Classifier mode
GET /admin/api/classifier-config / POST /admin/api/classifier-config are
a separate pair from the allowlist above, deliberately: classifier.mode
plus its companion block (cloud_primary/cloud_primary_auto or encoder)
is not one scalar, so it needs both fields written together or an
intermediate invalid state could land on disk. The GET response includes
resolved_primary — the live result of
routing.cheapest_classifier_candidate against the current catalog when
cloud_primary_auto is set, null otherwise.
mode in that response is the saved value. running_mode is what the
process is using, and restart_pending is true when the saved classifier
block differs from the running one in any field, or null when the saved
config no longer validates and the comparison cannot be made. The block is read
at import, so only a restart closes the gap.
Model availability overrides
POST /admin/api/models/{model_id:path}/{provider}/availability marks a model
as active, blocked, deprecated, or stale. DELETE on the same path
removes the override. The {model_id:path} converter accepts model ids that
contain slashes, so providers like OpenRouter with slash-bearing ids are
handled the same way as plain ids. Deprecation feeds into the routing hard
filters.
Not built yet. The Loops panel does not yet show the evidence columns an operator would act on — agent, model, top repeated target, calls, and $ since the last landed change — and there is no per-model stall rollup, no Block model button beside the evidence, and no Blocked list with reasons. Those are the remaining surface of North Star #4 in CLAUDE.md, whose first half (conversation identity and the watchdog) is shipped.






