Change the unit of session_cache.staleness from minutes to seconds so it can express finer-grained (sub-minute) staleness windows. This is a straight rename, not an additive/compat knob — no deprecated alias, per the project's convention of updating every consumer in the same change. New bounds: floor 5 seconds (was 1 minute), ceiling 7200 seconds (was 120 minutes). Default: 1200 seconds (was 20 minutes). The validator's reasoning is unit-independent and carries over: the floor is deliberately > 0 because 0 would make session_cache.get() miss every turn while put() still writes and the classifier-failure cascade's stale_read ignores staleness; the ceiling reasoning (unbounded window = never-expiring cache, 7200s still >> 840s real max run) also carries over in seconds. Every consumer updated in the same commit: - src/config.py: STALENESS_MINUTES_MIN/MAX -> STALENESS_SECONDS_MIN/MAX = 5/7200, staleness_minutes -> staleness_seconds: 1200, validator updated - src/dispatcher.py: drop the * 60 conversion (field is native seconds) - src/admin.py: _INT_KNOBS key/path/constants, _CONFIG_ALLOWLIST, _CONFIG_GET_ORDER, _runtime_state, error message template - admin/frontend/controls.html: note keys, tooltip, NUMBER_BOUNDS - config/config.yaml: staleness_seconds: 1200 - tests: test_admin_runtime/config/frontend/knob_coverage, plus stale comment in test_chat_completions - docs: admin-portal.md, evaluation.md, README.md config.local.yaml is gitignored and will be migrated separately.
451 lines
24 KiB
Markdown
451 lines
24 KiB
Markdown
> Deep dive into the admin web portal. Back: [README](../README.md).
|
||
|
||
### Admin web portal
|
||
|
||
`GET /admin/` serves a management portal from the running router on the same
|
||
loopback-only bind as the API. It has no auth layer yet, so like the other
|
||
endpoints it is reachable only from `127.0.0.1`.
|
||
|
||
The portal is a **six-page, glass dark-mode** dashboard built on Tabler/Bootstrap
|
||
with a few lighter inline SVG icons (see `plans/admin-design-standards.md`
|
||
for the visual language and `plans/admin-work-framework.md` for the method
|
||
used to evolve it). The six pages are Dashboard, Models, Profiles, Decisions,
|
||
Controls, and Providers. It is implemented in `admin.py`, with
|
||
`config/admin_schema.sql` for its future tables and `admin/frontend/` for the
|
||
browser UI.
|
||
|
||
**Dashboard** — `GET /admin/` aggregates the router at a glance: a quota chip
|
||
(burn against `objective.plan_kwh_per_period`), per-model usage bars, verdict
|
||
mix and category breakdown, five history mini-charts, and a recent-decisions
|
||
table. Clicking the chip opens a detail modal with per-provider balances and
|
||
runway; each line is labeled by its billing shape, so `"telemetry"` balances
|
||
read as "overage allowance" and `"polled"` balances read as "credits". Tiny
|
||
values keep their sign: a negative balance smaller than half a cent renders as
|
||
`~$0.00 (slight overage)`, a positive tiny balance renders as `<$0.01`, and an
|
||
exact zero balance renders as `$0.00`.
|
||
|
||
**Pinch savings** — the dashboard card shows pinch's effect: share pruned (as a
|
||
percent), tokens saved, median saved, and 30d dollars saved. When pinch is
|
||
disabled the card shows "Pinch is disabled". The TUI dashboard has the same
|
||
pinch rows (`share_pruned`, `total_tokens_saved`, `median_tokens_saved`,
|
||
`dollars_saved_usd_30d`) pulled from the snapshot. See [docs/pinch.md](pinch.md)
|
||
for the full model.
|
||
|
||
<div align="center">
|
||
<img src="../assets/screenshots/dashboard.png" alt="Admin dashboard" width="820">
|
||
</div>
|
||
|
||
**Models** — `GET /admin/models` is the model-availability table. Each row shows
|
||
the serving class, tier, status, and an override dropdown (`active` /
|
||
`deprecated` / `stale`) that writes through to the routing hard filters. By
|
||
default the table hides deprecated and stale rows; enable the **Show deprecated
|
||
/ stale** toggle to include them.
|
||
|
||
<div align="center">
|
||
<img src="../assets/screenshots/models.png" alt="Admin models page" width="820">
|
||
</div>
|
||
|
||
**Profiles** — `GET /admin/profiles` shows each routing profile and what it
|
||
admits under the live catalog.
|
||
|
||
The page uses a **two-pane layout** (not a card grid): a narrow master list on
|
||
the left carries one compact line per profile — name, icon, admitted/interactive
|
||
counts, default badge, and a category-coverage info icon for access-gated rows.
|
||
Clicking a line loads the detail pane on the right, which consumes the full
|
||
width. The page header has a **New profile** button (opens the create/edit modal
|
||
in `config.local.yaml`).
|
||
|
||
Built-in profiles are read-only; the detail pane shows a **Duplicate** button
|
||
for each — opening a create modal pre-filled from the built-in definition under a
|
||
new name, producing an ordinary config profile. Config profiles can be **edited
|
||
inline** in the detail pane, **duplicated**, and **deleted** (unless they are the
|
||
current `routing.default_profile`, in which case the delete button is disabled).
|
||
|
||
Admission is probed **per task category**, not once with `task_category=None`.
|
||
`eligible_categories` is a restrict-only gate, so a category-gated row is
|
||
excluded from a category-less question, and `locality` used to report
|
||
"admits 0 models" while its one active `ollama-local` row was serving the two
|
||
categories it exists for. A badge that calls a working profile broken teaches
|
||
operators to ignore the badge on the day it is real — and the day it is real
|
||
is [incident #3](incidents.md). The card now reads *"admits 1 model for 2 of
|
||
11 task categories"*, and the zero-admit warning is reserved for a profile
|
||
that admits nothing under any category.
|
||
|
||
<div align="center">
|
||
<img src="../assets/screenshots/profiles.png" alt="Admin profiles page" width="820">
|
||
</div>
|
||
|
||
**Proficiency** — `GET /admin/proficiency` is the model × category matrix of
|
||
`blended_score`. `proficiency_score` is the only category-dependent term in
|
||
the ranking, so this is the table that decides routing; before this page the
|
||
portal's entire surface for it was a top-N list for one category on the
|
||
dashboard.
|
||
|
||
The page's one load-bearing requirement is that **a score traffic earned must
|
||
not look like one that was copied in**. Colour carries provenance and nothing
|
||
else — value is left to the digits, because a heatmap of `blended_score` would
|
||
put the eye on the number and hide exactly the distinction the page is for:
|
||
|
||
| cell | meaning |
|
||
|---|---|
|
||
| green | measured here, from this row's own client outcomes |
|
||
| amber | not measured here — an inherited prior, or a copy |
|
||
| blue | benchmark, at or above `proficiency.self_eval_min_samples` |
|
||
| slate | benchmark, below that floor (thin) |
|
||
| corner wedge | `inherited_from` names the row it was copied from |
|
||
|
||
**Inheritance outranks `source`**, and that ordering matters: a `-flex` row
|
||
carries its family's number verbatim, `source` column included, so rows exist
|
||
reading `source='outcome_blended'` with `inherited_from` set. Colouring those
|
||
green would paint a copy as a measurement.
|
||
|
||
Every cell opens a detail popup with `leaderboard_score`, `self_eval_score`,
|
||
`outcome_score`, both sample counts, `inherited_from`, and `last_updated`.
|
||
Filter by category, source or model name; sort by model, mean score or samples.
|
||
|
||
The page is **read-only and there is no write endpoint**, deliberately.
|
||
Proficiency is derived from evaluation and client outcomes, so a hand-edited
|
||
score is a fabricated measurement — the same failure as the empty
|
||
`leaderboards.yaml` and the provider's `static_fallback` carbon constant this
|
||
project already excludes.
|
||
|
||
Its lower card lists recent `POST /outcome` reports — the only ground truth
|
||
the router gets — with each one's verdict, whether it was **counted or
|
||
excluded** (`model_attributable`), and whether `feedback.py` has folded it in
|
||
yet. That exclusion needed surfacing: degraded-classification outcomes are
|
||
recorded and then silently dropped from folding, and "recorded but never
|
||
surfaced" is a pattern this project has now hit four times.
|
||
|
||
<div align="center">
|
||
<img src="../assets/screenshots/proficiency.png" alt="Admin proficiency matrix" width="820">
|
||
</div>
|
||
|
||
**Controls** — `GET /admin/controls` holds the operational triggers, runtime
|
||
knobs, the classifier's mode config, and the persisted config editor (writes
|
||
land in `config/config.local.yaml`; see
|
||
[config-local-overlay](config-local-overlay.md)). Changes are marked as dirty
|
||
and written only on save.
|
||
|
||
It carries two dedicated cards, neither a row in the generic runtime-knob
|
||
list, because both have cross-field structure a flat scalar/boolean input
|
||
can't safely represent:
|
||
|
||
- **Local Compute** — "gaming mode". See [Gaming mode](#gaming-mode) below.
|
||
- **Classifier** — `classifier.mode`'s config (`cloud_llm` needs a pinned
|
||
primary or `auto`; `local_encoder` needs a model id). It shows the current
|
||
mode and, for `cloud_primary_auto`, the **live** resolved cheapest
|
||
candidate — computed with `routing.cheapest_classifier_candidate`, the same
|
||
function `classify()` calls, not a static echo of the config. An admin
|
||
panel that shows a config value instead of what dispatch actually does is a
|
||
real class of bug — the same lesson the profiles zero-admit badge above
|
||
already taught this project — and a "live" reading that is secretly stale
|
||
is worse than no reading at all. `POST /admin/api/classifier-config`
|
||
validates the same way config load does: an invalid combination (e.g.
|
||
`cloud_llm` with neither a pinned primary nor `auto`) is rejected before
|
||
anything reaches `config.local.yaml`, and mode plus its companion block are
|
||
written as one atomic change so an in-between invalid state is never even
|
||
written transiently.
|
||
|
||
<div align="center">
|
||
<img src="../assets/screenshots/controls.png" alt="Admin controls page" width="820">
|
||
</div>
|
||
|
||
**Providers** — `GET /admin/providers` is the dispatch-provider management page.
|
||
The left card adds a new provider (`name`, `base_url`, `api_key_env`,
|
||
`has_energy_telemetry`, `enabled`); the right card lists configured providers
|
||
with their source badge (base config vs overlay) and edit/delete actions. Base
|
||
providers are read-only; overlay providers can be edited or deleted. Creating,
|
||
editing, or deleting a provider writes to `config/config.local.yaml`, and a
|
||
service reload is required before dispatch sees the change.
|
||
|
||
Below the provider list is the **Provider allowlists** card. It lists every
|
||
provider and whether `require_allowlist` is enabled. For providers with an
|
||
allowlist, click **Manage** to see the current allowed model ids, remove
|
||
entries, and browse a live upstream catalog preview to add models with an
|
||
optional note. The catalog preview calls `poller.fetch_openrouter` directly, so
|
||
it is only available for OpenRouter providers that require an allowlist.
|
||
|
||
<div align="center">
|
||
<img src="../assets/screenshots/providers.png" alt="Admin providers page" width="820">
|
||
</div>
|
||
|
||
**Decisions** — `GET /admin/decisions` is the full decision log with
|
||
kind/category/tier filters and free-text search.
|
||
|
||
<div align="center">
|
||
<img src="../assets/screenshots/decisions.png" alt="Admin decisions log" width="820">
|
||
</div>
|
||
|
||
### The stale-page notice
|
||
|
||
An admin tab left open across a frontend change keeps running the JavaScript it
|
||
loaded, and the way that surfaces is unhelpful: the old code posts a field the
|
||
shipped code no longer sends, and the operator gets a 422 about a field that is
|
||
not in the file in front of them.
|
||
|
||
**This is not an HTTP caching bug, and no header fixes it.** Every frontend
|
||
route in `admin.py` already sets `Cache-Control: no-cache`, and `FileResponse`
|
||
supplies `ETag` and `Last-Modified`, so any page *load* revalidates and gets
|
||
fresh bytes. Verified against the live service:
|
||
|
||
```
|
||
GET /admin/profiles
|
||
cache-control: no-cache
|
||
etag: "27f7f7286a7435089bd3e372f447d4a5"
|
||
last-modified: Sat, 12 Sep 2026 16:02:48 GMT
|
||
```
|
||
|
||
The tab never loads again, so there is no request to attach a header to. The
|
||
page has to ask instead.
|
||
|
||
`GET /admin/api/frontend-version` returns `{"digest": "..."}` — a SHA-256 over
|
||
`(name, st_mtime_ns, st_size)` for every file the frontend routes serve. It
|
||
reads no file contents: the question is only whether the files under an open
|
||
page moved, and stat()ing them answers it. mtime is in the digest as well as
|
||
size because the case this exists for is a hand-edit during iteration, which
|
||
frequently does not change a file's length. A missing file contributes a marker
|
||
rather than raising.
|
||
|
||
`admin._FRONTEND_FILES` is the set that gets fingerprinted, and
|
||
`tests/test_admin_frontend.py` pins it against the directory listing. A page
|
||
added to the portal but not to that list would be invisible to the poll, which
|
||
reopens exactly the failure the notice exists to close.
|
||
|
||
`navbar.js` — one file, loaded by all eight pages — captures the digest on load
|
||
and re-checks it every 30s, plus whenever the tab becomes visible again, which
|
||
is the moment that matches the failure's shape. A change shows a fixed pill in
|
||
the bottom-right corner reading **This page is out of date**, with a **Reload**
|
||
button and a dismiss. It is `position: fixed` and hidden until it fires, so it
|
||
shifts nothing when it appears, and it is not a slot in the navbar cluster
|
||
because appearing there would shove the status dot and the profile switcher
|
||
sideways — and the navbar scrolls out of view while the reason to reload does
|
||
not. Dismissing hides it until the *next* change.
|
||
|
||
**It never reloads by itself.** An operator may be mid-edit in a profile modal
|
||
or holding a dirty config row, and discarding that silently costs more than the
|
||
staleness does. The offer is the feature.
|
||
|
||
The poll interval is a constant in `navbar.js`, not a config knob. The repo's
|
||
"every knob belongs in `config.yaml`" rule is about config the *router* reads;
|
||
nothing in the routing path depends on this number, and a key in `config.py`
|
||
plus `config.yaml` whose only consumer is a browser timer would be a knob
|
||
pretending to be policy.
|
||
|
||
### Gaming mode
|
||
|
||
`local_compute.enabled` (default `true`) is the outer gate over every call the
|
||
router makes to local hardware. Turn it off when you stop Ollama to give the
|
||
GPU back to something else, and the router **skips** the local path instead of
|
||
discovering the outage one `classifier.timeout_seconds` at a time — 120s per
|
||
request on this deployment — at each call site independently.
|
||
|
||
Skip, not fail. The gate sits above the client construction at every site:
|
||
|
||
| call site | behaviour with local compute off |
|
||
|---|---|
|
||
| classifier | skipped without dialling; goes straight to the cascade |
|
||
| `/health` probe | `classifier_reachable: null` — not asked, not unreachable |
|
||
| local verification | both the buffered and the streamed path decline |
|
||
| local vision fallback | declines |
|
||
| local dispatch rows | dropped in `load_candidates`, so a local row is never selected and then 503'd at dispatch |
|
||
| `/v1/models` | stops listing local rows; an explicit pin gets a 503 naming the flag |
|
||
|
||
It is **one flag the code reads, not a macro that writes five keys.** A macro
|
||
is hard to undo cleanly, drifts the moment a sixth call site appears, and
|
||
leaves nobody able to answer "why isn't the classifier running?" from one
|
||
place. `verification.local_llm_enabled`, `local_vision.enabled` and
|
||
`local_energy.enabled` keep their own meanings; this ANDs over them, so
|
||
turning it back on restores exactly the state you left.
|
||
|
||
**It refuses to engage without `classifier.cloud_fallback`** — on the runtime
|
||
knob (409) and at config load (a validation error). Skipping the local
|
||
classifier does not make classification remote; without a cloud classifier it
|
||
stops classifying, and every request falls through to a static guess recorded
|
||
as `general_chat`, a fully scored category that is indistinguishable from a
|
||
real classification afterwards. A refusal rather than a warning, because a
|
||
warning is what nobody reads while their game is loading. Nothing auto-writes
|
||
the `cloud_fallback` block — uncommenting the shipped example in
|
||
`config/config.yaml` is the intended setup path.
|
||
|
||
With the mode on, cascade steps 1 and 2 still run **ahead** of the cloud call.
|
||
A stale session classification is free and was a real classification of that
|
||
same session, so paying to re-derive an answer already held would be spending
|
||
money for nothing. "Force cloud" replaces the local *model*, not free correct
|
||
answers. Cloud classifications carry `source="classifier_cloud"`, which is in
|
||
the attributable set, so `POST /outcome` keeps training proficiency normally.
|
||
|
||
**Read-only dashboards** — `GET /admin/api/snapshot` exposes data for:
|
||
|
||
- `quota` — per-provider billing shape from `quota_accounts()`: `period` (billing window with start/next_reset/elapsed_fraction/source), `accounts[]` (list of per-provider dicts with `provider`, `shape`: `metered_plan | prepaid_credit | self_hosted | unmetered`, `spend_usd`, and type-specific blocks: `plan` for metered_plan, `pool` for prepaid_credit, `burn` for burn-rate metrics, `credit` when a balance URL was polled, `energy` for kWh/calls), `spend` (aggregate spend with `by_provider_usd`, `total_usd`, `estimated_usd`, and `estimate_ratio`), and `alarm` (plan-pace or stale-reading alert with kind/severity/headline). The old `by_provider` / `total_balance_usd` shape was removed.
|
||
- per-model usage from `energy_observations`
|
||
- live routing decisions from `route_decisions`
|
||
- verdict mix and scoring coverage
|
||
- history: `GET /admin/api/history?range=6h` (also `1h`, `24h`, `7d`, `30d`)
|
||
|
||
`GET /admin/api/proficiency` serves the proficiency matrix (every row with
|
||
`source`, both sample counts, `inherited_from`, `last_updated`, and a
|
||
server-computed `thin` flag) plus the recent client outcomes. Both queries live
|
||
in `metrics.py`, which must never import `dispatcher` — that separation is what
|
||
keeps them usable from `/health`, the TUI and the portal without a circular
|
||
import. There is no write counterpart.
|
||
|
||
`GET /admin/api/frontend-version` returns the frontend fingerprint that backs
|
||
the stale-page notice above. It stats the served files and reads none of them.
|
||
|
||
**Operational triggers** — async, fire-and-forget maintenance jobs:
|
||
|
||
- `POST /admin/api/refresh-catalog` runs the poller and tier pass
|
||
- `POST /admin/api/seed-energy?samples=N` starts a reference sweep
|
||
- `POST /admin/api/apply-feedback?dry_run=true` runs `feedback.py`
|
||
- `POST /admin/api/restart-service` restarts the running systemd unit
|
||
|
||
**Runtime toggles**
|
||
|
||
`GET /admin/api/runtime` shows persisted-vs-runtime values. `POST /admin/api/runtime/{knob}`
|
||
flips in-memory settings such as `log_route_decisions`, `local_llm_enabled`
|
||
and `local_compute_enabled`. Changes take effect immediately but reset on
|
||
restart. One knob can be refused rather than applied: see
|
||
[Gaming mode](#gaming-mode).
|
||
|
||
Most knobs are booleans (`_BOOL_KNOBS` in `admin.py`) and
|
||
`default_flex_preference` / `active_profile` are strings. Numbers get one table
|
||
per type, not one shared table:
|
||
|
||
| table | knob | range | blank |
|
||
|---|---|---|---|
|
||
| `_FLOAT_KNOBS` | `incumbent_challenger_cache_rate` | 0 - 1 | neutral |
|
||
| `_INT_KNOBS` | `session_cache_staleness_seconds` | 5 - 7200 | refused |
|
||
|
||
They are separate because the declared type has to survive the write.
|
||
`_FLOAT_KNOBS` coerces with `float(value)`, and `session_cache.staleness_seconds`
|
||
is declared `int` — a float there would leave the running `cfg` holding a value
|
||
`load_config` could never produce, and a fractional one would 422 the persisted
|
||
twin at load. The float tuple also carries a fourth element, `neutral_path`,
|
||
which exists only because the dial has a null-means-neutral semantic; an int
|
||
knob with no neutral would need a sentinel in it.
|
||
|
||
A runtime write bypasses every Pydantic validator — `StrictModel` sets
|
||
`extra="forbid"`, not `validate_assignment` — so the endpoint does the type and
|
||
range checking itself: a numeric knob refuses anything that is not a number
|
||
(booleans included, since `True` is an `int` in Python and would land as `1.0`
|
||
or a 1-second window) and anything outside its declared bounds. `_INT_KNOBS`
|
||
also refuses floats rather than truncating them, and imports its bounds from
|
||
`config.STALENESS_SECONDS_MIN`/`MAX` so the runtime path and the config
|
||
validator cannot drift into disagreeing about what the service will boot with.
|
||
|
||
**Persisted config edits**
|
||
|
||
`GET /admin/api/config` lists allowlisted keys. `POST /admin/api/config/{key}`
|
||
writes one allowlisted key back to `config/config.local.yaml` with a timestamped
|
||
backup and whole-config validation of the merged base + overlay. Arbitrary keys
|
||
are rejected.
|
||
|
||
**Incumbent cache pricing (Wave 2)**
|
||
|
||
Both knobs appear in *both* panels, as two adjacent rows:
|
||
|
||
| key | runtime knob | persisted key |
|
||
|---|---|---|
|
||
| gate | `incumbent_cache_pricing` | `objective.incumbent_cache_pricing` |
|
||
| challenger dial | `incumbent_challenger_cache_rate` | `objective.incumbent_challenger_cache_rate` |
|
||
|
||
The runtime pair is the one that matters for tuning. The dial exists so the
|
||
feature can be walked from neutral to full penalty **without a revert**, with a
|
||
re-measurement between each step (see `docs/routing.md` § incumbency pricing);
|
||
a loop that needs a tracked-file edit and a `systemctl --user restart` between
|
||
readings is a loop nobody walks. `dispatcher.py` reads `cfg.objective.*` per
|
||
request, so an in-memory write is live on the next one.
|
||
|
||
**Blank means neutral, never zero.** They are opposite ends of the same dial:
|
||
`null` follows `objective.assumed_cache_rate`, so a challenger pays nothing for
|
||
discarding the incumbent's cache, while `0.0` prices every challenger as a
|
||
fully cold prompt — the maximum incumbent advantage. So an empty field posts
|
||
`null`, not `0` (`Number('')` is `0` in JavaScript, which is exactly how an
|
||
operator clearing the field to switch the feature off would instead have
|
||
switched it to maximum). The persisted path writes `null` through to the
|
||
overlay and lets `Objective._resolve_challenger_cache_rate` resolve it at load;
|
||
the runtime path makes the same substitution itself, so `cfg` only ever holds
|
||
one representation of neutral. Both directions are pinned by tests in
|
||
`tests/test_admin_config.py` and `tests/test_admin_runtime.py`.
|
||
|
||
The dial is a cache **rate in [0, 1]**, not a percentage, and the UI does no
|
||
scaling — what is typed is what is stored. That is the
|
||
`classifier.confidence_threshold` lesson applied rather than relearned: that
|
||
control saved a raw `80` meaning 80% for a field the loader wanted as
|
||
`0.0`-`1.0`, and it was caught before the restart only by luck. Out-of-range
|
||
values are refused by the endpoint on the runtime path and by `RouterConfig`
|
||
validation of the merged config on the persisted path, before any byte reaches
|
||
disk.
|
||
|
||
Neither knob is enabled in tracked config: `incumbent_cache_pricing` ships
|
||
`false` and the dial ships blank.
|
||
|
||
**Session classification cache**
|
||
|
||
The same shape as the pair above — a switch and the window it gates, adjacent
|
||
in both panels:
|
||
|
||
| key | runtime knob | persisted key |
|
||
|---|---|---|
|
||
| switch | `session_cache_enabled` | `session_cache.enabled` |
|
||
| window | `session_cache_staleness_seconds` | `session_cache.staleness_seconds` |
|
||
|
||
The window decides how long **one** classification keeps steering routing, so
|
||
it is the knob that sets the size of the concession CLAUDE.md's north star rule
|
||
2 describes, not an implementation detail of the switch. Measured on live
|
||
traffic: 96.6% of all classifications are cache replays, and one classification
|
||
drove 107 consecutive turns across 840s. Before this control the only
|
||
available settings were a 1200-second window or no cache at all, and the answer
|
||
is almost certainly in between — which is why the runtime half matters.
|
||
`dispatcher.py` reads `cfg.session_cache.staleness_seconds` per request, so a
|
||
narrower window is live on the next one and the replay share can be re-measured
|
||
without a restart.
|
||
|
||
**Units are seconds and the UI does no scaling** — what is typed is what is
|
||
stored, the `classifier.confidence_threshold` lesson applied rather than
|
||
relearned. The field is an `int`, and a float is refused rather than truncated:
|
||
`20.5` quietly becoming `20` is a window nobody chose.
|
||
|
||
**Blank is refused, and so is 0.** They are not the same as switching the cache
|
||
off, and the difference is easy to miss:
|
||
|
||
- `session_cache.put` still writes on every turn, so the entry exists.
|
||
- The classifier-failure cascade reads it with `session_cache.stale_read`,
|
||
which **ignores staleness entirely** — so a 0-second window still replays a
|
||
session's label whenever the classifier is down.
|
||
|
||
So 0 is out of range (floor 5, the pre-existing `> 0` rule) and an empty field
|
||
posts `null`, which the runtime endpoint refuses with a message naming
|
||
`session_cache_enabled` as the switch the operator was reaching for. The
|
||
persisted path refuses both through whole-config validation of the merged
|
||
base + overlay, before any byte reaches disk.
|
||
|
||
The ceiling, 7200, is a judgement: an unbounded window is a cache that never
|
||
expires. It is 7200/1200 = 6x (same ratio) the shipped default and 7200/840 ≈
|
||
8.57x the longest single-classification
|
||
run measured, so it sits above every value there is a reason to try while still
|
||
guaranteeing a label cannot outlive the working session that produced it. The
|
||
bound lives in `config.py` (`STALENESS_SECONDS_MIN`/`MAX`) so a file edit, an
|
||
overlay write and a runtime POST all enforce the same range.
|
||
|
||
The shipped default is unchanged: `enabled: true`, `staleness_seconds: 1200`.
|
||
|
||
**Classifier mode**
|
||
|
||
`GET /admin/api/classifier-config` / `POST /admin/api/classifier-config` are
|
||
a separate pair from the allowlist above, deliberately: `classifier.mode`
|
||
plus its companion block (`cloud_primary`/`cloud_primary_auto` or `encoder`)
|
||
is not one scalar, so it needs both fields written together or an
|
||
intermediate invalid state could land on disk. The GET response includes
|
||
`resolved_primary` — the live result of
|
||
`routing.cheapest_classifier_candidate` against the current catalog when
|
||
`cloud_primary_auto` is set, `null` otherwise.
|
||
|
||
**Model availability overrides**
|
||
|
||
`POST /admin/api/models/{model_id:path}/{provider}/availability` marks a model
|
||
as `active`, `deprecated`, or `stale`. `DELETE` on the same path removes the
|
||
override. The `{model_id:path}` converter accepts model ids that contain
|
||
slashes, so providers like OpenRouter with slash-bearing ids are handled the
|
||
same way as plain ids. Deprecation feeds into the routing hard filters.
|