diff --git a/CLAUDE.md b/CLAUDE.md
index 3b972ab..fc3b62d 100644
--- a/CLAUDE.md
+++ b/CLAUDE.md
@@ -1139,6 +1139,66 @@ The same applies to an Ollama shared over a VPN — it has no auth either, so
`0.0.0.0`, which would publish it on whatever network the client happens to
be on.
+### A bare SIGTERM could hang the process forever — fixed 2026-08-29
+
+Observed live: the dispatcher went unreachable (opencode: `ConnectionError:
+... Connection refused`, retrying) while `systemctl --user status` still
+reported it `active (running)`. The process was alive — it had logged
+`Shutting down` / `Waiting for connections to close` and never got past that.
+Its listening socket was already closed, but the SSE clients holding
+`/events/decisions` open (the TUI, the admin dashboard) never disconnect, so
+uvicorn's graceful shutdown had nothing to wait *out*.
+
+That alone would just be slow. What made it fatal: there was no matching
+`Stopping Local LLM model router...` line from systemd before it — something
+sent SIGTERM straight to the PID rather than through `systemctl
+stop`/`restart`. Systemd only enforces `TimeoutStopUSec` (10s here) when *it*
+is the one running the stop job; a signal delivered outside that path leaves
+the unit `active (running)` forever — the main PID never exits, so
+`Restart=on-failure` never fires either. The process just sat there,
+permanently unreachable, with no supervisor noticing anything was wrong.
+
+The source of that bare SIGTERM is still unknown — nothing in
+`dispatcher.py`, `admin.py`, or `tui.py` sends a signal to the process, so it
+was most likely a manual `kill` / `pkill` / process-manager action in another
+terminal while iterating on the admin portal work. Worth checking for if it
+recurs.
+
+**Fixed the failure mode, not just the trigger.** `ExecStart` now passes
+`--timeout-graceful-shutdown 5` (`deploy/llm-router.service`), which caps
+uvicorn's own connection-draining wait at 5s *regardless of who sends the
+signal or whether systemd is tracking a stop job*. A bare SIGTERM now
+converges to a clean exit in 5s instead of hanging forever. This also fixes
+routine restarts: every ordinary `systemctl --user restart` was silently
+taking the full 10s-then-SIGKILL path already, because the same open SSE
+connections were blocking graceful shutdown there too — it just didn't
+matter, because systemd was tracking that stop job and enforcing the kill.
+
+Recovery if this happens again: `systemctl --user restart
+llm-router.service`. Since the hung process was never in a tracked stop job,
+this issues a fresh stop/start cycle that systemd *does* enforce the timeout
+on, so it reliably clears the hang.
+
+**Follow-up, same day: the graceful-shutdown fix exposed a second bug.**
+Once the process could exit cleanly on a bare SIGTERM instead of hanging,
+systemd's `Restart=on-failure` turned out to explicitly exclude clean
+termination by SIGTERM/SIGINT from auto-restart — its assumption is that
+SIGTERM always means someone deliberately asked the service to stop. That
+assumption is false for this bare signal, so the fixed process was now
+exiting cleanly and then just staying `inactive (dead)` until a human
+noticed. Changed `Restart=on-failure` → `Restart=always` in both
+`deploy/llm-router.service` and the live unit; `systemctl stop`/`restart`
+are still honored correctly regardless of that policy (systemd tracks
+deliberate stops separately from the Restart= decision), so this only adds
+self-healing for the unexplained case.
+
+The signal's actual source is still open — live investigation and audit
+trail in `code_plans/router-unreachable-signal-investigation.md`. Confirmed
+so far: it is not `systemctl`, not the admin portal's restart trigger, not
+suspend/resume, not the OOM killer, and — via `auditd` — not delivered
+through the `kill` or `tgkill` syscalls either, which is why the watch was
+extended to `pidfd_send_signal` and the `rt_*sigqueueinfo` syscalls.
+
## Pointing a coding agent at it
The `/v1` endpoints are OpenAI-compatible, so any normal client works —
diff --git a/README.md b/README.md
index 37d617f..8578ac9 100644
--- a/README.md
+++ b/README.md
@@ -1,4 +1,7 @@
-# Local LLM Model Router
+
+
+ 6krrt — Local LLM Model Router
+
This router sits between a coding agent (opencode, an SDK, plain curl — any
OpenAI-compatible client) and the cloud LLMs it calls. It uses a local model
@@ -41,6 +44,7 @@ measurement is how you check whether it still holds for you.
- [Installation](#installation)
- [Local (`venv`)](#local-venv)
- [As a systemd service](#as-a-systemd-service)
+ - [Where Ollama lives](#where-ollama-lives)
- [Usage](#usage)
- [Route without spending anything](#route-without-spending-anything)
- [Skip the classifier when you already know the shape](#skip-the-classifier-when-you-already-know-the-shape)
@@ -71,13 +75,11 @@ measurement is how you check whether it still holds for you.
- [What routing actually returns, and why it moves](#what-routing-actually-returns-and-why-it-moves)
- [API Endpoints](#api-endpoints)
- [Logging and Traceability](#logging-and-traceability)
-- [Scheduled Jobs (systemd)](#scheduled-jobs-systemd)
- [Self-Eval Harness (`eval_proficiency.py`)](#self-eval-harness-eval_proficiencypy)
- [Classifier Reliability Notes](#classifier-reliability-notes)
- [Pointing a Coding Agent at It](#pointing-a-coding-agent-at-it)
- [Testing](#testing)
-- [Setup](#setup)
- - [Where Ollama lives](#where-ollama-lives)
+- [Advanced: Fitting local models into 24GB VRAM](#advanced-fitting-local-models-into-24gb-vram)
- [Known Limitations & Open Items](#known-limitations--open-items)
## Requirements
@@ -142,6 +144,33 @@ This now feeds `eco` only. Cost is priced per request from catalog prices, so
routing no longer depends on the sweep at all; disabling this timer costs you
carbon figures, not routing quality.
+### Where Ollama lives
+
+```bash
+ollama pull mistral-nemo:12b # or whatever you set as classifier.model
+```
+
+It does not have to be on the machine running the router; the box with the
+GPU usually isn't the laptop. To use one across a VPN, point **both**
+endpoints at it:
+
+```yaml
+classifier:
+ base_url: "http://:11434/v1"
+verification:
+ base_url: "http://:11434" # same host, so `model` can stay null
+```
+
+and apply `deploy/ollama-over-vpn.conf` on the serving host — Ollama binds
+`127.0.0.1` by default and will otherwise refuse. Bind it to the VPN address
+rather than `0.0.0.0`: Ollama has no authentication, so anything reaching the
+port can run inference and enumerate your models.
+
+Both endpoints move together because the verifier speaks Ollama's *native*
+API and cannot follow the classifier to a cloud provider. Config load refuses
+the case where they are on different hosts and `verification.model` is null,
+because that combination fails silently.
+
## Usage
### Route without spending anything
@@ -1070,6 +1099,37 @@ config as arguments, so the suite runs offline on a clean checkout.
| `test_classifier_input.py` | Classifier framing: `_previous_context` scope, `_classifier_user_content` framing |
| `test_route_decisions.py` | `route_decisions` table, inline-create helper, config gate |
+## Advanced: Fitting local models into 24GB VRAM
+
+Ollama loads `mistral-nemo:12b` and `qwen3-vl:4b` at the context size embedded in their library tags unless a custom tag overrides it. On a 24GB card that default context is expensive: `mistral-nemo:12b` came up at ~14GB (`num_ctx=32768`), and `qwen3-vl:4b` came up at ~9.4GB (`num_ctx=32768`). Running both together left only ~1.9GB of headroom.
+
+You cannot fix this per request. The OpenAI-compatible `/v1/chat/completions` endpoint in Ollama 0.22.0 was verified live with every field shape that seemed plausible, and in every case it returned 200 while silently keeping the loaded context unchanged. The context size has to be baked into the model tag itself with a `Modelfile`:
+
+```bash
+# Router uses this for classification and verification (same resident model)
+printf 'FROM mistral-nemo:12b\nPARAMETER num_ctx 8192\n' > Modelfile.router
+ollama create mistral-nemo-router:12b -f Modelfile.router
+
+# Vision fallback only
+printf 'FROM qwen3-vl:4b\nPARAMETER num_ctx 16384\n' > Modelfile.vision
+ollama create qwen3-vl-router:4b -f Modelfile.vision
+```
+
+`8192` is deliberately conservative for classification and verification. It leaves about 2× headroom over the real working set, and it keeps the classifier and verifier on the same resident instance rather than risking two differently-sized copies of the same base model. `16384` is a conservative cut for the raw message list the vision fallback receives, because that path is called before context pruning.
+
+With these tags the combined resident footprint was measured at roughly **15.9GB** in the worst case: classifier and verifier share one `mistral-nemo-router:12b` instance at ~8.6GB, with the `qwen3-vl-router:4b` vision fallback loaded alongside it. That leaves real headroom on a 24GB card.
+
+`config.yaml` already points at these tags by default:
+
+```yaml
+classifier:
+ model: mistral-nemo-router:12b
+verification:
+ model: mistral-nemo-router:12b
+local_vision:
+ model: qwen3-vl-router:4b
+```
+
## Known Limitations & Open Items
- **Leaderboard priors unfilled** — `leaderboards.yaml` ships empty. `python leaderboard.py --check`
diff --git a/admin.py b/admin.py
new file mode 100644
index 0000000..4c9fc2b
--- /dev/null
+++ b/admin.py
@@ -0,0 +1,802 @@
+"""Admin endpoints for the router: /admin/api/health and /admin/api/snapshot.
+
+This module MUST NOT import ``dispatcher`` — it mirrors ``metrics.py``'s
+contract (explicit ``(conn, cfg)`` args, no module-level globals, no import of
+the service module) and is mounted onto the dispatcher app by
+``dispatcher.app.include_router(admin.build_router(cfg, _db))``.
+
+Like ``metrics.py``, it never touches the TUI, never widens the bind, and
+never exposes prompts, session_dir, or conversation text. It is loopback-bound
+exactly like /health and /metrics.
+
+``build_router`` takes a ``_db_callable`` factory (returning an
+``sqlite3.Connection`` with ``row_factory`` set) so the module owns nothing
+global; the caller supplies its own connection source at mount time.
+"""
+
+from __future__ import annotations
+
+import asyncio
+import os
+import shutil
+import sqlite3
+import subprocess
+import sys
+import threading
+import time
+import uuid
+from collections.abc import Callable
+from datetime import datetime, timezone
+from pathlib import Path
+from typing import Any, List, Optional
+
+from fastapi import APIRouter, BackgroundTasks, HTTPException
+from fastapi.responses import FileResponse
+from openai import OpenAI, OpenAIError
+from pydantic import BaseModel, ValidationError
+from ruamel.yaml import YAML
+
+import metrics
+from config import FlexPreference, RouterConfig, load_config
+
+# Named /api/history ranges -> (span_seconds, bucket_seconds). Buckets floor
+# an observed_at to index = unix_seconds // bucket, so the representative ts of
+# each bucket is a whole multiple of bucket_seconds.
+_ADMIN_HISTORY_RANGES = {
+ "1h": (3600, 300),
+ "6h": (21600, 900),
+ "24h": (86400, 3600),
+ "7d": (604800, 21600),
+ "30d": (2592000, 86400),
+}
+
+# Column expression for the non-count energy series. All are energy_observations
+# sums; the keys line up with the /api/history response series.
+_HISTORY_ENERGY_SERIES = {
+ "requests_per_bucket": "COUNT(*)",
+ "cost_per_bucket": "COALESCE(SUM(cost_usd), 0)",
+ "energy_per_bucket": "COALESCE(SUM(energy_kwh), 0)",
+ "carbon_per_bucket": "COALESCE(SUM(carbon_g_co2eq), 0)",
+}
+
+
+def ensure_admin_tables(conn):
+ """Idempotently create the admin_model_overrides table.
+
+ Mirrors ``ensure_route_decisions`` in dispatcher.py: CREATE TABLE IF NOT
+ EXISTS + CREATE INDEX IF NOT EXISTS, both safe to call repeatedly.
+ """
+ conn.execute(
+ "CREATE TABLE IF NOT EXISTS admin_model_overrides ("
+ " model_id TEXT NOT NULL,"
+ " provider TEXT NOT NULL,"
+ " availability TEXT NOT NULL,"
+ " reason TEXT,"
+ " updated_at TEXT NOT NULL,"
+ " PRIMARY KEY (model_id, provider)"
+ ")"
+ )
+ conn.execute(
+ "CREATE INDEX IF NOT EXISTS idx_admin_model_overrides_availability "
+ "ON admin_model_overrides (availability)"
+ )
+ conn.commit()
+
+
+def _history_series(
+ conn: sqlite3.Connection,
+ span: int,
+ bucket: int,
+ table: str,
+ aggregate: str,
+) -> List[list]:
+ """Bucketed GROUP BY over *table* with the given *aggregate* expression.
+
+ Returns ``[[unix_ts, value], ...]`` ascending, one row per bucket that has
+ rows inside the last *span* seconds; buckets with no rows are omitted, so an
+ empty table yields ``[]``. ``table`` and ``aggregate`` are only ever module
+ literals (never user input), so they are safe to interpolate.
+ """
+ rows = conn.execute(
+ f"""
+ SELECT CAST((julianday(observed_at) - julianday('1970-01-01')) * 86400.0
+ / ? AS INTEGER) * ? AS bucket_ts,
+ {aggregate} AS value
+ FROM {table}
+ WHERE julianday(observed_at) >= julianday('now') - ? / 86400.0
+ GROUP BY CAST((julianday(observed_at) - julianday('1970-01-01'))
+ * 86400.0 / ? AS INTEGER)
+ ORDER BY bucket_ts
+ """,
+ (bucket, bucket, span, bucket),
+ ).fetchall()
+ return [[row["bucket_ts"], row["value"]] for row in rows]
+
+
+def _classifier_reachable(cfg: Any) -> bool:
+ """Whether the classifier endpoint answers a models.list() probe.
+
+ Replicates dispatcher's /health logic: ``api_key_env`` unset means the
+ unauthenticated local Ollama case (SDK still needs a key, so "ollama" is
+ used), and any OpenAIError (including a connection refusal) is treated as
+ unreachable.
+ """
+ key_env = cfg.classifier.api_key_env
+ api_key = os.environ.get(key_env) if key_env else "ollama"
+ client = OpenAI(base_url=cfg.classifier.base_url, api_key=api_key, max_retries=0)
+ try:
+ client.models.list()
+ return True
+ except OpenAIError:
+ return False
+
+
+def _health_check(conn: sqlite3.Connection, cfg: Any) -> dict:
+ """The same diagnostics /health reports, as a dict for snapshot to embed."""
+ counts = dict(
+ conn.execute(
+ """
+ SELECT 'models', COUNT(*) FROM models
+ UNION ALL SELECT 'routable', COUNT(*) FROM models
+ WHERE access_level = 'public' AND availability = 'active'
+ UNION ALL SELECT 'proficiency', COUNT(*) FROM proficiency
+ UNION ALL SELECT 'energy_observations', COUNT(*) FROM energy_observations
+ """
+ ).fetchall()
+ )
+ return {
+ "status": "ok",
+ "counts": counts,
+ "scoring": metrics.scoring_coverage(conn, cfg),
+ "classifier_reachable": _classifier_reachable(cfg),
+ "classifier_model": cfg.classifier.model,
+ "tiers": cfg.tiers,
+ "providers": list(cfg.dispatch_providers),
+ "api_keys_present": {
+ name: bool(os.environ.get(p.api_key_env))
+ for name, p in cfg.dispatch_providers.items()
+ },
+ }
+
+
+# Scalar fields a model detail exposes to the admin UI, renamed from the SQLite
+# column names to the JSON key names baked into the response contract. The
+# explicit rename list (rather than ``dict(row)``) guarantees we never leak
+# internal columns (e.g. poller/bookkeeping fields) into the response.
+_MODEL_FIELDS: dict[str, str] = {
+ "model_id": "model_id",
+ "provider": "provider",
+ "base_model_id": "base_model_id",
+ "display_name": "display_name",
+ "availability": "availability",
+ "tier": "tier",
+ "context_window": "context_window",
+ "effective_context_window": "effective_context_window",
+ "latency_class": "latency_class",
+ "reasoning_mode": "reasoning_mode",
+ "context_variant": "context_variant",
+ "access_level": "access_level",
+ "supports_tools": "supports_tools",
+ "supports_json_mode": "supports_json_mode",
+ "supports_vision": "supports_vision",
+ "supports_reasoning": "supports_reasoning",
+ "reasoning_default_enabled": "reasoning_default_enabled",
+ "cost_per_1m_prompt": "cost_per_1m_prompt",
+ "cost_per_1m_completion": "cost_per_1m_completion",
+}
+
+
+def _model_rows(conn: sqlite3.Connection) -> List[dict]:
+ """All models with their proficiency map, in the admin JSON contract.
+
+ Proficiency is a LEFT JOIN from ``models`` to ``proficiency`` grouped by
+ category so a model with no proficiency rows still appears (its map is
+ empty). Overrides from ``admin_model_overrides`` are merged so every
+ returned model object carries ``effective_availability`` (which may differ
+ from the DB ``availability``), ``is_overridden`` (bool), and the raw
+ ``availability`` column. Internal columns never escape.
+ """
+ # Load overrides into a (model_id, provider) -> availability lookup.
+ overrides = {}
+ for row in conn.execute(
+ "SELECT model_id, provider, availability FROM admin_model_overrides"
+ ):
+ overrides[(row["model_id"], row["provider"])] = row["availability"]
+
+ rows = conn.execute(
+ """
+ SELECT m.model_id,
+ m.provider,
+ m.base_model_id,
+ m.display_name,
+ m.availability,
+ m.tier,
+ m.context_window,
+ m.effective_context_window,
+ m.latency_class,
+ m.reasoning_mode,
+ m.context_variant,
+ m.access_level,
+ m.supports_tools,
+ m.supports_json_mode,
+ m.supports_vision,
+ m.supports_reasoning,
+ m.reasoning_default_enabled,
+ m.cost_per_1m_prompt,
+ m.cost_per_1m_completion,
+ p.category,
+ p.blended_score
+ FROM models m
+ LEFT JOIN proficiency p
+ ON p.model_id = m.model_id AND p.provider = m.provider
+ AND p.blended_score IS NOT NULL
+ ORDER BY m.model_id, m.provider, p.category
+ """
+ ).fetchall()
+
+ models: dict[tuple[str, str], dict] = {}
+ order: list[tuple[str, str]] = []
+ for row in rows:
+ key = (row["model_id"], row["provider"])
+ if key not in models:
+ entry = {
+ json_key: row[col]
+ for col, json_key in _MODEL_FIELDS.items()
+ }
+ for bool_col in (
+ "supports_tools",
+ "supports_json_mode",
+ "supports_vision",
+ "supports_reasoning",
+ "reasoning_default_enabled",
+ ):
+ entry[bool_col] = bool(entry[bool_col])
+ entry["proficiency"] = {}
+ override_avail = overrides.get(key)
+ entry["effective_availability"] = (
+ override_avail if override_avail else row["availability"]
+ )
+ entry["is_overridden"] = key in overrides
+ models[key] = entry
+ order.append(key)
+ if row["category"] is not None:
+ models[key]["proficiency"][row["category"]] = row["blended_score"]
+ return [models[key] for key in order]
+
+
+# --- runtime toggle knobs ------------------------------------------------------
+# Each boolean knob maps to a dotted attribute path on RouterConfig. The GET
+# response keys differ from the POST knob names only for circuit_breaker, which
+# is reported as a nested {"enabled": ...} object. Values are NEVER persisted
+# to config.yaml and never touch secrets, model names, or URLs.
+_BOOL_KNOBS: dict[str, tuple[str, ...]] = {
+ "log_route_decisions": ("logging", "log_route_decisions"),
+ "log_energy_observations": ("logging", "log_energy_observations"),
+ "circuit_breaker_enabled": ("circuit_breaker", "enabled"),
+ "local_llm_enabled": ("verification", "local_llm_enabled"),
+ "session_cache_enabled": ("session_cache", "enabled"),
+ "pinch_enabled": ("pinch", "enabled"),
+ "pinch_relevance_enabled": ("pinch", "relevance", "enabled"),
+}
+
+_FLEX_KNOB = "default_flex_preference"
+_FLEX_PATH: tuple[str, ...] = ("routing", "default_flex_preference")
+_FLEX_VALUES = frozenset(v.value for v in FlexPreference)
+
+# --- persisted config allowlist -------------------------------------------------
+# Dotted config.yaml paths an operator is allowed to edit. Everything else —
+# classifier/verification/local_vision URLs and model names, api_key_env,
+# provider blocks, dispatch_providers, and secrets — is deliberately OFF this
+# list. Editing works only on the named scalars. Writes go through
+# comment-preserving ruamel.yaml round-trip and are validated via ``RouterConfig``
+# before any byte reaches disk; a backup is made first.
+_CONFIG_ALLOWLIST: dict[str, tuple[str, ...]] = {
+ "logging.level": ("logging", "level"),
+ "objective.quality_tolerance": ("objective", "quality_tolerance"),
+ "objective.max_energy_per_request": ("objective", "max_energy_per_request"),
+ "objective.plan_kwh_per_period": ("objective", "plan_kwh_per_period"),
+ "circuit_breaker.enabled": ("circuit_breaker", "enabled"),
+ "session_cache.enabled": ("session_cache", "enabled"),
+ "verification.local_llm_enabled": ("verification", "local_llm_enabled"),
+ "pinch.enabled": ("pinch", "enabled"),
+ "pinch.relevance.enabled": ("pinch", "relevance", "enabled"),
+ "routing.default_flex_preference": ("routing", "default_flex_preference"),
+}
+
+# Order preserves config.yaml layout for the GET response.
+_CONFIG_GET_ORDER: list[str] = [
+ "logging.level",
+ "objective.quality_tolerance",
+ "objective.max_energy_per_request",
+ "objective.plan_kwh_per_period",
+ "circuit_breaker.enabled",
+ "session_cache.enabled",
+ "verification.local_llm_enabled",
+ "pinch.enabled",
+ "pinch.relevance.enabled",
+ "routing.default_flex_preference",
+]
+
+
+# Serializes the load/validate/backup/write sequence of the persisted-config
+# endpoints so concurrent writes cannot interleave and truncate config.yaml.
+_config_write_lock = threading.Lock()
+
+
+def _dict_get_at(store: Any, path: tuple[str, ...]) -> Any:
+ """Walk a dotted *path* (tuple) through a plain-nested mapping."""
+ cur = store
+ for part in path:
+ cur = cur[part]
+ return cur
+
+
+def _dict_set_at(store: Any, path: tuple[str, ...], value: Any) -> None:
+ """Set *value* at dotted *path* in a plain-nested mapping."""
+ cur = store
+ for part in path[:-1]:
+ cur = cur[part]
+ cur[path[-1]] = value
+
+
+def load_config_store(config_path: Path) -> Any:
+ """Load *config_path* as a ruamel CommentedMap so comments survive a dump."""
+ yaml = YAML()
+ yaml.preserve_quotes = True
+ return yaml.load(config_path.read_text())
+
+
+def _persist_config_value(
+ config_path: Path, path: tuple[str, ...], value: Any
+) -> None:
+ """Atomically persist *value* at dotted *path* in *config_path*.
+
+ The load/validate/backup/write sequence runs under the module-level
+ ``_config_write_lock`` so concurrent writes cannot interleave. The write is
+ a ruamel.yaml round-trip (comments survive), the WHOLE config is re-
+ validated via ``RouterConfig`` before any byte touches disk, and the on-
+ disk replacement uses a ``.tmp`` file + ``os.replace`` so a reader never
+ observes a truncated config.yaml. Raises ``ValidationError`` (leaving no
+ backup on disk) if the candidate config is invalid.
+ """
+ with _config_write_lock:
+ store = load_config_store(config_path)
+ _dict_set_at(store, path, value)
+ RouterConfig(**store)
+ backup = config_path.with_name(
+ f"config.yaml.bak.{int(time.time())}"
+ )
+ shutil.copyfile(config_path, backup)
+ tmp = config_path.with_suffix(config_path.suffix + ".tmp")
+ with tmp.open("w") as fh:
+ YAML().dump(store, fh)
+ os.replace(tmp, config_path)
+
+
+
+class _AvailabilityBody(BaseModel):
+ availability: str
+ reason: Optional[str] = None
+
+
+class _ValueBody(BaseModel):
+ """A single runtime knob value: a bool for boolean knobs, a flex string
+ for ``default_flex_preference``. Type-checked at validation time."""
+
+ value: Any
+
+
+def _get_at(cfg: Any, path: tuple[str, ...]) -> Any:
+ cur = cfg
+ for part in path:
+ cur = getattr(cur, part)
+ return cur
+
+
+def _set_at(cfg: Any, path: tuple[str, ...], value: Any) -> None:
+ cur = cfg
+ for part in path[:-1]:
+ cur = getattr(cur, part)
+ setattr(cur, path[-1], value)
+
+
+def _runtime_state(cfg: Any) -> dict:
+ """Read every toggle knob's current in-memory value off ``cfg``."""
+ return {
+ "log_route_decisions": _get_at(cfg, _BOOL_KNOBS["log_route_decisions"]),
+ "log_energy_observations": _get_at(
+ cfg, _BOOL_KNOBS["log_energy_observations"]
+ ),
+ "circuit_breaker": {
+ "enabled": _get_at(cfg, _BOOL_KNOBS["circuit_breaker_enabled"])
+ },
+ "local_llm_enabled": _get_at(cfg, _BOOL_KNOBS["local_llm_enabled"]),
+ "session_cache_enabled": _get_at(cfg, _BOOL_KNOBS["session_cache_enabled"]),
+ "pinch_enabled": _get_at(cfg, _BOOL_KNOBS["pinch_enabled"]),
+ "pinch_relevance_enabled": _get_at(
+ cfg, _BOOL_KNOBS["pinch_relevance_enabled"]
+ ),
+ "default_flex_preference": _get_at(cfg, _FLEX_PATH).value,
+ }
+
+
+def build_router(
+ cfg: Any,
+ _db_callable: Callable[[], sqlite3.Connection],
+ base_dir: Optional[str] = None,
+) -> APIRouter:
+ """Build the admin APIRouter bound to the caller's config and DB factory.
+
+ ``base_dir`` is the repo root (the directory holding config.yaml and the
+ maintenance scripts). The persisted-config endpoints use it to locate
+ ``config.yaml`` for comment-preserving writes; the maintenance triggers use
+ it as their spawn CWD. Defaults to this module's parent (the repo root), so
+ the router is portable and testable without an explicit base_dir.
+ """
+ _module_dir = Path(__file__).resolve().parent
+ config_path = (
+ Path(base_dir) / "config.yaml"
+ if base_dir is not None
+ else _module_dir / "config.yaml"
+ )
+ router = APIRouter()
+
+ _admin_frontend = _module_dir / "admin" / "frontend" / "index.html"
+
+ @router.get("/")
+ def admin_index() -> FileResponse:
+ return FileResponse(_admin_frontend, media_type="text/html")
+
+ @router.get("/api/health")
+ def admin_health() -> dict:
+ return {"status": "ok"}
+
+ @router.get("/api/models")
+ def admin_models() -> list:
+ """All models with per-category proficiency, for the admin model table."""
+ conn = _db_callable()
+ try:
+ return _model_rows(conn)
+ finally:
+ conn.close()
+
+ @router.get("/api/models/{model_id}/{provider}")
+ def admin_model_detail(model_id: str, provider: str) -> dict:
+ """A single model's detail (same shape as one /api/models element)."""
+ conn = _db_callable()
+ try:
+ for row in _model_rows(conn):
+ if row["model_id"] == model_id and row["provider"] == provider:
+ return row
+ raise HTTPException(status_code=404, detail="model not found")
+ finally:
+ conn.close()
+
+ @router.post("/api/models/{model_id}/{provider}/availability")
+ def admin_set_availability(
+ model_id: str,
+ provider: str,
+ body: _AvailabilityBody,
+ ) -> dict:
+ """Upsert an admin override for a model's availability.
+
+ Validates that availability is one of active/deprecated/stale.
+ Returns 404 if the model_id/provider pair does not exist in the
+ models table, 422 for an invalid availability value.
+ """
+ # Validate availability value (Pydantic only typed as str; keep explicit
+ # allow-list so unknown values get a clear 422).
+ valid_avail = {"active", "deprecated", "stale"}
+ avail_val = body.availability
+ reason_val = body.reason
+ if avail_val not in valid_avail:
+ raise HTTPException(
+ status_code=422,
+ detail=f"availability must be one of {sorted(valid_avail)}, got {avail_val!r}",
+ )
+
+ conn = _db_callable()
+ try:
+ # Check model exists.
+ exists = conn.execute(
+ "SELECT 1 FROM models WHERE model_id=? AND provider=?",
+ (model_id, provider),
+ ).fetchone()
+ if not exists:
+ raise HTTPException(status_code=404, detail="model not found")
+
+ now = datetime.now(timezone.utc).isoformat()
+ conn.execute(
+ """
+ INSERT INTO admin_model_overrides
+ (model_id, provider, availability, reason, updated_at)
+ VALUES (?, ?, ?, ?, ?)
+ ON CONFLICT(model_id, provider) DO UPDATE SET
+ availability = excluded.availability,
+ reason = excluded.reason,
+ updated_at = excluded.updated_at
+ """,
+ (model_id, provider, avail_val, reason_val, now),
+ )
+ conn.commit()
+ # Return the updated model detail (now carries effective_availability).
+ for row in _model_rows(conn):
+ if row["model_id"] == model_id and row["provider"] == provider:
+ return row
+ # Should not happen since we already checked existence.
+ raise HTTPException(status_code=500, detail="unexpected: row lost after upsert")
+ finally:
+ conn.close()
+
+ @router.delete("/api/models/{model_id}/{provider}/availability")
+ def admin_delete_availability(model_id: str, provider: str) -> dict:
+ """Delete an admin availability override for a model.
+
+ Returns 404 if the model does not exist. If no override row exists
+ for this model, still returns the model detail (no-op from routing
+ perspective).
+ """
+ if isinstance(model_id, bytes):
+ model_id = model_id.decode()
+ if isinstance(provider, bytes):
+ provider = provider.decode()
+
+ conn = _db_callable()
+ try:
+ # Check model exists.
+ exists = conn.execute(
+ "SELECT 1 FROM models WHERE model_id=? AND provider=?",
+ (model_id, provider),
+ ).fetchone()
+ if not exists:
+ raise HTTPException(status_code=404, detail="model not found")
+
+ conn.execute(
+ "DELETE FROM admin_model_overrides WHERE model_id=? AND provider=?",
+ (model_id, provider),
+ )
+ conn.commit()
+ for row in _model_rows(conn):
+ if row["model_id"] == model_id and row["provider"] == provider:
+ return row
+ raise HTTPException(status_code=500, detail="unexpected: row lost")
+ finally:
+ conn.close()
+
+ @router.get("/api/snapshot")
+ def admin_snapshot() -> dict:
+ """Fused metrics + health + recent-history summary for the admin UI.
+
+ Covers analytics (quota, coverage, per-model aggregates, verdict mix,
+ top proficiency) plus a liveness/health snapshot, timestamped as
+ ``generated_at``.
+ """
+ conn = _db_callable()
+ try:
+ return {
+ "quota": metrics.quota_burn(conn, cfg),
+ "coverage": metrics.scoring_coverage(conn, cfg),
+ "recent_decisions": metrics.recent_decisions(conn),
+ "per_model": metrics.per_model(conn),
+ "verdict_mix": metrics.verdict_mix(conn),
+ "top_proficiency": metrics.top_proficiency(conn, "coding_general"),
+ "health": _health_check(conn, cfg),
+ "generated_at": datetime.now(timezone.utc).isoformat(),
+ }
+ finally:
+ conn.close()
+
+ @router.get("/api/history")
+ def admin_history(range: str = "24h") -> dict:
+ """Bucketed time-series over energy/decision tables for the admin UI.
+
+ ``range`` selects a (span, bucket) pair from ``_ADMIN_HISTORY_RANGES``.
+ Returns five series (decisions, requests, cost, energy, carbon) as
+ ``[unix_ts, value]`` pairs occupying buckets that had rows inside the
+ span; empty tables yield ``[]`` for every series. Raises 400 for any
+ unlisted range value.
+ """
+ if range not in _ADMIN_HISTORY_RANGES:
+ raise HTTPException(
+ status_code=400,
+ detail=f"unknown history range: {range!r}",
+ )
+ span, bucket = _ADMIN_HISTORY_RANGES[range]
+ conn = _db_callable()
+ try:
+ return {
+ "decisions_per_bucket": _history_series(
+ conn, span, bucket, "route_decisions", "COUNT(*)"
+ ),
+ **{
+ name: _history_series(
+ conn, span, bucket, "energy_observations", aggregate
+ )
+ for name, aggregate in _HISTORY_ENERGY_SERIES.items()
+ },
+ }
+ finally:
+ conn.close()
+
+ # --- operational triggers ------------------------------------------------
+ # Maintenance scripts run as child processes. Running them with
+ # ``asyncio.create_subprocess_exec`` (never ``subprocess.run`` inline) keeps
+ # the event loop responsive. Commands run from the repo root (the directory
+ # holding admin.py) so relative ``config.yaml`` / ``router.db`` resolve.
+ _repo_root = str(Path(__file__).resolve().parent)
+
+ def _job(command: str, status: str, returncode, output_tail: str) -> dict:
+ """The documented job object returned by every trigger endpoint."""
+ return {
+ "id": uuid.uuid4().hex[:12],
+ "command": command,
+ "status": status,
+ "returncode": returncode,
+ "output_tail": output_tail,
+ }
+
+ async def _run_steps(
+ steps: list[list[str]], timeout: float, cwd: str
+ ) -> dict:
+ """Run maintenance steps sequentially (``&&`` semantics), return a job.
+
+ Each step is spawned with ``asyncio.create_subprocess_exec``; output
+ (stdout merged with stderr) is capped to the last 2048 chars for
+ ``output_tail``. A single wall-clock deadline covers the whole chain,
+ so a slow poller cannot quietly eat the whole budget and then leave a
+ late tier step hanging.
+ """
+ command = " && ".join(" ".join(s) for s in steps)
+ loop = asyncio.get_running_loop()
+ deadline = loop.time() + timeout
+ collected: list[str] = []
+ proc = None
+
+ for step in steps:
+ try:
+ proc = await asyncio.create_subprocess_exec(
+ *step,
+ cwd=cwd,
+ stdout=asyncio.subprocess.PIPE,
+ stderr=asyncio.subprocess.STDOUT,
+ )
+ except Exception as exc: # noqa: BLE001 - top-level spawn boundary
+ return _job(
+ command, "failed", None, f"spawn failed ({step[0]}): {exc}"[:2048]
+ )
+
+ remaining = deadline - loop.time()
+ try:
+ out, _ = await asyncio.wait_for(
+ proc.communicate(), timeout=max(0.0, remaining)
+ )
+ except asyncio.TimeoutError:
+ proc.kill()
+ await proc.wait()
+ return _job(
+ command,
+ "timed_out",
+ None,
+ f"Command timed out after {int(timeout)}s",
+ )
+
+ if out:
+ collected.append(out.decode("utf-8", errors="replace"))
+ if proc.returncode != 0:
+ break # && semantics: stop on the first failing step
+
+ output = "".join(collected)
+ status = "success" if proc is not None and proc.returncode == 0 else "failed"
+ return _job(command, status, proc.returncode if proc else None, output[-2048:])
+
+ @router.post("/api/refresh-catalog")
+ async def admin_refresh_catalog() -> dict:
+ """Run ``poller.py && tier.py`` to refresh the model catalog."""
+ return await _run_steps(
+ [[sys.executable, "poller.py"], [sys.executable, "tier.py"]],
+ timeout=120.0,
+ cwd=_repo_root,
+ )
+
+ @router.post("/api/seed-energy")
+ async def admin_seed_energy(samples: int = 5) -> dict:
+ """Sweep the energy reference workload for ``samples`` per model."""
+ return await _run_steps(
+ [[sys.executable, "seed_energy.py", "--samples", str(samples)]],
+ timeout=600.0,
+ cwd=_repo_root,
+ )
+
+ @router.post("/api/apply-feedback")
+ async def admin_apply_feedback(dry_run: bool = False) -> dict:
+ """Fold verification outcomes back into proficiency (optionally dry)."""
+ argv = [sys.executable, "feedback.py"]
+ if dry_run:
+ argv.append("--dry-run")
+ return await _run_steps([argv], timeout=120.0, cwd=_repo_root)
+
+ @router.post("/api/restart-service")
+ def admin_restart_service(background: BackgroundTasks) -> dict:
+ """Schedule a systemctl restart and return immediately.
+
+ The systemctl call runs as a post-response BackgroundTask (fire-and-
+ forget) so the HTTP response flushes before the service process dies.
+ It is deliberately never awaited inline.
+ """
+ background.add_task(_restart_service)
+ return {"status": "restarting"}
+
+ def _restart_service() -> None:
+ subprocess.run(
+ ["systemctl", "--user", "restart", "llm-router.service"],
+ check=False,
+ capture_output=True,
+ )
+
+ @router.get("/api/runtime")
+ def admin_runtime_state() -> dict:
+ """Persisted (config.yaml) vs runtime (in-memory) value of each knob."""
+ persisted = _runtime_state(load_config("config.yaml"))
+ runtime = _runtime_state(cfg)
+ return {
+ key: {"persisted": persisted[key], "runtime": runtime[key]}
+ for key in persisted
+ }
+
+ @router.post("/api/runtime/{knob}")
+ def admin_set_runtime_knob(knob: str, body: _ValueBody) -> dict:
+ """Flip a single knob's in-memory value; config.yaml is never written."""
+ if knob in _BOOL_KNOBS:
+ if not isinstance(body.value, bool):
+ raise HTTPException(
+ status_code=422,
+ detail=f"{knob} expects a boolean value",
+ )
+ _set_at(cfg, _BOOL_KNOBS[knob], body.value)
+ return {"ok": True, knob: body.value}
+ if knob == _FLEX_KNOB:
+ if body.value not in _FLEX_VALUES:
+ raise HTTPException(
+ status_code=422,
+ detail=(
+ f"{knob} must be one of {sorted(_FLEX_VALUES)}, "
+ f"got {body.value!r}"
+ ),
+ )
+ _set_at(cfg, _FLEX_PATH, FlexPreference(body.value))
+ return {"ok": True, knob: body.value}
+ raise HTTPException(
+ status_code=400,
+ detail=f"unknown runtime knob: {knob}",
+ )
+
+ @router.get("/api/config")
+ def admin_config_get() -> dict:
+ """The allowlisted config.yaml values, as {dotted_key: value}."""
+ return {key: _dict_get_at(load_config_store(config_path), path)
+ for key, path in _CONFIG_ALLOWLIST.items()}
+
+ @router.post("/api/config/{key}")
+ def admin_config_write(key: str, body: _ValueBody) -> dict:
+ """Persist one allowlisted value to config.yaml (comment-preserving)."""
+ if key not in _CONFIG_ALLOWLIST:
+ raise HTTPException(
+ status_code=403,
+ detail=f"config key is not editable: {key}",
+ )
+ try:
+ _persist_config_value(
+ config_path, _CONFIG_ALLOWLIST[key], body.value
+ )
+ except ValidationError as exc:
+ raise HTTPException(
+ status_code=422,
+ detail=exc.errors()[0]["msg"],
+ ) from exc
+ return {
+ "key": key,
+ "value": body.value,
+ "message": "A restart is required for this change to take effect",
+ }
+
+ return router
diff --git a/admin/frontend/__init__.py b/admin/frontend/__init__.py
new file mode 100644
index 0000000..e69de29
diff --git a/admin/frontend/index.html b/admin/frontend/index.html
new file mode 100644
index 0000000..fd91224
--- /dev/null
+++ b/admin/frontend/index.html
@@ -0,0 +1,916 @@
+
+
+
+
+
+Admin Dashboard — LLM Router
+
+
+
+
+
+
+
+
admin@router ▸ dashboard
+
+ connecting…
+
+
—
+
+
+
+
+
+
+
⚡ Quota Meter
+
Loading…
+
+
+
◈ Model Availability
+
+
+
model
provider
tier
status
override
+
Loading…
+
+
+
+
+
+
+
+
⟁ Recent Decisions
+
+
+
time
kind
category
tier
model
cost
prof
+
+
+
+
+
▐ Per-Model Usage
+
+
Loading…
+
+
+
+
+
+
+
+
◉ Verdict Mix
+
+
+
+
+
◆ Category Breakdown
+
+
Loading…
+
+
+
+
⚠ Warnings
+
+
None
+
+
+
+
+
+
+
+
+ ◈ History
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
⚙ Controls
+
+
+
+
Operational Triggers — fire-and-forget maintenance jobs
+
+
+
+
+
+
+
+
+
+
+
+
Runtime Knobs — toggle in-memory; no config.yaml write
+
+
+
+
+
+
Persisted Config — allowlisted keys only; persisted to config.yaml
``
+- `addDecision`, the live SSE row-append path (`index.html:403`): identical
+ unescaped interpolation.
+
+Both are `tbody.innerHTML = ...` / `insertBefore` on a freshly-built
+`
`, so this is a direct HTML injection, not merely an attribute-quoting
+issue. The same field, from the same data source, *is* escaped one panel
+over in `renderCategoryBreakdown` (`index.html:577`,
+`escapeHtml(cat)`) — confirming this is a missed spot rather than a
+considered choice, since `escapeHtml` already exists in the file
+(`index.html:884-886`) and is used correctly in `renderModels`,
+`renderConfig`, and `renderWarnings`.
+
+**Concrete reproduction:**
+
+```bash
+curl -s localhost:8080/route -H 'content-type: application/json' -d '{
+ "task": "x", "task_tier": 1, "required_context_tokens": 0,
+ "task_category": ""
+}'
+```
+
+The next time an operator has `/admin/` open — including passively, since
+the live decisions table updates via SSE without a page reload — the
+`onerror` fires in the page that already holds every admin capability:
+`restart-service`, allowlisted `config.yaml` writes, and model-availability
+overrides, all same-origin `fetch()` calls away with no additional
+credential needed. This is why the auth/CSRF gap in item 2 above is worse
+than it looks in isolation: an `Origin` check on the write endpoints
+wouldn't stop this, because the malicious request originates from
+JavaScript already running on `localhost:8080` itself.
+
+No test in the 9 new test files touches HTML escaping or the dashboard's JS
+at all (`tests/test_admin_frontend.py` only checks the file is served and
+contains the Chart.js tag) — consistent with the orchestration report's "F3
+Real manual QA: APPROVE" not having exercised this path.
+
+**Fix**: run `task_category` (and `rejected_reason`, `runner_up_models`, and
+any other free-text decision field) through `escapeHtml` in `renderDecisions`
+and `addDecision`, matching the pattern already used three other places in
+the same file. Validating `task_category` server-side against
+`classifier.py`'s known category set would also close it and is probably
+worth doing anyway — an unvalidated free-form override string being usable
+to skip the classifier is a second, smaller thing worth a look, independent
+of the rendering fix.
+
+### Second new finding, caught live in production rather than in review: concurrent config writes actually crash, and can transiently truncate `config.yaml`
+
+The gap analysis's F10 ("concurrent config.yaml edits are a race condition")
+was accepted as deferred/"Recommended" on the theory that it's a rare,
+low-stakes race. It isn't rare: `saveAllConfig()`
+(`admin/frontend/index.html:778-803`) fires every allowlisted key as a
+**parallel** `Promise.all` of `POST /admin/api/config/{key}` calls, so the
+one-click "Save All Config" button — the documented way to use this
+feature — guarantees concurrent writes on every use, not just under load.
+
+Caught live on this host at 02:35:34-36 on 2026-08-29: two of the ten
+parallel writes came back `500 Internal Server Error`
+(`pinch.enabled`, `routing.default_flex_preference`), both
+`TypeError: 'NoneType' object is not subscriptable` at `admin.py:332`
+(`_dict_set_at`) — meaning `load_config_store` (`admin.py:336-340`) returned
+`None` because it read `config.yaml` while a concurrent request's
+`config_path.open("w")` (`admin.py:772`) had already truncated the file to
+zero bytes and hadn't finished writing yet. Direct evidence this isn't
+theoretical: one of the timestamped backups made during that window,
+`config.yaml.bak.1787985334`, is itself **0 bytes** — `shutil.copyfile`
+(`admin.py:771`) copied `config.yaml` mid-truncation, so the safety backup
+this feature exists to provide was itself corrupted for that request.
+`config.yaml` itself survived intact this time (the last writer to finish
+happened to write a complete, valid file), but nothing in the current code
+guarantees that outcome — the write is an in-place `open("w")` truncate,
+not a write-to-temp-file-then-rename, so a crash or a slower writer mid-race
+could leave `config.yaml` empty on disk with no valid backup to recover
+from.
+
+**Fix**: two independent things, either one closes most of the risk:
+(1) make the write atomic — write to `config.yaml.tmp` and `os.replace()`
+over the real path, so a reader/writer never observes a truncated file; (2)
+serialize config writes with a lock (a module-level `threading.Lock` in
+`admin.py` is enough for a single-process service) so overlapping requests
+queue instead of interleaving. The frontend's `saveAllConfig()` sending
+requests sequentially instead of via `Promise.all` would also avoid
+triggering this in the one place it's currently guaranteed to happen, but
+the server-side fix is the one that actually closes the bug — a future
+`curl` script or a second browser tab hitting `/admin/api/config/*`
+concurrently would reproduce it regardless of what the shipped frontend
+does.
+
+## Recommendation
+
+Ship the fix for the XSS finding before this goes live somewhere reachable
+by more than one trusted user — it's a small, mechanical change
+(two `escapeHtml` calls) with an exact reproduction above to verify against.
+Everything else checked against the plan review's required list is
+genuinely done, not just documented as done: the F1 routing-integration test
+in particular is exactly the test that review asked for, not a weaker
+substitute. The auth/CSRF decision is still owed a real paragraph (even one
+sentence: "accepted, personal single-operator tool, revisit if this becomes
+multi-user") rather than the pre-existing read-only-surface sentence carried
+forward unchanged — worth doing at the same time as the XSS fix since they're
+the same conversation.
diff --git a/code_reviews/router-admin-portal-plan-review.md b/code_reviews/router-admin-portal-plan-review.md
new file mode 100644
index 0000000..8d453a0
--- /dev/null
+++ b/code_reviews/router-admin-portal-plan-review.md
@@ -0,0 +1,215 @@
+# Review: `router-admin-portal` work plan (pre-implementation)
+
+**What it was reviewing:** `code_plans/.omo/plans/router-admin-portal.md` —
+a 10-todo plan to add an integrated `/admin` FastAPI web portal (dashboards
++ operator controls: refresh catalog, reseed energy, restart service,
+runtime toggles, allowlisted config writes, model-availability overrides).
+Reviewed before any implementation exists, against the actual current code
+rather than the plan's own descriptions, and cross-checked against
+`code_plans/admin-portal-gap-analysis.md` (a 15-finding critique of an
+earlier draft, `code_plans/.omo/drafts/router-admin-portal.md`) to see
+whether the final plan actually incorporated what that analysis found.
+
+## Verdict: citations are excellent; one likely functional gap and one unexamined security assumption before this should be approved as-is
+
+### Citation accuracy — verified, essentially flawless
+
+Checked every file:line reference that matters for correctness, not a
+sample:
+
+- `dispatcher.py:154` (`cfg`), `:278` (`_db`), `:1306`/`:1353`/`:1417`
+ (`/health`/`/metrics`/`/events/decisions`) — all exact.
+- `config.py:185` (`default_flex_preference: FlexPreference =
+ FlexPreference.auto`) — confirmed real field, and every one of its four
+ cited use sites (`dispatcher.py:815, 2364, 2446` plus the docstring
+ mention at `:186`) — exact.
+- Todo 5's entire runtime-toggle line list — eleven citations across seven
+ config knobs (`circuit_breaker.enabled` at 2556/2573/2820/2823,
+ `log_route_decisions` at 935, `log_energy_observations` at 1261,
+ `verification.local_llm_enabled` at 2653/2794, `session_cache.enabled` at
+ 2266/2336, `pinch.enabled` at 2247/2475, the `pinch.relevance` gate at
+ 1827) — every single one checked out exact via direct grep.
+
+One trivial slip: Todo 3 cites `dispatcher.py:2126-2125` for `/v1/models` —
+a backwards range (start line > end line); the actual decorator is a single
+line at `2126`. Cosmetic, not worth blocking on.
+
+### The admin override table may not actually affect routing — verify before approving
+
+Todo 7 creates `admin_model_overrides(model_id, provider, availability,
+reason, updated_at)` specifically to survive `poller.py` overwriting
+`models.availability` on every catalog refresh. That part is correctly
+diagnosed: `poller.py`'s upsert really does set `availability = "active"`
+on every routable row it sees, per the gap analysis's F1 (confirmed
+independently, not taken on the gap analysis's word).
+
+But Todo 7's own text only says `GET /admin/api/models` **merges** the
+override into what it *returns to the dashboard*. Nowhere across Todos
+1–10 is there a step that feeds `admin_model_overrides` into
+`routing.select_candidates` / `routing.rejection_reason` — the functions
+that actually decide what a live request routes to. As specified, a model
+marked "deprecated" through the admin UI would show as deprecated on the
+dashboard while the router keeps dispatching to it exactly as before,
+because the hard filters still only read `models.availability` /
+`models.deprecated`, which the override table never touches.
+
+This is the same shape of defect the last review caught in
+`circuit_breaker.record_success`: a feature that is structurally complete
+— table, API, tests — but never reaches the code path it exists to affect.
+It matters more here, because the stated purpose of this whole control
+("mark a flaky model out of rotation") is exactly what was done by hand
+with a raw `UPDATE models SET deprecated = 1 ...` earlier this session when
+`gemma-4-31b` started failing — if the admin UI's version of that same
+action doesn't reach `routing.py`, the feature doesn't do the thing it's
+being built for. Add an explicit todo (or fold into Todo 7's acceptance
+criteria) requiring `route()`'s candidate-filtering to consult the override
+table the same way `_open_circuits` already consults `circuit_breaker.is_down`
+for the circuit-breaker exclusion set, and add a test that actually routes
+a request against an overridden model and asserts it's excluded — not just
+that the API reports it correctly.
+
+### Loopback-only/no-auth is inherited from a read-only posture without being re-examined for a read-write one
+
+The plan states "loopback-only and unauthenticated, matching the router's
+current security posture" as an already-made decision. The *current*
+posture (`/metrics`, `/health`, SSE decisions) is read-only observability,
+and loopback-only is a reasonable ACL for that. This plan adds `POST
+/admin/api/restart-service`, on-disk config writes, and subprocess
+triggers — a materially larger blast radius reachable by the same
+unauthenticated loopback bind. Loopback-only does not mean "only the
+operator can reach it": with no auth and no CSRF protection specified
+anywhere in the plan, any process running as the same user, or any webpage
+the operator's browser visits while the router happens to be running, can
+POST to `localhost:8080/admin/api/restart-service` via a plain HTML form —
+no preflight required to block a same-origin-policy-naive form POST. This
+is a known category of local-dev-server attack (CSRF/DNS-rebinding against
+`localhost`) that a pure read-only surface never had to consider. Worth a
+deliberate decision — even if the answer stays "loopback-only is
+acceptable for a personal single-operator tool" — rather than carrying the
+old answer forward because it was already the answer for a different kind
+of surface.
+
+Separately, Todo 4's restart-service mechanism is stated aspirationally
+("is expected to terminate the serving process; return a `restarting`
+status before the process exits") without specifying how the HTTP response
+gets flushed to the client before `systemctl --user restart` causes SIGTERM
+to land on the same process handling that request. If the endpoint `await`s
+the systemctl subprocess inline, the response can't be sent before the
+process that would send it is torn down. This needs an explicit
+fire-and-forget mechanism (e.g., schedule the restart via `BackgroundTasks`
+so it runs after the response is already flushed) named in the plan, not
+left as an implementation detail to be discovered while writing Todo 4.
+
+### Three "Recommended" gap-analysis findings are silently absent, not declared as deferred
+
+F9 (no audit trail for admin actions), F10 (no lock against concurrent
+config.yaml writes), and F11 (no endpoint to clear circuit-breaker/session
+cache state) are all "Recommended" rather than "Must-fix: Yes" in the gap
+analysis's own severity table, and none appear in the final plan's
+Must-have list. That's a defensible call for a personal, single-operator
+tool — but none of the three are named in the plan's "Must NOT have"
+section either, which is where deliberate exclusions are supposed to live.
+As written, a reader of the final plan alone has no way to tell "considered
+and deferred" from "the gap analysis was never actually read." Move them
+there explicitly, even as a one-line "deferred: F9/F10/F11, personal-use
+tool, revisit if this gets multi-operator" note.
+
+### Cross-repo citations also check out — including the parts I could actually verify
+
+`daashbrd` exists locally (`/home/alee/Sources/daashbrd`), so the plan's
+references to it aren't unverifiable hand-waving. Checked directly:
+`app/history.py` and `app/frontend/index.html` (2303 lines, matching Todo
+12's "fine for a single-page <3k line file" claim in the gap analysis)
+exist as cited; `app/main.py:154` is exactly the `StaticFiles.mount(...)`
+call and `:303` is exactly the `FileResponse(...)` call the gap analysis's
+F12 cited; `tests/test_main.py:23` and `:29` are exactly
+`test_frontend_dir_is_absolute_and_exists` and
+`test_static_mount_points_to_frontend_dir`, matching "mount checks."
+`poller.py`'s `main()`/`__main__` at 298/319, `feedback.py` at 143/171,
+`tier.py` at 82/92, and `schema.sql`'s `models` table (line 7) and its
+three indexes (244-246) all match Todo 4's and Todo 7's citations exactly
+as well. At this point essentially every checkable citation in the entire
+plan has been verified — this is the most citation-accurate plan reviewed
+in this loop so far.
+
+### Todo 2's history store duplicates data the router already durably records
+
+`admin_history.py` is specified as a **separate SQLite file**
+(`router_admin_history.db`) fed by a **background task sampling every
+60s** — a pattern lifted directly from `daashbrd`, which doesn't have
+anything better to query. This router already isn't in that position:
+`energy_observations` records `energy_kwh`, `cost_usd`, and
+`carbon_g_co2eq` **per request, as it happens**, and `route_decisions`
+records one row per routing decision — both timestamped, both already
+durable. `metrics.py`'s existing `/metrics` endpoint already computes
+rolling aggregates from `energy_observations` on demand, no sampling loop
+required.
+
+A bucketed time-series `GET /admin/api/history` can almost certainly be a
+`GROUP BY`-bucketed SQL query against the tables that already exist,
+computed on request rather than sampled every 60s into a second database.
+That would avoid three problems the current spec creates and doesn't
+solve: (1) a background-task lifecycle question that Todo 2 doesn't
+actually answer — `admin.py` must never import `dispatcher` per Todo 1, so
+nothing in the plan specifies *what* schedules this recurring task against
+the running event loop or on which object's startup hook; (2) a second
+SQLite file whose freshness is bounded by a 60s sampling interval instead
+of being exact; (3) genuinely duplicated data. Given this router's real
+traffic volume (one personal deployment; `config.yaml`'s own comments
+mention total billed traffic to date as $0.07), query performance against
+the existing tables is not a concern that justifies pre-aggregation. Worth
+reconsidering before Wave 1, since it's foundational to Todo 2's whole
+shape.
+
+### `ruamel.yaml` is a new dependency with no todo to pin it
+
+Todo 6 commits to "ruamel.yaml round-trip preservation" for config writes,
+but `ruamel.yaml` is not in `requirements.txt` today (checked directly —
+zero matches) and no todo in the plan adds it. `CLAUDE.md`'s own README
+section is explicit and deliberate about this: dependencies are pinned
+specifically so "a service that restarts on boot shouldn't change its
+dependency tree underneath itself," and bumps are meant to be deliberate.
+Introducing a new library for a real reason (comment-preserving YAML
+writes is the right tool for what F2 needed) is a legitimate, deliberate
+addition — but it needs its own explicit line in `requirements.txt` with a
+pinned version, and a todo (or an amendment to Todo 6) that says so, rather
+than being assumed available.
+
+### Minor: "single self-contained HTML file" and a `/admin/static` mount say two different things
+
+Todo 8 describes "a single self-contained HTML file using Chart.js from
+CDN" (matching gap-analysis F12's option (a): one `HTMLResponse` string,
+no static mount needed). Todo 9 then separately mounts
+`admin/frontend/` under `/admin/static`. If the page is genuinely
+self-contained with only a CDN script tag, there's nothing left to serve
+under a static mount and it can be dropped; if there *are* separate static
+assets planned (a favicon, a stylesheet), "self-contained" should say so.
+Small inconsistency, easy to resolve either direction — flagging so
+whichever way it goes is a decision rather than a leftover from copying
+daashbrd's shape wholesale.
+
+## Recommendation
+
+Don't approve as-is. Four things need an answer before implementation
+starts, not after — none of them require redoing work already done well:
+
+1. Confirm — or add a todo requiring — that `admin_model_overrides`
+ actually reaches `routing.select_candidates`, since without that the
+ model-management feature doesn't do what it's for.
+2. Make an explicit, stated decision about auth/CSRF for the new
+ write-capable endpoints rather than inheriting the read-only surface's
+ answer unexamined, and specify the restart-service
+ response-before-death mechanism concretely.
+3. Reconsider Todo 2's separate sampled history database against just
+ querying `energy_observations`/`route_decisions` directly — it's
+ simpler, exact instead of 60s-stale, and doesn't leave the background-
+ task lifecycle question unanswered.
+4. Add the missing `requirements.txt` pin for `ruamel.yaml` (or a todo that
+ does), matching this project's own stated policy on deliberate,
+ pinned dependency changes.
+
+The citation work underneath all of this is excellent — genuinely the most
+accurate plan reviewed in this loop so far, including cross-repo references
+into `daashbrd` that all checked out — and doesn't need rework. This is a
+plan worth sending back for one more pass on the four points above, not a
+rewrite.
diff --git a/deploy/llm-router.service b/deploy/llm-router.service
index 548dad2..54aa58f 100644
--- a/deploy/llm-router.service
+++ b/deploy/llm-router.service
@@ -16,9 +16,9 @@ WorkingDirectory=%h/llm-router
# Holds NEURALWATT_API_KEY. Create it with:
# echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env
EnvironmentFile=%h/llm-router/.env
-ExecStart=%h/llm-router/.venv/bin/uvicorn dispatcher:app --host 127.0.0.1 --port 8080
+ExecStart=%h/llm-router/.venv/bin/uvicorn dispatcher:app --host 127.0.0.1 --port 8080 --timeout-graceful-shutdown 5
-Restart=on-failure
+Restart=always
RestartSec=5s
# Loopback only by default — this service holds a billable API key and has no
diff --git a/documentation_plans/readme-restructure.md b/documentation_plans/readme-restructure.md
new file mode 100644
index 0000000..2f0f11d
--- /dev/null
+++ b/documentation_plans/readme-restructure.md
@@ -0,0 +1,191 @@
+# Spec: restructure README.md around a pitch → features → install → per-command usage shape
+
+**Origin.** Modeled on the structure of
+[`dsplce-co/supabase-plus`'s README](https://raw.githubusercontent.com/dsplce-co/supabase-plus/refs/heads/master/README.md)
+— its section shape and pacing, not its content or voice. `supabase-plus` is
+a Rust CLI tool with crates.io badges, install-method sub-sections, and
+punchy, jokey per-command usage blurbs; none of that is this project. What's
+being borrowed is purely structural: a short pitch up front, a scannable
+features list, a Table of Contents for a doc long enough to need one,
+installation broken into named sub-methods, and — the part worth the most —
+usage organized **per command**, each with its own short "why you'd want
+this" motivation before the command itself, rather than one long undifferentiated
+block of curl examples.
+
+This project's own voice stays: dry, evidence-first, "measured live on
+[date]" citations, numbers with dates on them. Nothing here asks for new
+jokes, new claims, or new numbers — only reorganizing what's already true in
+the current 1039-line `README.md` into a shape a first-time reader can
+actually navigate.
+
+No README changes accompany this document — this is the spec opencode
+builds from.
+
+---
+
+## What's being borrowed, what's being skipped, and why
+
+| Reference element | Verdict | Reasoning |
+|---|---|---|
+| Org attribution banner (top link) | **Skip** | No org; this is a personal project. |
+| Badges row (crates.io, license, version) | **Skip, or minimal** | Nothing here is published to a package registry. A license badge only makes sense once a license exists (see Open Decisions below) — don't badge something that isn't true yet. |
+| One-line pitch + short elaboration paragraph | **Adopt** | The current README already has this content (see §1 below) — it's just buried under a heading instead of leading the file. |
+| Italic disclaimer right after the pitch | **Adopt, re-aimed** | The reference's is a trademark disclaimer. This project's real analog already exists and is more load-bearing: *"Numbers in this README are measurements, not specifications"* (current README, "What This Is"). Same structural slot — a caveat before anyone trusts a number — different content. |
+| Demo GIF | **Optional, not this pass** | `tui.py`'s live dashboard is the natural candidate, but capturing one is a manual terminal recording step, not something a text plan can produce. Leave a placeholder comment (``) rather than skip the idea entirely. |
+| `## Features` — punchy bulleted list | **Adopt, own voice** | See §2. The existing "At a Glance" table already IS a features summary, just in dry-table form instead of scannable bullets. |
+| `---` / Table of Contents | **Adopt, straightforwardly** | The current README has **no TOC at all** across ~25 major sections and 1039 lines. This is the single highest-value, lowest-risk item in this whole plan — pure navigation, no content decisions required. |
+| `## Installation` with named sub-methods | **Adopt, honestly scoped** | The reference has 6 real install methods (nix/cargo/homebrew/deb/apt/aur). This project has exactly one (`venv` + `pip`) plus a systemd deployment path. Structure as two sub-sections — "Local" and "As a systemd service" — not six fake ones. Do not invent install methods that don't exist to mimic the reference's breadth. |
+| `## Usage` — one sub-heading per command, motivation → command → bullets | **Adopt — the main point of this plan** | See §3. This is the biggest actual improvement available: the current "Quick Usage" section is 9 curl commands in one code block with almost no narrative, while the *reasons* for each one are already written elsewhere in the doc, disconnected from the command they justify. |
+| `## Requirements` | **Adopt** | Already exists as a sentence in "Setup" ("You need: a Neuralwatt API key, Python 3.10+...") — promote it to its own short section, matching the reference's terse bulleted form. |
+| `## Repo & Contributions` | **Needs a decision, see below** | The reference is a public GitHub project soliciting PRs. This repo's remote is a private, self-hosted Gitea instance (`git.adlee.work`) — "PRs welcome" isn't true here. |
+| `## License` | **Needs a decision, see below** | No `LICENSE` file exists in this repo (checked directly). The reference names MIT/Apache-2.0 because those are real, chosen licenses. Do not write a license section that names something that isn't actually true. |
+
+## Open decisions (yours, not opencode's to invent)
+
+Two sections in the reference structure map to facts that don't exist yet
+in this repo. Both should be **explicit decisions**, not filled in by
+guessing during implementation:
+
+1. **License.** Pick one (or explicitly decide "unlicensed / private, not
+ for redistribution") before a `## License` section gets written. Whatever
+ is decided, add the matching `LICENSE` file at the same time — a README
+ section naming a license with no `LICENSE` file in the repo would be the
+ exact "documented but not actually true" failure mode this project's own
+ `config.yaml` strictness rules exist to prevent elsewhere.
+2. **Repo & Contributions.** Given the remote is a private Gitea instance
+ rather than a public GitHub project, decide whether this section exists
+ at all, and if so what it actually says — a link to the internal remote
+ for your own reference, not an open invitation for outside contributions
+ that can't reasonably arrive here.
+
+If no decision is made, the plan's default is: **omit both sections** rather
+than have opencode fabricate a license or a contribution policy that isn't
+real.
+
+## Section-by-section mechanism
+
+### §1. Pitch + elaboration (replaces the top of "What This Is")
+
+Move the existing lead paragraph up to immediately follow the `# Local LLM
+Model Router` title, ahead of any subheading — this is exactly what the
+reference does (pitch line, then one elaborating paragraph, before any `##`).
+Source material already exists verbatim in the current README's opening
+paragraph and the "Numbers in this README are measurements, not
+specifications" caveat — this section is a move-and-reflow, not a rewrite.
+
+### §2. `## Features`
+
+Reference pattern: one bullet per capability, phrased as a real situation
+the reader has been in, followed by the command that fixes it. This
+project's own voice should replace "clever/jokey" with "concrete/measured" —
+the through-line of the whole existing document. Candidate bullets, pulled
+directly from existing content rather than invented:
+
+- `POST /route` — "Want to know what a task would cost before you spend
+ anything on it?" *(current README: "no provider call, no cost")*
+- Per-request energy/cost pricing — "List price ranks models backwards for
+ this workload; billing is per-kWh, and this router prices per request from
+ what's actually shaped like your traffic." *(current: "Weighted Scoring"
+ section)*
+- Local vision fallback — "Your cheapest coding model doesn't support
+ images. This one falls back to a local model instead of 422ing."
+ *(current: "Local Vision Fallback")*
+- Live TUI — "`python tui.py`, a live routing feed with no polling delay."
+ *(current: "Monitoring")*
+- Structural + local-LLM verification — "Every response gets checked for
+ free before routing ever learns from it." *(current: "Verification
+ Pipeline")*
+
+Five bullets, matching the reference's scope (3 headline + "and others
+like"), not an exhaustive re-listing of every feature — the full detail
+still lives in its own section further down, same as the reference's
+`## Usage` expands on its `## Features` teasers.
+
+### §3. `## Usage` — the main restructuring work
+
+Current "Quick Usage" is one code block, 9 `curl` commands, minimal
+narrative. Reference pattern is one `###` sub-heading per command: a short
+paragraph on *why* (often phrased as the problem the reader already has),
+then the command, then a bullet list of what it actually does.
+
+This project already has the "why" prose for nearly every command — it's
+just located in a different section than the command itself. Concrete
+mapping (existing source section → new usage sub-heading):
+
+| New `### ` sub-heading | Command | Source prose already written |
+|---|---|---|
+| Route without spending anything | `POST /route` | API Endpoints table: "no provider call, no cost" |
+| Skip the classifier when you already know the shape | `/route` with overrides | "Input to `/route` and `/dispatch` can include... overrides — these skip the classifier" |
+| Actually dispatch and log energy | `POST /dispatch` | API Endpoints table |
+| Point any OpenAI-compatible client at it | `/v1/chat/completions`, `/v1/models` | "Pointing a Coding Agent at It" |
+| Route overnight/batch work through flex rows | `latency_tolerance: batch` | "auto:batch admits flex rows for async work" |
+| Ask an image question | image_url request | "Local Vision Fallback" |
+| Force JSON output | `response_format` | Weighted Scoring, hard filter #6 |
+| Watch it live | `python tui.py`, `/metrics`, `/events/decisions` | "Monitoring" section, mostly verbatim |
+| Probe routing without spending | `router_cli.py` | "Monitoring" section |
+
+This is reorganization, not new writing: every cell in the right column
+already exists in the current README. The work is moving each explanation
+next to the command it explains, and adding a one-line "why" lead-in where
+the existing text is purely descriptive rather than motivating.
+
+### §4. `## Installation`
+
+Two named sub-sections, matching what's actually true:
+
+- **Local (`venv`)** — the existing "Setup" code block, unchanged.
+- **As a systemd service** — the existing "Scheduled Jobs (systemd)" table
+ and its install pointer to `deploy/README.md`.
+
+Do not add a third method. The reference's breadth (6 install paths) exists
+because that project targets a general audience installing a CLI tool from
+multiple ecosystems; this one has an owner, a GPU, and a fixed deployment
+shape.
+
+### §5. `## Requirements`
+
+Pull the existing "Setup" opening sentence — "a Neuralwatt API key, Python
+3.10+ (suite verified on 3.10 and 3.14), and an Ollama reachable from
+wherever this runs with a classifier model pulled" — into its own short
+bulleted section, ahead of Installation, matching the reference's
+`requirements`-before-`installation` ordering.
+
+### §6. Table of Contents
+
+Auto-derivable from the final heading structure once the above moves are
+made. Every existing `##`/`###` heading gets an entry; nest sub-headings
+(Usage's per-command sections, Installation's two methods) the way the
+reference nests its own.
+
+## What must NOT change
+
+This is a restructuring of entry points and navigation, not a rewrite of
+substance. Everything below stays exactly as it is, moved but not
+rewritten:
+
+- The architecture diagram, the module table, the full schema documentation
+ (`models`/`proficiency`/`energy_observations`/`verifications`/
+ `route_decisions`), the weighted-scoring mechanism, the classifier
+ reliability notes, and Known Limitations & Open Items — none of this has
+ a reference-README equivalent because `supabase-plus` doesn't carry
+ this much operational depth. It stays as reference material past the new
+ Usage/Installation front matter, in its current form.
+- No new measurements, dates, or claims. Every fact placed into a new
+ section must already exist verbatim (or near-verbatim, for reflow)
+ somewhere in the current `README.md`.
+- No jokes or voice imported from the reference. "Had buckets locally once,
+ never found them in prod" works for a Supabase CLI's audience; this
+ project's own established register — measured numbers, dated
+ observations, named caveats — is what every other document in this repo
+ already uses, and the README should not be the one place that departs
+ from it.
+
+## Recommendation
+
+Build it. The highest-value pieces (a Table of Contents for a 1039-line
+document that currently has none, and moving existing "why" prose next to
+the commands it explains) require no new facts and no voice risk — they're
+pure reorganization. The two sections that need a real answer (License,
+Repo & Contributions) should get one from you before opencode touches them;
+absent that, the plan's default is to leave them out rather than invent
+them.
diff --git a/documentation_reviews/readme-restructure-review.md b/documentation_reviews/readme-restructure-review.md
new file mode 100644
index 0000000..abe8380
--- /dev/null
+++ b/documentation_reviews/readme-restructure-review.md
@@ -0,0 +1,102 @@
+# Review: README restructure (commit `14a7653`)
+
+**What it was reviewing:** opencode's implementation of
+`documentation_plans/readme-restructure.md` — reshaping `README.md` around a
+pitch → features → install → per-command usage structure borrowed from
+`supabase-plus`'s README shape. Checked the actual diff (`git show
+14a7653`) line by line against the spec's explicit constraints, not just the
+new document's surface quality.
+
+## Verdict: the restructuring itself is well done; one section's content was silently dropped, and one TOC link is dead
+
+### What's correct, including the part most likely to go wrong
+
+The spec's two "needs a decision, don't invent" items — a `## License`
+section and a `## Repo & Contributions` section — are both correctly
+**absent** from the new README. Neither a LICENSE file nor an honest
+contributions policy exists yet, and the plan was explicit that opencode
+should not fabricate either. Checked directly (`grep -i
+"license\|contribut"` against every heading): nothing was invented. This
+was the single most important constraint in the spec and it held.
+
+Everything else structural matches: `## Features` uses the five bullets the
+plan proposed, pulled from existing claims rather than invented ones;
+`## Requirements` is pulled out as its own section; `## Installation` has
+exactly the two real sub-methods (`Local (venv)`, `As a systemd service`)
+rather than fabricated ones; `## Usage` is restructured into nine
+per-command sections matching the plan's mapping table, and spot-checking
+"Watch it live" and "Probe routing without spending" confirms the
+Monitoring section's content (all five tools: `/metrics`, `/events/decisions`,
+`tui.py`, `baseline_report.py`, `router_cli.py`) survived the reflow intact,
+just condensed. A Table of Contents now exists where none did before.
+
+### `## Setup`'s "Where Ollama lives" subsection was deleted, not moved
+
+**Confirmed by diff, not inference.** `git show 14a7653 -- README.md`
+shows this entire block removed with no corresponding addition anywhere in
+the new document:
+
+```
+-### Where Ollama lives
+-
+-```bash
+-ollama pull mistral-nemo:12b # or whatever you set as classifier.model
+-```
+-
+-...To use one across a VPN, point **both** endpoints at it:
+-...and apply `deploy/ollama-over-vpn.conf` on the serving host — Ollama binds
+-`127.0.0.1` by default and will otherwise refuse. Bind it to the VPN address
+-rather than `0.0.0.0`: Ollama has no authentication, so anything reaching the
+-port can run inference and enumerate your models.
+-
+-Both endpoints move together because the verifier speaks Ollama's *native*
+-API and cannot follow the classifier to a cloud provider...
+```
+
+Grepped the current `README.md` for `ollama-over-vpn`, `ollama pull
+mistral-nemo`, and `Both endpoints move together` — zero hits. This isn't a
+paraphrase living somewhere else; the content is gone.
+
+**Why this matters more than an ordinary trim.** This section carried a real
+security instruction, not just background — Ollama has no authentication of
+its own, and the doc's own words were the thing telling a reader to bind it
+to a VPN address rather than `0.0.0.0`. The plan's own "What must NOT
+change" section was explicit: *"Everything below stays exactly as it is,
+moved but not rewritten"* and *"No new measurements, dates, or claims"* —
+the inverse, dropping an existing one, was never authorized either. The most
+likely mechanical cause: the old "Setup" section's opening sentence moved to
+`## Requirements` and its venv steps moved to `## Installation`, and
+"Where Ollama lives" — a `###` subsection of the same old `## Setup` — seems
+to have been left behind in that split rather than carried to either
+destination.
+
+### Dead TOC link: "Scheduled Jobs (systemd)"
+
+The Table of Contents (line 74) still has an entry
+`- [Scheduled Jobs (systemd)](#scheduled-jobs-systemd)`. That heading no
+longer exists — the content it pointed to was correctly moved into `###
+As a systemd service` under Installation, which already has its own
+correct TOC entry at line 43. The old line was never removed, so it's a
+link to nowhere sitting in a document whose main improvement this pass was
+*adding working navigation*. One line to delete.
+
+A handful of other headings with em-dashes, backticks, or `&` (e.g.
+`` `models` — one row per served model variant``, `Known Limitations & Open
+Items`) produced anchor mismatches against a straightforward slugify check,
+but GitHub- and Gitea-flavored anchor generation both have their own
+non-obvious rules for those characters that a quick script can't be trusted
+to reproduce exactly — these are worth a manual click-through in whatever
+renderer this repo actually displays in (Gitea, at `git.adlee.work`), not
+something to fix on the strength of this review alone.
+
+## Recommendation
+
+Restore "Where Ollama lives" verbatim — it's sitting in `git show
+14a7653^:README.md` (the pre-restructure version) if a clean copy is
+needed — into `## Installation`, most naturally as a third subsection
+alongside `Local (venv)` and `As a systemd service` (it's setup guidance
+that applies to either), or folded into `Local (venv)` if a separate
+heading feels like too much for one paragraph plus a warning. Delete the
+dead `Scheduled Jobs (systemd)` TOC line. Then do one manual pass clicking
+every TOC link in the actual Gitea-rendered view, since that's the renderer
+that matters here and not something worth guessing at from a script.
diff --git a/tests/test_admin_config.py b/tests/test_admin_config.py
new file mode 100644
index 0000000..58b9b35
--- /dev/null
+++ b/tests/test_admin_config.py
@@ -0,0 +1,187 @@
+"""Tests for the /admin/api persisted-config endpoints.
+
+``admin.py`` exposes ``GET /admin/api/config`` (the allowlisted config.yaml
+values) and ``POST /admin/api/config/{key}`` (persist one allowlisted value to
+config.yaml with comment-preserving ruamel.yaml round-trip, a backup copy, and
+whole-config validation via ``RouterConfig`` BEFORE anything touches disk).
+
+Each test builds its own isolated router against a temp copy of config.yaml by
+passing ``base_dir`` (a temp dir) to ``build_router`` — the real repo
+``config.yaml`` is never written. It mounts the router on a fresh FastAPI
+TestClient at ``prefix="/admin"``.
+"""
+
+from __future__ import annotations
+
+import re
+import shutil
+import sqlite3
+import threading
+from pathlib import Path
+
+import pytest
+from fastapi import FastAPI
+from starlette.testclient import TestClient
+
+from admin import _CONFIG_ALLOWLIST, _persist_config_value, build_router
+
+ROOT = Path(__file__).resolve().parent.parent
+
+_SENTINEL = "# SENTINEL_PRESERVED_12345"
+
+
+def _insert_sentinel(config_yaml: Path) -> None:
+ """Prepend a unique marker comment just above the ``logging:`` map."""
+ text = config_yaml.read_text()
+ assert "logging:" in text
+ config_yaml.write_text(text.replace("logging:", f"{_SENTINEL}\nlogging:", 1))
+
+
+def _make_db(tmp_path: Path) -> sqlite3.Connection:
+ conn = sqlite3.connect(str(tmp_path / "admin.db"))
+ conn.row_factory = sqlite3.Row
+ return conn
+
+
+@pytest.fixture
+def client(tmp_path, monkeypatch):
+ """A TestClient for an isolated admin router over a temp config.yaml copy.
+
+ ``base_dir`` is tmp_path, so the ``/admin/api/config`` endpoints read and
+ write ``tmp_path/config.yaml`` — never the repo's copy.
+ """
+ config_yaml = tmp_path / "config.yaml"
+ shutil.copyfile(ROOT / "config.yaml", config_yaml)
+
+ def _db_factory() -> sqlite3.Connection:
+ return _make_db(tmp_path)
+
+ router = build_router(None, _db_factory, base_dir=str(tmp_path))
+ app = FastAPI()
+ app.include_router(router, prefix="/admin")
+ return TestClient(app), config_yaml
+
+
+def test_config_GET_returns_allowlisted_values(client):
+ """GET /admin/api/config returns every allowlisted dotted key."""
+ tc, _ = client
+ resp = tc.get("/admin/api/config")
+ assert resp.status_code == 200
+ body = resp.json()
+ for key in (
+ "logging.level",
+ "objective.quality_tolerance",
+ "objective.max_energy_per_request",
+ "objective.plan_kwh_per_period",
+ "circuit_breaker.enabled",
+ "session_cache.enabled",
+ "verification.local_llm_enabled",
+ "pinch.enabled",
+ "pinch.relevance.enabled",
+ "routing.default_flex_preference",
+ ):
+ assert key in body
+ assert body["logging.level"] == "info"
+ assert body["objective.quality_tolerance"] == 0.10
+ assert body["routing.default_flex_preference"] == "auto"
+
+
+def test_config_POST_preserves_comments_and_changes_value(client):
+ """A valid write keeps the file's comments AND updates the value."""
+ tc, config_yaml = client
+ _insert_sentinel(config_yaml)
+
+ resp = tc.post("/admin/api/config/logging.level", json={"value": "warning"})
+ assert resp.status_code == 200
+ body = resp.json()
+ assert body["key"] == "logging.level"
+ assert body["value"] == "warning"
+ assert "restart is required" in body["message"]
+
+ text = config_yaml.read_text()
+ assert _SENTINEL in text
+ assert re.search(r"^\s*level:\s*warning\s*$", text, re.MULTILINE) is not None
+
+
+def test_config_POST_rejects_non_allowlisted_key(client):
+ """classifier.base_url (and any off-allowlist key) is refused with 403."""
+ tc, _ = client
+ resp = tc.post("/admin/api/config/classifier.base_url", json={"value": "http://x"})
+ assert resp.status_code == 403
+
+
+def test_config_POST_invalid_value_leaves_file_unchanged(client):
+ """quality_tolerance=1.5 (>1) fails RouterConfig validation -> 422, no write."""
+ tc, config_yaml = client
+ before = config_yaml.read_text()
+
+ resp = tc.post(
+ "/admin/api/config/objective.quality_tolerance", json={"value": 1.5}
+ )
+ assert resp.status_code == 422
+
+ after = config_yaml.read_text()
+ assert after == before
+
+
+def test_config_POST_creates_backup_before_write(client, tmp_path):
+ """A successful write produces a ``config.yaml.bak.`` backup copy."""
+ tc, config_yaml = client
+ _insert_sentinel(config_yaml)
+
+ resp = tc.post("/admin/api/config/circuit_breaker.enabled", json={"value": True})
+ assert resp.status_code == 200
+
+ backups = sorted(tmp_path.glob("config.yaml.bak.*"))
+ assert len(backups) == 1
+ # The backup captured the sentinel comment and the PRE-write value.
+ backup_text = backups[0].read_text()
+ assert _SENTINEL in backup_text
+ assert re.search(r"^\s*enabled:\s*false\s*$", backup_text, re.MULTILINE) is not None
+
+
+def test_config_concurrent_writes_are_atomic_no_zero_byte_backups(tmp_path):
+ """Concurrent writes never truncate config.yaml or leave 0-byte backups."""
+ config_yaml = tmp_path / "config.yaml"
+ shutil.copyfile(ROOT / "config.yaml", config_yaml)
+ path = _CONFIG_ALLOWLIST["logging.level"]
+
+ stop_reader = threading.Event()
+
+ def reader() -> None:
+ while not stop_reader.is_set():
+ try:
+ text = config_yaml.read_text()
+ except FileNotFoundError:
+ continue
+ assert text.strip() != "", "config.yaml observed empty"
+ assert "logging:" in text
+
+ reader_thread = threading.Thread(target=reader)
+ reader_thread.start()
+
+ def writer(level: str) -> None:
+ _persist_config_value(config_yaml, path, level)
+
+ threads = [
+ threading.Thread(target=writer, args=("warning",)),
+ threading.Thread(target=writer, args=("error",)),
+ ]
+ for t in threads:
+ t.start()
+ for t in threads:
+ t.join()
+ stop_reader.set()
+ reader_thread.join()
+
+ final = config_yaml.read_text()
+ assert re.search(
+ r"^\s*level:\s*(warning|error)\s*$", final, re.MULTILINE
+ ) is not None
+
+ backups = list(tmp_path.glob("config.yaml.bak.*"))
+ assert backups, "expected at least one backup"
+ for b in backups:
+ assert b.stat().st_size > 0, f"zero-byte backup: {b}"
+ assert b.read_text().strip() != ""
+
diff --git a/tests/test_admin_frontend.py b/tests/test_admin_frontend.py
new file mode 100644
index 0000000..74f0d61
--- /dev/null
+++ b/tests/test_admin_frontend.py
@@ -0,0 +1,97 @@
+"""Tests for the admin frontend serving route GET /admin/.
+
+The frontend is served as a single FileResponse (no StaticFiles mount). Its
+path is derived from ``Path(__file__).resolve().parent`` inside ``admin.py`` —
+not from ``base_dir`` — so it resolves to the real repo file even when the
+router is built without a base_dir, mirroring the daashbrd H3 portability fix.
+"""
+
+from __future__ import annotations
+
+import sqlite3
+import subprocess
+import sys
+from pathlib import Path
+
+import pytest
+from starlette.testclient import TestClient
+
+import admin
+import dispatcher
+from config import load_config
+
+ROOT = Path(__file__).resolve().parent.parent
+CFG = load_config(str(ROOT / "config.yaml"))
+
+
+@pytest.fixture
+def admin_client(tmp_path, monkeypatch):
+ """The real dispatcher app, which mounts admin at the /admin prefix."""
+ db_path = str(tmp_path / "admin.db")
+ conn = sqlite3.connect(db_path)
+ conn.row_factory = sqlite3.Row
+ conn.executescript((ROOT / "schema.sql").read_text())
+ conn.close()
+
+ monkeypatch.setattr(dispatcher.cfg.database, "path", db_path)
+ monkeypatch.setattr(dispatcher.cfg.verification, "local_llm_enabled", False)
+ monkeypatch.setattr(dispatcher.cfg.routing, "require_vision", False)
+ monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
+
+ with TestClient(dispatcher.app) as client:
+ yield client
+
+
+def test_admin_index_returns_html_with_chartjs(admin_client):
+ """GET /admin/ returns 200, text/html, and contains the Chart.js tag."""
+ resp = admin_client.get("/admin/")
+ assert resp.status_code == 200
+ assert resp.headers["content-type"].startswith("text/html")
+ assert "chart.js" in resp.text
+
+
+def test_admin_index_file_exists_at_module_derived_path():
+ """The served file lives at Path(__file__).parent/admin/frontend/index.html."""
+ expected = (
+ Path(admin.__file__).resolve().parent / "admin" / "frontend" / "index.html"
+ )
+ assert expected.is_file(), f"frontend file missing at {expected}"
+
+
+def test_admin_frontend_path_does_not_depend_on_base_dir():
+ """build_router WITHOUT base_dir already serves '/' on the raw router.
+
+ The route must derive its file path from admin.py's own location, not from
+ the optional base_dir argument, so it stays portable and testable.
+ """
+ probe = (
+ "import sqlite3\n"
+ "import admin\n"
+ "from fastapi import FastAPI\n"
+ "cfg = admin.load_config('config.yaml')\n"
+ "def db():\n"
+ " c = sqlite3.connect(':memory:')\n"
+ " c.row_factory = sqlite3.Row\n"
+ " return c\n"
+ "app = FastAPI()\n"
+ "app.include_router(admin.build_router(cfg, db))\n"
+ "from starlette.testclient import TestClient\n"
+ "with TestClient(app) as client:\n"
+ " resp = client.get('/')\n"
+ " assert resp.status_code == 200, resp.status_code\n"
+ " assert resp.headers['content-type'].startswith('text/html')\n"
+ )
+ subprocess.run(
+ [sys.executable, "-c", probe],
+ check=True,
+ cwd=str(ROOT),
+ capture_output=True,
+ )
+
+
+def test_admin_unknown_api_returns_json_404_not_html(admin_client):
+ """GET /admin/api/nonexistent is a 404 JSON body, not the HTML page."""
+ resp = admin_client.get("/admin/api/nonexistent")
+ assert resp.status_code == 404
+ assert resp.headers["content-type"].startswith("application/json")
+ assert "chart.js" not in resp.text
diff --git a/tests/test_admin_health.py b/tests/test_admin_health.py
new file mode 100644
index 0000000..d124205
--- /dev/null
+++ b/tests/test_admin_health.py
@@ -0,0 +1,208 @@
+"""Tests for the /admin/api health and snapshot endpoints.
+
+``admin.py`` is a standalone FastAPI router that mirrors ``metrics.py``'s
+contract — never import ``dispatcher``, take ``(conn, cfg)`` explicitly — and
+is mounted onto the dispatcher app under the ``/admin`` prefix. These tests
+drive a real TestClient GET against the seeded temp DB (never a mock-call
+assertion), mirroring ``test_metrics_endpoint.py``.
+"""
+
+from __future__ import annotations
+
+import sqlite3
+from datetime import datetime, timedelta, timezone
+from pathlib import Path
+
+import pytest
+from starlette.testclient import TestClient
+
+import dispatcher
+from config import load_config
+
+ROOT = Path(__file__).resolve().parent.parent
+SCHEMA_SQL = (ROOT / "schema.sql").read_text()
+CFG = load_config(str(ROOT / "config.yaml"))
+
+
+def _now() -> datetime:
+ return datetime.now(timezone.utc)
+
+
+def _make_db(tmp_path: Path) -> sqlite3.Connection:
+ conn = sqlite3.connect(str(tmp_path / "test.db"))
+ conn.row_factory = sqlite3.Row
+ conn.executescript(SCHEMA_SQL)
+ return conn
+
+
+def _seed_models(conn: sqlite3.Connection) -> None:
+ for model_id, tier, context, cost, vision in (
+ ("cheap", 2, 262128, 0.30, 1),
+ ("dear", 2, 262128, 9.00, 0),
+ ("tiny", 1, 131072, 0.10, 1),
+ ):
+ conn.execute(
+ """
+ INSERT INTO models (
+ model_id, provider, base_model_id, tier, context_window,
+ effective_context_window, max_output_tokens,
+ cost_per_1m_prompt, cost_per_1m_completion,
+ supports_vision, supports_json_mode,
+ latency_class, reasoning_mode, context_variant,
+ access_level, availability, last_updated
+ ) VALUES (?, 'neuralwatt', ?, ?, ?, 192500, 16384, ?, ?,
+ ?, 1, 'standard', 'default', 'full', 'public', 'active',
+ '2026-08-22T00:00:00+00:00')
+ """,
+ (model_id, model_id, tier, context, cost, cost / 3, vision),
+ )
+ conn.commit()
+
+
+def _seed_decision(conn: sqlite3.Connection) -> None:
+ conn.execute(
+ """
+ INSERT INTO route_decisions (
+ observed_at, kind, task_category, task_tier, required_context_tokens,
+ confidence, classifier_ms, classification_source, latency_tolerance,
+ candidates_considered, selected_model, selected_provider,
+ runner_up_models, est_cost_usd, est_proficiency,
+ session_key, tools, images, json_mode, streamed,
+ flex_preference, flex_swapped, flex_forced
+ ) VALUES (?, 'route', 'coding_general', 2, 100, 0.95, 200,
+ 'classifier', 'interactive', 5, 'cheap', 'neuralwatt',
+ '[{"model_id":"dear","provider":"neuralwatt"}]',
+ 0.001, 0.9, 'abc123', 0, 0, 0, 0,
+ 'auto', 0, 1)
+ """,
+ (_now().isoformat(),),
+ )
+ conn.commit()
+
+
+def _seed_energy(conn: sqlite3.Connection) -> None:
+ now = _now()
+ conn.execute(
+ "INSERT INTO energy_observations "
+ "(model_id, provider, task_category, completion_tokens, energy_kwh, "
+ "cost_usd, carbon_g_co2eq, attribution_ratio, observed_at) "
+ "VALUES ('cheap', 'neuralwatt', 'coding_general', 100, 5.0e-05, 0.001, "
+ "2.4e-03, 0.25, ?)",
+ ((now - timedelta(days=2)).isoformat(),),
+ )
+ conn.commit()
+
+
+def _seed_verification(conn: sqlite3.Connection) -> None:
+ conn.execute(
+ "INSERT INTO verifications (model_id, provider, kind, verdict, observed_at) "
+ "VALUES ('cheap', 'neuralwatt', 'structural', 'ok', ?)",
+ (_now().isoformat(),),
+ )
+ conn.commit()
+
+
+def _seed_proficiency(conn: sqlite3.Connection) -> None:
+ conn.execute(
+ "INSERT INTO proficiency (model_id, provider, category, blended_score, "
+ "source, last_updated) "
+ "VALUES ('cheap', 'neuralwatt', 'coding_general', 0.9, "
+ "'self_eval_thin', '2026-01-01T00:00:00+00:00')",
+ )
+ conn.commit()
+
+
+@pytest.fixture
+def seeded_client(tmp_path, monkeypatch):
+ """A TestClient wired to a seeded temp DB, at /admin."""
+ conn = _make_db(tmp_path)
+ _seed_models(conn)
+ for _ in range(3):
+ _seed_decision(conn)
+ _seed_energy(conn)
+ _seed_verification(conn)
+ _seed_proficiency(conn)
+ conn.close()
+
+ monkeypatch.setattr(dispatcher.cfg.database, "path", str(tmp_path / "test.db"))
+ monkeypatch.setattr(dispatcher.cfg.verification, "local_llm_enabled", False)
+ monkeypatch.setattr(dispatcher.cfg.routing, "require_vision", False)
+ monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
+
+ with TestClient(dispatcher.app) as client:
+ yield client
+
+
+def test_admin_health_returns_ok(seeded_client):
+ """GET /admin/api/health returns 200 with {"status": "ok"}."""
+ resp = seeded_client.get("/admin/api/health")
+ assert resp.status_code == 200
+ assert resp.json() == {"status": "ok"}
+
+
+def test_admin_snapshot_has_all_top_level_keys(seeded_client):
+ """GET /admin/api/snapshot returns 200 with every required key."""
+ resp = seeded_client.get("/admin/api/snapshot")
+ assert resp.status_code == 200
+ data = resp.json()
+ for key in (
+ "quota",
+ "coverage",
+ "recent_decisions",
+ "per_model",
+ "verdict_mix",
+ "top_proficiency",
+ "health",
+ "generated_at",
+ ):
+ assert key in data, f"missing top-level key {key!r}"
+
+
+def test_admin_snapshot_health_has_expected_shape(seeded_client):
+ """Snapshot's health block carries the same shape as /health."""
+ data = seeded_client.get("/admin/api/snapshot").json()
+ health = data["health"]
+ assert "status" in health
+ assert "counts" in health
+ assert "classifier_reachable" in health and isinstance(
+ health["classifier_reachable"], bool
+ )
+ assert "tiers" in health
+ assert "providers" in health
+ assert "api_keys_present" in health
+
+
+def test_admin_snapshot_returns_seeded_data(seeded_client):
+ """quota / per_model / verdict_mix / top_proficiency reflect the seed."""
+ data = seeded_client.get("/admin/api/snapshot").json()
+ assert data["quota"] is not None
+ assert isinstance(data["per_model"], list)
+ assert any(r["model_id"] == "cheap" for r in data["per_model"])
+ assert data["verdict_mix"]["ok"] == 1
+ assert isinstance(data["top_proficiency"], list)
+ assert data["top_proficiency"][0]["model_id"] == "cheap"
+ assert data["health"]["counts"]["models"] == 3
+
+
+def test_admin_snapshot_contains_no_session_dir(seeded_client):
+ """The JSON must never name session_dir or expose conversation text."""
+ body = seeded_client.get("/admin/api/snapshot").text
+ assert "session_dir" not in body
+
+
+def test_admin_does_not_import_dispatcher():
+ """Importing admin.py must not pull dispatcher into sys.modules."""
+ import subprocess
+ import sys
+
+ probe = (
+ "import admin; import sys; "
+ "assert 'dispatcher' not in sys.modules, "
+ "'import admin transitively imported dispatcher'"
+ )
+ subprocess.run(
+ [sys.executable, "-c", probe],
+ check=True,
+ cwd=str(ROOT),
+ capture_output=True,
+ )
diff --git a/tests/test_admin_history.py b/tests/test_admin_history.py
new file mode 100644
index 0000000..dfe07e6
--- /dev/null
+++ b/tests/test_admin_history.py
@@ -0,0 +1,249 @@
+"""Tests for the /admin/api/history bucketed time-series endpoint.
+
+``/api/history`` runs a bucketed GROUP BY over two existing tables
+(``energy_observations`` and ``route_decisions``), keyed by a ``range`` query
+param that selects a (span_seconds, bucket_seconds) pair. These tests drive a
+real TestClient GET against a temp DB seeded at controlled timestamps, exactly
+like ``test_admin_health.py`` — never a mock-call assertion.
+"""
+
+from __future__ import annotations
+
+import sqlite3
+from datetime import datetime, timedelta, timezone
+from pathlib import Path
+
+import pytest
+from starlette.testclient import TestClient
+
+import dispatcher
+from config import load_config
+
+ROOT = Path(__file__).resolve().parent.parent
+SCHEMA_SQL = (ROOT / "schema.sql").read_text()
+CFG = load_config(str(ROOT / "config.yaml"))
+
+# The 1h bucket size used by range=24h; expected bucket ts are whole-second
+# hour boundaries, computed the same way the SQL floors to a bucket index.
+BUCKET_24H = 3600
+
+
+def _make_db(tmp_path: Path) -> sqlite3.Connection:
+ conn = sqlite3.connect(str(tmp_path / "test.db"))
+ conn.row_factory = sqlite3.Row
+ conn.executescript(SCHEMA_SQL)
+ return conn
+
+
+def _seed_models(conn: sqlite3.Connection) -> None:
+ for model_id, tier, context, cost, vision in (
+ ("cheap", 2, 262128, 0.30, 1),
+ ("dear", 2, 262128, 9.00, 0),
+ ):
+ conn.execute(
+ """
+ INSERT INTO models (
+ model_id, provider, base_model_id, tier, context_window,
+ effective_context_window, max_output_tokens,
+ cost_per_1m_prompt, cost_per_1m_completion,
+ supports_vision, supports_json_mode,
+ latency_class, reasoning_mode, context_variant,
+ access_level, availability, last_updated
+ ) VALUES (?, 'neuralwatt', ?, ?, ?, 192500, 16384, ?, ?,
+ ?, 1, 'standard', 'default', 'full', 'public', 'active',
+ '2026-08-22T00:00:00+00:00')
+ """,
+ (model_id, model_id, tier, context, cost, cost / 3, vision),
+ )
+ conn.commit()
+
+
+def _seed_energy(
+ conn: sqlite3.Connection,
+ at: datetime,
+ energy_kwh: float,
+ cost_usd: float,
+ carbon_g_co2eq: float,
+) -> None:
+ conn.execute(
+ "INSERT INTO energy_observations "
+ "(model_id, provider, task_category, completion_tokens, energy_kwh, "
+ "cost_usd, carbon_g_co2eq, attribution_ratio, observed_at) "
+ "VALUES ('cheap', 'neuralwatt', 'coding_general', 100, ?, ?, ?, 0.25, ?)",
+ (energy_kwh, cost_usd, carbon_g_co2eq, at.isoformat()),
+ )
+
+
+def _seed_decision(conn: sqlite3.Connection, at: datetime) -> None:
+ conn.execute(
+ """
+ INSERT INTO route_decisions (
+ observed_at, kind, task_category, task_tier, required_context_tokens,
+ confidence, classifier_ms, classification_source, latency_tolerance,
+ candidates_considered, selected_model, selected_provider,
+ runner_up_models, est_cost_usd, est_proficiency,
+ session_key, tools, images, json_mode, streamed,
+ flex_preference, flex_swapped, flex_forced
+ ) VALUES (?, 'route', 'coding_general', 2, 100, 0.95, 200,
+ 'classifier', 'interactive', 5, 'cheap', 'neuralwatt',
+ '[{"model_id":"dear","provider":"neuralwatt"}]',
+ 0.001, 0.9, 'abc123', 0, 0, 0, 0,
+ 'auto', 0, 1)
+ """,
+ (at.isoformat(),),
+ )
+
+
+def _hour_floor(dt: datetime) -> datetime:
+ """The start of the hour containing *dt*, on a whole second."""
+ return dt.replace(second=0, microsecond=0).replace(minute=0)
+
+
+@pytest.fixture
+def seeded_client(tmp_path, monkeypatch):
+ """A TestClient at /admin wired to a temp DB.
+
+ Seeds, relative to the current hour, three energy rows across two
+ one-hour buckets and three decisions across the same two buckets — all
+ inside a 24h window:
+
+ - bucket A (current hour): R1 energy=0.001 cost=0.01 carbon=0.5,
+ R2 energy=0.002 cost=0.02 carbon=1.0
+ - bucket B (previous hour): R3 energy=0.004 cost=0.04 carbon=2.0
+ - decisions: D1+D2 in bucket A, D3 in bucket B
+ """
+ conn = _make_db(tmp_path)
+ _seed_models(conn)
+
+ now = _now()
+ bucket_a = _hour_floor(now)
+ bucket_b = bucket_a - timedelta(hours=1)
+
+ _seed_energy(conn, bucket_a + timedelta(seconds=60), 0.001, 0.01, 0.5)
+ _seed_energy(conn, bucket_a + timedelta(seconds=120), 0.002, 0.02, 1.0)
+ _seed_energy(conn, bucket_b + timedelta(seconds=30), 0.004, 0.04, 2.0)
+
+ _seed_decision(conn, bucket_a + timedelta(seconds=60))
+ _seed_decision(conn, bucket_a + timedelta(seconds=200))
+ _seed_decision(conn, bucket_b + timedelta(seconds=30))
+ conn.commit()
+ conn.close()
+
+ monkeypatch.setattr(dispatcher.cfg.database, "path", str(tmp_path / "test.db"))
+ monkeypatch.setattr(dispatcher.cfg.verification, "local_llm_enabled", False)
+ monkeypatch.setattr(dispatcher.cfg.routing, "require_vision", False)
+ monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
+
+ with TestClient(dispatcher.app) as client:
+ yield client
+
+
+def _now() -> datetime:
+ return datetime.now(timezone.utc)
+
+
+def _bucket_ts(dt: datetime) -> int:
+ """The unix second of the 1h bucket containing *dt*."""
+ return int(dt.timestamp() // BUCKET_24H * BUCKET_24H)
+
+
+def test_admin_history_24h_returns_series(seeded_client):
+ """range=24h returns all five series, each a list of [ts, value] pairs."""
+ resp = seeded_client.get("/admin/api/history", params={"range": "24h"})
+ assert resp.status_code == 200
+ data = resp.json()
+ assert set(data) == {
+ "decisions_per_bucket",
+ "requests_per_bucket",
+ "cost_per_bucket",
+ "energy_per_bucket",
+ "carbon_per_bucket",
+ }
+ for series in data.values():
+ assert isinstance(series, list)
+ for point in series:
+ assert len(point) == 2
+ assert isinstance(point[0], (int, float))
+
+
+def test_admin_history_24h_decisions_per_bucket(seeded_client):
+ """decisions_per_bucket counts route_decisions per 1h bucket."""
+ now = _now()
+ bucket_a = _bucket_ts(_hour_floor(now))
+ bucket_b = bucket_a - BUCKET_24H
+ series = seeded_client.get("/admin/api/history", params={"range": "24h"}).json()[
+ "decisions_per_bucket"
+ ]
+ assert series == [[bucket_b, 1], [bucket_a, 2]]
+
+
+def test_admin_history_24h_requests_per_bucket(seeded_client):
+ """requests_per_bucket counts energy_observations per 1h bucket."""
+ now = _now()
+ bucket_a = _bucket_ts(_hour_floor(now))
+ bucket_b = bucket_a - BUCKET_24H
+ series = seeded_client.get("/admin/api/history", params={"range": "24h"}).json()[
+ "requests_per_bucket"
+ ]
+ assert series == [[bucket_b, 1], [bucket_a, 2]]
+
+
+def test_admin_history_24h_cost_per_bucket(seeded_client):
+ """cost_per_bucket sums cost_usd per 1h bucket."""
+ now = _now()
+ bucket_a = _bucket_ts(_hour_floor(now))
+ bucket_b = bucket_a - BUCKET_24H
+ series = seeded_client.get("/admin/api/history", params={"range": "24h"}).json()[
+ "cost_per_bucket"
+ ]
+ assert series == [[bucket_b, 0.04], [bucket_a, 0.03]]
+
+
+def test_admin_history_24h_energy_per_bucket(seeded_client):
+ """energy_per_bucket sums energy_kwh per 1h bucket."""
+ now = _now()
+ bucket_a = _bucket_ts(_hour_floor(now))
+ bucket_b = bucket_a - BUCKET_24H
+ series = seeded_client.get("/admin/api/history", params={"range": "24h"}).json()[
+ "energy_per_bucket"
+ ]
+ assert series == [[bucket_b, 0.004], [bucket_a, 0.003]]
+
+
+def test_admin_history_24h_carbon_per_bucket(seeded_client):
+ """carbon_per_bucket sums carbon_g_co2eq per 1h bucket."""
+ now = _now()
+ bucket_a = _bucket_ts(_hour_floor(now))
+ bucket_b = bucket_a - BUCKET_24H
+ series = seeded_client.get("/admin/api/history", params={"range": "24h"}).json()[
+ "carbon_per_bucket"
+ ]
+ assert series == [[bucket_b, 2.0], [bucket_a, 1.5]]
+
+
+def test_admin_history_bogus_range_returns_400(seeded_client):
+ """An unlisted range value is rejected with 400."""
+ resp = seeded_client.get("/admin/api/history", params={"range": "bogus"})
+ assert resp.status_code == 400
+
+
+@pytest.fixture
+def empty_client(tmp_path, monkeypatch):
+ """A TestClient at /admin wired to an empty (schema-only) temp DB."""
+ conn = _make_db(tmp_path)
+ conn.close()
+ monkeypatch.setattr(dispatcher.cfg.database, "path", str(tmp_path / "test.db"))
+ monkeypatch.setattr(dispatcher.cfg.verification, "local_llm_enabled", False)
+ monkeypatch.setattr(dispatcher.cfg.routing, "require_vision", False)
+ monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
+ with TestClient(dispatcher.app) as client:
+ yield client
+
+
+def test_admin_history_empty_tables_return_empty_series(empty_client):
+ """Empty tables yield 200 with every series as an empty list."""
+ resp = empty_client.get("/admin/api/history", params={"range": "24h"})
+ assert resp.status_code == 200
+ data = resp.json()
+ for series in data.values():
+ assert series == []
diff --git a/tests/test_admin_models.py b/tests/test_admin_models.py
new file mode 100644
index 0000000..30415c1
--- /dev/null
+++ b/tests/test_admin_models.py
@@ -0,0 +1,318 @@
+"""Tests for the /admin/api/models read surfaces.
+
+``admin.py`` mirrors ``metrics.py``'s contract — never import dispatcher, take
+``(conn, cfg)`` explicitly — and is mounted under the ``/admin`` prefix. These
+tests drive a real TestClient GET against the seeded temp DB, mirroring
+``test_admin_health.py``'s ``seeded_client`` fixture.
+
+The routes under test:
+ - GET /admin/api/models -> list of all models + proficiency
+ - GET /admin/api/models/{model_id}/{provider}-> single model detail (404 if absent)
+ - POST /admin/api/models/{model_id}/{provider}/availability -> upsert override
+ - DELETE /admin/api/models/{model_id}/{provider}/availability -> remove override
+"""
+
+from __future__ import annotations
+
+import sqlite3
+from pathlib import Path
+
+import pytest
+from starlette.testclient import TestClient
+
+import dispatcher
+from config import load_config
+
+ROOT = Path(__file__).resolve().parent.parent
+SCHEMA_SQL = (ROOT / "schema.sql").read_text()
+CFG = load_config(str(ROOT / "config.yaml"))
+
+_ADMIN_TABLE_SQL = """
+CREATE TABLE IF NOT EXISTS admin_model_overrides (
+ model_id TEXT NOT NULL,
+ provider TEXT NOT NULL,
+ availability TEXT NOT NULL,
+ reason TEXT,
+ updated_at TEXT NOT NULL,
+ PRIMARY KEY (model_id, provider)
+);
+CREATE INDEX IF NOT EXISTS idx_admin_model_overrides_availability
+ ON admin_model_overrides (availability);
+"""
+
+
+def _make_db(tmp_path: Path) -> sqlite3.Connection:
+ conn = sqlite3.connect(str(tmp_path / "test.db"))
+ conn.row_factory = sqlite3.Row
+ conn.executescript(SCHEMA_SQL)
+ conn.executescript(_ADMIN_TABLE_SQL)
+ return conn
+
+
+def _seed_models(conn: sqlite3.Connection, model_ids: tuple[str, ...]) -> None:
+ for model_id in model_ids:
+ conn.execute(
+ """
+ INSERT INTO models (
+ model_id, provider, base_model_id, display_name, tier,
+ context_window, effective_context_window, max_output_tokens,
+ cost_per_1m_prompt, cost_per_1m_completion,
+ supports_tools, supports_json_mode, supports_vision,
+ supports_reasoning, reasoning_default_enabled,
+ latency_class, reasoning_mode, context_variant,
+ access_level, availability, last_updated
+ ) VALUES (?, 'neuralwatt', ?, ?, ?, ?, 192500, 16384,
+ ?, ?, 1, 1, 1, 1, 1,
+ 'standard', 'default', 'full', 'public', 'active',
+ '2026-08-22T00:00:00+00:00')
+ """,
+ (
+ model_id,
+ model_id,
+ model_id,
+ 2,
+ 262128,
+ 0.30,
+ 0.10,
+ ),
+ )
+ conn.commit()
+
+
+def _seed_proficiency(
+ conn: sqlite3.Connection,
+ spec: tuple[tuple[str, str, float], ...],
+) -> None:
+ """Insert proficiency rows as (model_id, category, blended_score)."""
+ for model_id, category, blended in spec:
+ conn.execute(
+ "INSERT INTO proficiency (model_id, provider, category, "
+ "blended_score, source, last_updated) "
+ "VALUES (?, 'neuralwatt', ?, ?, 'self_eval_thin', "
+ "'2026-01-01T00:00:00+00:00')",
+ (model_id, category, blended),
+ )
+ conn.commit()
+
+
+@pytest.fixture
+def seeded_client(tmp_path, monkeypatch):
+ """A TestClient wired to a seeded temp DB, at /admin."""
+ conn = _make_db(tmp_path)
+ _seed_models(conn, ("cheap", "dear", "tiny"))
+ _seed_proficiency(
+ conn,
+ (
+ ("cheap", "coding_general", 0.9),
+ ("cheap", "debugging", 0.85),
+ ("dear", "coding_general", 0.95),
+ ),
+ )
+ conn.close()
+
+ monkeypatch.setattr(dispatcher.cfg.database, "path", str(tmp_path / "test.db"))
+ monkeypatch.setattr(dispatcher.cfg.verification, "local_llm_enabled", False)
+ monkeypatch.setattr(dispatcher.cfg.routing, "require_vision", False)
+ monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
+
+ with TestClient(dispatcher.app) as client:
+ yield client
+
+
+@pytest.fixture
+def empty_client(tmp_path, monkeypatch):
+ """A TestClient over an empty DB (no models rows) at /admin."""
+ conn = _make_db(tmp_path)
+ conn.close()
+
+ monkeypatch.setattr(dispatcher.cfg.database, "path", str(tmp_path / "test.db"))
+ monkeypatch.setattr(dispatcher.cfg.verification, "local_llm_enabled", False)
+ monkeypatch.setattr(dispatcher.cfg.routing, "require_vision", False)
+ monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
+
+ with TestClient(dispatcher.app) as client:
+ yield client
+
+
+def test_admin_models_returns_all_models_with_proficiency(seeded_client):
+ """GET /admin/api/models returns one object per seeded model, each with a
+ category -> blended_score proficiency map."""
+ resp = seeded_client.get("/admin/api/models")
+ assert resp.status_code == 200
+ rows = resp.json()
+ assert isinstance(rows, list)
+ assert len(rows) == 3
+
+ by_id = {r["model_id"]: r for r in rows}
+ assert set(by_id) == {"cheap", "dear", "tiny"}
+
+ cheap = by_id["cheap"]
+ assert cheap["proficiency"] == {"coding_general": 0.9, "debugging": 0.85}
+ assert by_id["dear"]["proficiency"] == {"coding_general": 0.95}
+ assert by_id["tiny"]["proficiency"] == {}
+
+
+def test_admin_models_row_shape(seeded_client):
+ """Each model object carries every required scalar field."""
+ row = seeded_client.get("/admin/api/models").json()[0]
+ for key in (
+ "model_id",
+ "provider",
+ "base_model_id",
+ "display_name",
+ "availability",
+ "tier",
+ "context_window",
+ "effective_context_window",
+ "latency_class",
+ "reasoning_mode",
+ "context_variant",
+ "access_level",
+ "supports_tools",
+ "supports_json_mode",
+ "supports_vision",
+ "supports_reasoning",
+ "reasoning_default_enabled",
+ "cost_per_1m_prompt",
+ "cost_per_1m_completion",
+ "proficiency",
+ ):
+ assert key in row, f"missing field {key!r}"
+
+ assert row["provider"] == "neuralwatt"
+ assert row["availability"] == "active"
+ assert row["tier"] == 2
+ assert row["context_window"] == 262128
+ assert row["supports_tools"] is True
+ assert row["supports_json_mode"] is True
+ assert row["supports_vision"] is True
+ assert row["supports_reasoning"] is True
+ assert row["reasoning_default_enabled"] is True
+
+
+def test_admin_model_detail_returns_single_object(seeded_client):
+ """GET /admin/api/models/{model_id}/{provider} returns one object, not a list."""
+ resp = seeded_client.get("/admin/api/models/cheap/neuralwatt")
+ assert resp.status_code == 200
+ data = resp.json()
+ assert isinstance(data, dict)
+ assert data["model_id"] == "cheap"
+ assert data["provider"] == "neuralwatt"
+ assert data["proficiency"] == {"coding_general": 0.9, "debugging": 0.85}
+
+
+def test_admin_model_detail_unknown_model_returns_404(seeded_client):
+ """A model_id that does not exist -> 404."""
+ resp = seeded_client.get("/admin/api/models/nope/neuralwatt")
+ assert resp.status_code == 404
+
+
+def test_admin_model_detail_unknown_provider_returns_404(seeded_client):
+ """A provider that does not exist for a known model -> 404."""
+ resp = seeded_client.get("/admin/api/models/cheap/nope")
+ assert resp.status_code == 404
+
+
+def test_admin_models_empty_table_returns_empty_list(empty_client):
+ """GET /admin/api/models with zero rows -> 200 + empty list."""
+ resp = empty_client.get("/admin/api/models")
+ assert resp.status_code == 200
+ assert resp.json() == []
+
+
+def test_admin_models_never_expose_session_dir(seeded_client):
+ """No response object may name session_dir or carry a prompt/conversation key."""
+ rows = seeded_client.get("/admin/api/models").json()
+ for row in rows:
+ for key in ("session_dir", "prompt", "conversation"):
+ assert key not in row, f"leaked {key!r} in model row"
+
+
+def test_admin_model_detail_never_expose_session_dir(seeded_client):
+ """The single-model JSON must never name session_dir or conversation keys."""
+ data = seeded_client.get("/admin/api/models/cheap/neuralwatt").json()
+ for key in ("session_dir", "prompt", "conversation"):
+ assert key not in data, f"leaked {key!r} in model detail"
+
+# --- admin model overrides --------------------------------------------------
+
+
+def test_post_override_sets_effective_availability_deprecated(seeded_client):
+ """POST override -> effective_availability flips to deprecated, is_overridden
+ is True, raw availability stays 'active'."""
+ resp = seeded_client.post(
+ "/admin/api/models/cheap/neuralwatt/availability",
+ json={"availability": "deprecated", "reason": "failing verification"},
+ )
+ assert resp.status_code == 200
+ data = resp.json()
+ assert data["is_overridden"] is True
+ assert data["effective_availability"] == "deprecated"
+ assert data["availability"] == "active"
+
+
+def test_post_override_sets_effective_availability_stale(seeded_client):
+ """POST override with stale availability."""
+ resp = seeded_client.post(
+ "/admin/api/models/cheap/neuralwatt/availability",
+ json={"availability": "stale", "reason": "last seen long ago"},
+ )
+ assert resp.status_code == 200
+ data = resp.json()
+ assert data["is_overridden"] is True
+ assert data["effective_availability"] == "stale"
+
+
+def test_post_override_invalid_availability_returns_422(seeded_client):
+ """Bad availability value -> 422 validation error."""
+ resp = seeded_client.post(
+ "/admin/api/models/cheap/neuralwatt/availability",
+ json={"availability": "retired", "reason": "nope"},
+ )
+ assert resp.status_code == 422
+
+
+def test_post_override_nonexistent_model_returns_404(seeded_client):
+ """Trying to override a model that doesn't exist -> 404."""
+ resp = seeded_client.post(
+ "/admin/api/models/nonexistent/neuralwatt/availability",
+ json={"availability": "deprecated", "reason": "test"},
+ )
+ assert resp.status_code == 404
+
+
+def test_delete_override_reverts_to_db_value(seeded_client):
+ """POST then DELETE -> is_overridden becomes False, effective_availability
+ reverts to the DB value ('active')."""
+ # Set override
+ seeded_client.post(
+ "/admin/api/models/cheap/neuralwatt/availability",
+ json={"availability": "deprecated", "reason": "test"},
+ )
+ # Delete it
+ resp = seeded_client.delete(
+ "/admin/api/models/cheap/neuralwatt/availability"
+ )
+ assert resp.status_code == 200
+ data = resp.json()
+ assert data["is_overridden"] is False
+ assert data["effective_availability"] == "active"
+ assert data["availability"] == "active"
+
+
+def test_model_list_reflects_effective_availability(seeded_client):
+ """GET /admin/api/models includes effective_availability and is_overridden."""
+ seeded_client.post(
+ "/admin/api/models/cheap/neuralwatt/availability",
+ json={"availability": "deprecated", "reason": "test"},
+ )
+ rows = seeded_client.get("/admin/api/models").json()
+ by_id = {r["model_id"]: r for r in rows}
+ cheap = by_id["cheap"]
+ assert "effective_availability" in cheap
+ assert "is_overridden" in cheap
+ # dear and tiny should not be overridden
+ assert cheap["is_overridden"] is True
+ assert cheap["effective_availability"] == "deprecated"
+ assert by_id["dear"]["is_overridden"] is False
+
diff --git a/tests/test_admin_routing_override.py b/tests/test_admin_routing_override.py
new file mode 100644
index 0000000..39fcd1b
--- /dev/null
+++ b/tests/test_admin_routing_override.py
@@ -0,0 +1,286 @@
+"""Integration test: admin_model_overrides wire into routing.
+
+Verifies that POST /admin/api/models/{model_id}/{provider}/availability
+actually removes the model from /route candidates. The wired exclude set
+(_admin_deprecated_models) must intersect with select_candidates'
+exclude_models filter so the overridden model never appears in the
+route response.
+
+Uses the pattern from test_route_decisions.py: temp DB, seeded models with
+energy + proficiency so a specific model WOULD win, then override + /route
+assertion.
+"""
+
+from __future__ import annotations
+
+import json
+import sqlite3
+import time
+from pathlib import Path
+from typing import Any
+
+import pytest
+from openai import OpenAI
+from starlette.testclient import TestClient
+
+import dispatcher
+from config import load_config
+
+ROOT = Path(__file__).resolve().parent.parent
+SCHEMA_SQL = (ROOT / "schema.sql").read_text()
+CFG = load_config(str(ROOT / "config.yaml"))
+
+ADMIN_TABLE_SQL = """
+CREATE TABLE IF NOT EXISTS admin_model_overrides (
+ model_id TEXT NOT NULL,
+ provider TEXT NOT NULL,
+ availability TEXT NOT NULL,
+ reason TEXT,
+ updated_at TEXT NOT NULL,
+ PRIMARY KEY (model_id, provider)
+);
+CREATE INDEX IF NOT EXISTS idx_admin_model_overrides_availability
+ ON admin_model_overrides (availability);
+"""
+
+# A model we will seed with best proficiency + energy so it WOULD be picked.
+WINNER_MODEL = "premium"
+# A second model that would be runner-up.
+RUNNER_MODEL = "mid"
+
+
+def _completion(model_id: str) -> dict:
+ """An OpenAI-compatible completion body for a *fake* provider response."""
+ return {
+ "id": f"chatcmpl-{model_id}",
+ "object": "chat.completion",
+ "model": model_id,
+ "created": int(time.time()),
+ "choices": [{"index": 0, "finish_reason": "stop", "message": {"role": "assistant", "content": "done"}}],
+ }
+
+
+class FakeResponse:
+ """Minimal fake for requests.post() and httpx.Response."""
+
+ def __init__(
+ self,
+ body: dict | None = None,
+ lines: list[str] | None = None,
+ headers: dict | None = None,
+ ) -> None:
+ self.body = body or _completion("dummy")
+ self.lines = lines or []
+ self.headers = headers or {"content-type": "application/json"}
+ self.status_code = 200
+
+ @property
+ def text(self) -> str:
+ return json.dumps(self.body)
+
+ @property
+ def content(self) -> bytes:
+ return json.dumps(self.body).encode()
+
+ def json(self) -> dict:
+ return self.body
+
+
+@pytest.fixture
+def admin_override_router(tmp_path: Path, monkeypatch) -> tuple[TestClient, Path]:
+ """A TestClient with a temp DB that has admin_model_overrides table,
+ seeded with two models (WINNER_MODEL and RUNNER_MODEL) where WINNER
+ has best proficiency + energy."""
+
+ import admin
+
+ db_path = tmp_path / "test.db"
+ conn = sqlite3.connect(str(db_path))
+ conn.executescript(SCHEMA_SQL)
+ conn.executescript(ADMIN_TABLE_SQL)
+ dispatcher.ensure_route_decisions(conn)
+
+ # Seed two models: WINNER (tier 2, best) and RUNNER (tier 2).
+ for mid, cost_prompt, cost_compl, prof_score in [
+ (WINNER_MODEL, 0.50, 0.30, 0.95),
+ (RUNNER_MODEL, 0.30, 0.15, 0.60),
+ ]:
+ conn.execute(
+ """
+ INSERT INTO models (
+ model_id, provider, base_model_id, display_name, tier,
+ context_window, effective_context_window, max_output_tokens,
+ cost_per_1m_prompt, cost_per_1m_completion,
+ supports_tools, supports_json_mode, supports_vision,
+ supports_reasoning, reasoning_default_enabled,
+ latency_class, reasoning_mode, context_variant,
+ access_level, availability, last_updated
+ ) VALUES (?, 'neuralwatt', ?, ?, ?, 262128, 192500, 16384,
+ ?, ?, 1, 1, 1, 1, 1,
+ 'standard', 'default', 'full', 'public', 'active',
+ '2026-08-22T00:00:00+00:00')
+ """,
+ (
+ mid,
+ mid,
+ mid,
+ 2,
+ cost_prompt,
+ cost_compl,
+ ),
+ )
+
+ # Seed proficiency: WINNER has the best score for coding_general.
+ conn.execute(
+ "INSERT INTO proficiency (model_id, provider, category, "
+ "blended_score, source, last_updated) "
+ "VALUES (?, 'neuralwatt', 'coding_general', ?, 'self_eval_thin', "
+ "'2026-01-01T00:00:00+00:00')",
+ (WINNER_MODEL, 0.95),
+ )
+ conn.execute(
+ "INSERT INTO proficiency (model_id, provider, category, "
+ "blended_score, source, last_updated) "
+ "VALUES (?, 'neuralwatt', 'coding_general', ?, 'self_eval_thin', "
+ "'2026-01-01T00:00:00+00:00')",
+ (RUNNER_MODEL, 0.60),
+ )
+
+ # Seed one energy observation so scoring works.
+ now = "2026-08-22T00:00:00+00:00"
+ for mid in (WINNER_MODEL, RUNNER_MODEL):
+ conn.execute(
+ """
+ INSERT INTO energy_observations (
+ model_id, provider, prompt_tokens, completion_tokens,
+ energy_kwh, carbon_g_co2eq, cost_usd, observed_at
+ ) VALUES (?, 'neuralwatt', 200, 500, 0.005, 2.0, 0.10, ?)
+ """,
+ (mid, now),
+ )
+
+ conn.commit()
+ conn.close()
+
+ # Redirect the dispatcher to the temp DB.
+ monkeypatch.setattr(dispatcher.cfg.database, "path", str(db_path))
+ monkeypatch.setattr(dispatcher.cfg.verification, "local_llm_enabled", False)
+ monkeypatch.setattr(dispatcher.cfg.freshness, "exclude_stale", True)
+ monkeypatch.setattr(dispatcher.cfg.freshness, "exclude_deprecated", True)
+ monkeypatch.setattr(dispatcher.cfg.routing, "require_vision", False)
+ monkeypatch.setattr(dispatcher.cfg.logging, "log_route_decisions", True)
+ monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
+
+ # Ensure admin tables are wired into the startup migration.
+ from admin import ensure_admin_tables
+
+ temp_conn = sqlite3.connect(str(db_path))
+ temp_conn.row_factory = sqlite3.Row
+ try:
+ ensure_admin_tables(conn=temp_conn)
+ except Exception:
+ pass # may already exist
+ finally:
+ temp_conn.close()
+
+ # Stub the classifier — always returns coding_general, tier 2.
+ monkeypatch.setattr(
+ dispatcher,
+ "classify",
+ lambda task, context: dispatcher.Classification(
+ task_category="coding_general",
+ task_tier=2,
+ required_context_tokens=100,
+ confidence=0.9,
+ ),
+ )
+
+ # Stub the provider call so the /route endpoint doesn't actually call
+ # NeuralWatt. It just returns a minimal response.
+ def fake_post(url: str, headers: Any = None, json: Any = None,
+ stream: bool = False, timeout: int = 600):
+ if stream:
+ return FakeResponse(lines=["data: ..."])
+ return FakeResponse(_completion(json["model"] if json else "dummy"))
+
+ monkeypatch.setattr(dispatcher.requests, "post", fake_post)
+
+ app = dispatcher.app
+
+ with TestClient(app) as client:
+ yield client, db_path
+
+
+def test_admin_override_excludes_model_from_route(admin_override_router):
+ """When an admin override marks WINNER_MODEL as deprecated, POST /route
+ must NOT pick it. The selected model should be RUNNER_MODEL instead,
+ and candidates_considered should be 1 (not 2).
+
+ This is the CRITICAL integration test: override → router exclusion.
+ """
+ client, db_path = admin_override_router
+
+ # First, verify that WITHOUT an override, the router picks WINNER.
+ resp = client.post("/route", json={"task": "write a python function"})
+ assert resp.status_code == 200
+ body = resp.json()
+ assert body["selected"]["model_id"] == WINNER_MODEL
+ assert body["candidates_considered"] == 2
+
+ # Now mark WINNER as deprecated via admin override.
+ override_resp = client.post(
+ f"/admin/api/models/{WINNER_MODEL}/neuralwatt/availability",
+ json={"availability": "deprecated", "reason": "failing verification"},
+ )
+ assert override_resp.status_code == 200
+ override_data = override_resp.json()
+ assert override_data["is_overridden"] is True
+ assert override_data["effective_availability"] == "deprecated"
+
+ # POST /route again: the router must NOT pick the overridden model.
+ resp = client.post("/route", json={"task": "write a python function"})
+ assert resp.status_code == 200
+ body = resp.json()
+ assert body["selected"]["model_id"] == RUNNER_MODEL
+ assert body["candidates_considered"] == 1
+ # Verify the override model is in the excluded set via the decision log.
+ conn = sqlite3.connect(str(db_path))
+ conn.row_factory = sqlite3.Row
+ decision_row = conn.execute(
+ "SELECT * FROM route_decisions ORDER BY id DESC LIMIT 1"
+ ).fetchone()
+ conn.close()
+ assert decision_row["selected_model"] == RUNNER_MODEL
+
+
+def test_admin_override_revert_includes_model_again(admin_override_router):
+ """After DELETE on the admin override, the model must become routable
+ again and be picked if it still has best scores."""
+ client, db_path = admin_override_router
+
+ # Mark as deprecated.
+ client.post(
+ f"/admin/api/models/{WINNER_MODEL}/neuralwatt/availability",
+ json={"availability": "deprecated", "reason": "test"},
+ )
+
+ # Verify route picks RUNNER.
+ resp = client.post("/route", json={"task": "write a python function"})
+ assert resp.json()["selected"]["model_id"] == RUNNER_MODEL
+
+ # Delete override.
+ delete_resp = client.delete(
+ f"/admin/api/models/{WINNER_MODEL}/neuralwatt/availability"
+ )
+ assert delete_resp.status_code == 200
+ del_data = delete_resp.json()
+ assert del_data["is_overridden"] is False
+
+ # Route again: WINNER should be back on the radar.
+ resp = client.post("/route", json={"task": "write a python function"})
+ assert resp.status_code == 200
+ body = resp.json()
+ assert body["selected"]["model_id"] == WINNER_MODEL
+ assert body["candidates_considered"] == 2
+
+
diff --git a/tests/test_admin_runtime.py b/tests/test_admin_runtime.py
new file mode 100644
index 0000000..51db76c
--- /dev/null
+++ b/tests/test_admin_runtime.py
@@ -0,0 +1,180 @@
+"""Tests for the /admin/api runtime config toggle endpoints.
+
+``admin.py`` exposes GET /admin/api/runtime (persisted + in-memory state for
+each toggle knob) and POST /admin/api/runtime/{knob} (flip the in-memory cfg
+value only, never config.yaml). These tests drive a real TestClient against the
+seeded temp DB, mirroring ``test_admin_health.py``'s ``seeded_client`` fixture.
+
+The key contract under test: a POST flips the RUNTIME value but must leave the
+PERSISTED (config.yaml) value untouched.
+"""
+
+from __future__ import annotations
+
+import sqlite3
+from pathlib import Path
+
+import pytest
+from starlette.testclient import TestClient
+
+import dispatcher
+from config import load_config
+
+ROOT = Path(__file__).resolve().parent.parent
+SCHEMA_SQL = (ROOT / "schema.sql").read_text()
+CFG = load_config(str(ROOT / "config.yaml"))
+
+
+def _make_db(tmp_path: Path) -> sqlite3.Connection:
+ conn = sqlite3.connect(str(tmp_path / "test.db"))
+ conn.row_factory = sqlite3.Row
+ conn.executescript(SCHEMA_SQL)
+ return conn
+
+
+def _seed_models(conn: sqlite3.Connection) -> None:
+ for model_id, tier, context, cost, vision in (
+ ("cheap", 2, 262128, 0.30, 1),
+ ("dear", 2, 262128, 9.00, 0),
+ ("tiny", 1, 131072, 0.10, 1),
+ ):
+ conn.execute(
+ """
+ INSERT INTO models (
+ model_id, provider, base_model_id, tier, context_window,
+ effective_context_window, max_output_tokens,
+ cost_per_1m_prompt, cost_per_1m_completion,
+ supports_vision, supports_json_mode,
+ latency_class, reasoning_mode, context_variant,
+ access_level, availability, last_updated
+ ) VALUES (?, 'neuralwatt', ?, ?, ?, 192500, 16384, ?, ?,
+ ?, 1, 'standard', 'default', 'full', 'public', 'active',
+ '2026-08-22T00:00:00+00:00')
+ """,
+ (model_id, model_id, tier, context, cost, cost / 3, vision),
+ )
+ conn.commit()
+
+
+@pytest.fixture
+def seeded_client(tmp_path, monkeypatch):
+ """A TestClient wired to a seeded temp DB, at /admin."""
+ conn = _make_db(tmp_path)
+ _seed_models(conn)
+ conn.close()
+
+ monkeypatch.setattr(dispatcher.cfg.database, "path", str(tmp_path / "test.db"))
+ monkeypatch.setattr(dispatcher.cfg.verification, "local_llm_enabled", False)
+ monkeypatch.setattr(dispatcher.cfg.routing, "require_vision", False)
+ monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
+
+ with TestClient(dispatcher.app) as client:
+ yield client
+
+
+def test_runtime_GET_reports_every_knob(seeded_client):
+ """GET /admin/api/runtime reports persisted + runtime for all 8 knobs."""
+ resp = seeded_client.get("/admin/api/runtime")
+ assert resp.status_code == 200
+ body = resp.json()
+ for knob in (
+ "log_route_decisions",
+ "log_energy_observations",
+ "circuit_breaker",
+ "local_llm_enabled",
+ "session_cache_enabled",
+ "pinch_enabled",
+ "pinch_relevance_enabled",
+ "default_flex_preference",
+ ):
+ assert knob in body
+ assert set(body[knob]) == {"persisted", "runtime"}
+
+ # The runtime value mirrors the live dispatcher.cfg for the boolean knobs;
+ # the persisted value mirrors config.yaml.
+ assert body["log_route_decisions"]["runtime"] == dispatcher.cfg.logging.log_route_decisions
+ assert body["log_route_decisions"]["persisted"] == CFG.logging.log_route_decisions
+ assert body["circuit_breaker"] == {
+ "persisted": {"enabled": CFG.circuit_breaker.enabled},
+ "runtime": {"enabled": dispatcher.cfg.circuit_breaker.enabled},
+ }
+ assert body["default_flex_preference"] == {
+ "persisted": CFG.routing.default_flex_preference.value,
+ "runtime": dispatcher.cfg.routing.default_flex_preference.value,
+ }
+
+
+def test_toggle_log_route_decisions_flips_runtime_not_persisted(
+ seeded_client, monkeypatch
+):
+ """POST a boolean knob flips the runtime value; config.yaml is untouched."""
+ # Guard shared global state: record the original so it restores after this
+ # test, since the POST mutates dispatcher.cfg in place.
+ monkeypatch.setattr(
+ dispatcher.cfg.logging,
+ "log_route_decisions",
+ dispatcher.cfg.logging.log_route_decisions,
+ )
+ initial = seeded_client.get("/admin/api/runtime").json()
+ assert initial["log_route_decisions"]["runtime"] is True
+
+ resp = seeded_client.post(
+ "/admin/api/runtime/log_route_decisions", json={"value": False}
+ )
+ assert resp.status_code == 200
+
+ after = seeded_client.get("/admin/api/runtime").json()
+ # Runtime flipped ...
+ assert after["log_route_decisions"]["runtime"] is False
+ # ... but persisted is unchanged and still matches config.yaml.
+ assert after["log_route_decisions"]["persisted"] == CFG.logging.log_route_decisions
+ assert after["log_route_decisions"]["persisted"] == initial["log_route_decisions"][
+ "persisted"
+ ]
+
+
+def test_post_invalid_flex_value_returns_422(seeded_client):
+ """A flex preference outside the 4 allowed values is rejected with 422."""
+ resp = seeded_client.post(
+ "/admin/api/runtime/default_flex_preference", json={"value": "banana"}
+ )
+ assert resp.status_code == 422
+ # The in-memory value must not have changed.
+ assert (
+ seeded_client.get("/admin/api/runtime").json()["default_flex_preference"][
+ "runtime"
+ ]
+ == "auto"
+ )
+
+
+def test_post_valid_flex_value_updates_runtime(seeded_client, monkeypatch):
+ """A valid flex preference is accepted and reflected in runtime state."""
+ monkeypatch.setattr(
+ dispatcher.cfg.routing,
+ "default_flex_preference",
+ dispatcher.cfg.routing.default_flex_preference,
+ )
+ resp = seeded_client.post(
+ "/admin/api/runtime/default_flex_preference", json={"value": "force-flex"}
+ )
+ assert resp.status_code == 200
+ assert seeded_client.get("/admin/api/runtime").json()["default_flex_preference"][
+ "runtime"
+ ] == "force-flex"
+
+
+def test_post_unknown_knob_returns_400(seeded_client):
+ """POSTing an unknown knob name is rejected with 400."""
+ resp = seeded_client.post(
+ "/admin/api/runtime/not_a_real_knob", json={"value": True}
+ )
+ assert resp.status_code == 400
+
+
+def test_post_non_boolean_for_bool_knob_returns_422(seeded_client):
+ """POSTing a non-boolean value for a boolean knob is rejected with 422."""
+ resp = seeded_client.post(
+ "/admin/api/runtime/session_cache_enabled", json={"value": "yes"}
+ )
+ assert resp.status_code == 422
diff --git a/tests/test_admin_snapshot.py b/tests/test_admin_snapshot.py
new file mode 100644
index 0000000..3287b30
--- /dev/null
+++ b/tests/test_admin_snapshot.py
@@ -0,0 +1,176 @@
+"""Tests for GET /admin/api/snapshot: the full top-level key contract.
+
+The snapshot is the admin dashboard's main data payload. This file locks the
+complete key set (not just a hand-picked subset) so a dropped or renamed key
+fails the suite instead of silently missing from the UI. It drives a real
+TestClient against a seeded temp DB, mirroring ``test_admin_health.py``.
+"""
+
+from __future__ import annotations
+
+import sqlite3
+from datetime import datetime, timedelta, timezone
+from pathlib import Path
+
+import pytest
+from starlette.testclient import TestClient
+
+import dispatcher
+from config import load_config
+
+ROOT = Path(__file__).resolve().parent.parent
+SCHEMA_SQL = (ROOT / "schema.sql").read_text()
+CFG = load_config(str(ROOT / "config.yaml"))
+
+EXHAUSTIVE_KEYS = (
+ "quota",
+ "coverage",
+ "recent_decisions",
+ "per_model",
+ "verdict_mix",
+ "top_proficiency",
+ "health",
+ "generated_at",
+)
+
+
+def _now() -> datetime:
+ return datetime.now(timezone.utc)
+
+
+def _make_db(tmp_path: Path) -> sqlite3.Connection:
+ conn = sqlite3.connect(str(tmp_path / "test.db"))
+ conn.row_factory = sqlite3.Row
+ conn.executescript(SCHEMA_SQL)
+ return conn
+
+
+def _seed_models(conn: sqlite3.Connection) -> None:
+ for model_id, tier, context, cost, vision in (
+ ("cheap", 2, 262128, 0.30, 1),
+ ("dear", 2, 262128, 9.00, 0),
+ ):
+ conn.execute(
+ """
+ INSERT INTO models (
+ model_id, provider, base_model_id, tier, context_window,
+ effective_context_window, max_output_tokens,
+ cost_per_1m_prompt, cost_per_1m_completion,
+ supports_vision, supports_json_mode,
+ latency_class, reasoning_mode, context_variant,
+ access_level, availability, last_updated
+ ) VALUES (?, 'neuralwatt', ?, ?, ?, 192500, 16384, ?, ?,
+ ?, 1, 'standard', 'default', 'full', 'public', 'active',
+ '2026-08-22T00:00:00+00:00')
+ """,
+ (model_id, model_id, tier, context, cost, cost / 3, vision),
+ )
+ conn.commit()
+
+
+def _seed_decision(conn: sqlite3.Connection) -> None:
+ conn.execute(
+ """
+ INSERT INTO route_decisions (
+ observed_at, kind, task_category, task_tier, required_context_tokens,
+ confidence, classifier_ms, classification_source, latency_tolerance,
+ candidates_considered, selected_model, selected_provider,
+ runner_up_models, est_cost_usd, est_proficiency,
+ session_key, tools, images, json_mode, streamed,
+ flex_preference, flex_swapped, flex_forced
+ ) VALUES (?, 'route', 'coding_general', 2, 100, 0.95, 200,
+ 'classifier', 'interactive', 5, 'cheap', 'neuralwatt',
+ '[{"model_id":"dear","provider":"neuralwatt"}]',
+ 0.001, 0.9, 'abc123', 0, 0, 0, 0,
+ 'auto', 0, 1)
+ """,
+ (_now().isoformat(),),
+ )
+ conn.commit()
+
+
+def _seed_energy(conn: sqlite3.Connection) -> None:
+ conn.execute(
+ "INSERT INTO energy_observations "
+ "(model_id, provider, task_category, completion_tokens, energy_kwh, "
+ "cost_usd, carbon_g_co2eq, attribution_ratio, observed_at) "
+ "VALUES ('cheap', 'neuralwatt', 'coding_general', 100, 5.0e-05, 0.001, "
+ "2.4e-03, 0.25, ?)",
+ ((_now() - timedelta(days=2)).isoformat(),),
+ )
+ conn.commit()
+
+
+def _seed_verification(conn: sqlite3.Connection) -> None:
+ conn.execute(
+ "INSERT INTO verifications (model_id, provider, kind, verdict, observed_at) "
+ "VALUES ('cheap', 'neuralwatt', 'structural', 'ok', ?)",
+ (_now().isoformat(),),
+ )
+ conn.commit()
+
+
+def _seed_proficiency(conn: sqlite3.Connection) -> None:
+ conn.execute(
+ "INSERT INTO proficiency (model_id, provider, category, blended_score, "
+ "source, last_updated) "
+ "VALUES ('cheap', 'neuralwatt', 'coding_general', 0.9, "
+ "'self_eval_thin', '2026-01-01T00:00:00+00:00')",
+ )
+ conn.commit()
+
+
+@pytest.fixture
+def seeded_client(tmp_path, monkeypatch):
+ conn = _make_db(tmp_path)
+ _seed_models(conn)
+ for _ in range(3):
+ _seed_decision(conn)
+ _seed_energy(conn)
+ _seed_verification(conn)
+ _seed_proficiency(conn)
+ conn.close()
+
+ monkeypatch.setattr(dispatcher.cfg.database, "path", str(tmp_path / "test.db"))
+ monkeypatch.setattr(dispatcher.cfg.verification, "local_llm_enabled", False)
+ monkeypatch.setattr(dispatcher.cfg.routing, "require_vision", False)
+ monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
+
+ with TestClient(dispatcher.app) as client:
+ yield client
+
+
+def test_admin_snapshot_returns_exhaustive_top_level_keys(seeded_client):
+ """Every documented top-level snapshot key is present in the response."""
+ resp = seeded_client.get("/admin/api/snapshot")
+ assert resp.status_code == 200
+ data = resp.json()
+ for key in EXHAUSTIVE_KEYS:
+ assert key in data, f"missing top-level snapshot key {key!r}"
+
+
+def test_admin_snapshot_generated_at_is_iso(seeded_client):
+ """generated_at is an ISO-8601 timestamp and freshly generated."""
+ data = seeded_client.get("/admin/api/snapshot").json()
+ generated = datetime.fromisoformat(data["generated_at"])
+ assert generated.tzinfo is not None
+ delta = abs((_now() - generated).total_seconds())
+ assert delta < 60, "snapshot generated_at is stale"
+
+
+def test_admin_snapshot_empty_db_still_returns_all_keys(tmp_path, monkeypatch):
+ """An unseeded DB returns the same key set, with empty collections."""
+ conn = _make_db(tmp_path)
+ conn.close()
+ monkeypatch.setattr(dispatcher.cfg.database, "path", str(tmp_path / "test.db"))
+ monkeypatch.setattr(dispatcher.cfg.verification, "local_llm_enabled", False)
+ monkeypatch.setattr(dispatcher.cfg.routing, "require_vision", False)
+ monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
+
+ with TestClient(dispatcher.app) as client:
+ data = client.get("/admin/api/snapshot").json()
+ for key in EXHAUSTIVE_KEYS:
+ assert key in data, f"empty-DB snapshot missing key {key!r}"
+ assert data["per_model"] == []
+ assert data["recent_decisions"] == []
+ assert data["top_proficiency"] == []
diff --git a/tests/test_admin_triggers.py b/tests/test_admin_triggers.py
new file mode 100644
index 0000000..041327e
--- /dev/null
+++ b/tests/test_admin_triggers.py
@@ -0,0 +1,194 @@
+"""Tests for the /admin/api operational trigger endpoints.
+
+refresh-catalog, seed-energy, and apply-feedback run repo maintenance scripts
+via ``asyncio.create_subprocess_exec``; restart-service schedules systemctl
+through FastAPI BackgroundTasks. These tests monkeypatch
+``asyncio.create_subprocess_exec`` (and ``subprocess.run``) so no real process
+or network is ever touched — they assert the endpoint maps query params and
+CLI args onto the spawned command and returns the documented job shape.
+"""
+
+from __future__ import annotations
+
+import sqlite3
+import sys
+from pathlib import Path
+
+import pytest
+from starlette.testclient import TestClient
+
+import admin
+import dispatcher
+from config import load_config
+
+ROOT = Path(__file__).resolve().parent.parent
+CFG = load_config(str(ROOT / "config.yaml"))
+
+
+class FakeProcess:
+ """A stand-in for ``asyncio.subprocess.Process`` with a known outcome."""
+
+ def __init__(self, stdout: bytes = b"ok\n", returncode: int = 0):
+ self._stdout = stdout
+ self.returncode = returncode
+ self.killed = False
+
+ async def communicate(self):
+ return self._stdout, None
+
+ async def wait(self):
+ return self.returncode
+
+ def kill(self):
+ self.killed = True
+
+
+class _Recorder:
+ """Captures every ``create_subprocess_exec`` call as ``(args, kwargs)``."""
+
+ def __init__(self):
+ self.calls: list[tuple[tuple, dict]] = []
+
+
+@pytest.fixture
+def fake_spawn(monkeypatch):
+ """Replace ``asyncio.create_subprocess_exec`` with a success-bound fake."""
+ recorder = _Recorder()
+
+ async def _fake(*args, **kwargs):
+ recorder.calls.append((list(args), kwargs))
+ return FakeProcess(stdout=b"ok\n", returncode=0)
+
+ monkeypatch.setattr(admin.asyncio, "create_subprocess_exec", _fake)
+ return recorder
+
+
+@pytest.fixture
+def failing_spawn(monkeypatch):
+ """A fake that makes any spawned command fail with returncode 3."""
+ recorder = _Recorder()
+
+ async def _fake(*args, **kwargs):
+ recorder.calls.append((list(args), kwargs))
+ return FakeProcess(stdout=b"boom\n", returncode=3)
+
+ monkeypatch.setattr(admin.asyncio, "create_subprocess_exec", _fake)
+ return recorder
+
+
+@pytest.fixture
+def seeded_client(tmp_path, monkeypatch):
+ """A TestClient wired to dispatcher.app with a temp DB (mounts /admin)."""
+ conn = sqlite3.connect(str(tmp_path / "test.db"))
+ conn.executescript((ROOT / "schema.sql").read_text())
+ conn.close()
+ monkeypatch.setattr(dispatcher.cfg.database, "path", str(tmp_path / "test.db"))
+ monkeypatch.setenv("NEURALWATT_API_KEY", "test-key")
+ with TestClient(dispatcher.app) as client:
+ yield client
+
+
+# --- refresh-catalog --------------------------------------------------------
+
+
+def test_refresh_catalog_runs_poller_then_tier(seeded_client, fake_spawn):
+ """POST /admin/api/refresh-catalog spawns poller then tier, &&-semantics."""
+ resp = seeded_client.post("/admin/api/refresh-catalog")
+ assert resp.status_code == 200
+ job = resp.json()
+ assert job["status"] == "success"
+ assert job["returncode"] == 0
+ assert "poller.py" in job["command"]
+ assert "tier.py" in job["command"]
+
+ # One spawn per step, in order, with the venv python as the interpreter.
+ argv = [args for args, _ in fake_spawn.calls]
+ assert argv == [
+ [sys.executable, "poller.py"],
+ [sys.executable, "tier.py"],
+ ]
+ for _, kwargs in fake_spawn.calls:
+ assert kwargs["cwd"] == str(ROOT)
+
+
+def test_refresh_catalog_job_shape(seeded_client, fake_spawn):
+ """The job object carries id/command/status/returncode/output_tail."""
+ job = seeded_client.post("/admin/api/refresh-catalog").json()
+ assert isinstance(job["id"], str) and job["id"]
+ assert "command" in job
+ assert job["status"] == "success"
+ assert job["returncode"] == 0
+ # Two steps each emit "ok\n", so the merged tail carries both.
+ assert job["output_tail"] == "ok\nok\n"
+
+
+def test_failed_command_reports_failure(seeded_client, failing_spawn):
+ """A non-zero returncode is reported as failed with that returncode."""
+ job = seeded_client.post("/admin/api/apply-feedback").json()
+ assert job["status"] == "failed"
+ assert job["returncode"] == 3
+ assert job["output_tail"] == "boom\n"
+
+
+# --- seed-energy ------------------------------------------------------------
+
+
+def test_seed_energy_default_samples(seeded_client, fake_spawn):
+ """Without ?samples=, seed-energy defaults to ``--samples 5``."""
+ seeded_client.post("/admin/api/seed-energy")
+ argv = fake_spawn.calls[0][0]
+ assert argv == [sys.executable, "seed_energy.py", "--samples", "5"]
+
+
+def test_seed_energy_accepts_samples_query(seeded_client, fake_spawn):
+ """POST /admin/api/seed-energy?samples=3 forwards ``--samples 3``."""
+ resp = seeded_client.post("/admin/api/seed-energy?samples=3")
+ assert resp.status_code == 200
+ assert resp.json()["status"] == "success"
+ argv = fake_spawn.calls[0][0]
+ assert argv == [sys.executable, "seed_energy.py", "--samples", "3"]
+
+
+# --- apply-feedback ---------------------------------------------------------
+
+
+def test_apply_feedback_apply_mode(seeded_client, fake_spawn):
+ """Without dry_run, feedback.py runs with no extra flag."""
+ resp = seeded_client.post("/admin/api/apply-feedback")
+ assert resp.status_code == 200
+ assert resp.json()["status"] == "success"
+ argv = fake_spawn.calls[0][0]
+ assert argv == [sys.executable, "feedback.py"]
+
+
+def test_apply_feedback_dry_run(seeded_client, fake_spawn):
+ """?dry_run=true appends ``--dry-run`` to the feedback command."""
+ seeded_client.post("/admin/api/apply-feedback?dry_run=true")
+ argv = fake_spawn.calls[0][0]
+ assert argv == [sys.executable, "feedback.py", "--dry-run"]
+
+
+# --- restart-service --------------------------------------------------------
+
+
+def test_restart_service_returns_immediately(seeded_client, monkeypatch):
+ """POST /admin/api/restart-service returns {"status": "restarting"} 200.
+
+ The systemctl call is scheduled via BackgroundTasks, never awaited inline,
+ so the response body is the immediate "restarting" status and the actual
+ restart fires as a post-response background task.
+ """
+ calls: list[tuple] = []
+ monkeypatch.setattr(
+ admin.subprocess, "run", lambda *a, **k: calls.append((a, k))
+ )
+
+ resp = seeded_client.post("/admin/api/restart-service")
+ assert resp.status_code == 200
+ assert resp.json() == {"status": "restarting"}
+
+ # The BackgroundTask ran after the response was produced; it must have
+ # scheduled exactly the systemctl restart command.
+ assert calls, "BackgroundTask never fired systemctl"
+ spawned = calls[0][0][0]
+ assert spawned == ["systemctl", "--user", "restart", "llm-router.service"]