diff --git a/.gitignore b/.gitignore
index d23c91e..4de9287 100644
--- a/.gitignore
+++ b/.gitignore
@@ -7,3 +7,11 @@ router.log
.venv/
venv/
.omo/
+config.yaml.bak.*
+node_modules/
+package.json
+package-lock.json
+.playwright-mcp/
+.mypy_cache/
+.pytest_cache/
+.ruff_cache/
diff --git a/AGENTS.md b/AGENTS.md
index 568b657..2388947 100644
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -18,10 +18,10 @@ code.
| Language | Python 3.10+ (3.10 floor tested; 3.14 also verified) |
| Framework | FastAPI + uvicorn |
| Database | SQLite (`router.db`) |
-| Config | `config.yaml` + Pydantic (`config.py`), `extra="forbid"` |
+| Config | `config/config.yaml` + Pydantic (`src/config.py`), `extra="forbid"` |
| HTTP client | `requests` (pinned in `requirements.txt`) — do NOT add httpx2/aiohttp without a requirements bump |
| TUI | `textual==8.2.8` — imported only by `tui*.py` modules, never by the dispatch path |
-| Testing | `pytest`, 562 tests, all offline (no provider or local-model calls) |
+| Testing | `pytest`, 733 tests, all offline (no provider or local-model calls) |
| Dependencies | Pinned. Bump deliberately, never use `>=` |
## Module map and file boundaries
@@ -48,6 +48,19 @@ code.
| `proficiency.py` | Score blending: leaderboard + self-eval → weighted composite |
| `iteration.py` | Retry budget per tier, matching retry to failure kind |
+### Admin portal
+
+| File | Role |
+|---|---|
+| `admin.py` | FastAPI sub-router mounted by dispatcher at `/admin`; serves API + static HTML |
+| `admin/frontend/index.html` | Dashboard: quota chip, per-model usage, verdict mix, category breakdown, history mini-charts, recent decisions |
+| `admin/frontend/models.html` | Model availability table with override dropdown |
+| `admin/frontend/decisions.html` | Full decision log with filter + search |
+| `admin/frontend/controls.html` | Operational triggers, runtime knobs, persisted config editor |
+| `admin_schema.sql` | RBAC/audit schema ready for future auth work (config/admin_schema.sql) |
+
+When editing the admin UI, follow **`plans/admin-design-standards.md`** — design tokens, glass styling, responsive rules, and hard-won gotchas (e.g. navbar z-index, Chart.js canvas reuse, scroll-context bug). Follow **`plans/admin-work-framework.md`** for the *method* — surface-level diagnosis, grammar-first rule extraction, the pytest + Playwright-1400/800 + REFRESH_MS-idle evidence bar, and writing the commit as a standalone learning input. Key transferable rules: a settings list is one CSS-grid row (key | meta | fixed 176px control) — never restate a control's own value in a meta column, and a single-line button card goes full width with buttons in `.card-actions`, not into a `col-xl-6` beside a tall card.
+
### TUI modules (import `textual`; never imported by the dispatch path)
| File | Role |
@@ -77,7 +90,7 @@ code.
- No `# type: ignore`, no `as any`, no `@ts-ignore` equivalent.
- `except Exception` is acceptable at top-level boundaries with
`# noqa: BLE001` comment (see `dispatcher._refresh`, `persist_route_decision`).
-- Config is strict: every knob belongs in `config.yaml`, not only in a
+- Config is strict: every knob belongs in `config/config.yaml`, not only in a
Pydantic default. A default the file never mentions is invisible to a tuner.
- Named constants use `Final` in new code (`events.py`); existing code is
inconsistent — don't refactor just for this.
@@ -95,7 +108,7 @@ code.
- SQLite, `PRAGMA foreign_keys = ON`.
- `_db()` in `dispatcher.py` returns a `sqlite3.Connection` with
`row_factory = sqlite3.Row`.
-- Schema in `schema.sql` uses `CREATE TABLE IF NOT EXISTS` — safe to re-run.
+- Schema in `config/schema.sql` uses `CREATE TABLE IF NOT EXISTS` — safe to re-run.
- Code-side table creation (`ensure_route_decisions`, `proficiency_store.ensure_columns`)
mirrors the schema for live DBs that predate a feature.
@@ -116,25 +129,36 @@ code.
# Setup
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
-sqlite3 router.db < schema.sql
-cp .env.example .env # fill in NEURALWATT_API_KEY
+sqlite3 router.db < config/schema.sql
+cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
-# Start the service (binds 127.0.0.1:8080)
-python -m uvicorn dispatcher:app --reload
-# Or via systemd:
-systemctl --user start llm-router.service
+# Port 8080 is the systemd-managed PRODUCTION instance. opencode's own
+# model traffic goes through it (see opencode.json's baseURL) and
+# `Restart=always` resurrects it ~5s after any kill — so never
+# `pkill`/`kill` anything matching uvicorn/dispatcher/8080 to "free the
+# port". That fights the supervisor, and on this repo it can cut off your
+# own inference mid-task. See CLAUDE.md's "A bare SIGTERM could hang the
+# process forever" section for the incident this note comes from.
+#
+# To pick up a dispatcher.py change on the real instance:
+systemctl --user restart llm-router.service
+
+# For an ad hoc/manual run — iterating with --reload, a throwaway instance
+# for Playwright smoke tests against the admin frontend, anything that
+# isn't "use the real router" — bind a different port so it can't collide:
+PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
# Run the TUI (service must be running)
-python tui.py
+PYTHONPATH=src python -m tui
# Run the full test suite
python -m pytest
# Quick routing probe (no spend)
-python router_cli.py "Refactor this Django view"
+PYTHONPATH=src python -m router_cli "Refactor this Django view"
# Populate the catalog
-python poller.py && python tier.py
+PYTHONPATH=src python -m poller && PYTHONPATH=src python -m tier
```
## What's NOT built yet (open items)
@@ -149,9 +173,9 @@ See `CLAUDE.md` → "What's NOT built yet — pick up here" for the full list.
## After a code change
-- Run `python -m pytest` — 562 tests, ~30s.
+- Run `python -m pytest` — 733 tests, ~35s.
- If you changed `dispatcher.py`, restart the systemd service:
`systemctl --user restart llm-router.service` (it doesn't auto-reload code).
-- If you changed the TUI, run `python tui.py` to verify it starts.
+- If you changed the TUI, run `PYTHONPATH=src python -m tui` to verify it starts.
- Check `lsp_diagnostics` on changed files.
- Match the existing commit-message style: `fix:`, `feat:`, `docs:`, `test:`.
diff --git a/CLAUDE.md b/CLAUDE.md
index fc3b62d..1b61355 100644
--- a/CLAUDE.md
+++ b/CLAUDE.md
@@ -196,15 +196,15 @@ rather than from months of history.
## What's built and working
-- `schema.sql` — `models`, `proficiency`, `energy_observations`. Applies
- cleanly (`sqlite3 router.db < schema.sql`). `models` carries the serving
+- `config/schema.sql` — `models`, `proficiency`, `energy_observations`. Applies
+ cleanly (`sqlite3 router.db < config/schema.sql`). `models` carries the serving
class columns (below); `energy_observations` carries real carbon/cost.
- `poller.py` — fetches Neuralwatt's `/models` endpoint (public,
unauthenticated), normalizes, upserts, marks stale rows. **Verified
against the live API**: 19 models, and the `metadata.pricing` /
`metadata.capabilities` / `metadata.limits` field mappings are confirmed
correct.
-- `config.yaml` / `config.py` — weights, thresholds, provider settings,
+- `config/config.yaml` / `src/config.py` — weights, thresholds, provider settings,
Pydantic-validated.
- `scoring.py` — one `normalize_inverted` (cost and eco normalize
identically; they differ only in what is fed to them) plus the weighted
@@ -212,7 +212,7 @@ rather than from months of history.
- `seed_energy.py` — runs a fixed reference task N times per routable model
and writes `energy_observations` rows tagged `seed_reference`. This is what
makes `cost` and `eco` real numbers instead of the neutral 0.5. Re-run it
- after the catalog gains models: `python seed_energy.py --samples 5`
+ after the catalog gains models: `PYTHONPATH=src python -m seed_energy --samples 5`
(13 models x 5 = 65 calls, and the whole sweep cost **under a cent**).
- `tiering.py` / `tier.py` — pure tier resolver + the DB pass that applies it.
- `routing.py` — pure hard filters and ranking.
@@ -258,7 +258,7 @@ rather than from months of history.
`route_decisions`, per-model aggregates over `energy_observations`,
verification verdict mix, and top proficiency by category.
- `tui.py` — Textual terminal dashboard over `GET /metrics` and
- `GET /events/decisions`. A foreground entrypoint (`python tui.py`), not a
+ `GET /events/decisions`. A foreground entrypoint (`PYTHONPATH=src python -m tui`), not a
service. `textual` is imported only in the TUI modules (`tui.py`,
`tui_screens.py`, `tui_sse.py`), so the router's dispatch path has no UI
dependency. The dashboard has a live routing-decisions feed (via the SSE
@@ -270,7 +270,19 @@ rather than from months of history.
- `router_cli.py` — one-shot `/route` probe. Posts a task to the running router
and prints the decision tree, or emits raw JSON with `--json`. Spends no
quota because it only routes.
-- `tests/` — 562 tests across 27 files, all passing, all offline. Verified on
+- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — `/admin` management
+ portal served by the running router, loopback-only, no auth. Liquid-glass dark UI
+ (see `plans/admin-design-standards.md`). Pages: dashboard (`index.html`)
+ with quota chip, per-model usage, verdict-mix/category-breakdown bars, history
+ mini-charts and recent decisions; models (`models.html`) with availability
+ overrides; decisions (`decisions.html`) log with filter/search; controls
+ (`controls.html`) for operational triggers, runtime knobs, and persisted config
+ edits. SSE status dot + warnings bell are shared chrome. Operational triggers:
+ refresh catalog, seed energy, apply feedback, restart service. Runtime toggles
+ reset on restart. Persisted config edits limited to an allowlist. Model
+ availability overrides feed routing hard filters. Bucketed history endpoint over
+ `energy_observations` and `route_decisions`.
+- `tests/` — 733 tests across 41 files, all passing, all offline. Verified on
Python 3.10 and 3.14; nothing declares `requires-python`, so 3.10 is the
tested floor rather than a promised one.
@@ -289,7 +301,7 @@ failed write is logged and swallowed, because a decision record is worth
having but never worth failing or slowing a request for. The table is created
from code at module load (`_ensure_route_decisions_table`) and again on every
write (`ensure_route_decisions` inside `persist_route_decision`), mirroring
-the `proficiency_store.ensure_columns` migration pattern: `schema.sql` is
+ the `proficiency_store.ensure_columns` migration pattern: `config/schema.sql` is
`CREATE TABLE IF NOT EXISTS`, but a live `router.db` predating this table
needs the code-side migration. Both the module-load hook and the write-path
guarantee are idempotent and leave existing rows intact.
@@ -309,7 +321,7 @@ request body. Two new gates live in `routing.rejection_reason` and
`json_schema` and `routing.require_json_mode` is true. A model passes only if
`supports_json_mode = 1`; `NULL` also fails closed, for the same reason.
-Both gates default to **on** in `config.yaml`, because a wrong guess produces a
+Both gates default to **on** in `config/config.yaml`, because a wrong guess produces a
400. This is a deliberate asymmetry against the tool-proficiency gate below:
capability **flags** fail closed on unknown, while quality **measurements**
(tool proficiency, energy) admit on absent evidence ("unproven, not bad").
@@ -336,11 +348,11 @@ so failing early is better than an opaque provider 400.
### Local Ollama vision fallback
-`local_vision:` in `config.yaml` configures a fallback path for image requests
+`local_vision:` in `config/config.yaml` configures a fallback path for image requests
that find no cloud vision candidate — the cloud catalog excludes the cost
leader (deepseek) on vision, so without this fallback every image request
that would otherwise have routed there 422s instead. It is a **core feature**,
-**enabled by default** (`enabled: true`, both in `config.yaml` and in
+**enabled by default** (`enabled: true`, both in `config/config.yaml` and in
`LocalVisionConfig`'s own default, so a config that omits the section still
gets it). Disable it explicitly (`enabled: false`) on a host with no local
Ollama, or one that hasn't pulled the vision model.
@@ -1043,21 +1055,21 @@ in the same pass — it was declared, never read, and shadowed the
```bash
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
-sqlite3 router.db < schema.sql
-cp .env.example .env # fill in NEURALWATT_API_KEY
-python poller.py # populate the catalog
-python tier.py # resolve tiers
-python config.py # sanity-check config loads
-python -m uvicorn dispatcher:app --reload
+sqlite3 router.db < config/schema.sql
+cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
+PYTHONPATH=src python -m poller # populate the catalog
+PYTHONPATH=src python -m tier # resolve tiers
+PYTHONPATH=src python -m config # sanity-check config loads
+PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
```
-Then set what is deployment-specific in `config.yaml`: `classifier.model` and
+Then set what is deployment-specific in `config/config.yaml`: `classifier.model` and
`classifier.base_url` for your Ollama, and `objective.plan_kwh_per_period` to
your own plan's quota (it is reported in `/health` as burn against the
allowance; it does not gate anything).
Ollama must be reachable with the classifier model pulled — the name must
-match `classifier.model` in `config.yaml`, which points at a **Modelfile-tagged
+match `classifier.model` in `config/config.yaml`, which points at a **Modelfile-tagged
variant**, not the base library tag. Ollama loads a model at its library
Modelfile's default context unless told otherwise, and the base tag never
was — measured live on a 24GB card, `mistral-nemo:12b` alone came up at
@@ -1192,12 +1204,51 @@ are still honored correctly regardless of that policy (systemd tracks
deliberate stops separately from the Restart= decision), so this only adds
self-healing for the unexplained case.
-The signal's actual source is still open — live investigation and audit
-trail in `code_plans/router-unreachable-signal-investigation.md`. Confirmed
-so far: it is not `systemctl`, not the admin portal's restart trigger, not
-suspend/resume, not the OOM killer, and — via `auditd` — not delivered
-through the `kill` or `tgkill` syscalls either, which is why the watch was
-extended to `pidfd_send_signal` and the `rt_*sigqueueinfo` syscalls.
+**Resolved 2026-08-29 — the source was opencode, killing its own supply
+line.** Full audit trail in
+`plans/router-unreachable-signal-investigation.md`; the syscall-level
+extension to `pidfd_send_signal` (past `kill`/`tgkill`, both cleared by
+`auditd`) is what finally caught it. `sudo ausearch -k routerkill` matched
+five incidents in one afternoon to `pidfd_send_signal(..., SIGTERM)` /
+`SIGKILL`-after-escalation calls from a non-interactive `zsh -c "pkill ..."`
+(never in `~/.zsh_history`, since it's not a login shell), and
+`~/.local/share/opencode/log/opencode.log` matched every one of those
+timestamps, to the millisecond, to one opencode run testing the admin
+frontend: `pkill -f "uvicorn dispatcher:app"` (or `.*dispatcher`, or plain
+`"8080"`), then `python -m uvicorn dispatcher:app --host 127.0.0.1 --port
+8080` for its own throwaway instance. When the port came back occupied 5s
+later (`Restart=always` resurrecting the real service), the agent read that
+as "the kill didn't work" and escalated to `pkill -9` — the one case
+(16:01:00) that arrived as a bare `SIGKILL`, `status=9/KILL`, rather than a
+caught `SIGTERM`.
+
+This was self-inflicted in a sharper way than it looks: `opencode.json`
+points opencode's *own* model traffic at `http://127.0.0.1:8080/v1` — the
+same production instance it was killing to test against. Every kill briefly
+cut off the agent's own inference supply.
+
+The actual bug was in `AGENTS.md`, not in this service: its "How to run
+things" section told an agent to bring the router up with a bare
+`python -m uvicorn dispatcher:app --reload` (implicitly on 8080, no port
+flag) as an equally-valid alternative to `systemctl --user start`, with no
+warning that 8080 is normally already held by the supervised instance. An
+agent following that instruction and finding the port taken has no way to
+know the right move is `systemctl --user restart` (which the doc *does* say
+two sections later, for the "changed dispatcher.py" case, but not for "I
+want to smoke-test against a running instance"). Fixed there: manual/ad hoc
+runs now bind `--port 8081` explicitly, and the section says outright not to
+`pkill`/`kill` anything matching `uvicorn`/`dispatcher`/`8080` — that's the
+systemd-managed instance, `Restart=always` will fight you, and on this repo
+it may be your own model access.
+
+**Convention going forward: 8080 is production, always.** It's the port
+baked into `opencode.json`, every curl example in this file, the systemd
+unit, and the admin frontend's own fetches — moving it would touch more
+surface than the problem is worth. A throwaway instance (manual iteration,
+Playwright smoke tests against the admin frontend, anything that isn't "use
+the real router") binds **8081** instead, and nothing should ever send a
+kill signal to a process matched by name/port rather than by a PID it
+started itself.
## Pointing a coding agent at it
diff --git a/README.md b/README.md
index 8578ac9..4e22881 100644
--- a/README.md
+++ b/README.md
@@ -30,9 +30,14 @@ measurement is how you check whether it still holds for you.
- **Fall back to local vision when no cloud row supports images.** If no
vision-capable catalog candidate survives the hard filters, the router proxies
the request to a local Ollama vision model instead of returning 422.
-- **Watch decisions arrive live.** `python tui.py` opens a terminal dashboard
+- **Watch decisions arrive live.** `PYTHONPATH=src python -m tui` opens a terminal dashboard
that follows `/events/decisions` as decisions are recorded, with no polling
delay.
+- **Manage it from a browser.** `GET /admin/` serves a four-page glass dark-mode
+ portal from the running router: a dashboard (quota burn, per-model usage,
+ verdict mix, history), a model-availability table with routing overrides, a
+ searchable decision log, and a controls page for operational triggers, runtime
+ toggles, and allowlisted `config/config.yaml` edits. Loopback-only, no auth.
- **Verify before learning.** Every routed response is structurally parsed in
the background; larger prose answers get an async local-LLM spot-check, and
failures fold back into per-model proficiency through `feedback.py`.
@@ -54,6 +59,7 @@ measurement is how you check whether it still holds for you.
- [Ask an image question](#ask-an-image-question)
- [Force JSON output](#force-json-output)
- [Watch it live](#watch-it-live)
+ - [Admin web portal](#admin-web-portal)
- [Probe routing without spending](#probe-routing-without-spending)
- [At a Glance](#at-a-glance)
- [Verification Pipeline](#verification-pipeline)
@@ -97,15 +103,15 @@ Nothing else is assumed about the host — routing itself is SQLite and arithmet
```bash
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
-sqlite3 router.db < schema.sql
-cp .env.example .env # fill in NEURALWATT_API_KEY
-python poller.py # populate the catalog
-python tier.py # resolve tiers
-python config.py # sanity-check config loads
-python -m uvicorn dispatcher:app --reload
+sqlite3 router.db < config/schema.sql
+cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
+PYTHONPATH=src python -m poller # populate the catalog
+PYTHONPATH=src python -m tier # resolve tiers
+PYTHONPATH=src python -m config # sanity-check config loads
+PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
```
-Then edit `config.yaml` for your own setup — at minimum:
+Then edit `config/config.yaml` for your own setup — at minimum:
| Key | Why |
|---|---|
@@ -115,7 +121,7 @@ Then edit `config.yaml` for your own setup — at minimum:
| `objective.assumed_cache_rate` | 0.917 was measured from one client's traffic (40.7M tokens). Check yours against the provider's per-session cache-hit figures |
| `session_cache.enabled` | off by default; caches category/tier per session for `staleness_minutes` to skip repeat classifier round-trips on long agent sessions |
-`python seed_energy.py` is optional. It sweeps a fixed reference workload to
+`PYTHONPATH=src python -m seed_energy` is optional. It sweeps a fixed reference workload to
populate `eco`, which is logged but is not an objective — routing works
without it. It costs real money and quota, so it is not in the path above.
@@ -127,9 +133,9 @@ Five user units cover continuous dispatch, catalog polling, and periodic energy
|---|---|---|
| `llm-router.service` | Continuous | FastAPI dispatcher |
| `llm-router-poller.timer` | 2 min after boot, then every 2 h | triggers the poller unit |
-| `llm-router-poller.service` | oneshot | `poller.py` → `tier.py` |
+| `llm-router-poller.service` | oneshot | `python -m poller` → `python -m tier` |
| `llm-router-seed.timer` | Every 6 h | triggers the seed unit |
-| `llm-router-seed.service` | oneshot | small `seed_energy.py` sweep |
+| `llm-router-seed.service` | oneshot | small `python -m seed_energy` sweep |
**The poller timer is load-bearing, not optional.** `freshness.stale_after_days`
is 3 with `exclude_stale: true` — an unpolled catalog marks every row stale
@@ -250,7 +256,7 @@ curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application
- Only inline `data:` URIs are accepted; remote `http(s)` image URLs are declined to avoid SSRF.
- Image count and total payload size are bounded before the local call is made.
-- Configure the fallback in `config.yaml` under `local_vision:`.
+- Configure the fallback in `config/config.yaml` under `local_vision:`.
### Force JSON output
@@ -277,23 +283,100 @@ curl -s localhost:8080/metrics | python -m json.tool
curl -s localhost:8080/events/decisions
# Terminal dashboard with live routing feed
-python tui.py
+PYTHONPATH=src python -m tui
```
- **`GET /metrics`** returns quota burn, coverage, recent decisions, per-model totals, verdict mix, and top proficiency. Loopback-only, no auth.
- **`GET /events/decisions`** is a Server-Sent Events stream of routing decisions. It replays recent decisions, then streams new ones as they happen; `:heartbeat` keepalive comments keep the connection alive between events.
-- **`python tui.py`** is a Textual dashboard with a live routing feed, a category → model breakdown panel, a detail popup (press Enter or `e`), and quota/per-model/verdict/warnings panels. It is a separate entrypoint, not a systemd unit.
-- **`baseline_report.py --since 2026-08-01`** is a read-only retrospective that replays recent decisions against two trivial counterfactuals — always pick the cheapest eligible candidate and always pick the best-proficiency eligible candidate — and reports total/mean cost and proficiency plus a dominance share. Add `--category coding_refactor` to scope it, or `--csv` for parseable output.
+- **`PYTHONPATH=src python -m tui`** is a Textual dashboard with a live routing feed, a category → model breakdown panel, a detail popup (press Enter or `e`), and quota/per-model/verdict/warnings panels. It is a separate entrypoint, not a systemd unit.
+- **`baseline_report.py --since 2026-08-01`** is a read-only retrospective (`PYTHONPATH=src python -m baseline_report`) that replays recent decisions against two trivial counterfactuals — always pick the cheapest eligible candidate and always pick the best-proficiency eligible candidate — and reports total/mean cost and proficiency plus a dominance share. Add `--category coding_refactor` to scope it, or `--csv` for parseable output.
`textual` is pinned in `requirements.txt` solely for the TUI modules (`tui.py`, `tui_screens.py`, `tui_sse.py`). It is imported only by these modules; the FastAPI service dispatch path never touches it, so the router itself has no UI dependency.
+### Admin web portal
+
+`GET /admin/` serves a management portal from the running router on the same
+loopback-only bind as the API. It has no auth layer yet, so like the other
+endpoints it is reachable only from `127.0.0.1`.
+
+The portal is a **four-page, glass dark-mode** dashboard built on Tabler/Bootstrap
+with a few lighter inline SVG icons (see `plans/admin-design-standards.md`
+for the visual language and `plans/admin-work-framework.md` for the method
+used to evolve it). It is implemented in `admin.py`, with `config/admin_schema.sql` for
+its future tables and `admin/frontend/` for the browser UI.
+
+**Dashboard** — `GET /admin/` aggregates the router at a glance: a quota chip
+(burn against `objective.plan_kwh_per_period`), per-model usage bars, verdict
+mix and category breakdown, five history mini-charts, and a recent-decisions
+table.
+
+
+
+
+
+**Models** — `GET /admin/models` is the model-availability table. Each row shows
+the serving class, tier, status, and an override dropdown (`active` /
+`deprecated` / `stale`) that writes through to the routing hard filters.
+
+
+
+
+
+**Controls** — `GET /admin/controls` holds the operational triggers, runtime
+knobs, and the persisted `config/config.yaml` editor. Changes are marked as dirty and
+written only on save.
+
+
+
+
+
+**Decisions** — `GET /admin/decisions` is the full decision log with
+kind/category/tier filters and free-text search.
+
+
+
+
+
+**Read-only dashboards** — `GET /admin/api/snapshot` exposes data for:
+
+- quota burn against `objective.plan_kwh_per_period`
+- per-model usage from `energy_observations`
+- live routing decisions from `route_decisions`
+- verdict mix and scoring coverage
+- history: `GET /admin/api/history?range=6h` (also `1h`, `24h`, `7d`, `30d`)
+
+**Operational triggers** — async, fire-and-forget maintenance jobs:
+
+- `POST /admin/api/refresh-catalog` runs the poller and tier pass
+- `POST /admin/api/seed-energy?samples=N` starts a reference sweep
+- `POST /admin/api/apply-feedback?dry_run=true` runs `feedback.py`
+- `POST /admin/api/restart-service` restarts the running systemd unit
+
+**Runtime toggles**
+
+`GET /admin/api/runtime` shows persisted-vs-runtime values. `POST /admin/api/runtime/{knob}`
+flips in-memory settings such as `log_route_decisions` and `local_llm_enabled`.
+Changes take effect immediately but reset on restart.
+
+**Persisted config edits**
+
+`GET /admin/api/config` lists allowlisted keys. `POST /admin/api/config/{key}`
+writes one allowlisted key back to `config/config.yaml` with a timestamped backup
+and whole-config validation. Arbitrary keys are rejected.
+
+**Model availability overrides**
+
+`POST /admin/api/models/{model_id}/{provider}/availability` marks a model as
+`active`, `deprecated`, or `stale`. `DELETE` on the same path removes the override.
+Deprecation feeds into the routing hard filters.
+
### Probe routing without spending
`router_cli.py` is a one-shot shell probe that POSTs to `/route` once and prints the full decision tree.
```bash
-python router_cli.py "Refactor this Django view into service objects"
-python router_cli.py "Summarize this diff" --category summarization --tier 2
+PYTHONPATH=src python -m router_cli "Refactor this Django view into service objects"
+PYTHONPATH=src python -m router_cli "Summarize this diff" --category summarization --tier 2
```
- Prints the selected model, candidates, estimated cost, and estimated proficiency.
@@ -377,8 +460,8 @@ Observation → learning. `feedback.py` folds verification failures into
23-task benchmark:
```bash
-python feedback.py --dry-run # preview what would change
-python feedback.py # apply
+PYTHONPATH=src python -m feedback --dry-run # preview what would change
+PYTHONPATH=src python -m feedback # apply
```
Key behaviors:
@@ -458,11 +541,11 @@ Key behaviors:
| **Local Classification** | Ollama, OpenAI-compatible — `localhost:11434/v1` or an Ollama across your VPN |
| **Local Model** | `classifier.model` — `mistral-nemo:12b` by default; any Ollama model works |
| **Cloud Provider** | Neuralwatt only |
-| **Config** | `config.yaml` loaded & validated by Pydantic (`config.py`) |
+| **Config** | `config/config.yaml` loaded & validated by Pydantic (`src/config.py`) |
| **OpenAI Client** | `openai==3.0.0` (official SDK) |
| **HTTP** | `requests` for poller, `httpx` (via openai/uvicorn) |
-| **Testing** | `pytest` — 562 tests across 27 files, all offline |
-| **Config Files** | `config.yaml`, `leaderboards.yaml`, `evals/tasks.yaml` |
+| **Testing** | `pytest` — 733 tests across 41 files, all offline |
+| **Config Files** | `config/config.yaml`, `config/leaderboards.yaml`, `evals/tasks.yaml` |
| **Deployment** | systemd user units (`.service` + `.timer` files in `deploy/`) |
| **Integration** | `opencode.json` in the repo routes through it by default; any OpenAI-compatible client works |
@@ -505,6 +588,7 @@ restarts on boot shouldn't change its dependency tree underneath itself.
| **`config.py`** | YAML loader + Pydantic validators (blend weights sum to 1, valid tiers, endpoints separately addressable) | Yes (file) |
| **`metrics.py`** | Read-only aggregations for `/health` and `GET /metrics`: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency | Yes (DB) |
| **`events.py`** | In-memory decision-event broker for the TUI's live feed: bounded ring buffer + thread-safe fan-out to SSE subscribers | Pure |
+| **`admin.py`** | `/admin` management portal: serves the four frontend pages + read/write API (snapshot, history, models/availability, runtime knobs, allowlisted config edits, operational triggers) | Yes (DB, network, file) |
| **`tui.py`** | Textual terminal dashboard over `GET /metrics` and `GET /events/decisions`; live routing feed, detail popup, category breakdown. Foreground tool, not a service | Yes (network) |
| **`tui_model.py`** | Pure data layer for the TUI: `build_model`, `build_category_breakdown`, `decision_row` — no Textual import, testable without a terminal | Pure |
| **`tui_sse.py`** | Background-thread SSE consumer for the TUI: reconnects on failure, marshals live decisions onto the UI thread | Yes (network) |
@@ -570,7 +654,7 @@ parses them into `access_level` and `routing.allowed_access_levels` (default
| `inherited_from` | TEXT | Model this row was copied from, NULL if measured directly |
| `last_updated` | TEXT | ISO8601 |
-**Category set** (9 categories, defined in `config.yaml`):
+**Category set** (9 categories, defined in `config/config.yaml`):
| Category | Example | Scoring type |
|---|---|---|
@@ -741,7 +825,7 @@ rows with `supports_vision = 1` survive the hard filters. If **no** cloud
candidate survives, the router can fall back to a local vision model instead
of returning 422.
-`local_vision:` in `config.yaml` controls this path:
+`local_vision:` in `config/config.yaml` controls this path:
| Key | Default | Purpose |
|---|---|---|
@@ -753,7 +837,7 @@ of returning 422.
| `max_images` | `4` | Refuse requests with more image parts |
| `max_image_bytes` | `9437184` (9 MiB) | Refuse requests whose image payload exceeds this |
-The fallback is **enabled by default** both in `config.yaml` and in
+The fallback is **enabled by default** both in `config/config.yaml` and in
`LocalVisionConfig`, so omitting the section still turns it on. Disable it
explicitly (`enabled: false`) on a host with no local Ollama or one that has
not pulled the vision model.
@@ -883,6 +967,7 @@ allowance.
| `POST` | `/dispatch` | Same as `/route`, plus complete the provider call, stream response, log observation |
| `GET` | `/v1/models` | OpenAI-compatible model list (router virtual models + catalog) |
| `POST` | `/v1/chat/completions` | OpenAI-compatible completions — routes then proxies, **streaming supported** |
+| `GET` | `/admin` | Loopback-only web management portal (read-only dashboards, operational triggers, runtime toggles, allowlisted config edits) |
`/metrics` returns a single JSON object with these top-level keys:
@@ -951,7 +1036,7 @@ rank id=r9116d9 pos=0 model=deepseek-v4-flash prof=1 est_usd=0.00016296
**No conversation text is logged at any level**, prompt or answer — prompts
here run 60k–150k tokens and the journal is on disk. A test enforces it.
-Set the level in `config.yaml` (`logging.level`), or override it without
+Set the level in `config/config.yaml` (`logging.level`), or override it without
touching a tracked file:
```bash
@@ -976,10 +1061,10 @@ out clean.
## Self-Eval Harness (`eval_proficiency.py`)
```bash
-python eval_proficiency.py # every routable model × every task
-python eval_proficiency.py --models kimi-k3 # subset of models
-python eval_proficiency.py --categories coding_general
-python eval_proficiency.py --dry-run # plan only
+PYTHONPATH=src python -m eval_proficiency # every routable model × every task
+PYTHONPATH=src python -m eval_proficiency --models kimi-k3 # subset of models
+PYTHONPATH=src python -m eval_proficiency --categories coding_general
+PYTHONPATH=src python -m eval_proficiency --dry-run # plan only
```
- Runs every task through the target provider, scores it, writes to `proficiency`.
@@ -1037,7 +1122,7 @@ Several settings keep it from cascading failures:
unparseable output. A coding agent would rather have a mid-tier answer than
an error. Escalation deliberately skips fallbacks so an unavailable local
model doesn't silently promote every request to the frontier tier.
-- **Session classification cache** (`session_cache:` in `config.yaml`,
+- **Session classification cache** (`session_cache:` in `config/config.yaml`,
off by default) remembers the last `task_category`/`task_tier` decision per
session for `staleness_minutes` (default 20), so a long agent session skips
the classifier round-trip on every turn. It only ever short-circuits the
@@ -1068,7 +1153,7 @@ carries the same modality block.
## Testing
```bash
-python -m pytest # 562 tests
+python -m pytest # 733 tests
python -m pytest --cov # with coverage
```
@@ -1093,6 +1178,7 @@ config as arguments, so the suite runs offline on a clean checkout.
| `test_session_identity.py` | Outcome attribution: session matching, ambiguity refusal |
| `test_config_endpoints.py` | Classifier and verifier are separately addressable; guards on the split |
| `test_metrics_endpoint.py` | `/metrics` endpoint, SSE `/events/decisions` headers + replay/stream behavior |
+| `test_admin_frontend.py` | Admin portal serves the four HTML pages with their expected markers and Chart.js asset |
| `test_events.py` | Decision-event broker: publish, subscribe/replay, unsubscribe, full-subscriber eviction |
| `test_tui.py` | TUI data model, category breakdown, detail popup, live SSE decision handling, keyboard controls |
| `test_context_prune.py` | Context pruning: image_url handling, structured content, recency guards, stats accuracy |
@@ -1119,7 +1205,7 @@ ollama create qwen3-vl-router:4b -f Modelfile.vision
With these tags the combined resident footprint was measured at roughly **15.9GB** in the worst case: classifier and verifier share one `mistral-nemo-router:12b` instance at ~8.6GB, with the `qwen3-vl-router:4b` vision fallback loaded alongside it. That leaves real headroom on a 24GB card.
-`config.yaml` already points at these tags by default:
+`config/config.yaml` already points at these tags by default:
```yaml
classifier:
diff --git a/admin/frontend/controls.html b/admin/frontend/controls.html
new file mode 100644
index 0000000..9746cb6
--- /dev/null
+++ b/admin/frontend/controls.html
@@ -0,0 +1,652 @@
+
+
+
+
+
+Controls · LLM Router Admin
+
+
+
+
+
+
+
+
+
+
+
+
+
Available models, their tier, and live availability overrides
+
+
+
+
+
+
+
+
+
+
+
+
+
Model Availability
+
+
Change the override dropdown to mark a model active, deprecated, or stale
+
+
+
+
+
+
Model
+
Provider
+
Tier
+
Status
+
Override
+
+
+
+
Loading…
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
diff --git a/assets/6krrt-logo-outlined.svg b/assets/6krrt-logo-outlined.svg
new file mode 100644
index 0000000..6f7645f
--- /dev/null
+++ b/assets/6krrt-logo-outlined.svg
@@ -0,0 +1,18 @@
+
+
\ No newline at end of file
diff --git a/assets/screenshots/controls.png b/assets/screenshots/controls.png
new file mode 100644
index 0000000..47533cf
Binary files /dev/null and b/assets/screenshots/controls.png differ
diff --git a/assets/screenshots/dashboard.png b/assets/screenshots/dashboard.png
new file mode 100644
index 0000000..7408a51
Binary files /dev/null and b/assets/screenshots/dashboard.png differ
diff --git a/assets/screenshots/decisions.png b/assets/screenshots/decisions.png
new file mode 100644
index 0000000..0387595
Binary files /dev/null and b/assets/screenshots/decisions.png differ
diff --git a/assets/screenshots/models.png b/assets/screenshots/models.png
new file mode 100644
index 0000000..d50f76c
Binary files /dev/null and b/assets/screenshots/models.png differ
diff --git a/admin_schema.sql b/config/admin_schema.sql
similarity index 100%
rename from admin_schema.sql
rename to config/admin_schema.sql
diff --git a/config.yaml b/config/config.yaml
similarity index 98%
rename from config.yaml
rename to config/config.yaml
index c536550..b0559ce 100644
--- a/config.yaml
+++ b/config/config.yaml
@@ -18,7 +18,7 @@ objective:
# currently rest on 2-3 samples per category, so a 0.05 gap is
# indistinguishable from sampling variation and paying for it buys noise.
# Narrow it as samples accumulate.
- quality_tolerance: 0.10
+ quality_tolerance: 0.1
# Cost is priced per-request from catalog prices, NOT from a benchmark
# sweep. A fixed 400-token reference task ranked glm-5.2-fast 3.2x cheaper
@@ -52,7 +52,7 @@ objective:
# wall you hit mid-task. So this is the cost mandate stated as a guarantee.
# For scale: the reference task runs ~5e-06 kWh on the cheapest model and
# ~2.2e-04 on the most expensive.
- max_energy_per_request: null
+ max_energy_per_request:
# The subscription's kWh allowance per billing period, for reporting burn in
# /health. Set to match your plan; null disables the report. NeuralWatt also
@@ -115,15 +115,15 @@ proficiency:
leaderboard_weight: 0.3
self_eval_weight: 0.7
categories:
- - coding_general
- - coding_refactor
- - debugging
- - docs_writing
- - summarization
- - translation
- - reasoning_math
- - tool_use_agentic
- - general_chat
+ - coding_general
+ - coding_refactor
+ - debugging
+ - docs_writing
+ - summarization
+ - translation
+ - reasoning_math
+ - tool_use_agentic
+ - general_chat
escalation:
enabled: true
@@ -203,7 +203,7 @@ session_cache:
# Off by default, matching every other new-and-unproven knob in this
# project: ship it, watch route_decisions.source="cached" on real traffic,
# then decide the right default.
- enabled: false
+ enabled: true
staleness_minutes: 20
circuit_breaker:
@@ -224,7 +224,7 @@ routing:
# routing excludes anything not listed here. Add 'preview'/'canary' only if
# the account actually holds the grant — otherwise dispatch earns a 403.
allowed_access_levels:
- - public
+ - public
# '-flex' rows are held server-side during peak until a capacity gap opens.
# That's correct for overnight/batch agent work and wrong for anything
@@ -274,7 +274,7 @@ routing:
# off, let POST /outcome report real pass/fail, and compare
# tool_use_agentic proficiency for deepseek before and after. That is the
# one signal here that knows whether the work actually worked.
- min_tool_proficiency: null
+ min_tool_proficiency:
tool_use_category: tool_use_agentic
# Request-side capability gates. These read the request body (image parts,
@@ -299,7 +299,7 @@ local_vision:
# and the model must be pulled (`ollama pull qwen3-vl:4b`) on that host.
enabled: true
base_url: "http://localhost:11434/v1"
- api_key_env: null
+ api_key_env:
# A Modelfile-tagged variant of qwen3-vl:4b, not the base library tag.
# Measured live: the base tag comes up at Ollama's own default num_ctx
# (32768) and costs 9.4GB loaded — resident alongside the classifier's
@@ -398,7 +398,7 @@ classifier:
base_url: "http://localhost:11434/v1"
# Unset means unauthenticated, which is the Ollama case. Name the env var
# holding the key when the endpoint actually checks one.
- api_key_env: null
+ api_key_env:
# A Modelfile-tagged variant of mistral-nemo:12b, not the base library tag
# — must match a model `ollama list` reports. Ollama loads a model at its
# library Modelfile's default context unless told otherwise, and the base
@@ -517,3 +517,4 @@ logging:
# systemctl --user edit llm-router # Environment="LLM_ROUTER_LOG_LEVEL=debug"
# systemctl --user restart llm-router
level: info
+
diff --git a/leaderboards.yaml b/config/leaderboards.yaml
similarity index 100%
rename from leaderboards.yaml
rename to config/leaderboards.yaml
diff --git a/schema.sql b/config/schema.sql
similarity index 100%
rename from schema.sql
rename to config/schema.sql
diff --git a/deploy/README.md b/deploy/README.md
index 8bea6d8..0ac26c1 100644
--- a/deploy/README.md
+++ b/deploy/README.md
@@ -9,9 +9,9 @@ all.
| file | what it does |
|---|---|
| `llm-router.service` | the FastAPI dispatcher, on `127.0.0.1:8080` |
-| `llm-router-poller.service` | one-shot: `poller.py` then `tier.py` |
+| `llm-router-poller.service` | one-shot: `PYTHONPATH=src python -m poller` then `PYTHONPATH=src python -m tier` |
| `llm-router-poller.timer` | fires the poller 2 min after boot, then every 2 h |
-| `llm-router-seed.service` | one-shot: a small `seed_energy.py` reference sweep |
+| `llm-router-seed.service` | one-shot: a small `PYTHONPATH=src python -m seed_energy` reference sweep |
| `llm-router-seed.timer` | every 6 h — energy attribution drifts with pool load across hours, so the median has to span time rather than one sweep |
These are **user** units — no root, and they run as you with your own
@@ -37,6 +37,10 @@ done
systemctl --user daemon-reload
systemctl --user enable --now llm-router.service llm-router-poller.timer llm-router-seed.timer
+# Note: the units rely on `Environment=PYTHONPATH=%h/llm-router/src` (rewritten
+# by the same sed to your REPO) so the modules under `src/` are importable
+# without an editable install. .env is read from the repo root and stays there.
+
# 3. Survive logout/reboot (user units stop with your session otherwise)
loginctl enable-linger "$USER"
@@ -54,17 +58,17 @@ journalctl --user -u 'llm-router*' -f # everything, live (quote the g
journalctl --user -u llm-router -f -o cat # the request log, message only
journalctl --user -u llm-router -p warning # fallbacks, retries, refusals
journalctl --user -u llm-router-poller.service # catalog refreshes
-systemctl --user restart llm-router.service # after editing config.yaml
+systemctl --user restart llm-router.service # after editing config/config.yaml
systemctl --user start llm-router-poller.service # force a refresh now
```
-`config.yaml` is read once at startup, so weight and threshold changes need a
+`config/config.yaml` is read once at startup, so weight and threshold changes need a
restart. The catalog is read per-request, so a poller run takes effect
immediately.
### Turning up the logs
-`logging.level` in config.yaml is the documented setting, but flipping it means
+`logging.level` in config/config.yaml is the documented setting, but flipping it means
editing a tracked file. For a running service use a drop-in instead:
```bash
@@ -91,7 +95,7 @@ priorities when systemd owns its stderr (`SyslogLevelPrefix` is on by default).
A foreground `uvicorn` prints them clean, so the same binary is readable either
way.
-**The oneshot units buffer.** `poller.py` and `seed_energy.py` print progress
+**The oneshot units buffer.** `python -m poller` and `python -m seed_energy` print progress
with plain `print()`, and Python block-buffers stdout when it is not a
terminal, so their output arrives in one dump at exit rather than
progressively. Add `Environment="PYTHONUNBUFFERED=1"` to those units if you
@@ -123,7 +127,7 @@ sudo nano /etc/systemd/system/ollama.service.d/override.conf
sudo systemctl daemon-reload && sudo systemctl restart ollama
```
-On the **client** host, in `config.yaml`:
+On the **client** host, in `config/config.yaml`:
```yaml
classifier:
diff --git a/deploy/llm-router-poller.service b/deploy/llm-router-poller.service
index 39dd4f7..3a09f30 100644
--- a/deploy/llm-router-poller.service
+++ b/deploy/llm-router-poller.service
@@ -7,12 +7,13 @@ Wants=network-online.target
[Service]
Type=oneshot
WorkingDirectory=%h/llm-router
+Environment=PYTHONPATH=%h/llm-router/src
EnvironmentFile=%h/llm-router/.env
-# poller.py refreshes the catalog; tier.py re-resolves tiers from it. Tiers
+# poller refreshes the catalog; tier re-resolves tiers from it. Tiers
# are derived from cost and reasoning fields the poll may have changed, so
-# they always run as a pair.
-ExecStart=%h/llm-router/.venv/bin/python poller.py
-ExecStart=%h/llm-router/.venv/bin/python tier.py
+# they always run as a pair. PYTHONPATH above resolves src/ for `-m` imports.
+ExecStart=%h/llm-router/.venv/bin/python -m poller
+ExecStart=%h/llm-router/.venv/bin/python -m tier
NoNewPrivileges=true
PrivateTmp=true
diff --git a/deploy/llm-router-seed.service b/deploy/llm-router-seed.service
index db7e1a7..b78ea26 100644
--- a/deploy/llm-router-seed.service
+++ b/deploy/llm-router-seed.service
@@ -7,10 +7,12 @@ Wants=network-online.target
[Service]
Type=oneshot
WorkingDirectory=%h/llm-router
+Environment=PYTHONPATH=%h/llm-router/src
EnvironmentFile=%h/llm-router/.env
# Fewer samples per run than a manual sweep, because the point is coverage
-# across TIME rather than depth at one moment — see the timer.
-ExecStart=%h/llm-router/.venv/bin/python seed_energy.py --samples 3
+# across TIME rather than depth at one moment — see the timer. PYTHONPATH
+# above resolves src/ for the `-m` import.
+ExecStart=%h/llm-router/.venv/bin/python -m seed_energy --samples 3
NoNewPrivileges=true
PrivateTmp=true
diff --git a/deploy/llm-router.service b/deploy/llm-router.service
index 54aa58f..8d76add 100644
--- a/deploy/llm-router.service
+++ b/deploy/llm-router.service
@@ -10,9 +10,12 @@ Wants=network-online.target
[Service]
Type=exec
-# config.yaml, router.db and router.log are all referenced as relative paths,
-# so this has to be the repo root.
+# router.db and router.log are referenced as relative paths, so this has to be
+# the repo root. config.yaml now lives under config/ and the Python modules
+# under src/, so PYTHONPATH points at src/ for the dispatcher import to resolve.
WorkingDirectory=%h/llm-router
+# Resolves `dispatcher` (and any other src/ module) for the ExecStart below.
+Environment=PYTHONPATH=%h/llm-router/src
# Holds NEURALWATT_API_KEY. Create it with:
# echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env
EnvironmentFile=%h/llm-router/.env
diff --git a/code_plans/.playwright-mcp/console-2026-08-29T06-11-18-590Z.log b/plans/.playwright-mcp/console-2026-08-29T06-11-18-590Z.log
similarity index 100%
rename from code_plans/.playwright-mcp/console-2026-08-29T06-11-18-590Z.log
rename to plans/.playwright-mcp/console-2026-08-29T06-11-18-590Z.log
diff --git a/code_plans/.playwright-mcp/console-2026-08-29T07-07-39-971Z.log b/plans/.playwright-mcp/console-2026-08-29T07-07-39-971Z.log
similarity index 100%
rename from code_plans/.playwright-mcp/console-2026-08-29T07-07-39-971Z.log
rename to plans/.playwright-mcp/console-2026-08-29T07-07-39-971Z.log
diff --git a/code_plans/.playwright-mcp/console-2026-08-29T07-17-18-482Z.log b/plans/.playwright-mcp/console-2026-08-29T07-17-18-482Z.log
similarity index 100%
rename from code_plans/.playwright-mcp/console-2026-08-29T07-17-18-482Z.log
rename to plans/.playwright-mcp/console-2026-08-29T07-17-18-482Z.log
diff --git a/code_plans/.playwright-mcp/console-2026-08-29T14-42-18-565Z.log b/plans/.playwright-mcp/console-2026-08-29T14-42-18-565Z.log
similarity index 100%
rename from code_plans/.playwright-mcp/console-2026-08-29T14-42-18-565Z.log
rename to plans/.playwright-mcp/console-2026-08-29T14-42-18-565Z.log
diff --git a/code_plans/.playwright-mcp/console-2026-08-29T14-47-56-594Z.log b/plans/.playwright-mcp/console-2026-08-29T14-47-56-594Z.log
similarity index 100%
rename from code_plans/.playwright-mcp/console-2026-08-29T14-47-56-594Z.log
rename to plans/.playwright-mcp/console-2026-08-29T14-47-56-594Z.log
diff --git a/code_plans/.playwright-mcp/console-2026-08-29T15-00-11-582Z.log b/plans/.playwright-mcp/console-2026-08-29T15-00-11-582Z.log
similarity index 100%
rename from code_plans/.playwright-mcp/console-2026-08-29T15-00-11-582Z.log
rename to plans/.playwright-mcp/console-2026-08-29T15-00-11-582Z.log
diff --git a/code_plans/.playwright-mcp/console-2026-08-29T15-01-33-174Z.log b/plans/.playwright-mcp/console-2026-08-29T15-01-33-174Z.log
similarity index 100%
rename from code_plans/.playwright-mcp/console-2026-08-29T15-01-33-174Z.log
rename to plans/.playwright-mcp/console-2026-08-29T15-01-33-174Z.log
diff --git a/code_plans/.playwright-mcp/page-2026-08-29T06-11-19-085Z.yml b/plans/.playwright-mcp/page-2026-08-29T06-11-19-085Z.yml
similarity index 100%
rename from code_plans/.playwright-mcp/page-2026-08-29T06-11-19-085Z.yml
rename to plans/.playwright-mcp/page-2026-08-29T06-11-19-085Z.yml
diff --git a/code_plans/.playwright-mcp/page-2026-08-29T06-11-46-660Z.yml b/plans/.playwright-mcp/page-2026-08-29T06-11-46-660Z.yml
similarity index 100%
rename from code_plans/.playwright-mcp/page-2026-08-29T06-11-46-660Z.yml
rename to plans/.playwright-mcp/page-2026-08-29T06-11-46-660Z.yml
diff --git a/code_plans/.playwright-mcp/page-2026-08-29T07-07-40-410Z.yml b/plans/.playwright-mcp/page-2026-08-29T07-07-40-410Z.yml
similarity index 100%
rename from code_plans/.playwright-mcp/page-2026-08-29T07-07-40-410Z.yml
rename to plans/.playwright-mcp/page-2026-08-29T07-07-40-410Z.yml
diff --git a/code_plans/.playwright-mcp/page-2026-08-29T07-17-18-936Z.yml b/plans/.playwright-mcp/page-2026-08-29T07-17-18-936Z.yml
similarity index 100%
rename from code_plans/.playwright-mcp/page-2026-08-29T07-17-18-936Z.yml
rename to plans/.playwright-mcp/page-2026-08-29T07-17-18-936Z.yml
diff --git a/code_plans/.playwright-mcp/page-2026-08-29T14-42-18-860Z.yml b/plans/.playwright-mcp/page-2026-08-29T14-42-18-860Z.yml
similarity index 100%
rename from code_plans/.playwright-mcp/page-2026-08-29T14-42-18-860Z.yml
rename to plans/.playwright-mcp/page-2026-08-29T14-42-18-860Z.yml
diff --git a/code_plans/.playwright-mcp/page-2026-08-29T14-42-58-029Z.yml b/plans/.playwright-mcp/page-2026-08-29T14-42-58-029Z.yml
similarity index 100%
rename from code_plans/.playwright-mcp/page-2026-08-29T14-42-58-029Z.yml
rename to plans/.playwright-mcp/page-2026-08-29T14-42-58-029Z.yml
diff --git a/code_plans/.playwright-mcp/page-2026-08-29T14-44-04-172Z.yml b/plans/.playwright-mcp/page-2026-08-29T14-44-04-172Z.yml
similarity index 100%
rename from code_plans/.playwright-mcp/page-2026-08-29T14-44-04-172Z.yml
rename to plans/.playwright-mcp/page-2026-08-29T14-44-04-172Z.yml
diff --git a/code_plans/.playwright-mcp/page-2026-08-29T14-47-57-000Z.yml b/plans/.playwright-mcp/page-2026-08-29T14-47-57-000Z.yml
similarity index 100%
rename from code_plans/.playwright-mcp/page-2026-08-29T14-47-57-000Z.yml
rename to plans/.playwright-mcp/page-2026-08-29T14-47-57-000Z.yml
diff --git a/code_plans/.playwright-mcp/page-2026-08-29T14-48-31-579Z.yml b/plans/.playwright-mcp/page-2026-08-29T14-48-31-579Z.yml
similarity index 100%
rename from code_plans/.playwright-mcp/page-2026-08-29T14-48-31-579Z.yml
rename to plans/.playwright-mcp/page-2026-08-29T14-48-31-579Z.yml
diff --git a/code_plans/.playwright-mcp/page-2026-08-29T15-00-11-987Z.yml b/plans/.playwright-mcp/page-2026-08-29T15-00-11-987Z.yml
similarity index 100%
rename from code_plans/.playwright-mcp/page-2026-08-29T15-00-11-987Z.yml
rename to plans/.playwright-mcp/page-2026-08-29T15-00-11-987Z.yml
diff --git a/code_plans/.playwright-mcp/page-2026-08-29T15-01-33-298Z.yml b/plans/.playwright-mcp/page-2026-08-29T15-01-33-298Z.yml
similarity index 100%
rename from code_plans/.playwright-mcp/page-2026-08-29T15-01-33-298Z.yml
rename to plans/.playwright-mcp/page-2026-08-29T15-01-33-298Z.yml
diff --git a/code_plans/.playwright-mcp/task-8-admin-visual-fixes-v2-full.png b/plans/.playwright-mcp/task-8-admin-visual-fixes-v2-full.png
similarity index 100%
rename from code_plans/.playwright-mcp/task-8-admin-visual-fixes-v2-full.png
rename to plans/.playwright-mcp/task-8-admin-visual-fixes-v2-full.png
diff --git a/code_plans/.playwright-mcp/task-8-admin-visual-fixes-v2.png b/plans/.playwright-mcp/task-8-admin-visual-fixes-v2.png
similarity index 100%
rename from code_plans/.playwright-mcp/task-8-admin-visual-fixes-v2.png
rename to plans/.playwright-mcp/task-8-admin-visual-fixes-v2.png
diff --git a/code_plans/.playwright-mcp/task-8-v2-console-clean.log b/plans/.playwright-mcp/task-8-v2-console-clean.log
similarity index 100%
rename from code_plans/.playwright-mcp/task-8-v2-console-clean.log
rename to plans/.playwright-mcp/task-8-v2-console-clean.log
diff --git a/code_plans/.playwright-mcp/task-8-v2-console-final.log b/plans/.playwright-mcp/task-8-v2-console-final.log
similarity index 100%
rename from code_plans/.playwright-mcp/task-8-v2-console-final.log
rename to plans/.playwright-mcp/task-8-v2-console-final.log
diff --git a/code_plans/.playwright-mcp/task-8-v2-console-full.log b/plans/.playwright-mcp/task-8-v2-console-full.log
similarity index 100%
rename from code_plans/.playwright-mcp/task-8-v2-console-full.log
rename to plans/.playwright-mcp/task-8-v2-console-full.log
diff --git a/plans/admin-controls-relayout.jpg b/plans/admin-controls-relayout.jpg
new file mode 100644
index 0000000..6448c02
Binary files /dev/null and b/plans/admin-controls-relayout.jpg differ
diff --git a/plans/admin-design-standards.md b/plans/admin-design-standards.md
new file mode 100644
index 0000000..ff14e92
--- /dev/null
+++ b/plans/admin-design-standards.md
@@ -0,0 +1,364 @@
+# Admin Portal Design Standards
+
+This document captures the design system used by the Claude-built admin portal uplift (commits around `23b74a3`–`906c9e6`) so future edits to `admin/frontend/` stay visually consistent.
+
+Intended readers: any agent or human touching `admin/frontend/*.html`, `admin.py`, or the admin API surface.
+
+**Supplementary sources** (read these too before touching the admin UI):
+- `.omo/notepads/admin-facelift/learnings.md` — the running log of every visual QA finding, root cause, and fix across the whole 10-wave facelift. This is the single richest record of what went wrong and how it was fixed.
+- `plans/admin-facelift.md` — the original spec (survival guardrails S1–S5, device contracts).
+- `plans/admin-visual-fixes-review.md` and `plans/router-admin-portal-implementation-review.md` — post-hoc audits that caught regressions the QA pass missed.
+
+## Philosophy
+
+- **Liquid-glass dashboard, not a marketing page.** Background is a subtle gradient wash; cards are translucent, blurred, and softly lit from the top. The goal is information density with low visual fatigue.
+- **One visual language everywhere.** Use the same frosted-card treatment, pill buttons, and glass badges on every page so the portal feels like a single app.
+- **Subtle cues over neon alerts.** Status colors are desaturated translucent tints. Warnings live in the navbar bell, not in a full card.
+
+## Theme
+
+- Always dark mode. Root is ``.
+- **Base background** (body):
+ - Three overlapping `radial-gradient` blobs:
+ - amber (`rgba(245,158,11,.10)`) top-left
+ - blue (`rgba(59,130,246,.10)`) top-right
+ - purple (`rgba(139,92,246,.07)`) bottom
+ - Base color `#111827`.
+ - `background-attachment: fixed` so it stays put during scroll.
+
+## Layout shell (every page)
+
+1. Bootstrap 5 / Tabler page skeleton:
+ ```html
+ …
+
+
+
…
+
…
+
+
+
+ ```
+2. Scrolling context: **`` scrolls, `` does not.** This works around a Chrome/Linux phantom-margin bug caused by the HTML element also being the scrollbar container.
+ ```css
+ html { overflow:hidden; height:100%; margin:0!important; padding:0!important }
+ body { overflow-y:auto; height:100%; margin:0!important; padding:0!important }
+ ```
+3. Container width: `container-xl` inside `.page-body`.
+4. Cards grid: `.row.row-cards > .col-xl-6` or `.col-12`.
+
+## Navbar
+
+- Glass background: `rgba(31,41,55,.6)` with `backdrop-filter: blur(16px) saturate(160%)`.
+- Bottom border: `1px solid rgba(255,255,255,.06)`.
+- Explicit `position: relative; z-index: 1030;` on `header.navbar` so its dropdowns paint above later backdrop-filter stacking contexts (e.g. the quota chip).
+- Container padding pinned to exactly `20px` left/right so the brand edge stays aligned across breakpoints:
+ ```css
+ header.navbar > .container-fluid { padding-left:20px!important; padding-right:20px!important; }
+ ```
+- Brand wordmark: **6krrt LLM Router**, in `'Quicksand', var(--tblr-font-sans-serif)` at `font-weight: 700; letter-spacing: 0.01em`. The leading `6` is `.brand-six` — brand amber, `1.14em` — because the mascot is itself a 6. Keep the wordmark inside a single `.brand-word` span: the anchor is `inline-flex` with a `gap`, so a bare span around the `6` becomes its own flex item and that gap opens up between `6` and `krrt`.
+- Logo SVG inline; force `filter:none!important` and size `64x64px` to defeat Tabler's autodark invert filter. The size lives in three places that must agree — the CSS rule, and the inline `width`/`height` attributes **and** `style="width:64px;height:64px"` on the `