Root reorg: src/ config/ plans/ layout + admin facelift #9

Merged
alee merged 18 commits from neuralwatt-router-service into main 2026-08-30 04:52:49 +00:00
153 changed files with 4275 additions and 840 deletions

8
.gitignore vendored
View File

@@ -7,3 +7,11 @@ router.log
.venv/
venv/
.omo/
config.yaml.bak.*
node_modules/
package.json
package-lock.json
.playwright-mcp/
.mypy_cache/
.pytest_cache/
.ruff_cache/

View File

@@ -18,10 +18,10 @@ code.
| Language | Python 3.10+ (3.10 floor tested; 3.14 also verified) |
| Framework | FastAPI + uvicorn |
| Database | SQLite (`router.db`) |
| Config | `config.yaml` + Pydantic (`config.py`), `extra="forbid"` |
| Config | `config/config.yaml` + Pydantic (`src/config.py`), `extra="forbid"` |
| HTTP client | `requests` (pinned in `requirements.txt`) — do NOT add httpx2/aiohttp without a requirements bump |
| TUI | `textual==8.2.8` — imported only by `tui*.py` modules, never by the dispatch path |
| Testing | `pytest`, 562 tests, all offline (no provider or local-model calls) |
| Testing | `pytest`, 733 tests, all offline (no provider or local-model calls) |
| Dependencies | Pinned. Bump deliberately, never use `>=` |
## Module map and file boundaries
@@ -48,6 +48,19 @@ code.
| `proficiency.py` | Score blending: leaderboard + self-eval → weighted composite |
| `iteration.py` | Retry budget per tier, matching retry to failure kind |
### Admin portal
| File | Role |
|---|---|
| `admin.py` | FastAPI sub-router mounted by dispatcher at `/admin`; serves API + static HTML |
| `admin/frontend/index.html` | Dashboard: quota chip, per-model usage, verdict mix, category breakdown, history mini-charts, recent decisions |
| `admin/frontend/models.html` | Model availability table with override dropdown |
| `admin/frontend/decisions.html` | Full decision log with filter + search |
| `admin/frontend/controls.html` | Operational triggers, runtime knobs, persisted config editor |
| `admin_schema.sql` | RBAC/audit schema ready for future auth work (config/admin_schema.sql) |
When editing the admin UI, follow **`plans/admin-design-standards.md`** — design tokens, glass styling, responsive rules, and hard-won gotchas (e.g. navbar z-index, Chart.js canvas reuse, scroll-context bug). Follow **`plans/admin-work-framework.md`** for the *method* — surface-level diagnosis, grammar-first rule extraction, the pytest + Playwright-1400/800 + REFRESH_MS-idle evidence bar, and writing the commit as a standalone learning input. Key transferable rules: a settings list is one CSS-grid row (key | meta | fixed 176px control) — never restate a control's own value in a meta column, and a single-line button card goes full width with buttons in `.card-actions`, not into a `col-xl-6` beside a tall card.
### TUI modules (import `textual`; never imported by the dispatch path)
| File | Role |
@@ -77,7 +90,7 @@ code.
- No `# type: ignore`, no `as any`, no `@ts-ignore` equivalent.
- `except Exception` is acceptable at top-level boundaries with
`# noqa: BLE001` comment (see `dispatcher._refresh`, `persist_route_decision`).
- Config is strict: every knob belongs in `config.yaml`, not only in a
- Config is strict: every knob belongs in `config/config.yaml`, not only in a
Pydantic default. A default the file never mentions is invisible to a tuner.
- Named constants use `Final` in new code (`events.py`); existing code is
inconsistent — don't refactor just for this.
@@ -95,7 +108,7 @@ code.
- SQLite, `PRAGMA foreign_keys = ON`.
- `_db()` in `dispatcher.py` returns a `sqlite3.Connection` with
`row_factory = sqlite3.Row`.
- Schema in `schema.sql` uses `CREATE TABLE IF NOT EXISTS` — safe to re-run.
- Schema in `config/schema.sql` uses `CREATE TABLE IF NOT EXISTS` — safe to re-run.
- Code-side table creation (`ensure_route_decisions`, `proficiency_store.ensure_columns`)
mirrors the schema for live DBs that predate a feature.
@@ -116,25 +129,36 @@ code.
# Setup
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
sqlite3 router.db < schema.sql
cp .env.example .env # fill in NEURALWATT_API_KEY
sqlite3 router.db < config/schema.sql
cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
# Start the service (binds 127.0.0.1:8080)
python -m uvicorn dispatcher:app --reload
# Or via systemd:
systemctl --user start llm-router.service
# Port 8080 is the systemd-managed PRODUCTION instance. opencode's own
# model traffic goes through it (see opencode.json's baseURL) and
# `Restart=always` resurrects it ~5s after any kill — so never
# `pkill`/`kill` anything matching uvicorn/dispatcher/8080 to "free the
# port". That fights the supervisor, and on this repo it can cut off your
# own inference mid-task. See CLAUDE.md's "A bare SIGTERM could hang the
# process forever" section for the incident this note comes from.
#
# To pick up a dispatcher.py change on the real instance:
systemctl --user restart llm-router.service
# For an ad hoc/manual run — iterating with --reload, a throwaway instance
# for Playwright smoke tests against the admin frontend, anything that
# isn't "use the real router" — bind a different port so it can't collide:
PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
# Run the TUI (service must be running)
python tui.py
PYTHONPATH=src python -m tui
# Run the full test suite
python -m pytest
# Quick routing probe (no spend)
python router_cli.py "Refactor this Django view"
PYTHONPATH=src python -m router_cli "Refactor this Django view"
# Populate the catalog
python poller.py && python tier.py
PYTHONPATH=src python -m poller && PYTHONPATH=src python -m tier
```
## What's NOT built yet (open items)
@@ -149,9 +173,9 @@ See `CLAUDE.md` → "What's NOT built yet — pick up here" for the full list.
## After a code change
- Run `python -m pytest` — 562 tests, ~30s.
- Run `python -m pytest` — 733 tests, ~35s.
- If you changed `dispatcher.py`, restart the systemd service:
`systemctl --user restart llm-router.service` (it doesn't auto-reload code).
- If you changed the TUI, run `python tui.py` to verify it starts.
- If you changed the TUI, run `PYTHONPATH=src python -m tui` to verify it starts.
- Check `lsp_diagnostics` on changed files.
- Match the existing commit-message style: `fix:`, `feat:`, `docs:`, `test:`.

View File

@@ -196,15 +196,15 @@ rather than from months of history.
## What's built and working
- `schema.sql` — `models`, `proficiency`, `energy_observations`. Applies
cleanly (`sqlite3 router.db < schema.sql`). `models` carries the serving
- `config/schema.sql` — `models`, `proficiency`, `energy_observations`. Applies
cleanly (`sqlite3 router.db < config/schema.sql`). `models` carries the serving
class columns (below); `energy_observations` carries real carbon/cost.
- `poller.py` — fetches Neuralwatt's `/models` endpoint (public,
unauthenticated), normalizes, upserts, marks stale rows. **Verified
against the live API**: 19 models, and the `metadata.pricing` /
`metadata.capabilities` / `metadata.limits` field mappings are confirmed
correct.
- `config.yaml` / `config.py` — weights, thresholds, provider settings,
- `config/config.yaml` / `src/config.py` — weights, thresholds, provider settings,
Pydantic-validated.
- `scoring.py` — one `normalize_inverted` (cost and eco normalize
identically; they differ only in what is fed to them) plus the weighted
@@ -212,7 +212,7 @@ rather than from months of history.
- `seed_energy.py` — runs a fixed reference task N times per routable model
and writes `energy_observations` rows tagged `seed_reference`. This is what
makes `cost` and `eco` real numbers instead of the neutral 0.5. Re-run it
after the catalog gains models: `python seed_energy.py --samples 5`
after the catalog gains models: `PYTHONPATH=src python -m seed_energy --samples 5`
(13 models x 5 = 65 calls, and the whole sweep cost **under a cent**).
- `tiering.py` / `tier.py` — pure tier resolver + the DB pass that applies it.
- `routing.py` — pure hard filters and ranking.
@@ -258,7 +258,7 @@ rather than from months of history.
`route_decisions`, per-model aggregates over `energy_observations`,
verification verdict mix, and top proficiency by category.
- `tui.py` — Textual terminal dashboard over `GET /metrics` and
`GET /events/decisions`. A foreground entrypoint (`python tui.py`), not a
`GET /events/decisions`. A foreground entrypoint (`PYTHONPATH=src python -m tui`), not a
service. `textual` is imported only in the TUI modules (`tui.py`,
`tui_screens.py`, `tui_sse.py`), so the router's dispatch path has no UI
dependency. The dashboard has a live routing-decisions feed (via the SSE
@@ -270,7 +270,19 @@ rather than from months of history.
- `router_cli.py` — one-shot `/route` probe. Posts a task to the running router
and prints the decision tree, or emits raw JSON with `--json`. Spends no
quota because it only routes.
- `tests/` — 562 tests across 27 files, all passing, all offline. Verified on
- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — `/admin` management
portal served by the running router, loopback-only, no auth. Liquid-glass dark UI
(see `plans/admin-design-standards.md`). Pages: dashboard (`index.html`)
with quota chip, per-model usage, verdict-mix/category-breakdown bars, history
mini-charts and recent decisions; models (`models.html`) with availability
overrides; decisions (`decisions.html`) log with filter/search; controls
(`controls.html`) for operational triggers, runtime knobs, and persisted config
edits. SSE status dot + warnings bell are shared chrome. Operational triggers:
refresh catalog, seed energy, apply feedback, restart service. Runtime toggles
reset on restart. Persisted config edits limited to an allowlist. Model
availability overrides feed routing hard filters. Bucketed history endpoint over
`energy_observations` and `route_decisions`.
- `tests/` — 733 tests across 41 files, all passing, all offline. Verified on
Python 3.10 and 3.14; nothing declares `requires-python`, so 3.10 is the
tested floor rather than a promised one.
@@ -289,7 +301,7 @@ failed write is logged and swallowed, because a decision record is worth
having but never worth failing or slowing a request for. The table is created
from code at module load (`_ensure_route_decisions_table`) and again on every
write (`ensure_route_decisions` inside `persist_route_decision`), mirroring
the `proficiency_store.ensure_columns` migration pattern: `schema.sql` is
the `proficiency_store.ensure_columns` migration pattern: `config/schema.sql` is
`CREATE TABLE IF NOT EXISTS`, but a live `router.db` predating this table
needs the code-side migration. Both the module-load hook and the write-path
guarantee are idempotent and leave existing rows intact.
@@ -309,7 +321,7 @@ request body. Two new gates live in `routing.rejection_reason` and
`json_schema` and `routing.require_json_mode` is true. A model passes only if
`supports_json_mode = 1`; `NULL` also fails closed, for the same reason.
Both gates default to **on** in `config.yaml`, because a wrong guess produces a
Both gates default to **on** in `config/config.yaml`, because a wrong guess produces a
400. This is a deliberate asymmetry against the tool-proficiency gate below:
capability **flags** fail closed on unknown, while quality **measurements**
(tool proficiency, energy) admit on absent evidence ("unproven, not bad").
@@ -336,11 +348,11 @@ so failing early is better than an opaque provider 400.
### Local Ollama vision fallback
`local_vision:` in `config.yaml` configures a fallback path for image requests
`local_vision:` in `config/config.yaml` configures a fallback path for image requests
that find no cloud vision candidate — the cloud catalog excludes the cost
leader (deepseek) on vision, so without this fallback every image request
that would otherwise have routed there 422s instead. It is a **core feature**,
**enabled by default** (`enabled: true`, both in `config.yaml` and in
**enabled by default** (`enabled: true`, both in `config/config.yaml` and in
`LocalVisionConfig`'s own default, so a config that omits the section still
gets it). Disable it explicitly (`enabled: false`) on a host with no local
Ollama, or one that hasn't pulled the vision model.
@@ -1043,21 +1055,21 @@ in the same pass — it was declared, never read, and shadowed the
```bash
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
sqlite3 router.db < schema.sql
cp .env.example .env # fill in NEURALWATT_API_KEY
python poller.py # populate the catalog
python tier.py # resolve tiers
python config.py # sanity-check config loads
python -m uvicorn dispatcher:app --reload
sqlite3 router.db < config/schema.sql
cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
PYTHONPATH=src python -m poller # populate the catalog
PYTHONPATH=src python -m tier # resolve tiers
PYTHONPATH=src python -m config # sanity-check config loads
PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
```
Then set what is deployment-specific in `config.yaml`: `classifier.model` and
Then set what is deployment-specific in `config/config.yaml`: `classifier.model` and
`classifier.base_url` for your Ollama, and `objective.plan_kwh_per_period` to
your own plan's quota (it is reported in `/health` as burn against the
allowance; it does not gate anything).
Ollama must be reachable with the classifier model pulled — the name must
match `classifier.model` in `config.yaml`, which points at a **Modelfile-tagged
match `classifier.model` in `config/config.yaml`, which points at a **Modelfile-tagged
variant**, not the base library tag. Ollama loads a model at its library
Modelfile's default context unless told otherwise, and the base tag never
was — measured live on a 24GB card, `mistral-nemo:12b` alone came up at
@@ -1192,12 +1204,51 @@ are still honored correctly regardless of that policy (systemd tracks
deliberate stops separately from the Restart= decision), so this only adds
self-healing for the unexplained case.
The signal's actual source is still open — live investigation and audit
trail in `code_plans/router-unreachable-signal-investigation.md`. Confirmed
so far: it is not `systemctl`, not the admin portal's restart trigger, not
suspend/resume, not the OOM killer, and — via `auditd` — not delivered
through the `kill` or `tgkill` syscalls either, which is why the watch was
extended to `pidfd_send_signal` and the `rt_*sigqueueinfo` syscalls.
**Resolved 2026-08-29 — the source was opencode, killing its own supply
line.** Full audit trail in
`plans/router-unreachable-signal-investigation.md`; the syscall-level
extension to `pidfd_send_signal` (past `kill`/`tgkill`, both cleared by
`auditd`) is what finally caught it. `sudo ausearch -k routerkill` matched
five incidents in one afternoon to `pidfd_send_signal(..., SIGTERM)` /
`SIGKILL`-after-escalation calls from a non-interactive `zsh -c "pkill ..."`
(never in `~/.zsh_history`, since it's not a login shell), and
`~/.local/share/opencode/log/opencode.log` matched every one of those
timestamps, to the millisecond, to one opencode run testing the admin
frontend: `pkill -f "uvicorn dispatcher:app"` (or `.*dispatcher`, or plain
`"8080"`), then `python -m uvicorn dispatcher:app --host 127.0.0.1 --port
8080` for its own throwaway instance. When the port came back occupied 5s
later (`Restart=always` resurrecting the real service), the agent read that
as "the kill didn't work" and escalated to `pkill -9` — the one case
(16:01:00) that arrived as a bare `SIGKILL`, `status=9/KILL`, rather than a
caught `SIGTERM`.
This was self-inflicted in a sharper way than it looks: `opencode.json`
points opencode's *own* model traffic at `http://127.0.0.1:8080/v1` — the
same production instance it was killing to test against. Every kill briefly
cut off the agent's own inference supply.
The actual bug was in `AGENTS.md`, not in this service: its "How to run
things" section told an agent to bring the router up with a bare
`python -m uvicorn dispatcher:app --reload` (implicitly on 8080, no port
flag) as an equally-valid alternative to `systemctl --user start`, with no
warning that 8080 is normally already held by the supervised instance. An
agent following that instruction and finding the port taken has no way to
know the right move is `systemctl --user restart` (which the doc *does* say
two sections later, for the "changed dispatcher.py" case, but not for "I
want to smoke-test against a running instance"). Fixed there: manual/ad hoc
runs now bind `--port 8081` explicitly, and the section says outright not to
`pkill`/`kill` anything matching `uvicorn`/`dispatcher`/`8080` — that's the
systemd-managed instance, `Restart=always` will fight you, and on this repo
it may be your own model access.
**Convention going forward: 8080 is production, always.** It's the port
baked into `opencode.json`, every curl example in this file, the systemd
unit, and the admin frontend's own fetches — moving it would touch more
surface than the problem is worth. A throwaway instance (manual iteration,
Playwright smoke tests against the admin frontend, anything that isn't "use
the real router") binds **8081** instead, and nothing should ever send a
kill signal to a process matched by name/port rather than by a PID it
started itself.
## Pointing a coding agent at it

152
README.md
View File

@@ -30,9 +30,14 @@ measurement is how you check whether it still holds for you.
- **Fall back to local vision when no cloud row supports images.** If no
vision-capable catalog candidate survives the hard filters, the router proxies
the request to a local Ollama vision model instead of returning 422.
- **Watch decisions arrive live.** `python tui.py` opens a terminal dashboard
- **Watch decisions arrive live.** `PYTHONPATH=src python -m tui` opens a terminal dashboard
that follows `/events/decisions` as decisions are recorded, with no polling
delay.
- **Manage it from a browser.** `GET /admin/` serves a four-page glass dark-mode
portal from the running router: a dashboard (quota burn, per-model usage,
verdict mix, history), a model-availability table with routing overrides, a
searchable decision log, and a controls page for operational triggers, runtime
toggles, and allowlisted `config/config.yaml` edits. Loopback-only, no auth.
- **Verify before learning.** Every routed response is structurally parsed in
the background; larger prose answers get an async local-LLM spot-check, and
failures fold back into per-model proficiency through `feedback.py`.
@@ -54,6 +59,7 @@ measurement is how you check whether it still holds for you.
- [Ask an image question](#ask-an-image-question)
- [Force JSON output](#force-json-output)
- [Watch it live](#watch-it-live)
- [Admin web portal](#admin-web-portal)
- [Probe routing without spending](#probe-routing-without-spending)
- [At a Glance](#at-a-glance)
- [Verification Pipeline](#verification-pipeline)
@@ -97,15 +103,15 @@ Nothing else is assumed about the host — routing itself is SQLite and arithmet
```bash
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
sqlite3 router.db < schema.sql
cp .env.example .env # fill in NEURALWATT_API_KEY
python poller.py # populate the catalog
python tier.py # resolve tiers
python config.py # sanity-check config loads
python -m uvicorn dispatcher:app --reload
sqlite3 router.db < config/schema.sql
cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
PYTHONPATH=src python -m poller # populate the catalog
PYTHONPATH=src python -m tier # resolve tiers
PYTHONPATH=src python -m config # sanity-check config loads
PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
```
Then edit `config.yaml` for your own setup — at minimum:
Then edit `config/config.yaml` for your own setup — at minimum:
| Key | Why |
|---|---|
@@ -115,7 +121,7 @@ Then edit `config.yaml` for your own setup — at minimum:
| `objective.assumed_cache_rate` | 0.917 was measured from one client's traffic (40.7M tokens). Check yours against the provider's per-session cache-hit figures |
| `session_cache.enabled` | off by default; caches category/tier per session for `staleness_minutes` to skip repeat classifier round-trips on long agent sessions |
`python seed_energy.py` is optional. It sweeps a fixed reference workload to
`PYTHONPATH=src python -m seed_energy` is optional. It sweeps a fixed reference workload to
populate `eco`, which is logged but is not an objective — routing works
without it. It costs real money and quota, so it is not in the path above.
@@ -127,9 +133,9 @@ Five user units cover continuous dispatch, catalog polling, and periodic energy
|---|---|---|
| `llm-router.service` | Continuous | FastAPI dispatcher |
| `llm-router-poller.timer` | 2 min after boot, then every 2 h | triggers the poller unit |
| `llm-router-poller.service` | oneshot | `poller.py` → `tier.py` |
| `llm-router-poller.service` | oneshot | `python -m poller` → `python -m tier` |
| `llm-router-seed.timer` | Every 6 h | triggers the seed unit |
| `llm-router-seed.service` | oneshot | small `seed_energy.py` sweep |
| `llm-router-seed.service` | oneshot | small `python -m seed_energy` sweep |
**The poller timer is load-bearing, not optional.** `freshness.stale_after_days`
is 3 with `exclude_stale: true` — an unpolled catalog marks every row stale
@@ -250,7 +256,7 @@ curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application
- Only inline `data:` URIs are accepted; remote `http(s)` image URLs are declined to avoid SSRF.
- Image count and total payload size are bounded before the local call is made.
- Configure the fallback in `config.yaml` under `local_vision:`.
- Configure the fallback in `config/config.yaml` under `local_vision:`.
### Force JSON output
@@ -277,23 +283,100 @@ curl -s localhost:8080/metrics | python -m json.tool
curl -s localhost:8080/events/decisions
# Terminal dashboard with live routing feed
python tui.py
PYTHONPATH=src python -m tui
```
- **`GET /metrics`** returns quota burn, coverage, recent decisions, per-model totals, verdict mix, and top proficiency. Loopback-only, no auth.
- **`GET /events/decisions`** is a Server-Sent Events stream of routing decisions. It replays recent decisions, then streams new ones as they happen; `:heartbeat` keepalive comments keep the connection alive between events.
- **`python tui.py`** is a Textual dashboard with a live routing feed, a category → model breakdown panel, a detail popup (press Enter or `e`), and quota/per-model/verdict/warnings panels. It is a separate entrypoint, not a systemd unit.
- **`baseline_report.py --since 2026-08-01`** is a read-only retrospective that replays recent decisions against two trivial counterfactuals — always pick the cheapest eligible candidate and always pick the best-proficiency eligible candidate — and reports total/mean cost and proficiency plus a dominance share. Add `--category coding_refactor` to scope it, or `--csv` for parseable output.
- **`PYTHONPATH=src python -m tui`** is a Textual dashboard with a live routing feed, a category → model breakdown panel, a detail popup (press Enter or `e`), and quota/per-model/verdict/warnings panels. It is a separate entrypoint, not a systemd unit.
- **`baseline_report.py --since 2026-08-01`** is a read-only retrospective (`PYTHONPATH=src python -m baseline_report`) that replays recent decisions against two trivial counterfactuals — always pick the cheapest eligible candidate and always pick the best-proficiency eligible candidate — and reports total/mean cost and proficiency plus a dominance share. Add `--category coding_refactor` to scope it, or `--csv` for parseable output.
`textual` is pinned in `requirements.txt` solely for the TUI modules (`tui.py`, `tui_screens.py`, `tui_sse.py`). It is imported only by these modules; the FastAPI service dispatch path never touches it, so the router itself has no UI dependency.
### Admin web portal
`GET /admin/` serves a management portal from the running router on the same
loopback-only bind as the API. It has no auth layer yet, so like the other
endpoints it is reachable only from `127.0.0.1`.
The portal is a **four-page, glass dark-mode** dashboard built on Tabler/Bootstrap
with a few lighter inline SVG icons (see `plans/admin-design-standards.md`
for the visual language and `plans/admin-work-framework.md` for the method
used to evolve it). It is implemented in `admin.py`, with `config/admin_schema.sql` for
its future tables and `admin/frontend/` for the browser UI.
**Dashboard** — `GET /admin/` aggregates the router at a glance: a quota chip
(burn against `objective.plan_kwh_per_period`), per-model usage bars, verdict
mix and category breakdown, five history mini-charts, and a recent-decisions
table.
<div align="center">
<img src="assets/screenshots/dashboard.png" alt="Admin dashboard" width="820">
</div>
**Models** — `GET /admin/models` is the model-availability table. Each row shows
the serving class, tier, status, and an override dropdown (`active` /
`deprecated` / `stale`) that writes through to the routing hard filters.
<div align="center">
<img src="assets/screenshots/models.png" alt="Admin models page" width="820">
</div>
**Controls** — `GET /admin/controls` holds the operational triggers, runtime
knobs, and the persisted `config/config.yaml` editor. Changes are marked as dirty and
written only on save.
<div align="center">
<img src="assets/screenshots/controls.png" alt="Admin controls page" width="820">
</div>
**Decisions** — `GET /admin/decisions` is the full decision log with
kind/category/tier filters and free-text search.
<div align="center">
<img src="assets/screenshots/decisions.png" alt="Admin decisions log" width="820">
</div>
**Read-only dashboards** — `GET /admin/api/snapshot` exposes data for:
- quota burn against `objective.plan_kwh_per_period`
- per-model usage from `energy_observations`
- live routing decisions from `route_decisions`
- verdict mix and scoring coverage
- history: `GET /admin/api/history?range=6h` (also `1h`, `24h`, `7d`, `30d`)
**Operational triggers** — async, fire-and-forget maintenance jobs:
- `POST /admin/api/refresh-catalog` runs the poller and tier pass
- `POST /admin/api/seed-energy?samples=N` starts a reference sweep
- `POST /admin/api/apply-feedback?dry_run=true` runs `feedback.py`
- `POST /admin/api/restart-service` restarts the running systemd unit
**Runtime toggles**
`GET /admin/api/runtime` shows persisted-vs-runtime values. `POST /admin/api/runtime/{knob}`
flips in-memory settings such as `log_route_decisions` and `local_llm_enabled`.
Changes take effect immediately but reset on restart.
**Persisted config edits**
`GET /admin/api/config` lists allowlisted keys. `POST /admin/api/config/{key}`
writes one allowlisted key back to `config/config.yaml` with a timestamped backup
and whole-config validation. Arbitrary keys are rejected.
**Model availability overrides**
`POST /admin/api/models/{model_id}/{provider}/availability` marks a model as
`active`, `deprecated`, or `stale`. `DELETE` on the same path removes the override.
Deprecation feeds into the routing hard filters.
### Probe routing without spending
`router_cli.py` is a one-shot shell probe that POSTs to `/route` once and prints the full decision tree.
```bash
python router_cli.py "Refactor this Django view into service objects"
python router_cli.py "Summarize this diff" --category summarization --tier 2
PYTHONPATH=src python -m router_cli "Refactor this Django view into service objects"
PYTHONPATH=src python -m router_cli "Summarize this diff" --category summarization --tier 2
```
- Prints the selected model, candidates, estimated cost, and estimated proficiency.
@@ -377,8 +460,8 @@ Observation → learning. `feedback.py` folds verification failures into
23-task benchmark:
```bash
python feedback.py --dry-run # preview what would change
python feedback.py # apply
PYTHONPATH=src python -m feedback --dry-run # preview what would change
PYTHONPATH=src python -m feedback # apply
```
Key behaviors:
@@ -458,11 +541,11 @@ Key behaviors:
| **Local Classification** | Ollama, OpenAI-compatible — `localhost:11434/v1` or an Ollama across your VPN |
| **Local Model** | `classifier.model` — `mistral-nemo:12b` by default; any Ollama model works |
| **Cloud Provider** | Neuralwatt only |
| **Config** | `config.yaml` loaded & validated by Pydantic (`config.py`) |
| **Config** | `config/config.yaml` loaded & validated by Pydantic (`src/config.py`) |
| **OpenAI Client** | `openai==3.0.0` (official SDK) |
| **HTTP** | `requests` for poller, `httpx` (via openai/uvicorn) |
| **Testing** | `pytest` — 562 tests across 27 files, all offline |
| **Config Files** | `config.yaml`, `leaderboards.yaml`, `evals/tasks.yaml` |
| **Testing** | `pytest` — 733 tests across 41 files, all offline |
| **Config Files** | `config/config.yaml`, `config/leaderboards.yaml`, `evals/tasks.yaml` |
| **Deployment** | systemd user units (`.service` + `.timer` files in `deploy/`) |
| **Integration** | `opencode.json` in the repo routes through it by default; any OpenAI-compatible client works |
@@ -505,6 +588,7 @@ restarts on boot shouldn't change its dependency tree underneath itself.
| **`config.py`** | YAML loader + Pydantic validators (blend weights sum to 1, valid tiers, endpoints separately addressable) | Yes (file) |
| **`metrics.py`** | Read-only aggregations for `/health` and `GET /metrics`: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency | Yes (DB) |
| **`events.py`** | In-memory decision-event broker for the TUI's live feed: bounded ring buffer + thread-safe fan-out to SSE subscribers | Pure |
| **`admin.py`** | `/admin` management portal: serves the four frontend pages + read/write API (snapshot, history, models/availability, runtime knobs, allowlisted config edits, operational triggers) | Yes (DB, network, file) |
| **`tui.py`** | Textual terminal dashboard over `GET /metrics` and `GET /events/decisions`; live routing feed, detail popup, category breakdown. Foreground tool, not a service | Yes (network) |
| **`tui_model.py`** | Pure data layer for the TUI: `build_model`, `build_category_breakdown`, `decision_row` — no Textual import, testable without a terminal | Pure |
| **`tui_sse.py`** | Background-thread SSE consumer for the TUI: reconnects on failure, marshals live decisions onto the UI thread | Yes (network) |
@@ -570,7 +654,7 @@ parses them into `access_level` and `routing.allowed_access_levels` (default
| `inherited_from` | TEXT | Model this row was copied from, NULL if measured directly |
| `last_updated` | TEXT | ISO8601 |
**Category set** (9 categories, defined in `config.yaml`):
**Category set** (9 categories, defined in `config/config.yaml`):
| Category | Example | Scoring type |
|---|---|---|
@@ -741,7 +825,7 @@ rows with `supports_vision = 1` survive the hard filters. If **no** cloud
candidate survives, the router can fall back to a local vision model instead
of returning 422.
`local_vision:` in `config.yaml` controls this path:
`local_vision:` in `config/config.yaml` controls this path:
| Key | Default | Purpose |
|---|---|---|
@@ -753,7 +837,7 @@ of returning 422.
| `max_images` | `4` | Refuse requests with more image parts |
| `max_image_bytes` | `9437184` (9 MiB) | Refuse requests whose image payload exceeds this |
The fallback is **enabled by default** both in `config.yaml` and in
The fallback is **enabled by default** both in `config/config.yaml` and in
`LocalVisionConfig`, so omitting the section still turns it on. Disable it
explicitly (`enabled: false`) on a host with no local Ollama or one that has
not pulled the vision model.
@@ -883,6 +967,7 @@ allowance.
| `POST` | `/dispatch` | Same as `/route`, plus complete the provider call, stream response, log observation |
| `GET` | `/v1/models` | OpenAI-compatible model list (router virtual models + catalog) |
| `POST` | `/v1/chat/completions` | OpenAI-compatible completions — routes then proxies, **streaming supported** |
| `GET` | `/admin` | Loopback-only web management portal (read-only dashboards, operational triggers, runtime toggles, allowlisted config edits) |
`/metrics` returns a single JSON object with these top-level keys:
@@ -951,7 +1036,7 @@ rank id=r9116d9 pos=0 model=deepseek-v4-flash prof=1 est_usd=0.00016296
**No conversation text is logged at any level**, prompt or answer — prompts
here run 60k–150k tokens and the journal is on disk. A test enforces it.
Set the level in `config.yaml` (`logging.level`), or override it without
Set the level in `config/config.yaml` (`logging.level`), or override it without
touching a tracked file:
```bash
@@ -976,10 +1061,10 @@ out clean.
## Self-Eval Harness (`eval_proficiency.py`)
```bash
python eval_proficiency.py # every routable model × every task
python eval_proficiency.py --models kimi-k3 # subset of models
python eval_proficiency.py --categories coding_general
python eval_proficiency.py --dry-run # plan only
PYTHONPATH=src python -m eval_proficiency # every routable model × every task
PYTHONPATH=src python -m eval_proficiency --models kimi-k3 # subset of models
PYTHONPATH=src python -m eval_proficiency --categories coding_general
PYTHONPATH=src python -m eval_proficiency --dry-run # plan only
```
- Runs every task through the target provider, scores it, writes to `proficiency`.
@@ -1037,7 +1122,7 @@ Several settings keep it from cascading failures:
unparseable output. A coding agent would rather have a mid-tier answer than
an error. Escalation deliberately skips fallbacks so an unavailable local
model doesn't silently promote every request to the frontier tier.
- **Session classification cache** (`session_cache:` in `config.yaml`,
- **Session classification cache** (`session_cache:` in `config/config.yaml`,
off by default) remembers the last `task_category`/`task_tier` decision per
session for `staleness_minutes` (default 20), so a long agent session skips
the classifier round-trip on every turn. It only ever short-circuits the
@@ -1068,7 +1153,7 @@ carries the same modality block.
## Testing
```bash
python -m pytest # 562 tests
python -m pytest # 733 tests
python -m pytest --cov # with coverage
```
@@ -1093,6 +1178,7 @@ config as arguments, so the suite runs offline on a clean checkout.
| `test_session_identity.py` | Outcome attribution: session matching, ambiguity refusal |
| `test_config_endpoints.py` | Classifier and verifier are separately addressable; guards on the split |
| `test_metrics_endpoint.py` | `/metrics` endpoint, SSE `/events/decisions` headers + replay/stream behavior |
| `test_admin_frontend.py` | Admin portal serves the four HTML pages with their expected markers and Chart.js asset |
| `test_events.py` | Decision-event broker: publish, subscribe/replay, unsubscribe, full-subscriber eviction |
| `test_tui.py` | TUI data model, category breakdown, detail popup, live SSE decision handling, keyboard controls |
| `test_context_prune.py` | Context pruning: image_url handling, structured content, recency guards, stats accuracy |
@@ -1119,7 +1205,7 @@ ollama create qwen3-vl-router:4b -f Modelfile.vision
With these tags the combined resident footprint was measured at roughly **15.9GB** in the worst case: classifier and verifier share one `mistral-nemo-router:12b` instance at ~8.6GB, with the `qwen3-vl-router:4b` vision fallback loaded alongside it. That leaves real headroom on a 24GB card.
`config.yaml` already points at these tags by default:
`config/config.yaml` already points at these tags by default:
```yaml
classifier:

View File

@@ -0,0 +1,652 @@
<!DOCTYPE html>
<html lang="en" data-bs-theme="dark">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Controls &middot; LLM Router Admin</title>
<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/@tabler/core@1.4.0/dist/css/tabler.min.css">
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=Quicksand:wght@600;700&display=swap">
<link rel="icon" href="data:image/svg+xml,%3Csvg%20xmlns%3D%22http%3A//www.w3.org/2000/svg%22%20viewBox%3D%220%200%20256%20256%22%3E%3Ccircle%20cx%3D%22146%22%20cy%3D%2284%22%20r%3D%2245.6%22%20fill%3D%22%23fff%22/%3E%3Cpath%20d%3D%22M124%20110%20Q90%20140%2090%20176%20Q90%20212%20122%20212%20Q156%20212%20172%20186%22%20fill%3D%22none%22%20stroke%3D%22%23fff%22%20stroke-width%3D%2239.2%22%20stroke-linecap%3D%22round%22/%3E%3Ccircle%20cx%3D%22146%22%20cy%3D%2284%22%20r%3D%2242%22%20fill%3D%22%23f59e0b%22/%3E%3Cpath%20d%3D%22M124%20110%20Q90%20140%2090%20176%20Q90%20212%20122%20212%20Q156%20212%20172%20186%22%20fill%3D%22none%22%20stroke%3D%22%23f59e0b%22%20stroke-width%3D%2232%22%20stroke-linecap%3D%22round%22/%3E%3Ccircle%20cx%3D%22128%22%20cy%3D%2278%22%20r%3D%229%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22164%22%20cy%3D%2278%22%20r%3D%229%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22128%22%20cy%3D%2278%22%20r%3D%226%22%20fill%3D%22%231e293b%22/%3E%3Ccircle%20cx%3D%22164%22%20cy%3D%2278%22%20r%3D%226%22%20fill%3D%22%231e293b%22/%3E%3Ccircle%20cx%3D%22131%22%20cy%3D%2274%22%20r%3D%222.2%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22167%22%20cy%3D%2274%22%20r%3D%222.2%22%20fill%3D%22%23fff%22/%3E%3Cpath%20d%3D%22M134%20100%20Q146%20112%20158%20100%22%20fill%3D%22none%22%20stroke%3D%22%23fff%22%20stroke-width%3D%227%22%20stroke-linecap%3D%22round%22/%3E%3Cpath%20d%3D%22M134%20100%20Q146%20112%20158%20100%22%20fill%3D%22none%22%20stroke%3D%22%231e293b%22%20stroke-width%3D%223.5%22%20stroke-linecap%3D%22round%22/%3E%3Cline%20x1%3D%22150.47%22%20y1%3D%2247.59%22%20x2%3D%22156.01%22%20y2%3D%2227.13%22%20stroke%3D%22%23fff%22%20stroke-width%3D%2210.73%22%20stroke-linecap%3D%22round%22/%3E%3Cline%20x1%3D%22150.47%22%20y1%3D%2247.59%22%20x2%3D%22155.08%22%20y2%3D%2230.54%22%20stroke%3D%22%231e293b%22%20stroke-width%3D%226.26%22%20stroke-linecap%3D%22round%22/%3E%3Ccircle%20cx%3D%2295%22%20cy%3D%22226%22%20r%3D%2214%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22150%22%20cy%3D%22226%22%20r%3D%2214%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%2295%22%20cy%3D%22226%22%20r%3D%229%22%20fill%3D%22%231e293b%22/%3E%3Ccircle%20cx%3D%22150%22%20cy%3D%22226%22%20r%3D%229%22%20fill%3D%22%231e293b%22/%3E%3C/svg%3E">
<script src="https://cdn.jsdelivr.net/npm/@tabler/core@1.4.0/dist/js/tabler.min.js"></script>
<style>
/* ── Liquid glass: soft gradient wash behind everything so the blur has
something to catch, then translucent+blurred navbar/cards on top. Kept
subtle — this is a working dashboard, not a marketing page. ── */
body{
background:
radial-gradient(900px circle at 6% -8%, rgba(245,158,11,.10), transparent 55%),
radial-gradient(800px circle at 94% -6%, rgba(59,130,246,.10), transparent 50%),
radial-gradient(760px circle at 50% 108%, rgba(139,92,246,.07), transparent 55%),
#111827;
background-attachment: fixed;
}
header.navbar{
background: rgba(31,41,55,.6)!important;
backdrop-filter: blur(16px) saturate(160%);
-webkit-backdrop-filter: blur(16px) saturate(160%);
border-bottom: 1px solid rgba(255,255,255,.06);
}
.card{
background: rgba(31,41,55,.55);
backdrop-filter: blur(18px) saturate(140%);
-webkit-backdrop-filter: blur(18px) saturate(140%);
border: 1px solid rgba(255,255,255,.07);
box-shadow: 0 10px 30px -14px rgba(0,0,0,.55), inset 0 1px 0 rgba(255,255,255,.04);
}
.card-header{border-bottom-color: rgba(255,255,255,.06)}
/* ── The "gap only at wide windows" bug, root-caused: this Chrome/Linux
build reports a phantom left margin on <html> exactly equal to the
scrollbar's width whenever <html> itself is the scrolling element.
Moving the scroll context down to <body> sidesteps it. ── */
html{overflow:hidden;height:100%;margin:0!important;padding:0!important}
body{overflow-y:auto;height:100%;margin:0!important;padding:0!important}
/* ── Buttons: glass pills instead of flat saturated fills. ── */
.btn-success{background:rgba(34,197,94,.16)!important;border-color:rgba(74,222,128,.45)!important;color:#4ade80!important}
.btn-success:hover{background:rgba(34,197,94,.28)!important;border-color:rgba(74,222,128,.7)!important;color:#86efac!important}
.btn-danger{background:rgba(239,68,68,.16)!important;border-color:rgba(248,113,113,.45)!important;color:#f87171!important}
.btn-danger:hover{background:rgba(239,68,68,.28)!important;border-color:rgba(248,113,113,.7)!important;color:#fca5a5!important}
.btn{backdrop-filter:blur(6px);-webkit-backdrop-filter:blur(6px)}
/* ── Badges: same glass treatment. ── */
.badge.bg-success{background:rgba(34,197,94,.20)!important;color:#4ade80!important}
.badge.bg-warning{background:rgba(245,158,11,.20)!important;color:#fbbf24!important}
.badge.bg-danger{background:rgba(239,68,68,.20)!important;color:#f87171!important}
.badge.bg-info{background:rgba(59,130,246,.20)!important;color:#93c5fd!important}
.badge.bg-secondary{background:rgba(148,163,184,.20)!important;color:#e2e8f0!important}
.badge.bg-purple{background:rgba(168,85,247,.20)!important;color:#d8b4fe!important}
/* ── Progress bars: translucent track + glow instead of flat solid. ── */
.progress{background:rgba(255,255,255,.08)!important;border-radius:999px;overflow:hidden}
.progress-bar{box-shadow:0 0 6px 0 currentColor;filter:saturate(1.25)}
header.navbar>.container-fluid{padding-left:20px!important;padding-right:20px!important}
.navbar-brand, .page-title{font-family:'Quicksand',var(--tblr-font-sans-serif,ui-sans-serif,system-ui,sans-serif);font-weight:700;letter-spacing:.01em}
.navbar-brand a{text-decoration:none!important;color:inherit}
/* ── Brand: Tabler's own rule is [data-bs-theme=dark] .navbar-brand-autodark
.navbar-brand-image{filter:brightness(0) invert(1)} — attribute selector +
2 classes outranks a plain 2-class override, so `filter:none` alone loses
the cascade regardless of source order. !important is what actually beats
it. That filter was stripping the mascot's orange color, and a sub-36px
cap shrank his face to sub-pixel dots. Restore color + force an explicit size (64px, below). ── */
.navbar-brand-autodark .navbar-brand-image{filter:none!important;height:64px;width:64px}
/* ── 64px mascot, same bar height. His face is only ~40% of his height (the
rest is tail and wheels), so at 45px the eyes were ~4px and the whole point
of having a mascot was lost. The 19px comes out of padding, not out of the
navbar: the brand h1 carries 8px top and bottom, and the navbar itself 4px.
Zero the first and halve the second and the bar lands at 69px — 1px shorter
than the 45px version it replaces. Measure before changing either number. ── */
.navbar-brand{padding-top:0!important;padding-bottom:0!important}
header.navbar{padding-top:2px!important;padding-bottom:2px!important}
/* ── Speed lines: he is a wheeled rover, so give him a trail. Drawn as their
own inline SVG rather than baked into the logo, so assets/*.svg stays the
plain mascot and the lines can drop out on narrow screens (d-none d-md-flex).
The gradient fades AWAY from him (transparent at the far left, amber where
he is), which is what reads as speed rather than as three floating dashes;
it is userSpaceOnUse because an objectBoundingBox gradient on a horizontal
line has a zero-height bbox and is undefined. ── */
.navbar-brand a{display:inline-flex;align-items:center;gap:.15rem}
/* The mascot IS a 6, so the wordmark's leading 6 echoes him in the brand amber.
The whole wordmark is wrapped in .brand-word first: the anchor is inline-flex
with a gap, so a bare <span> around the 6 would make it its own flex item and
open that gap between "6" and "krrt". */
.brand-six{color:#f59e0b;font-size:1.14em;line-height:1}
/* -20px: the logo's own viewBox carries ~18px of empty space to the left of
his tail at 64px, so the pull has to cover that before it buys any real
closeness. Measured gap from the middle line's tip to his back: ~4.5px. */
.brand-speedlines{align-items:center;line-height:0;margin-right:-20px}
.brand-speedlines svg{filter:none!important}
/* Motion only on hover — a dashboard that twitches at rest is a dashboard you
stop looking at. */
@keyframes brand-dash{
0%{transform:translateX(-5px);opacity:.2}
55%{opacity:1}
100%{transform:translateX(4px);opacity:0}
}
.navbar-brand a:hover .brand-speedlines line{animation:brand-dash .75s ease-in infinite}
.navbar-brand a:hover .brand-speedlines line:nth-child(2){animation-delay:.12s}
.navbar-brand a:hover .brand-speedlines line:nth-child(3){animation-delay:.24s}
@media (prefers-reduced-motion: reduce){
.navbar-brand a:hover .brand-speedlines line{animation:none}
}
/* ── Settings rows: ONE row grammar shared by Runtime Knobs and Persisted
Config — `key ......... [meta] [control]` on a 3-track grid with a fixed
control track, so every switch/input/select lands on the same right edge
and the eye scans a single column instead of chasing floating controls.
The old layout used Bootstrap .row col-5/col-3/col-4, which put the
control mid-row and spent a whole column restating the value the control
already shows. ── */
/* Negative margin == row padding, so the keys line up with the card body's
own content edge while the hover tint still bleeds to the card's inner edge. */
.settings-list{margin:0 -.5rem}
.setting-row{
display:grid;
grid-template-columns:minmax(0,1fr) auto 176px;
align-items:center;
gap:.75rem;
/* Fixed row height: a switch and a text input are ~10px apart in natural
height, which made the list look ragged when the two alternate. */
min-height:38px;
padding:.25rem .5rem;
border-left:2px solid transparent;
border-radius:6px;
}
.setting-row+.setting-row{border-top:1px solid rgba(255,255,255,.05)}
.setting-row:hover{background:rgba(255,255,255,.025)}
.setting-row.is-dirty{border-left-color:#fbbf24;background:rgba(245,158,11,.06)}
.setting-key{
font-size:.82rem;line-height:1.25;
overflow:hidden;text-overflow:ellipsis;white-space:nowrap;
}
.setting-key .scope{color:var(--tblr-secondary);opacity:.7}
.setting-meta{font-size:.72rem;white-space:nowrap;color:var(--tblr-secondary);font-variant-numeric:tabular-nums}
.setting-control{display:flex;justify-content:flex-end;align-items:center;min-width:0}
.setting-control .form-control,
.setting-control .form-select{width:100%;font-variant-numeric:tabular-nums}
/* Bootstrap floats .form-check-input and reserves 1.5em of left padding;
both fight right-alignment inside the control track. */
.setting-control .form-check{margin:0;padding:0;min-height:0;display:flex;justify-content:flex-end}
.setting-control .form-check-input{margin:0;float:none}
.settings-hint{font-size:.75rem;margin-bottom:.5rem}
/* Card bodies that are pure lists: let rows own the vertical rhythm. */
.card-body.settings-body{padding:.75rem 1rem 1rem}
</style>
</head>
<body>
<!-- ═══ Navbar ═══ -->
<header class="navbar navbar-expand-md navbar-dark" data-bs-theme="dark">
<div class="container-fluid">
<button class="navbar-toggler" type="button" data-bs-toggle="collapse" data-bs-target="#navbar-menu" aria-controls="navbar-menu" aria-expanded="false" aria-label="Toggle navigation"><span class="navbar-toggler-icon"></span></button>
<h1 class="navbar-brand navbar-brand-autodark d-none-navbar-horizontal pe-0 pe-md-3">
<a href="/admin/">
<span class="brand-speedlines d-none d-md-flex" aria-hidden="true">
<svg viewBox="0 0 40 64" width="40" height="64" xmlns="http://www.w3.org/2000/svg">
<defs>
<linearGradient id="speedfade" gradientUnits="userSpaceOnUse" x1="0" y1="0" x2="40" y2="0">
<stop offset="0" stop-color="#f59e0b" stop-opacity="0"/>
<stop offset="1" stop-color="#f59e0b" stop-opacity=".85"/>
</linearGradient>
</defs>
<g stroke="url(#speedfade)" stroke-linecap="round" fill="none">
<line x1="14" y1="20" x2="38" y2="20" stroke-width="5"/>
<line x1="2" y1="33" x2="36" y2="33" stroke-width="6"/>
<line x1="18" y1="46" x2="34" y2="46" stroke-width="4.5"/>
</g>
</svg>
</span>
<svg class="navbar-brand-image" viewBox="0 0 256 256" width="64" height="64" style="width:64px;height:64px" role="img" aria-hidden="true" xmlns="http://www.w3.org/2000/svg">
<circle cx="146" cy="84" r="45.6" fill="#fff"/>
<path d="M124 110 Q 90 140 90 176 Q 90 212 122 212 Q 156 212 172 186" fill="none" stroke="#fff" stroke-width="39.2" stroke-linecap="round"/>
<circle cx="146" cy="84" r="42" fill="#f59e0b"/>
<path d="M124 110 Q 90 140 90 176 Q 90 212 122 212 Q 156 212 172 186" fill="none" stroke="#f59e0b" stroke-width="32" stroke-linecap="round"/>
<circle cx="128" cy="78" r="9" fill="#fff"/><circle cx="164" cy="78" r="9" fill="#fff"/>
<circle cx="128" cy="78" r="6" fill="#1e293b"/><circle cx="164" cy="78" r="6" fill="#1e293b"/>
<circle cx="131" cy="74" r="2.2" fill="#fff"/><circle cx="167" cy="74" r="2.2" fill="#fff"/>
<path d="M134 100 Q146 112 158 100" fill="none" stroke="#fff" stroke-width="7" stroke-linecap="round"/>
<path d="M134 100 Q146 112 158 100" fill="none" stroke="#1e293b" stroke-width="3.5" stroke-linecap="round"/>
<line x1="150.47" y1="47.59" x2="156.01" y2="27.13" stroke="#ffffff" stroke-width="10.73" stroke-linecap="round"/>
<line x1="150.47" y1="47.59" x2="155.08" y2="30.54" stroke="#1e293b" stroke-width="6.26" stroke-linecap="round"/>
<ellipse cx="157.68" cy="-18.58" fill="#fff" transform="matrix(0.96737,0.25337,-0.26099,0.96534,0,0)" rx="9.05" ry="8.83"/>
<ellipse cx="157.68" cy="-18.58" fill="#1e293b" transform="matrix(0.96737,0.25337,-0.26099,0.96534,0,0)" rx="5.79" ry="5.65"/>
<circle cx="95" cy="226" r="14" fill="#fff"/><circle cx="150" cy="226" r="14" fill="#fff"/>
<circle cx="95" cy="226" r="9" fill="#1e293b"/><circle cx="150" cy="226" r="9" fill="#1e293b"/>
</svg>
<span class="brand-word"><span class="brand-six">6</span>krrt LLM Router</span>
</a>
</h1>
<div class="navbar-nav flex-row order-md-last">
<div class="d-flex align-items-center gap-3">
<span id="sse-status" class="badge bg-warning">connecting&hellip;</span>
<span id="sse-text" class="visually-hidden"></span>
</div>
</div>
<div class="collapse navbar-collapse" id="navbar-menu">
<div class="d-flex flex-column flex-md-row flex-fill align-items-stretch align-items-md-center">
<ul class="navbar-nav">
<li class="nav-item"><a class="nav-link" href="/admin/">Dashboard</a></li>
<li class="nav-item"><a class="nav-link" href="/admin/models">Models</a></li>
<li class="nav-item"><a class="nav-link" href="/admin/decisions">Decisions</a></li>
<li class="nav-item active"><a class="nav-link" href="/admin/controls">Controls</a></li>
</ul>
</div>
</div>
</div>
</header>
<!-- ═══ Page shell ═══ -->
<div class="page">
<div class="page-wrapper">
<div class="page-header d-print-none">
<div class="container-xl">
<div class="row g-2 align-items-center">
<div class="col">
<h2 class="page-title">Controls</h2>
<div class="text-muted mt-1">Maintenance triggers, runtime knobs, and persisted configuration</div>
</div>
</div>
</div>
</div>
<div class="page-body">
<div class="container-xl">
<div class="row row-cards align-items-start">
<!-- Row 1: Operational Triggers as a full-width strip. It is one
line of buttons, so as a half-width card beside the much taller
Runtime Knobs it left a column of dead space no matter what
align-items said. Buttons live in .card-actions; the body is a
single description/status line. -->
<div class="col-12">
<div class="card">
<div class="card-header">
<h3 class="card-title"><span class="me-2" data-icon="zap"></span>Operational Triggers</h3>
<div class="card-actions d-flex gap-2 flex-wrap">
<button class="btn btn-success btn-sm" onclick="triggerJob('refresh-catalog')">Refresh Catalog</button>
<button class="btn btn-success btn-sm" onclick="triggerJob('seed-energy')">Seed Energy</button>
<button class="btn btn-success btn-sm" onclick="triggerJob('apply-feedback')">Apply Feedback</button>
<button class="btn btn-danger btn-sm" onclick="triggerJob('restart-service')">Restart Service</button>
</div>
</div>
<div class="card-body py-2 d-flex flex-wrap align-items-center justify-content-between gap-2">
<span class="text-muted" style="font-size:.78rem">Fire-and-forget maintenance jobs.</span>
<span id="job-status" class="text-muted text-end" style="font-size:0.75rem"></span>
</div>
</div>
</div>
<!-- Row 2: the two settings panels side by side. Same row grammar,
same control column, so they read as one table split in half:
left is what the process is running now, right is what the file
says on next start. -->
<div class="col-xl-6">
<div class="card">
<div class="card-header"><h3 class="card-title"><span class="me-2" data-icon="settings"></span>Runtime Knobs</h3></div>
<div class="card-body settings-body">
<p class="text-muted settings-hint">In-memory only &mdash; reverts on restart, no <code>config.yaml</code> write.</p>
<div class="settings-list" id="runtime-list"></div>
</div>
</div>
</div>
<div class="col-xl-6">
<div class="card">
<div class="card-header">
<h3 class="card-title"><span class="me-2" data-icon="database"></span>Persisted Config</h3>
<div class="card-actions d-flex align-items-center gap-2">
<span id="config-save-status" class="text-muted" style="font-size:.72rem"></span>
<button class="btn btn-primary btn-sm" id="config-save-btn" onclick="saveAllConfig()" disabled>Save</button>
</div>
</div>
<div class="card-body settings-body">
<p class="text-muted settings-hint">Allowlisted keys, written to <code>config.yaml</code>.</p>
<div class="settings-list" id="config-list">
<div class="text-muted fst-italic small px-2 py-1">Loading config&hellip;</div>
</div>
</div>
</div>
</div>
</div><!-- /row row-cards -->
</div><!-- /container-xl -->
</div><!-- /page-body -->
<footer class="footer footer-transparent d-print-none">
<div class="container-xl">
&copy; 2026 adLee &middot; 6krrt, local LLM model router &middot; admin
</div>
</footer>
</div>
</div>
<!-- Toast (Bootstrap native) -->
<div id="toast" class="toast align-items-center text-bg-info border-0 position-fixed bottom-0 end-0 p-3" role="alert" aria-live="assertive" aria-atomic="true"></div>
<script>
/* ═══════════════════════════════════════════════════
admin/frontend/controls.html — single-file controls page
API base: relative (works under /admin/)
═══════════════════════════════════════════════════ */
const API = ''; // relative to /admin/
const SSE_URL = '/events/decisions'; // root-level endpoint
/* ═─ Inline icon helper (no tabler-icons webfont) ═─ */
function icon(name, size) {
const s = size || 18;
const svgs = {
zap: `<svg xmlns="http://www.w3.org/2000/svg" width="${s}" height="${s}" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polygon points="13 2 3 14 12 14 11 22 21 10 12 10 13 2"/></svg>`,
settings: `<svg xmlns="http://www.w3.org/2000/svg" width="${s}" height="${s}" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="3"/><path d="M19.4 15a1.65 1.65 0 0 0 .33 1.82l.06.06a2 2 0 0 1-2.83 2.83l-.06-.06a1.65 1.65 0 0 0-1.82-.33 1.65 1.65 0 0 0-1 1.51V21a2 2 0 0 1-4 0v-.09A1.65 1.65 0 0 0 9 19.4a1.65 1.65 0 0 0-1.82.33l-.06.06a2 2 0 0 1-2.83-2.83l.06-.06A1.65 1.65 0 0 0 4.68 15a1.65 1.65 0 0 0-1.51-1H3a2 2 0 0 1 0-4h.09A1.65 1.65 0 0 0 4.6 9a1.65 1.65 0 0 0-.33-1.82l-.06-.06a2 2 0 0 1 2.83-2.83l.06.06A1.65 1.65 0 0 0 9 4.68a1.65 1.65 0 0 0 1-1.51V3a2 2 0 0 1 4 0v.09a1.65 1.65 0 0 0 1 1.51 1.65 1.65 0 0 0 1.82-.33l.06-.06a2 2 0 0 1 2.83 2.83l-.06.06A1.65 1.65 0 0 0 19.4 9a1.65 1.65 0 0 0 1.51 1H21a2 2 0 0 1 0 4h-.09a1.65 1.65 0 0 0-1.51 1z"/></svg>`,
database: `<svg xmlns="http://www.w3.org/2000/svg" width="${s}" height="${s}" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><ellipse cx="12" cy="5" rx="9" ry="3"/><path d="M21 12c0 1.66-4 3-9 3s-9-1.34-9-3"/><path d="M3 5v14c0 1.66 4 3 9 3s9-1.34 9-3V5"/></svg>`,
default: `<svg xmlns="http://www.w3.org/2000/svg" width="${s}" height="${s}" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="1"/></svg>`,
};
return svgs[name] || svgs.default;
}
function renderStaticIcons() {
document.querySelectorAll('[data-icon]').forEach(el => {
el.innerHTML = icon(el.dataset.icon);
});
}
/* ── Toast (Bootstrap native via tabler wrapper) ── */
let _toastEl = null;
function toast(msg, type = 'info') {
const bgClass = { success: 'text-bg-success', error: 'text-bg-danger', info: 'text-bg-info' }[type] || 'text-bg-info';
if (!_toastEl) {
_toastEl = document.getElementById('toast');
_toastEl.innerHTML = `<div class="d-flex">
<div class="toast-body"></div>
<button type="button" class="btn-close btn-close-white me-2 m-auto" data-bs-dismiss="toast" aria-label="Close"></button>
</div>`;
_toastEl._bt = new tabler.Toast(_toastEl, { delay: 4000 });
}
_toastEl.querySelector('.toast-body').textContent = msg;
_toastEl.className = `toast align-items-center ${bgClass} border-0 position-fixed bottom-0 end-0 p-3`;
_toastEl._bt.show();
}
/* ── Fetch Wrapper ── */
async function apiFetch(url, opts = {}) {
try {
const resp = await fetch(url, opts);
if (!resp.ok) throw new Error(`${resp.status} ${resp.statusText}`);
return await resp.json();
} catch (e) {
console.warn('API call failed:', url, e);
return null;
}
}
/* ═══════════════════════════════════════
SSE LIVE STREAM
═══════════════════════════════════════ */
let sseConn = null;
function connectSSE() {
try {
if (sseConn) { sseConn.close(); }
sseConn = new EventSource(SSE_URL);
const statusDot = document.getElementById('sse-status');
const statusText = document.getElementById('sse-text');
statusDot.className = 'badge bg-warning';
statusDot.textContent = 'connecting…';
statusText.textContent = 'connecting…';
sseConn.onopen = () => {
statusDot.className = 'badge bg-success';
statusDot.textContent = 'live';
statusText.textContent = 'live';
};
sseConn.onerror = () => {
statusDot.className = 'badge bg-danger';
statusDot.textContent = 'reconnecting…';
statusText.textContent = 'reconnecting…';
};
} catch (e) {
document.getElementById('sse-status').className = 'badge bg-danger';
document.getElementById('sse-status').textContent = 'offline';
document.getElementById('sse-text').textContent = 'offline';
}
}
/* ═══════════════════════════════════════
OPERATIONAL TRIGGERS
═══════════════════════════════════════ */
async function triggerJob(action) {
const statusEl = document.getElementById('job-status');
let url;
switch(action) {
case 'refresh-catalog': url = `${API}api/refresh-catalog`; break;
case 'seed-energy': url = `${API}api/seed-energy?samples=5`; break;
case 'apply-feedback': url = `${API}api/apply-feedback?dry_run=true`; break;
case 'restart-service': url = `${API}api/restart-service`; break;
default: return;
}
statusEl.innerHTML = `<span style="color:var(--tblr-warning)">Running ${action}&hellip;</span>`;
const resp = await apiFetch(url, { method:'POST' });
if (resp) {
if (resp.status === 'restarting') {
statusEl.innerHTML = `<span style="color:var(--tblr-danger)">Service restarting, page will reload automatically</span>`;
toast('Service restart triggered', 'info');
setTimeout(() => location.reload(), 5000);
} else {
statusEl.innerHTML = `<span style="color:var(--tblr-success)">${action}: ${resp.status || 'done'}</span>`;
toast(`${action} completed`, 'success');
}
} else {
statusEl.innerHTML = `<span style="color:var(--tblr-danger)">${action} failed</span>`;
toast(`${action} failed`, 'error');
}
}
/* ═══════════════════════════════════════
RUNTIME KNOBS
═══════════════════════════════════════ */
function knobDisplayValue(v) {
if (v !== null && typeof v === 'object') {
if ('enabled' in v) return String(v.enabled);
try { return JSON.stringify(v); } catch (_) { return String(v); }
}
if (typeof v === 'boolean') return String(v);
return String(v);
}
// Knobs (and Persisted Config keys, below) with only a few legal values get
// a <select> instead of a free-text input — the values come straight from
// the Pydantic enum/validator on the backend, so this list stays in sync
// with config.py by hand, not by inspection.
const ENUM_VALUES = {
default_flex_preference: ['no-flex', 'auto', 'prefer-flex', 'force-flex'],
'routing.default_flex_preference': ['no-flex', 'auto', 'prefer-flex', 'force-flex'],
'logging.level': ['debug', 'info', 'warning', 'error'],
};
function enumSelect(key, value, { dataAttr = 'data-knob', onchangeExpr = '' } = {}) {
const opts = ENUM_VALUES[key].map(v =>
`<option value="${v}"${v === value ? ' selected' : ''}>${v}</option>`
).join('');
const onchange = onchangeExpr ? ` onchange="${onchangeExpr}"` : '';
return `<select class="form-select form-select-sm toggle-input" ${dataAttr}="${key}"${onchange}>${opts}</select>`;
}
// A dotted key reads as scope + leaf; dimming the scope makes a column of
// `objective.*` / `pinch.*` keys scannable by the part that differs.
function keyHtml(key) {
const cut = key.lastIndexOf('.');
if (cut === -1) return escapeHtml(key);
return `<span class="scope">${escapeHtml(key.slice(0, cut + 1))}</span>${escapeHtml(key.slice(cut + 1))}`;
}
// Units the config file expresses but the key name doesn't.
const UNITS = {
'objective.max_energy_per_request': 'kWh',
'objective.plan_kwh_per_period': 'kWh',
};
function settingRow({ key, meta = '', control, dirtyAttrs = '' }) {
return `<div class="setting-row"${dirtyAttrs}>
<span class="setting-key" title="${escapeHtml(key)}">${keyHtml(key)}</span>
<span class="setting-meta">${meta}</span>
<span class="setting-control">${control}</span>
</div>`;
}
function renderRuntime(state) {
const el = document.getElementById('runtime-list');
const knobs = state || {};
const html = Object.keys(knobs).map(key => {
const knob = knobs[key];
const persisted = knob.persisted;
const runtime = knob.runtime;
const isDiff = JSON.stringify(persisted) !== JSON.stringify(runtime);
const runtimeStr = knobDisplayValue(runtime);
let control;
if (typeof persisted === 'boolean') {
control = `<div class="form-check form-switch">
<input class="form-check-input toggle-input" type="checkbox" data-knob="${key}" ${runtime ? 'checked' : ''} onchange="toggleKnob('${key}', this.checked)">
</div>`;
} else if (ENUM_VALUES[key]) {
control = enumSelect(key, runtime, { onchangeExpr: `toggleKnob('${key}', this.value)` });
} else if (typeof persisted === 'string') {
control = `<input class="form-control form-control-sm toggle-input" type="text" data-knob="${key}" value="${escapeHtml(runtimeStr)}" onchange="toggleKnob('${key}', this.value)">`;
} else {
control = `<span class="text-muted small">${escapeHtml(runtimeStr)}</span>`;
}
// The control already shows the live value, so the only fact worth a
// second column is disagreement with the file.
const meta = isDiff
? `<span class="badge bg-warning-subtle text-warning" title="config.yaml says ${escapeHtml(knobDisplayValue(persisted))}; running with ${escapeHtml(runtimeStr)}">file: ${escapeHtml(knobDisplayValue(persisted))}</span>`
: '';
return settingRow({ key, meta, control });
}).join('');
el.innerHTML = html || '<div class="text-muted fst-italic small px-2 py-1">No runtime knobs</div>';
}
async function toggleKnob(knob, value) {
const body = { value };
if (typeof value === 'string') body.value = value;
const resp = await apiFetch(
`${API}api/runtime/${encodeURIComponent(knob)}`,
{ method:'POST', headers:{'Content-Type':'application/json'}, body:JSON.stringify(body) }
);
if (resp?.ok) toast(`${knob} toggled`, 'success');
else toast(`Failed to toggle ${knob}`, 'error');
}
/* ═══════════════════════════════════════
PERSISTED CONFIG EDITOR
═══════════════════════════════════════ */
function renderConfig(config) {
const list = document.getElementById('config-list');
if (!config || !Object.keys(config).length) {
list.innerHTML = '<div class="text-muted fst-italic small px-2 py-1">No config data</div>';
return;
}
const html = Object.entries(config).map(([key, val]) => {
let control;
if (typeof val === 'boolean') {
// A switch, matching Runtime Knobs — the bare checkbox here was the one
// control on the page that didn't look like the others.
control = `<div class="form-check form-switch">
<input class="form-check-input config-input toggle-input" type="checkbox" ${val ? 'checked' : ''} data-config-input>
</div>`;
} else if (ENUM_VALUES[key]) {
control = enumSelect(key, val, { dataAttr: 'data-config-input' });
} else {
control = `<input class="form-control form-control-sm config-input toggle-input" type="text" value="${val === null ? '' : escapeHtml(String(val))}" placeholder="${val === null ? 'null' : ''}" data-config-input>`;
}
return settingRow({
key,
meta: UNITS[key] ? escapeHtml(UNITS[key]) : '',
control,
dirtyAttrs: ` data-key="${escapeHtml(key)}" data-orig="${escapeHtml(String(val === null ? '' : val))}"`,
});
}).join('');
list.innerHTML = html;
markDirty();
}
/* The inputs hold the values, so the only thing worth flagging is an edit that
hasn't been written yet: an amber rail on the row, a live count on Save. */
function rowValue(row) {
const input = row.querySelector('[data-config-input]');
return input.type === 'checkbox' ? String(input.checked) : input.value;
}
function markDirty() {
const rows = document.querySelectorAll('#config-list .setting-row[data-key]');
let dirty = 0;
rows.forEach(row => {
const changed = rowValue(row) !== row.getAttribute('data-orig');
row.classList.toggle('is-dirty', changed);
if (changed) dirty += 1;
});
const btn = document.getElementById('config-save-btn');
btn.disabled = dirty === 0;
btn.textContent = dirty ? `Save ${dirty} change${dirty === 1 ? '' : 's'}` : 'Save';
return dirty;
}
async function saveAllConfig() {
const rows = [...document.querySelectorAll('#config-list .setting-row[data-key]')]
.filter(row => rowValue(row) !== row.getAttribute('data-orig'));
const statusEl = document.getElementById('config-save-status');
if (!rows.length) return;
let ok = 0;
let failed = 0;
for (const row of rows) {
const key = row.getAttribute('data-key');
const input = row.querySelector('[data-config-input]');
const val = input.type === 'checkbox' ? input.checked : input.value;
let parsed = val;
if (typeof val === 'string') {
if (val.trim() === '') {
parsed = null;
} else {
const num = Number(val);
if (!isNaN(num)) parsed = num;
}
}
statusEl.textContent = `saving ${key}…`;
statusEl.style.color = 'var(--tblr-secondary)';
const res = await apiFetch(
`${API}api/config/${encodeURIComponent(key)}`,
{ method:'POST', headers:{'Content-Type':'application/json'}, body:JSON.stringify({ value: parsed }) }
);
if (res === null) {
failed += 1;
row.style.color = 'var(--tblr-danger)';
} else {
ok += 1;
row.style.color = '';
}
}
const allOk = failed === 0;
statusEl.textContent = allOk ? `saved ${ok}` : `${failed} write(s) failed, ${ok} saved`;
statusEl.style.color = allOk ? 'var(--tblr-success)' : 'var(--tblr-danger)';
if (allOk) toast('Config updated', 'success');
setTimeout(loadControls, 1500);
}
/* ═══════════════════════════════════════
DATA FETCHING
═══════════════════════════════════════ */
async function loadControls() {
const runtime = await apiFetch(`${API}api/runtime`);
if (runtime) renderRuntime(runtime);
const config = await apiFetch(`${API}api/config`);
if (config) renderConfig(config);
}
/* ═══════════════════════════════════════
UTILITIES
═══════════════════════════════════════ */
function escapeHtml(s) {
return String(s).replace(/&/g,'&amp;').replace(/</g,'&lt;').replace(/>/g,'&gt;').replace(/"/g,'&quot;');
}
/* ═══════════════════════════════════════
INIT
═══════════════════════════════════════ */
function init() {
renderStaticIcons();
// Delegated so it survives every re-render of the list.
const configList = document.getElementById('config-list');
configList.addEventListener('input', markDirty);
configList.addEventListener('change', markDirty);
loadControls();
connectSSE();
}
init();
</script>
</body>
</html>

View File

@@ -0,0 +1,553 @@
<!DOCTYPE html>
<html lang="en" data-bs-theme="dark">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Decisions &middot; LLM Router Admin</title>
<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/@tabler/core@1.4.0/dist/css/tabler.min.css">
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=Quicksand:wght@600;700&display=swap">
<link rel="icon" href="data:image/svg+xml,%3Csvg%20xmlns%3D%22http%3A//www.w3.org/2000/svg%22%20viewBox%3D%220%200%20256%20256%22%3E%3Ccircle%20cx%3D%22146%22%20cy%3D%2284%22%20r%3D%2245.6%22%20fill%3D%22%23fff%22/%3E%3Cpath%20d%3D%22M124%20110%20Q90%20140%2090%20176%20Q90%20212%20122%20212%20Q156%20212%20172%20186%22%20fill%3D%22none%22%20stroke%3D%22%23fff%22%20stroke-width%3D%2239.2%22%20stroke-linecap%3D%22round%22/%3E%3Ccircle%20cx%3D%22146%22%20cy%3D%2284%22%20r%3D%2242%22%20fill%3D%22%23f59e0b%22/%3E%3Cpath%20d%3D%22M124%20110%20Q90%20140%2090%20176%20Q90%20212%20122%20212%20Q156%20212%20172%20186%22%20fill%3D%22none%22%20stroke%3D%22%23f59e0b%22%20stroke-width%3D%2232%22%20stroke-linecap%3D%22round%22/%3E%3Ccircle%20cx%3D%22128%22%20cy%3D%2278%22%20r%3D%229%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22164%22%20cy%3D%2278%22%20r%3D%229%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22128%22%20cy%3D%2278%22%20r%3D%226%22%20fill%3D%22%231e293b%22/%3E%3Ccircle%20cx%3D%22164%22%20cy%3D%2278%22%20r%3D%226%22%20fill%3D%22%231e293b%22/%3E%3Ccircle%20cx%3D%22131%22%20cy%3D%2274%22%20r%3D%222.2%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22167%22%20cy%3D%2274%22%20r%3D%222.2%22%20fill%3D%22%23fff%22/%3E%3Cpath%20d%3D%22M134%20100%20Q146%20112%20158%20100%22%20fill%3D%22none%22%20stroke%3D%22%23fff%22%20stroke-width%3D%227%22%20stroke-linecap%3D%22round%22/%3E%3Cpath%20d%3D%22M134%20100%20Q146%20112%20158%20100%22%20fill%3D%22none%22%20stroke%3D%22%231e293b%22%20stroke-width%3D%223.5%22%20stroke-linecap%3D%22round%22/%3E%3Cline%20x1%3D%22150.47%22%20y1%3D%2247.59%22%20x2%3D%22156.01%22%20y2%3D%2227.13%22%20stroke%3D%22%23fff%22%20stroke-width%3D%2210.73%22%20stroke-linecap%3D%22round%22/%3E%3Cline%20x1%3D%22150.47%22%20y1%3D%2247.59%22%20x2%3D%22155.08%22%20y2%3D%2230.54%22%20stroke%3D%22%231e293b%22%20stroke-width%3D%226.26%22%20stroke-linecap%3D%22round%22/%3E%3Ccircle%20cx%3D%2295%22%20cy%3D%22226%22%20r%3D%2214%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22150%22%20cy%3D%22226%22%20r%3D%2214%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%2295%22%20cy%3D%22226%22%20r%3D%229%22%20fill%3D%22%231e293b%22/%3E%3Ccircle%20cx%3D%22150%22%20cy%3D%22226%22%20r%3D%229%22%20fill%3D%22%231e293b%22/%3E%3C/svg%3E">
<script src="https://cdn.jsdelivr.net/npm/@tabler/core@1.4.0/dist/js/tabler.min.js"></script>
<style>
/* ── Liquid glass: soft gradient wash behind everything so the blur has
something to catch, then translucent+blurred navbar/cards on top. Kept
subtle — this is a working dashboard, not a marketing page. ── */
body{
background:
radial-gradient(900px circle at 6% -8%, rgba(245,158,11,.10), transparent 55%),
radial-gradient(800px circle at 94% -6%, rgba(59,130,246,.10), transparent 50%),
radial-gradient(760px circle at 50% 108%, rgba(139,92,246,.07), transparent 55%),
#111827;
background-attachment: fixed;
}
header.navbar{
background: rgba(31,41,55,.6)!important;
backdrop-filter: blur(16px) saturate(160%);
-webkit-backdrop-filter: blur(16px) saturate(160%);
border-bottom: 1px solid rgba(255,255,255,.06);
position: relative;
z-index: 1030;
}
.card{
background: rgba(31,41,55,.55);
backdrop-filter: blur(18px) saturate(140%);
-webkit-backdrop-filter: blur(18px) saturate(140%);
border: 1px solid rgba(255,255,255,.07);
box-shadow: 0 10px 30px -14px rgba(0,0,0,.55), inset 0 1px 0 rgba(255,255,255,.04);
}
.card-header{border-bottom-color: rgba(255,255,255,.06)}
/* ── The "gap only at wide windows" bug, root-caused: this Chrome/Linux
build reports a phantom left margin on <html> exactly equal to the
scrollbar's width whenever <html> itself is the scrolling element.
Moving the scroll context down to <body> sidesteps it. ── */
html{overflow:hidden;height:100%;margin:0!important;padding:0!important}
body{overflow-y:auto;height:100%;margin:0!important;padding:0!important}
/* ── Buttons: glass pills instead of flat saturated fills. ── */
.btn-success{background:rgba(34,197,94,.16)!important;border-color:rgba(74,222,128,.45)!important;color:#4ade80!important}
.btn-success:hover{background:rgba(34,197,94,.28)!important;border-color:rgba(74,222,128,.7)!important;color:#86efac!important}
.btn-danger{background:rgba(239,68,68,.16)!important;border-color:rgba(248,113,113,.45)!important;color:#f87171!important}
.btn-danger:hover{background:rgba(239,68,68,.28)!important;border-color:rgba(248,113,113,.7)!important;color:#fca5a5!important}
.btn{backdrop-filter:blur(6px);-webkit-backdrop-filter:blur(6px)}
/* ── Badges: same glass treatment. ── */
.badge.bg-success{background:rgba(34,197,94,.20)!important;color:#4ade80!important}
.badge.bg-warning{background:rgba(245,158,11,.20)!important;color:#fbbf24!important}
.badge.bg-danger{background:rgba(239,68,68,.20)!important;color:#f87171!important}
.badge.bg-info{background:rgba(59,130,246,.20)!important;color:#93c5fd!important}
.badge.bg-secondary{background:rgba(148,163,184,.20)!important;color:#e2e8f0!important}
.badge.bg-purple{background:rgba(168,85,247,.20)!important;color:#d8b4fe!important}
/* ── Progress bars: translucent track + glow instead of flat solid. ── */
.progress{background:rgba(255,255,255,.08)!important;border-radius:999px;overflow:hidden}
.progress-bar{box-shadow:0 0 6px 0 currentColor;filter:saturate(1.25)}
header.navbar>.container-fluid{padding-left:20px!important;padding-right:20px!important}
.navbar-brand, .page-title{font-family:'Quicksand',var(--tblr-font-sans-serif,ui-sans-serif,system-ui,sans-serif);font-weight:700;letter-spacing:.01em}
.navbar-brand a{text-decoration:none!important;color:inherit}
/* ── Brand: Tabler's own rule is [data-bs-theme=dark] .navbar-brand-autodark
.navbar-brand-image{filter:brightness(0) invert(1)} — attribute selector +
2 classes outranks a plain 2-class override, so `filter:none` alone loses
the cascade regardless of source order. !important is what actually beats
it. That filter was stripping the mascot's orange color, and a sub-45px
cap shrank his face to sub-pixel dots. Restore color + force an explicit size (64px, below). ── */
.navbar-brand-autodark .navbar-brand-image{filter:none!important;height:64px;width:64px}
/* ── 64px mascot, same bar height. His face is only ~40% of his height (the
rest is tail and wheels), so at 45px the eyes were ~4px and the whole point
of having a mascot was lost. The 19px comes out of padding, not out of the
navbar: the brand h1 carries 8px top and bottom, and the navbar itself 4px.
Zero the first and halve the second and the bar lands at 69px — 1px shorter
than the 45px version it replaces. Measure before changing either number. ── */
.navbar-brand{padding-top:0!important;padding-bottom:0!important}
header.navbar{padding-top:2px!important;padding-bottom:2px!important}
/* ── Speed lines: he is a wheeled rover, so give him a trail. Drawn as their
own inline SVG rather than baked into the logo, so assets/*.svg stays the
plain mascot and the lines can drop out on narrow screens (d-none d-md-flex).
The gradient fades AWAY from him (transparent at the far left, amber where
he is), which is what reads as speed rather than as three floating dashes;
it is userSpaceOnUse because an objectBoundingBox gradient on a horizontal
line has a zero-height bbox and is undefined. ── */
.navbar-brand a{display:inline-flex;align-items:center;gap:.15rem}
/* The mascot IS a 6, so the wordmark's leading 6 echoes him in the brand amber.
The whole wordmark is wrapped in .brand-word first: the anchor is inline-flex
with a gap, so a bare <span> around the 6 would make it its own flex item and
open that gap between "6" and "krrt". */
.brand-six{color:#f59e0b;font-size:1.14em;line-height:1}
/* -20px: the logo's own viewBox carries ~18px of empty space to the left of
his tail at 64px, so the pull has to cover that before it buys any real
closeness. Measured gap from the middle line's tip to his back: ~4.5px. */
.brand-speedlines{align-items:center;line-height:0;margin-right:-20px}
.brand-speedlines svg{filter:none!important}
/* Motion only on hover — a dashboard that twitches at rest is a dashboard you
stop looking at. */
@keyframes brand-dash{
0%{transform:translateX(-5px);opacity:.2}
55%{opacity:1}
100%{transform:translateX(4px);opacity:0}
}
.navbar-brand a:hover .brand-speedlines line{animation:brand-dash .75s ease-in infinite}
.navbar-brand a:hover .brand-speedlines line:nth-child(2){animation-delay:.12s}
.navbar-brand a:hover .brand-speedlines line:nth-child(3){animation-delay:.24s}
@media (prefers-reduced-motion: reduce){
.navbar-brand a:hover .brand-speedlines line{animation:none}
}
.tier-1{color:var(--tblr-success)}
.tier-2{color:var(--tblr-warning)}
.tier-3{color:var(--tblr-danger)}
.tier-?{color:var(--tblr-secondary)}
/* ── Filter bar ── */
.dec-filters{display:flex;flex-wrap:wrap;gap:8px;align-items:center}
.dec-filters .form-select,.dec-filters .form-control{width:auto;min-width:120px}
.dec-filters .form-control{min-width:220px;flex:1 1 220px}
.dec-count{font-size:.8rem;color:var(--tblr-secondary);white-space:nowrap}
/* ── Row density: this table's whole point is scanning a lot of rows at
once, so it trades card-table's normal comfortable spacing for a dense
grid — smaller type plus tighter padding, not just one or the other. ── */
#dec-tbody td, thead th{padding-top:.2rem!important;padding-bottom:.2rem!important}
#dec-tbody{font-size:.72rem}
#dec-tbody td{line-height:1.25}
/* ── Visual cues: kind already reads as a colored badge — category, tier
and source now match that language instead of sitting as plain text,
so the whole row scans by color instead of by reading every cell. ── */
.dec-badge{display:inline-block;padding:.15rem .5rem;border-radius:999px;font-size:.72rem;font-weight:600;white-space:nowrap}
.flag-chip{display:inline-flex;align-items:center;justify-content:center;width:18px;height:18px;border-radius:5px;font-size:.62rem;font-weight:700;margin-right:3px;background:rgba(255,255,255,.08);color:var(--tblr-secondary);cursor:default}
.flag-chip.on{background:rgba(59,130,246,.22);color:#93c5fd}
</style>
</head>
<body>
<!-- ═══ Navbar ═══ -->
<header class="navbar navbar-expand-md navbar-dark" data-bs-theme="dark">
<div class="container-fluid">
<button class="navbar-toggler" type="button" data-bs-toggle="collapse" data-bs-target="#navbar-menu" aria-controls="navbar-menu" aria-expanded="false" aria-label="Toggle navigation"><span class="navbar-toggler-icon"></span></button>
<h1 class="navbar-brand navbar-brand-autodark d-none-navbar-horizontal pe-0 pe-md-3">
<a href="/admin/" style="display:inline-flex;align-items:center;gap:8px">
<span class="brand-speedlines d-none d-md-flex" aria-hidden="true">
<svg viewBox="0 0 40 64" width="40" height="64" xmlns="http://www.w3.org/2000/svg">
<defs>
<linearGradient id="speedfade" gradientUnits="userSpaceOnUse" x1="0" y1="0" x2="40" y2="0">
<stop offset="0" stop-color="#f59e0b" stop-opacity="0"/>
<stop offset="1" stop-color="#f59e0b" stop-opacity=".85"/>
</linearGradient>
</defs>
<g stroke="url(#speedfade)" stroke-linecap="round" fill="none">
<line x1="14" y1="20" x2="38" y2="20" stroke-width="5"/>
<line x1="2" y1="33" x2="36" y2="33" stroke-width="6"/>
<line x1="18" y1="46" x2="34" y2="46" stroke-width="4.5"/>
</g>
</svg>
</span>
<svg class="navbar-brand-image" viewBox="0 0 256 256" width="64" height="64" style="width:64px;height:64px" role="img" aria-hidden="true" xmlns="http://www.w3.org/2000/svg">
<circle cx="146" cy="84" r="45.6" fill="#fff"/>
<path d="M124 110 Q 90 140 90 176 Q 90 212 122 212 Q 156 212 172 186" fill="none" stroke="#fff" stroke-width="39.2" stroke-linecap="round"/>
<circle cx="146" cy="84" r="42" fill="#f59e0b"/>
<path d="M124 110 Q 90 140 90 176 Q 90 212 122 212 Q 156 212 172 186" fill="none" stroke="#f59e0b" stroke-width="32" stroke-linecap="round"/>
<circle cx="128" cy="78" r="9" fill="#fff"/><circle cx="164" cy="78" r="9" fill="#fff"/>
<circle cx="128" cy="78" r="6" fill="#1e293b"/><circle cx="164" cy="78" r="6" fill="#1e293b"/>
<circle cx="131" cy="74" r="2.2" fill="#fff"/><circle cx="167" cy="74" r="2.2" fill="#fff"/>
<path d="M134 100 Q146 112 158 100" fill="none" stroke="#fff" stroke-width="7" stroke-linecap="round"/>
<path d="M134 100 Q146 112 158 100" fill="none" stroke="#1e293b" stroke-width="3.5" stroke-linecap="round"/>
<line x1="150.47" y1="47.59" x2="156.01" y2="27.13" stroke="#ffffff" stroke-width="10.73" stroke-linecap="round"/>
<line x1="150.47" y1="47.59" x2="155.08" y2="30.54" stroke="#1e293b" stroke-width="6.26" stroke-linecap="round"/>
<ellipse cx="157.68" cy="-18.58" fill="#fff" transform="matrix(0.96737,0.25337,-0.26099,0.96534,0,0)" rx="9.05" ry="8.83"/>
<ellipse cx="157.68" cy="-18.58" fill="#1e293b" transform="matrix(0.96737,0.25337,-0.26099,0.96534,0,0)" rx="5.79" ry="5.65"/>
<circle cx="95" cy="226" r="14" fill="#fff"/><circle cx="150" cy="226" r="14" fill="#fff"/>
<circle cx="95" cy="226" r="9" fill="#1e293b"/><circle cx="150" cy="226" r="9" fill="#1e293b"/>
</svg>
<span class="brand-word"><span class="brand-six">6</span>krrt LLM Router</span>
</a>
</h1>
<div class="navbar-nav flex-row order-md-last">
<div class="d-flex align-items-center gap-3">
<span id="sse-status" class="badge bg-warning">connecting&hellip;</span>
<span id="sse-text" class="visually-hidden"></span>
<span id="generated-at" class="text-muted small">&mdash;</span>
</div>
</div>
<div class="collapse navbar-collapse" id="navbar-menu">
<div class="d-flex flex-column flex-md-row flex-fill align-items-stretch align-items-md-center">
<ul class="navbar-nav">
<li class="nav-item"><a class="nav-link" href="/admin/">Dashboard</a></li>
<li class="nav-item"><a class="nav-link" href="/admin/models">Models</a></li>
<li class="nav-item active"><a class="nav-link" href="/admin/decisions">Decisions</a></li>
<li class="nav-item"><a class="nav-link" href="/admin/controls">Controls</a></li>
</ul>
</div>
</div>
</div>
</header>
<!-- ═══ Page shell ═══ -->
<div class="page">
<div class="page-wrapper">
<div class="page-header d-print-none">
<div class="container-xl">
<div class="row g-2 align-items-center">
<div class="col">
<h2 class="page-title">Decisions</h2>
<div class="text-muted mt-1">Full routing decision history &mdash; filter, search, and inspect</div>
</div>
</div>
</div>
</div>
<div class="page-body">
<div class="container-xl">
<div class="row row-cards">
<div class="col-12">
<div class="card">
<div class="card-header flex-wrap gap-2">
<div class="dec-filters">
<select class="form-select form-select-sm" id="f-kind">
<option value="">All kinds</option>
</select>
<select class="form-select form-select-sm" id="f-category">
<option value="">All categories</option>
</select>
<select class="form-select form-select-sm" id="f-tier">
<option value="">All tiers</option>
<option value="1">Tier 1</option>
<option value="2">Tier 2</option>
<option value="3">Tier 3</option>
</select>
<input type="text" class="form-control form-control-sm" id="f-search" placeholder="Search model, session, rejected reason&hellip;">
<select class="form-select form-select-sm" id="f-page-size" style="min-width:auto">
<option value="50">50 / page</option>
<option value="100" selected>100 / page</option>
<option value="250">250 / page</option>
</select>
<button type="button" class="btn btn-outline-secondary btn-sm" id="f-clear">Clear</button>
</div>
<span class="dec-count ms-auto" id="dec-count"></span>
</div>
<div class="table-responsive">
<table class="table table-vcenter table-hover card-table">
<thead>
<tr>
<th>Time</th><th>Kind</th><th>Category</th><th>Tier</th>
<th>Source</th><th>Model</th><th>Flags</th><th>Cost</th><th>Prof</th><th>Rejected</th>
</tr>
</thead>
<tbody id="dec-tbody">
<tr><td colspan="10" class="text-muted">Loading&hellip;</td></tr>
</tbody>
</table>
</div>
<div class="card-footer d-flex align-items-center justify-content-between">
<span class="text-muted small" id="dec-page-info"></span>
<div class="btn-group btn-group-sm">
<button type="button" class="btn btn-outline-secondary" id="dec-prev">&larr; Prev</button>
<button type="button" class="btn btn-outline-secondary" id="dec-next">Next &rarr;</button>
</div>
</div>
</div>
</div>
</div><!-- /row row-cards -->
</div><!-- /container-xl -->
</div><!-- /page-body -->
<footer class="footer footer-transparent d-print-none">
<div class="container-xl">
&copy; 2026 adLee &middot; 6krrt, local LLM model router &middot; admin
</div>
</footer>
</div>
</div>
<script>
/* ═══════════════════════════════════════════════════
admin/frontend/decisions.html — standalone decision log
API base: relative (works under /admin/)
═══════════════════════════════════════════════════ */
const API = ''; // relative to /admin/
const SSE_URL = '/events/decisions'; // root-level endpoint
const MAX_ROWS = 2000;
let _allDecisions = []; // newest first
let _knownKinds = new Set();
let _knownCategories = new Set();
let _pageSize = 100;
let _currentPage = 1; // 1-indexed
/* ── Fetch wrapper ── */
async function apiFetch(url, opts = {}) {
try {
const resp = await fetch(url, opts);
if (!resp.ok) throw new Error(`${resp.status} ${resp.statusText}`);
return await resp.json();
} catch (e) {
console.warn('API call failed:', url, e);
return null;
}
}
function escapeHtml(s) {
return String(s).replace(/&/g,'&amp;').replace(/</g,'&lt;').replace(/>/g,'&gt;').replace(/"/g,'&quot;');
}
function formatTs(ts) {
if (!ts) return '&mdash;';
try {
const d = new Date(ts);
return d.toLocaleString([], { month: 'short', day: 'numeric', hour: '2-digit', minute: '2-digit', second: '2-digit' });
} catch { return String(ts); }
}
function badgeForKind(kind) {
return {
route: 'bg-info', dispatch: 'bg-success', chat: 'bg-purple',
'local-vision': 'bg-cyan', passthrough: 'bg-info'
}[kind] || 'bg-secondary';
}
// Category has no small fixed set like kind does, so its badge color is a
// deterministic hash-of-the-string hue instead of a lookup table — same
// category always lands on the same color without a color map to maintain
// as new task_category values show up.
function hashCode(s) {
let h = 0;
for (let i = 0; i < s.length; i++) { h = (h << 5) - h + s.charCodeAt(i); h |= 0; }
return h;
}
function categoryStyle(label) {
const hue = Math.abs(hashCode(String(label).toLowerCase())) % 360;
return `background:hsla(${hue},70%,55%,.20);color:hsl(${hue},85%,74%)`;
}
const SOURCE_STYLES = {
classifier: 'background:rgba(59,130,246,.18);color:#93c5fd',
cached: 'background:rgba(34,197,94,.18);color:#4ade80',
fallback: 'background:rgba(239,68,68,.18);color:#f87171',
};
function sourceStyle(source) {
return SOURCE_STYLES[source] || 'background:rgba(148,163,184,.18);color:#cbd5e1';
}
function flagChips(d) {
// Only the flags that are actually ON — showing all four dimmed on every
// row for hundreds of rows was noise, not a cue. Nothing on -> empty cell.
const flags = [
['T', 'tools', d.tools],
['I', 'images', d.images],
['J', 'json mode', d.json_mode],
['S', 'streamed', d.streamed],
];
return flags.filter(([, , on]) => on)
.map(([letter, label]) => `<span class="flag-chip on" title="${escapeHtml(label)}">${letter}</span>`)
.join('');
}
/* ═══════════════════════════════════════
SSE LIVE STREAM
═══════════════════════════════════════ */
let sseConn = null;
function connectSSE() {
try {
if (sseConn) { sseConn.close(); }
sseConn = new EventSource(SSE_URL);
const statusDot = document.getElementById('sse-status');
const statusText = document.getElementById('sse-text');
statusDot.className = 'badge bg-warning';
statusDot.textContent = 'connecting…';
statusText.textContent = 'connecting…';
sseConn.onopen = () => {
statusDot.className = 'badge bg-success';
statusDot.textContent = 'live';
statusText.textContent = 'live';
};
sseConn.onmessage = (ev) => {
try {
const data = JSON.parse(ev.data);
addDecision(data);
} catch (_) { /* heartbeat, ignore */ }
};
sseConn.onerror = () => {
statusDot.className = 'badge bg-danger';
statusDot.textContent = 'reconnecting…';
statusText.textContent = 'reconnecting…';
};
} catch (e) {
document.getElementById('sse-status').className = 'badge bg-danger';
document.getElementById('sse-status').textContent = 'offline';
document.getElementById('sse-text').textContent = 'offline';
}
}
function addDecision(dec) {
_allDecisions.unshift(dec);
if (_allDecisions.length > MAX_ROWS) _allDecisions.length = MAX_ROWS;
refreshFilterOptions([dec]);
renderTable();
}
/* ═══════════════════════════════════════
FILTERS
═══════════════════════════════════════ */
function refreshFilterOptions(rows) {
const kindSel = document.getElementById('f-kind');
const catSel = document.getElementById('f-category');
let changed = false;
for (const d of rows) {
if (d.kind && !_knownKinds.has(d.kind)) { _knownKinds.add(d.kind); changed = true; }
if (d.task_category && !_knownCategories.has(d.task_category)) { _knownCategories.add(d.task_category); changed = true; }
}
if (!changed) return;
const kindVal = kindSel.value, catVal = catSel.value;
kindSel.innerHTML = '<option value="">All kinds</option>' +
[...(_knownKinds)].sort().map(k => `<option value="${escapeHtml(k)}">${escapeHtml(k)}</option>`).join('');
catSel.innerHTML = '<option value="">All categories</option>' +
[...(_knownCategories)].sort().map(c => `<option value="${escapeHtml(c)}">${escapeHtml(c)}</option>`).join('');
kindSel.value = kindVal;
catSel.value = catVal;
}
function currentFilters() {
return {
kind: document.getElementById('f-kind').value,
category: document.getElementById('f-category').value,
tier: document.getElementById('f-tier').value,
search: document.getElementById('f-search').value.trim().toLowerCase(),
};
}
function matchesFilters(d, f) {
if (f.kind && d.kind !== f.kind) return false;
if (f.category && d.task_category !== f.category) return false;
if (f.tier && String(d.task_tier || '') !== f.tier) return false;
if (f.search) {
const haystack = [
d.selected_model, d.selected_provider, d.session_key,
d.rejected_reason, d.task_category, d.kind,
].filter(Boolean).join(' ').toLowerCase();
if (!haystack.includes(f.search)) return false;
}
return true;
}
function resetPageAndRender() {
_currentPage = 1;
renderTable();
}
function setupFilters() {
const ids = ['f-kind', 'f-category', 'f-tier'];
ids.forEach(id => document.getElementById(id).addEventListener('change', resetPageAndRender));
document.getElementById('f-search').addEventListener('input', resetPageAndRender);
document.getElementById('f-clear').addEventListener('click', () => {
document.getElementById('f-kind').value = '';
document.getElementById('f-category').value = '';
document.getElementById('f-tier').value = '';
document.getElementById('f-search').value = '';
resetPageAndRender();
});
}
function setupPagination() {
document.getElementById('f-page-size').addEventListener('change', (e) => {
_pageSize = Number(e.target.value) || 100;
resetPageAndRender();
});
document.getElementById('dec-prev').addEventListener('click', () => {
if (_currentPage > 1) { _currentPage--; renderTable(); }
});
document.getElementById('dec-next').addEventListener('click', () => {
_currentPage++; // clamped inside renderTable
renderTable();
});
}
/* ═══════════════════════════════════════
TABLE
═══════════════════════════════════════ */
function renderTable() {
const tbody = document.getElementById('dec-tbody');
const countEl = document.getElementById('dec-count');
const pageInfoEl = document.getElementById('dec-page-info');
const prevBtn = document.getElementById('dec-prev');
const nextBtn = document.getElementById('dec-next');
const f = currentFilters();
const rows = _allDecisions.filter(d => matchesFilters(d, f));
countEl.textContent = `${rows.length} of ${_allDecisions.length}`;
if (!rows.length) {
tbody.innerHTML = `<tr><td colspan="10" class="text-muted text-center">No decisions match these filters</td></tr>`;
pageInfoEl.textContent = '';
prevBtn.disabled = true;
nextBtn.disabled = true;
return;
}
const totalPages = Math.max(1, Math.ceil(rows.length / _pageSize));
_currentPage = Math.min(Math.max(1, _currentPage), totalPages);
const start = (_currentPage - 1) * _pageSize;
const pageRows = rows.slice(start, start + _pageSize);
pageInfoEl.textContent = `Page ${_currentPage} of ${totalPages} (rows ${start + 1}–${start + pageRows.length})`;
prevBtn.disabled = _currentPage <= 1;
nextBtn.disabled = _currentPage >= totalPages;
tbody.innerHTML = pageRows.map(d => {
const badgeClass = badgeForKind(d.kind);
const modelStr = `${d.selected_model || '&mdash;'} / ${d.selected_provider || ''}`;
const category = d.task_category || '—';
const tier = d.task_tier != null ? String(d.task_tier) : '?';
const source = d.classification_source || '—';
return `<tr>
<td class="text-nowrap">${formatTs(d.observed_at || '')}</td>
<td><span class="badge ${badgeClass}">${escapeHtml(String(d.kind || '?'))}</span></td>
<td title="${escapeHtml(d.task_category || '')}"><span class="dec-badge" style="${categoryStyle(category)}">${escapeHtml(category)}</span></td>
<td class="tier-${tier}">${escapeHtml(tier)}</td>
<td><span class="dec-badge" style="${sourceStyle(source)}">${escapeHtml(source)}</span></td>
<td title="${escapeHtml(modelStr)}">${escapeHtml(modelStr)}</td>
<td class="text-nowrap">${flagChips(d)}</td>
<td>${d.est_cost_usd != null ? '$' + Number(d.est_cost_usd).toFixed(4) : '—'}</td>
<td>${d.est_proficiency != null ? Number(d.est_proficiency).toFixed(2) : '—'}</td>
<td class="text-muted small">${escapeHtml(d.rejected_reason || '')}</td>
</tr>`;
}).join('');
}
async function loadDecisions() {
const data = await apiFetch(`${API}api/decisions?limit=1000`);
if (data) {
_allDecisions = data;
refreshFilterOptions(data);
renderTable();
}
}
/* ═══════════════════════════════════════
INIT
═══════════════════════════════════════ */
function init() {
setupFilters();
setupPagination();
loadDecisions();
connectSSE();
}
init();
</script>
</body>
</html>

File diff suppressed because it is too large Load Diff

374
admin/frontend/models.html Normal file
View File

@@ -0,0 +1,374 @@
<!DOCTYPE html>
<html lang="en" data-bs-theme="dark">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Models &middot; LLM Router Admin</title>
<script src="https://cdn.jsdelivr.net/npm/chart.js@4.4.7/dist/chart.umd.min.js"></script>
<script src="https://cdn.jsdelivr.net/npm/chartjs-adapter-date-fns@3.0.0/dist/chartjs-adapter-date-fns.bundle.min.js"></script>
<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/@tabler/core@1.4.0/dist/css/tabler.min.css">
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=Quicksand:wght@600;700&display=swap">
<link rel="icon" href="data:image/svg+xml,%3Csvg%20xmlns%3D%22http%3A//www.w3.org/2000/svg%22%20viewBox%3D%220%200%20256%20256%22%3E%3Ccircle%20cx%3D%22146%22%20cy%3D%2284%22%20r%3D%2245.6%22%20fill%3D%22%23fff%22/%3E%3Cpath%20d%3D%22M124%20110%20Q90%20140%2090%20176%20Q90%20212%20122%20212%20Q156%20212%20172%20186%22%20fill%3D%22none%22%20stroke%3D%22%23fff%22%20stroke-width%3D%2239.2%22%20stroke-linecap%3D%22round%22/%3E%3Ccircle%20cx%3D%22146%22%20cy%3D%2284%22%20r%3D%2242%22%20fill%3D%22%23f59e0b%22/%3E%3Cpath%20d%3D%22M124%20110%20Q90%20140%2090%20176%20Q90%20212%20122%20212%20Q156%20212%20172%20186%22%20fill%3D%22none%22%20stroke%3D%22%23f59e0b%22%20stroke-width%3D%2232%22%20stroke-linecap%3D%22round%22/%3E%3Ccircle%20cx%3D%22128%22%20cy%3D%2278%22%20r%3D%229%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22164%22%20cy%3D%2278%22%20r%3D%229%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22128%22%20cy%3D%2278%22%20r%3D%226%22%20fill%3D%22%231e293b%22/%3E%3Ccircle%20cx%3D%22164%22%20cy%3D%2278%22%20r%3D%226%22%20fill%3D%22%231e293b%22/%3E%3Ccircle%20cx%3D%22131%22%20cy%3D%2274%22%20r%3D%222.2%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22167%22%20cy%3D%2274%22%20r%3D%222.2%22%20fill%3D%22%23fff%22/%3E%3Cpath%20d%3D%22M134%20100%20Q146%20112%20158%20100%22%20fill%3D%22none%22%20stroke%3D%22%23fff%22%20stroke-width%3D%227%22%20stroke-linecap%3D%22round%22/%3E%3Cpath%20d%3D%22M134%20100%20Q146%20112%20158%20100%22%20fill%3D%22none%22%20stroke%3D%22%231e293b%22%20stroke-width%3D%223.5%22%20stroke-linecap%3D%22round%22/%3E%3Cline%20x1%3D%22150.47%22%20y1%3D%2247.59%22%20x2%3D%22156.01%22%20y2%3D%2227.13%22%20stroke%3D%22%23fff%22%20stroke-width%3D%2210.73%22%20stroke-linecap%3D%22round%22/%3E%3Cline%20x1%3D%22150.47%22%20y1%3D%2247.59%22%20x2%3D%22155.08%22%20y2%3D%2230.54%22%20stroke%3D%22%231e293b%22%20stroke-width%3D%226.26%22%20stroke-linecap%3D%22round%22/%3E%3Ccircle%20cx%3D%2295%22%20cy%3D%22226%22%20r%3D%2214%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%22150%22%20cy%3D%22226%22%20r%3D%2214%22%20fill%3D%22%23fff%22/%3E%3Ccircle%20cx%3D%2295%22%20cy%3D%22226%22%20r%3D%229%22%20fill%3D%22%231e293b%22/%3E%3Ccircle%20cx%3D%22150%22%20cy%3D%22226%22%20r%3D%229%22%20fill%3D%22%231e293b%22/%3E%3C/svg%3E">
<script src="https://cdn.jsdelivr.net/npm/@tabler/core@1.4.0/dist/js/tabler.min.js"></script>
<style>
/* ── Liquid glass: soft gradient wash behind everything so the blur has
something to catch, then translucent+blurred navbar/cards on top. Kept
subtle — this is a working dashboard, not a marketing page. ── */
body{
background:
radial-gradient(900px circle at 6% -8%, rgba(245,158,11,.10), transparent 55%),
radial-gradient(800px circle at 94% -6%, rgba(59,130,246,.10), transparent 50%),
radial-gradient(760px circle at 50% 108%, rgba(139,92,246,.07), transparent 55%),
#111827;
background-attachment: fixed;
}
header.navbar{
background: rgba(31,41,55,.6)!important;
backdrop-filter: blur(16px) saturate(160%);
-webkit-backdrop-filter: blur(16px) saturate(160%);
border-bottom: 1px solid rgba(255,255,255,.06);
}
.card{
background: rgba(31,41,55,.55);
backdrop-filter: blur(18px) saturate(140%);
-webkit-backdrop-filter: blur(18px) saturate(140%);
border: 1px solid rgba(255,255,255,.07);
box-shadow: 0 10px 30px -14px rgba(0,0,0,.55), inset 0 1px 0 rgba(255,255,255,.04);
}
.card-header{border-bottom-color: rgba(255,255,255,.06)}
/* ── The "gap only at wide windows" bug, root-caused: this Chrome/Linux
build reports a phantom left margin on <html> exactly equal to the
scrollbar's width whenever <html> itself is the scrolling element.
Moving the scroll context down to <body> sidesteps it. ── */
html{overflow:hidden;height:100%;margin:0!important;padding:0!important}
body{overflow-y:auto;height:100%;margin:0!important;padding:0!important}
/* ── Buttons: glass pills instead of flat saturated fills. ── */
.btn-success{background:rgba(34,197,94,.16)!important;border-color:rgba(74,222,128,.45)!important;color:#4ade80!important}
.btn-success:hover{background:rgba(34,197,94,.28)!important;border-color:rgba(74,222,128,.7)!important;color:#86efac!important}
.btn-danger{background:rgba(239,68,68,.16)!important;border-color:rgba(248,113,113,.45)!important;color:#f87171!important}
.btn-danger:hover{background:rgba(239,68,68,.28)!important;border-color:rgba(248,113,113,.7)!important;color:#fca5a5!important}
.btn{backdrop-filter:blur(6px);-webkit-backdrop-filter:blur(6px)}
/* ── Badges: same glass treatment. ── */
.badge.bg-success{background:rgba(34,197,94,.20)!important;color:#4ade80!important}
.badge.bg-warning{background:rgba(245,158,11,.20)!important;color:#fbbf24!important}
.badge.bg-danger{background:rgba(239,68,68,.20)!important;color:#f87171!important}
.badge.bg-info{background:rgba(59,130,246,.20)!important;color:#93c5fd!important}
.badge.bg-secondary{background:rgba(148,163,184,.20)!important;color:#e2e8f0!important}
.badge.bg-purple{background:rgba(168,85,247,.20)!important;color:#d8b4fe!important}
/* ── Progress bars: translucent track + glow instead of flat solid. ── */
.progress{background:rgba(255,255,255,.08)!important;border-radius:999px;overflow:hidden}
.progress-bar{box-shadow:0 0 6px 0 currentColor;filter:saturate(1.25)}
header.navbar>.container-fluid{padding-left:20px!important;padding-right:20px!important}
.navbar-brand, .page-title{font-family:'Quicksand',var(--tblr-font-sans-serif,ui-sans-serif,system-ui,sans-serif);font-weight:700;letter-spacing:.01em}
.navbar-brand a{text-decoration:none!important;color:inherit}
/* ── Brand: Tabler's own rule is [data-bs-theme=dark] .navbar-brand-autodark
.navbar-brand-image{filter:brightness(0) invert(1)} — attribute selector +
2 classes outranks a plain 2-class override, so `filter:none` alone loses
the cascade regardless of source order. !important is what actually beats
it. That filter was stripping the mascot's orange color, and the 2rem
.navbar-brand-image cap shrank his face to sub-pixel dots.
Restore color + force an explicit size (64px, below). ── */
.navbar-brand-autodark .navbar-brand-image{filter:none!important;height:64px;width:64px}
/* ── 64px mascot, same bar height. His face is only ~40% of his height (the
rest is tail and wheels), so at 45px the eyes were ~4px and the whole point
of having a mascot was lost. The 19px comes out of padding, not out of the
navbar: the brand h1 carries 8px top and bottom, and the navbar itself 4px.
Zero the first and halve the second and the bar lands at 69px — 1px shorter
than the 45px version it replaces. Measure before changing either number. ── */
.navbar-brand{padding-top:0!important;padding-bottom:0!important}
header.navbar{padding-top:2px!important;padding-bottom:2px!important}
/* ── Speed lines: he is a wheeled rover, so give him a trail. Drawn as their
own inline SVG rather than baked into the logo, so assets/*.svg stays the
plain mascot and the lines can drop out on narrow screens (d-none d-md-flex).
The gradient fades AWAY from him (transparent at the far left, amber where
he is), which is what reads as speed rather than as three floating dashes;
it is userSpaceOnUse because an objectBoundingBox gradient on a horizontal
line has a zero-height bbox and is undefined. ── */
.navbar-brand a{display:inline-flex;align-items:center;gap:.15rem}
/* The mascot IS a 6, so the wordmark's leading 6 echoes him in the brand amber.
The whole wordmark is wrapped in .brand-word first: the anchor is inline-flex
with a gap, so a bare <span> around the 6 would make it its own flex item and
open that gap between "6" and "krrt". */
.brand-six{color:#f59e0b;font-size:1.14em;line-height:1}
/* -20px: the logo's own viewBox carries ~18px of empty space to the left of
his tail at 64px, so the pull has to cover that before it buys any real
closeness. Measured gap from the middle line's tip to his back: ~4.5px. */
.brand-speedlines{align-items:center;line-height:0;margin-right:-20px}
.brand-speedlines svg{filter:none!important}
/* Motion only on hover — a dashboard that twitches at rest is a dashboard you
stop looking at. */
@keyframes brand-dash{
0%{transform:translateX(-5px);opacity:.2}
55%{opacity:1}
100%{transform:translateX(4px);opacity:0}
}
.navbar-brand a:hover .brand-speedlines line{animation:brand-dash .75s ease-in infinite}
.navbar-brand a:hover .brand-speedlines line:nth-child(2){animation-delay:.12s}
.navbar-brand a:hover .brand-speedlines line:nth-child(3){animation-delay:.24s}
@media (prefers-reduced-motion: reduce){
.navbar-brand a:hover .brand-speedlines line{animation:none}
}
/* ── Tier colors on the availability table ── */
.tier-1{color:var(--tblr-success)}
.tier-2{color:var(--tblr-warning)}
.tier-3{color:var(--tblr-danger)}
.tier-?{color:var(--tblr-secondary)}
</style>
</head>
<body>
<!-- ═══ Navbar ═══ -->
<header class="navbar navbar-expand-md navbar-dark" data-bs-theme="dark">
<div class="container-fluid">
<button class="navbar-toggler" type="button" data-bs-toggle="collapse" data-bs-target="#navbar-menu" aria-controls="navbar-menu" aria-expanded="false" aria-label="Toggle navigation"><span class="navbar-toggler-icon"></span></button>
<h1 class="navbar-brand navbar-brand-autodark d-none-navbar-horizontal pe-0 pe-md-3">
<a href="/admin/" style="display:inline-flex;align-items:center;gap:8px">
<span class="brand-speedlines d-none d-md-flex" aria-hidden="true">
<svg viewBox="0 0 40 64" width="40" height="64" xmlns="http://www.w3.org/2000/svg">
<defs>
<linearGradient id="speedfade" gradientUnits="userSpaceOnUse" x1="0" y1="0" x2="40" y2="0">
<stop offset="0" stop-color="#f59e0b" stop-opacity="0"/>
<stop offset="1" stop-color="#f59e0b" stop-opacity=".85"/>
</linearGradient>
</defs>
<g stroke="url(#speedfade)" stroke-linecap="round" fill="none">
<line x1="14" y1="20" x2="38" y2="20" stroke-width="5"/>
<line x1="2" y1="33" x2="36" y2="33" stroke-width="6"/>
<line x1="18" y1="46" x2="34" y2="46" stroke-width="4.5"/>
</g>
</svg>
</span>
<svg class="navbar-brand-image" viewBox="0 0 256 256" width="64" height="64" style="width:64px;height:64px" role="img" aria-hidden="true" xmlns="http://www.w3.org/2000/svg">
<circle cx="146" cy="84" r="45.6" fill="#fff"/>
<path d="M124 110 Q 90 140 90 176 Q 90 212 122 212 Q 156 212 172 186" fill="none" stroke="#fff" stroke-width="39.2" stroke-linecap="round"/>
<circle cx="146" cy="84" r="42" fill="#f59e0b"/>
<path d="M124 110 Q 90 140 90 176 Q 90 212 122 212 Q 156 212 172 186" fill="none" stroke="#f59e0b" stroke-width="32" stroke-linecap="round"/>
<circle cx="128" cy="78" r="9" fill="#fff"/><circle cx="164" cy="78" r="9" fill="#fff"/>
<circle cx="128" cy="78" r="6" fill="#1e293b"/><circle cx="164" cy="78" r="6" fill="#1e293b"/>
<circle cx="131" cy="74" r="2.2" fill="#fff"/><circle cx="167" cy="74" r="2.2" fill="#fff"/>
<path d="M134 100 Q146 112 158 100" fill="none" stroke="#fff" stroke-width="7" stroke-linecap="round"/>
<path d="M134 100 Q146 112 158 100" fill="none" stroke="#1e293b" stroke-width="3.5" stroke-linecap="round"/>
<line x1="150.47" y1="47.59" x2="156.01" y2="27.13" stroke="#ffffff" stroke-width="10.73" stroke-linecap="round"/>
<line x1="150.47" y1="47.59" x2="155.08" y2="30.54" stroke="#1e293b" stroke-width="6.26" stroke-linecap="round"/>
<ellipse cx="157.68" cy="-18.58" fill="#fff" transform="matrix(0.96737,0.25337,-0.26099,0.96534,0,0)" rx="9.05" ry="8.83"/>
<ellipse cx="157.68" cy="-18.58" fill="#1e293b" transform="matrix(0.96737,0.25337,-0.26099,0.96534,0,0)" rx="5.79" ry="5.65"/>
<circle cx="95" cy="226" r="14" fill="#fff"/><circle cx="150" cy="226" r="14" fill="#fff"/>
<circle cx="95" cy="226" r="9" fill="#1e293b"/><circle cx="150" cy="226" r="9" fill="#1e293b"/>
</svg>
<span class="brand-word"><span class="brand-six">6</span>krrt LLM Router</span>
</a>
</h1>
<div class="navbar-nav flex-row order-md-last">
<div class="d-flex align-items-center gap-3">
<span id="sse-status" class="badge bg-warning">connecting&hellip;</span>
<span id="sse-text" class="visually-hidden"></span>
<span id="generated-at" class="text-muted small">&mdash;</span>
</div>
</div>
<div class="collapse navbar-collapse" id="navbar-menu">
<div class="d-flex flex-column flex-md-row flex-fill align-items-stretch align-items-md-center">
<ul class="navbar-nav">
<li class="nav-item"><a class="nav-link" href="/admin/">Dashboard</a></li>
<li class="nav-item active"><a class="nav-link" href="/admin/models">Models</a></li>
<li class="nav-item"><a class="nav-link" href="/admin/decisions">Decisions</a></li>
<li class="nav-item"><a class="nav-link" href="/admin/controls">Controls</a></li>
</ul>
</div>
</div>
</div>
</header>
<!-- ═══ Page shell ═══ -->
<div class="page">
<div class="page-wrapper">
<div class="page-header d-print-none">
<div class="container-xl">
<div class="row g-2 align-items-center">
<div class="col">
<h2 class="page-title">Models</h2>
<div class="text-muted mt-1">Available models, their tier, and live availability overrides</div>
</div>
</div>
</div>
</div>
<div class="page-body">
<div class="container-xl">
<div class="row row-cards">
<div class="col-12">
<div class="card">
<div class="card-header">
<div class="d-flex align-items-center gap-2">
<svg xmlns="http://www.w3.org/2000/svg" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><ellipse cx="12" cy="5" rx="9" ry="3"/><path d="M21 12c0 1.66-4 3-9 3s-9-1.34-9-3"/><path d="M3 5v14c0 1.66 4 3 9 3s9-1.34 9-3V5"/></svg>
<h3 class="card-title mb-0">Model Availability</h3>
</div>
<div class="card-subtitle text-muted ms-2">Change the override dropdown to mark a model active, deprecated, or stale</div>
</div>
<div class="table-responsive">
<table class="table table-vcenter table-hover card-table">
<thead>
<tr>
<th>Model</th>
<th>Provider</th>
<th>Tier</th>
<th>Status</th>
<th>Override</th>
</tr>
</thead>
<tbody id="model-tbody">
<tr><td colspan="5" class="text-muted">Loading&hellip;</td></tr>
</tbody>
</table>
</div>
</div>
</div>
</div><!-- /row row-cards -->
</div><!-- /container-xl -->
</div><!-- /page-body -->
<footer class="footer footer-transparent d-print-none">
<div class="container-xl">
&copy; 2026 adLee &middot; 6krrt, local LLM model router &middot; admin
</div>
</footer>
</div>
</div>
<!-- Toast (Bootstrap native) -->
<div id="toast" class="toast align-items-center text-bg-info border-0 position-fixed bottom-0 end-0 p-3" role="alert" aria-live="assertive" aria-atomic="true"></div>
<script>
/* ═══════════════════════════════════════════════════
admin/frontend/models.html — model availability page
API base: relative (works under /admin/)
═══════════════════════════════════════════════════ */
const API = ''; // relative to /admin/
const SSE_URL = '/events/decisions'; // root-level endpoint
/* ── Toast (Bootstrap native via tabler wrapper) ── */
let _toastEl = null;
function toast(msg, type = 'info') {
const bgClass = { success: 'text-bg-success', error: 'text-bg-danger', info: 'text-bg-info' }[type] || 'text-bg-info';
if (!_toastEl) {
_toastEl = document.getElementById('toast');
_toastEl.innerHTML = `<div class="d-flex">
<div class="toast-body"></div>
<button type="button" class="btn-close btn-close-white me-2 m-auto" data-bs-dismiss="toast" aria-label="Close"></button>
</div>`;
_toastEl._bt = new tabler.Toast(_toastEl, { delay: 4000 });
}
_toastEl.querySelector('.toast-body').textContent = msg;
_toastEl.className = `toast align-items-center ${bgClass} border-0 position-fixed bottom-0 end-0 p-3`;
_toastEl._bt.show();
}
/* ── Fetch Wrapper ── */
async function apiFetch(url, opts = {}) {
try {
const resp = await fetch(url, opts);
if (!resp.ok) throw new Error(`${resp.status} ${resp.statusText}`);
return await resp.json();
} catch (e) {
console.warn('API call failed:', url, e);
return null;
}
}
/* ═══════════════════════════════════════
SSE LIVE STREAM
═══════════════════════════════════════ */
let sseConn = null;
function connectSSE() {
try {
if (sseConn) { sseConn.close(); }
sseConn = new EventSource(SSE_URL);
const statusDot = document.getElementById('sse-status');
const statusText = document.getElementById('sse-text');
statusDot.className = 'badge bg-warning';
statusDot.textContent = 'connecting…';
statusText.textContent = 'connecting…';
sseConn.onopen = () => {
statusDot.className = 'badge bg-success';
statusDot.textContent = 'live';
statusText.textContent = 'live';
};
sseConn.onerror = () => {
statusDot.className = 'badge bg-danger';
statusDot.textContent = 'reconnecting…';
statusText.textContent = 'reconnecting…';
};
} catch (e) {
document.getElementById('sse-status').className = 'badge bg-danger';
document.getElementById('sse-status').textContent = 'offline';
document.getElementById('sse-text').textContent = 'offline';
}
}
/* ═══════════════════════════════════════
MODEL AVAILABILITY TABLE
═══════════════════════════════════════ */
function renderModels(models) {
const tbody = document.getElementById('model-tbody');
if (!models || !models.length) {
tbody.innerHTML = '<tr><td colspan="5" class="text-muted text-center">No models</td></tr>';
return;
}
const html = models.filter(m => (m.access_level === 'public') || !m.access_level).map(m => {
const availClass = { active: 'var(--tblr-success)', deprecated: 'var(--tblr-danger)', stale: 'var(--tblr-warning)' }[m.effective_availability] || 'var(--tblr-secondary)';
const selectOpts = ['active','deprecated','stale'].map(a =>
`<option value="${a}"${m.effective_availability===a?' selected':''}>${a}</option>`
).join('');
const tierClass = `tier-${m.tier || '?'}`;
return `<tr data-model="${m.model_id}" data-provider="${m.provider}">
<td>${escapeHtml(m.model_id)}</td>
<td>${escapeHtml(m.provider)}</td>
<td class="${tierClass}">${m.tier || '?'}</td>
<td style="color:${availClass}"><span class="badge bg-secondary-subtle text-secondary">${escapeHtml(String(m.effective_availability))}</span></td>
<td><select class="form-select form-select-sm" style="width:130px" onchange="updateAvailability('${escapeHtml(m.model_id)}','${escapeHtml(m.provider)}',this.value)"${m.is_overridden?' disabled':''}>${selectOpts}</select></td>
</tr>`;
}).join('');
tbody.innerHTML = html;
}
async function updateAvailability(modelId, provider, avail) {
// Don't post if already that value via dropdown
const resp = await apiFetch(
`${API}api/models/${encodeURIComponent(modelId)}/${encodeURIComponent(provider)}/availability`,
{ method:'POST', headers:{'Content-Type':'application/json'}, body:JSON.stringify({ availability: avail }) }
);
if (resp) toast('Availability updated', 'success');
else toast('Update failed', 'error');
}
async function loadModels() {
const models = await apiFetch(`${API}api/models`);
if (models) renderModels(models);
}
/* ═══════════════════════════════════════
UTILITIES
═══════════════════════════════════════ */
function escapeHtml(s) {
return String(s).replace(/&/g,'&amp;').replace(/</g,'&lt;').replace(/>/g,'&gt;').replace(/"/g,'&quot;');
}
/* ═══════════════════════════════════════
INIT
═══════════════════════════════════════ */
function init() {
loadModels();
connectSSE();
}
init();
</script>
</body>
</html>

View File

@@ -0,0 +1,18 @@
<svg class="navbar-brand-image" viewBox="0 0 256 256" width="64" height="64" style="width:64px;height:64px" role="img" aria-hidden="true" xmlns="http://www.w3.org/2000/svg">
<circle cx="146" cy="84" r="45.6" fill="#fff"></circle>
<path d="M124 110 Q 90 140 90 176 Q 90 212 122 212 Q 156 212 172 186" fill="none" stroke="#fff" stroke-width="39.2" stroke-linecap="round"></path>
<circle cx="146" cy="84" r="42" fill="#f59e0b"></circle>
<path d="M124 110 Q 90 140 90 176 Q 90 212 122 212 Q 156 212 172 186" fill="none" stroke="#f59e0b" stroke-width="32" stroke-linecap="round"></path>
<circle cx="128" cy="78" r="9" fill="#fff"></circle><circle cx="164" cy="78" r="9" fill="#fff"></circle>
<circle cx="128" cy="78" r="6" fill="#1e293b"></circle><circle cx="164" cy="78" r="6" fill="#1e293b"></circle>
<circle cx="131" cy="74" r="2.2" fill="#fff"></circle><circle cx="167" cy="74" r="2.2" fill="#fff"></circle>
<path d="M134 100 Q146 112 158 100" fill="none" stroke="#fff" stroke-width="7" stroke-linecap="round"></path>
<path d="M134 100 Q146 112 158 100" fill="none" stroke="#1e293b" stroke-width="3.5" stroke-linecap="round"></path>
<line x1="150.47" y1="47.59" x2="156.01" y2="27.13" stroke="#ffffff" stroke-width="10.73" stroke-linecap="round"></line>
<line x1="150.47" y1="47.59" x2="155.08" y2="30.54" stroke="#1e293b" stroke-width="6.26" stroke-linecap="round"></line>
<ellipse cx="157.68" cy="-18.58" fill="#fff" transform="matrix(0.96737,0.25337,-0.26099,0.96534,0,0)" rx="9.05" ry="8.83"></ellipse>
<ellipse cx="157.68" cy="-18.58" fill="#1e293b" transform="matrix(0.96737,0.25337,-0.26099,0.96534,0,0)" rx="5.79" ry="5.65"></ellipse>
<circle cx="95" cy="226" r="14" fill="#fff"></circle><circle cx="150" cy="226" r="14" fill="#fff"></circle>
<circle cx="95" cy="226" r="9" fill="#1e293b"></circle><circle cx="150" cy="226" r="9" fill="#1e293b"></circle>
</svg>

After

Width:  |  Height:  |  Size: 2.0 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 604 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 642 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 839 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 428 KiB

View File

@@ -18,7 +18,7 @@ objective:
# currently rest on 2-3 samples per category, so a 0.05 gap is
# indistinguishable from sampling variation and paying for it buys noise.
# Narrow it as samples accumulate.
quality_tolerance: 0.10
quality_tolerance: 0.1
# Cost is priced per-request from catalog prices, NOT from a benchmark
# sweep. A fixed 400-token reference task ranked glm-5.2-fast 3.2x cheaper
@@ -52,7 +52,7 @@ objective:
# wall you hit mid-task. So this is the cost mandate stated as a guarantee.
# For scale: the reference task runs ~5e-06 kWh on the cheapest model and
# ~2.2e-04 on the most expensive.
max_energy_per_request: null
max_energy_per_request:
# The subscription's kWh allowance per billing period, for reporting burn in
# /health. Set to match your plan; null disables the report. NeuralWatt also
@@ -203,7 +203,7 @@ session_cache:
# Off by default, matching every other new-and-unproven knob in this
# project: ship it, watch route_decisions.source="cached" on real traffic,
# then decide the right default.
enabled: false
enabled: true
staleness_minutes: 20
circuit_breaker:
@@ -274,7 +274,7 @@ routing:
# off, let POST /outcome report real pass/fail, and compare
# tool_use_agentic proficiency for deepseek before and after. That is the
# one signal here that knows whether the work actually worked.
min_tool_proficiency: null
min_tool_proficiency:
tool_use_category: tool_use_agentic
# Request-side capability gates. These read the request body (image parts,
@@ -299,7 +299,7 @@ local_vision:
# and the model must be pulled (`ollama pull qwen3-vl:4b`) on that host.
enabled: true
base_url: "http://localhost:11434/v1"
api_key_env: null
api_key_env:
# A Modelfile-tagged variant of qwen3-vl:4b, not the base library tag.
# Measured live: the base tag comes up at Ollama's own default num_ctx
# (32768) and costs 9.4GB loaded — resident alongside the classifier's
@@ -398,7 +398,7 @@ classifier:
base_url: "http://localhost:11434/v1"
# Unset means unauthenticated, which is the Ollama case. Name the env var
# holding the key when the endpoint actually checks one.
api_key_env: null
api_key_env:
# A Modelfile-tagged variant of mistral-nemo:12b, not the base library tag
# — must match a model `ollama list` reports. Ollama loads a model at its
# library Modelfile's default context unless told otherwise, and the base
@@ -517,3 +517,4 @@ logging:
# systemctl --user edit llm-router # Environment="LLM_ROUTER_LOG_LEVEL=debug"
# systemctl --user restart llm-router
level: info

View File

@@ -9,9 +9,9 @@ all.
| file | what it does |
|---|---|
| `llm-router.service` | the FastAPI dispatcher, on `127.0.0.1:8080` |
| `llm-router-poller.service` | one-shot: `poller.py` then `tier.py` |
| `llm-router-poller.service` | one-shot: `PYTHONPATH=src python -m poller` then `PYTHONPATH=src python -m tier` |
| `llm-router-poller.timer` | fires the poller 2 min after boot, then every 2 h |
| `llm-router-seed.service` | one-shot: a small `seed_energy.py` reference sweep |
| `llm-router-seed.service` | one-shot: a small `PYTHONPATH=src python -m seed_energy` reference sweep |
| `llm-router-seed.timer` | every 6 h — energy attribution drifts with pool load across hours, so the median has to span time rather than one sweep |
These are **user** units — no root, and they run as you with your own
@@ -37,6 +37,10 @@ done
systemctl --user daemon-reload
systemctl --user enable --now llm-router.service llm-router-poller.timer llm-router-seed.timer
# Note: the units rely on `Environment=PYTHONPATH=%h/llm-router/src` (rewritten
# by the same sed to your REPO) so the modules under `src/` are importable
# without an editable install. .env is read from the repo root and stays there.
# 3. Survive logout/reboot (user units stop with your session otherwise)
loginctl enable-linger "$USER"
@@ -54,17 +58,17 @@ journalctl --user -u 'llm-router*' -f # everything, live (quote the g
journalctl --user -u llm-router -f -o cat # the request log, message only
journalctl --user -u llm-router -p warning # fallbacks, retries, refusals
journalctl --user -u llm-router-poller.service # catalog refreshes
systemctl --user restart llm-router.service # after editing config.yaml
systemctl --user restart llm-router.service # after editing config/config.yaml
systemctl --user start llm-router-poller.service # force a refresh now
```
`config.yaml` is read once at startup, so weight and threshold changes need a
`config/config.yaml` is read once at startup, so weight and threshold changes need a
restart. The catalog is read per-request, so a poller run takes effect
immediately.
### Turning up the logs
`logging.level` in config.yaml is the documented setting, but flipping it means
`logging.level` in config/config.yaml is the documented setting, but flipping it means
editing a tracked file. For a running service use a drop-in instead:
```bash
@@ -91,7 +95,7 @@ priorities when systemd owns its stderr (`SyslogLevelPrefix` is on by default).
A foreground `uvicorn` prints them clean, so the same binary is readable either
way.
**The oneshot units buffer.** `poller.py` and `seed_energy.py` print progress
**The oneshot units buffer.** `python -m poller` and `python -m seed_energy` print progress
with plain `print()`, and Python block-buffers stdout when it is not a
terminal, so their output arrives in one dump at exit rather than
progressively. Add `Environment="PYTHONUNBUFFERED=1"` to those units if you
@@ -123,7 +127,7 @@ sudo nano /etc/systemd/system/ollama.service.d/override.conf
sudo systemctl daemon-reload && sudo systemctl restart ollama
```
On the **client** host, in `config.yaml`:
On the **client** host, in `config/config.yaml`:
```yaml
classifier:

View File

@@ -7,12 +7,13 @@ Wants=network-online.target
[Service]
Type=oneshot
WorkingDirectory=%h/llm-router
Environment=PYTHONPATH=%h/llm-router/src
EnvironmentFile=%h/llm-router/.env
# poller.py refreshes the catalog; tier.py re-resolves tiers from it. Tiers
# poller refreshes the catalog; tier re-resolves tiers from it. Tiers
# are derived from cost and reasoning fields the poll may have changed, so
# they always run as a pair.
ExecStart=%h/llm-router/.venv/bin/python poller.py
ExecStart=%h/llm-router/.venv/bin/python tier.py
# they always run as a pair. PYTHONPATH above resolves src/ for `-m` imports.
ExecStart=%h/llm-router/.venv/bin/python -m poller
ExecStart=%h/llm-router/.venv/bin/python -m tier
NoNewPrivileges=true
PrivateTmp=true

View File

@@ -7,10 +7,12 @@ Wants=network-online.target
[Service]
Type=oneshot
WorkingDirectory=%h/llm-router
Environment=PYTHONPATH=%h/llm-router/src
EnvironmentFile=%h/llm-router/.env
# Fewer samples per run than a manual sweep, because the point is coverage
# across TIME rather than depth at one moment — see the timer.
ExecStart=%h/llm-router/.venv/bin/python seed_energy.py --samples 3
# across TIME rather than depth at one moment — see the timer. PYTHONPATH
# above resolves src/ for the `-m` import.
ExecStart=%h/llm-router/.venv/bin/python -m seed_energy --samples 3
NoNewPrivileges=true
PrivateTmp=true

View File

@@ -10,9 +10,12 @@ Wants=network-online.target
[Service]
Type=exec
# config.yaml, router.db and router.log are all referenced as relative paths,
# so this has to be the repo root.
# router.db and router.log are referenced as relative paths, so this has to be
# the repo root. config.yaml now lives under config/ and the Python modules
# under src/, so PYTHONPATH points at src/ for the dispatcher import to resolve.
WorkingDirectory=%h/llm-router
# Resolves `dispatcher` (and any other src/ module) for the ExecStart below.
Environment=PYTHONPATH=%h/llm-router/src
# Holds NEURALWATT_API_KEY. Create it with:
# echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env
EnvironmentFile=%h/llm-router/.env

View File

Before

Width:  |  Height:  |  Size: 348 KiB

After

Width:  |  Height:  |  Size: 348 KiB

View File

Before

Width:  |  Height:  |  Size: 70 KiB

After

Width:  |  Height:  |  Size: 70 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 40 KiB

View File

@@ -0,0 +1,364 @@
# Admin Portal Design Standards
This document captures the design system used by the Claude-built admin portal uplift (commits around `23b74a3`–`906c9e6`) so future edits to `admin/frontend/` stay visually consistent.
Intended readers: any agent or human touching `admin/frontend/*.html`, `admin.py`, or the admin API surface.
**Supplementary sources** (read these too before touching the admin UI):
- `.omo/notepads/admin-facelift/learnings.md` — the running log of every visual QA finding, root cause, and fix across the whole 10-wave facelift. This is the single richest record of what went wrong and how it was fixed.
- `plans/admin-facelift.md` — the original spec (survival guardrails S1–S5, device contracts).
- `plans/admin-visual-fixes-review.md` and `plans/router-admin-portal-implementation-review.md` — post-hoc audits that caught regressions the QA pass missed.
## Philosophy
- **Liquid-glass dashboard, not a marketing page.** Background is a subtle gradient wash; cards are translucent, blurred, and softly lit from the top. The goal is information density with low visual fatigue.
- **One visual language everywhere.** Use the same frosted-card treatment, pill buttons, and glass badges on every page so the portal feels like a single app.
- **Subtle cues over neon alerts.** Status colors are desaturated translucent tints. Warnings live in the navbar bell, not in a full card.
## Theme
- Always dark mode. Root is `<html lang="en" data-bs-theme="dark">`.
- **Base background** (body):
- Three overlapping `radial-gradient` blobs:
- amber (`rgba(245,158,11,.10)`) top-left
- blue (`rgba(59,130,246,.10)`) top-right
- purple (`rgba(139,92,246,.07)`) bottom
- Base color `#111827`.
- `background-attachment: fixed` so it stays put during scroll.
## Layout shell (every page)
1. Bootstrap 5 / Tabler page skeleton:
```html
<header class="navbar navbar-expand-md navbar-dark" data-bs-theme="dark">…</header>
<div class="page">
<div class="page-wrapper">
<div class="page-header d-print-none">…</div>
<div class="page-body">…</div>
<footer class="footer footer-transparent d-print-none">…</footer>
</div>
</div>
```
2. Scrolling context: **`<body>` scrolls, `<html>` does not.** This works around a Chrome/Linux phantom-margin bug caused by the HTML element also being the scrollbar container.
```css
html { overflow:hidden; height:100%; margin:0!important; padding:0!important }
body { overflow-y:auto; height:100%; margin:0!important; padding:0!important }
```
3. Container width: `container-xl` inside `.page-body`.
4. Cards grid: `.row.row-cards > .col-xl-6` or `.col-12`.
## Navbar
- Glass background: `rgba(31,41,55,.6)` with `backdrop-filter: blur(16px) saturate(160%)`.
- Bottom border: `1px solid rgba(255,255,255,.06)`.
- Explicit `position: relative; z-index: 1030;` on `header.navbar` so its dropdowns paint above later backdrop-filter stacking contexts (e.g. the quota chip).
- Container padding pinned to exactly `20px` left/right so the brand edge stays aligned across breakpoints:
```css
header.navbar > .container-fluid { padding-left:20px!important; padding-right:20px!important; }
```
- Brand wordmark: **6krrt LLM Router**, in `'Quicksand', var(--tblr-font-sans-serif)` at `font-weight: 700; letter-spacing: 0.01em`. The leading `6` is `.brand-six` — brand amber, `1.14em` — because the mascot is itself a 6. Keep the wordmark inside a single `.brand-word` span: the anchor is `inline-flex` with a `gap`, so a bare span around the `6` becomes its own flex item and that gap opens up between `6` and `krrt`.
- Logo SVG inline; force `filter:none!important` and size `64x64px` to defeat Tabler's autodark invert filter. The size lives in three places that must agree — the CSS rule, and the inline `width`/`height` attributes **and** `style="width:64px;height:64px"` on the `<svg>` (that inline style outranks the stylesheet, so changing only the CSS silently does nothing).
- **64px brand, 69px bar.** The mascot's face is ~40% of his height, so at 45px his eyes were ~4px. The extra 19px comes out of padding rather than out of the bar: `.navbar-brand` gives up its 8px top/bottom, and `header.navbar` drops 4px to 2px. Net bar height 69px, 1px *shorter* than the 45px version. Re-measure `header.navbar.offsetHeight` before touching either number.
### Logo
Two variants in `assets/`:
| File | Use |
|---|---|
| `6krrt-logo.svg` | plain mascot (Inkscape source) |
| `6krrt-logo-outlined.svg` | + white keyline; **this is what the admin portal uses** |
- The keyline is a 3.6-unit white outline around the head and tail, so the amber
silhouette separates from `#111827` and from darker browser tab bars.
- It is **not** a stroke on the amber shapes — a stroke centers on the path and
would eat half its width into the body. It is the same two shapes drawn first,
larger, in white: `circle r 42 -> 45.6`, tail `stroke-width 32 -> 39.2`. Draw
order is the whole mechanism; the white pair must stay above the amber pair.
- Eyes, smile, antenna and wheels already carry their own white rings and need
nothing added.
- Keep the keyline well under 9 units: by then the white has grown across the
counter of the "6" and closed the hole that makes the silhouette a digit.
- The favicon is the same artwork as a percent-encoded `data:image/svg+xml` URI
in each page's `<link rel="icon">`, so any logo change has to be applied to
both copies in all four HTML files.
### Speed lines
Three tapered amber lines trail the mascot inside the brand link
(`.brand-speedlines`), because he is a wheeled rover and should look like he is
going somewhere.
- Their own inline SVG, **not** part of the logo asset — `assets/*.svg` stays the
plain mascot, and the lines can drop out on narrow screens (`d-none d-md-flex`).
- The gradient fades *away* from him — transparent at the far left, amber at his
back. That direction is what reads as speed; reverse it and you get three
floating dashes.
- `gradientUnits="userSpaceOnUse"`. The default `objectBoundingBox` is undefined
on a horizontal line, whose bbox has zero height.
- `margin-right: -20px` on the wrapper. The logo's viewBox carries ~18px of empty
space left of his tail at 64px, so the pull spends most of itself cancelling
that before it buys any actual closeness. Target gap ≈ 4.5px.
- Motion is **hover-only**, and off under `prefers-reduced-motion`. A dashboard
that twitches at rest is a dashboard you stop looking at.
- Right cluster: warnings bell (dashboard + decisions only), SSE status badge, optional generated-at timestamp.
## Cards
```css
.card {
background: rgba(31,41,55,.55);
backdrop-filter: blur(18px) saturate(140%);
border: 1px solid rgba(255,255,255,.07);
box-shadow: 0 10px 30px -14px rgba(0,0,0,.55),
inset 0 1px 0 rgba(255,255,255,.04);
}
.card-header { border-bottom-color: rgba(255,255,255,.06); }
```
- Titles use `.card-title` with a leading inline SVG icon injected via `[data-icon]` + the inline `icon()` helper.
- Header actions align with flex utilities (`d-flex align-items-center justify-content-between`).
## Buttons
Override the solid Bootstrap semantic buttons to glass pills:
```css
.btn-success {
background: rgba(34,197,94,.16)!important;
border-color: rgba(74,222,128,.45)!important;
color: #4ade80!important;
}
.btn-success:hover {
background: rgba(34,197,94,.28)!important;
border-color: rgba(74,222,128,.7)!important;
color: #86efac!important;
}
.btn-danger { /* analogous red */ }
.btn { backdrop-filter: blur(6px); }
```
- Outlined secondary buttons are used for neutral actions (clear, view-all).
- Keep it small: `btn-sm` in headers and tables.
## Badges
All `.badge.bg-*` are glass tints:
| Class | Background | Text |
|---|---|---|
| `.bg-success` | `rgba(34,197,94,.20)` | `#4ade80` |
| `.bg-warning` | `rgba(245,158,11,.20)` | `#fbbf24` |
| `.bg-danger` | `rgba(239,68,68,.20)` | `#f87171` |
| `.bg-info` | `rgba(59,130,246,.20)` | `#93c5fd` |
| `.bg-secondary` | `rgba(148,163,184,.20)` | `#e2e8f0` |
| `.bg-purple` | `rgba(168,85,247,.20)` | `#d8b4fe` |
## Progress bars
Used for per-model usage, verdict mix, and category breakdown.
```css
.progress {
background: rgba(255,255,255,.08)!important;
border-radius: 999px;
overflow: hidden;
}
.progress-bar {
box-shadow: 0 0 6px 0 currentColor;
filter: saturate(1.25);
}
```
## Color palette
- Brand amber: `#f59e0b` / `#fbbf24`
- Success green: `#4ade80` / `#2ecc71`
- Danger red: `#f87171` / `#e74c3c`
- Warning yellow: `#fbbf24` / `#f1c40f`
- Info blue: `#93c5fd` / `#3498db`
- Purple accent: `#d8b4fe` / `#9b59b6`
- Cyan accent: used for `local-vision` kind badges (`--tblr-cyan`)
- Muted text: `var(--tblr-secondary)` (#94a3b8 region)
- Surface: `rgba(31,41,55,…)` (`#1f2937`)
- Background: `#111827`
Functional color mapping:
- Tier 1 → success green
- Tier 2 → warning amber
- Tier 3 → danger red
- Tier unknown → secondary
## Typography
- Body / UI text: Tabler default (`--tblr-font-sans-serif`).
- Brand + page titles: `'Quicksand', var(--tblr-font-sans-serif)`, weight 700, `letter-spacing: 0.01em`.
- Numeric data: `font-variant-numeric: tabular-nums` so numbers don't jitter when updating.
- Small meta text: `.text-muted small`, `font-size` around `0.75–0.85rem`.
## Icons
- **Do not** use the Tabler icons webfont ( avoided to prevent FOUT / network dependency).
- Use the inline `icon(name, size=18)` helper and `[data-icon]` placeholders. Each page ships a local `svgs` dictionary with only the icons it needs.
- Stroke-based SVGs, `width="18"`, `stroke-width="2"`, consistent with Feather/Tabler style.
## Charts
Dashboard uses Chart.js 4.4.7 with `chartjs-adapter-date-fns`.
### Line/area mini charts (History)
- One chart per metric; title rendered as plain text above the canvas.
- Common Chart.js options:
- `responsive: true`, `maintainAspectRatio: false`
- Legend hidden (`plugins.legend.display: false`)
- Grid: `rgba(255,255,255,0.04)`
- Y axis: `beginAtZero: true`, `maxTicksLimit: 6`
- X axis: `type: 'time'` with unit derived from range
- `pointRadius: 0`, `tension: 0.2`, `fill: false`
- Metric colors:
| Metric | Border | Fill bg |
|---|---|---|
| decisions | `#3498db` | `rgba(52,152,219,0.1)` |
| requests | `#9b59b6` | `rgba(155,89,182,0.1)` |
| cost | `#2ecc71` | `rgba(46,204,113,0.1)` |
| energy | `#e67e22` | `rgba(230,126,34,0.1)` |
| carbon | `#e74c3c` | `rgba(231,76,60,0.1)` |
### Verdict Mix / Category Breakdown bars
- Pure HTML/CSS, no Chart.js.
- Row layout: `.verdict-row { label | progress | count (share%) }`.
- Colors shared via `categoryColor(label)` so verdict bars and category dots line up.
## The quota chip
- Lives in `.page-header` right side, **not** in a dashboard card.
- Amber/blue gradient glass pill:
```css
background: linear-gradient(135deg, rgba(245,158,11,.12), rgba(59,130,246,.08));
border: 1px solid rgba(245,158,11,.22);
backdrop-filter: blur(14px) saturate(150%);
```
- Opens a modal with the full Quota Meter data.
## Modals
```css
.modal-content {
background: rgba(31,41,55,.75);
backdrop-filter: blur(24px) saturate(150%);
border: 1px solid rgba(255,255,255,.09);
box-shadow: 0 20px 50px -20px rgba(0,0,0,.6);
}
.modal-backdrop { background: #000; opacity: .6; }
```
## Warnings navbar bell
- Hidden if no warnings.
- Amber glass dot count; dropdown is a glass menu with dismissable small alerts.
- Dismissal is per-browser `localStorage`, not server ack.
## Warning-signal / API contract details (from learnings.md)
- **Never use `${icon(...)}` in static HTML.** The `icon()` helper only works in
JS-generated markup. In static card headers, use `<span data-icon="..."></span>`
and let `renderStaticIcons()` fill it on init. Using the literal template in
static markup renders the raw `${icon('x')}` text into the page — a real
regression that shipped once.
- **Tabler exposes Bootstrap components under the `tabler.*` global, NOT
`bootstrap.*`.** `tabler.min.js` does not expose a `bootstrap` global.
Always `new tabler.Toast(...)`, never `bootstrap.Toast(...)` (would throw
`ReferenceError`).
- **Keep a Chart.js `<script>` tag** even on pages that no longer use charts —
`tests/test_admin_frontend.py` asserts the string `"chart.js"` appears in
served page HTML. (Note: current pages load Chart.js only where actually
used; the test targets pages that historically included it — verify which
assertion applies to the page you touch.)
- **SSE status uses a Tabler `.badge bg-*` pill**, not a dot: classes are
`bg-warning` (connecting) → `bg-success` (live) → `bg-danger`
(reconnecting). Both `#sse-status` and `#sse-text` are updated together;
`#sse-text` carries the text for screen readers.
- **`var(--red)` / `var(--green)` / `var(--yellow)` / `var(--text-dim)` are
dead custom tokens** — replaced by Tabler's `--tblr-danger/success/warning/secondary`.
Never reintroduce the old names.
## Tables
- Use `.table.table-vcenter.card-table` or `.table-vcenter.table-hover.card-table`.
- Fixed column widths via utility classes where needed.
- Tier column uses `.tier-{1,2,3,?}` classes.
## Forms / controls
- Boolean knobs: `.form-check.form-switch` — **including** persisted-config
booleans. A bare `.form-check-input` in one panel and a switch in the other
reads as two different apps.
- Model availability: `.form-select.form-select-sm`.
### Settings rows (Controls page)
Runtime Knobs and Persisted Config share one row grammar, `.setting-row`:
```
[ key ................... ] [ meta ] [ control ]
minmax(0,1fr) auto 176px
```
- CSS grid, not `.row`/`.col-*`. A fixed control track is the point: every
switch, input and select lands on the same right edge, so the eye scans one
column. The Bootstrap-grid version put controls mid-row at three different
widths.
- `min-height: 38px` on the row — a switch and a text input differ by ~10px of
natural height, and alternating them looked ragged.
- `.settings-list` carries a negative inline margin equal to the row padding,
so keys align with the card body's content edge while the hover tint bleeds
to the card's inner edge.
- Dotted keys render as `<span class="scope">objective.</span>quality_tolerance`
— dim the scope, keep the leaf bright.
- **Do not add a column that restates the control's own value.** The meta track
is for facts the control cannot show: a unit (`kWh`), or a
`.badge.bg-warning-subtle.text-warning` reading `file: <value>` when a runtime
knob disagrees with `config.yaml`.
- Unsaved persisted-config edits are marked on the row (`.is-dirty`: amber left
rail + tint) and counted on the save button (`Save 2 changes`, disabled at
zero), rather than being invisible until a blind save-everything.
### Card-level layout
A card whose content is a single line of buttons goes **full width** with the
buttons in `.card-actions`, not into a `col-xl-6` beside a tall card — no
`align-items` value fixes the dead column that creates.
## Data patterns
- API base is `''` (relative to `/admin/`); SSE is `/events/decisions` (root level).
- Every page connects the same SSE stream for live status dot/bell updates.
- Escape everything inserted via `innerHTML` with `escapeHtml()`.
- Use `apiFetch()` wrapper that returns `null` on failure; render empty states accordingly.
- Use Bootstrap native toast (`tabler.Toast`) for user feedback, bottom-right, 4s delay.
## Hard-won rules
1. Backdrop-filter stacking contexts need explicit `position: relative; z-index: 1030;` on the navbar or dropdowns sink below later glass elements.
2. Always assign the return value of `new Chart(...)` to a variable, or polling will recreate charts and throw "Canvas is already in use".
3. Keep `<body>` the scroll context, not `<html>`, or Chrome/Linux can add a phantom left margin equal to scrollbar width.
4. Defeat Tabler's `filter: brightness(0) invert(1)` on `.navbar-brand-image` with `filter:none!important` and explicit 45px sizing.
5. Use `!important` on badge/button glass overrides because Bootstrap's attribute-selectored dark rules have higher specificity.
6. Chart.js 4 does **not** stack values within a single dataset. To build a stacked bar (verdict mix), emit **one dataset per segment**.
7. `renderStaticIcons()` must be called first in `init()` — static `[data-icon]` placeholders render undefined otherwise.
8. History range button selectors must match where the buttons actually live. A selector assuming they sit in a `.history-rows` ancestor broke once because the buttons were moved into `.card-title`. Prefer `button[data-range]` globally.
9. Verify rendered DOM with screenshots, not just grep — static-vs-JS template interpolation is easy to get wrong (see rule 6/7 in learnings).
10. After a change to `/admin` pages, run `python -m pytest` (733 tests) AND a Playwright load + one `REFRESH_MS` (30s) idle wait, confirming zero console errors. A fresh-load screenshot alone has missed canvas-reuse and selector bugs twice in this project's history.
## File inventory
| File | Purpose |
|---|---|
| `admin/frontend/index.html` | Dashboard (quota chip, per-model usage, verdict mix, category breakdown, history mini-charts, recent decisions) |
| `admin/frontend/models.html` | Model availability table with override dropdown |
| `admin/frontend/decisions.html` | Full decision log with filter + search |
| `admin/frontend/controls.html` | Operational triggers, runtime knobs, persisted config editor |
| `admin.py` | FastAPI sub-router mounted at `/admin` |
| `admin_schema.sql` | RBAC/audit tables (not yet wired in UI) |

91
plans/admin-facelift.md Normal file
View File

@@ -0,0 +1,91 @@
# Spec: admin dashboard facelift — standardize on Tabler (dark mode)
**Origin.** The admin dashboard (`admin/frontend/index.html`) is a hand-rolled,
single-file dark theme — custom CSS variables, a monospace terminal
aesthetic, and vanilla-JS panels with no component library. The ask is to
replace that visual language with [Tabler](https://github.com/tabler/tabler)
— a fully open-source (MIT) Bootstrap 5 admin template with dark mode built
into the core product, not gated behind a paid tier — while keeping every
existing behavior intact. (An earlier pass at this same ask targeted
Themesberg's Volt Dashboard; switched to Tabler specifically because Volt's
dark theme may be Pro-only, a risk Tabler doesn't carry.)
## Outcome (post-lift)
Spec became reality in commits `23b74a3`–`906c9e6`. The portal is now four
pages with a consistent liquid-glass dark theme. The consolidated design
system is captured in **
## Folders/files now in the admin portal
- `admin.py` (~800 LOC) — FastAPI sub-router with API + static HTML
- `admin_schema.sql`
- `admin/frontend/index.html` — dashboard
- `admin/frontend/models.html` — model availability
- `admin/frontend/decisions.html` — decision log
- `admin/frontend/controls.html` — triggers, knobs, persisted config
## Selected architecture decisions
- **CDN for Tabler + Chart.js + Google Fonts, everything else inline.** The
pages still contain their own `<style>` and `<script>` blocks — the original
"one self-contained file" architecture from F12 was relaxed, but no static
mount was added. Tabler core CSS/JS, Chart.js, and `Quicksand` load from
jsDelivr/fonts.googleapis.com.
- **Inline SVG icon helper.** No Tabler Icons webfont; each page carries a
minimal `icon(name, size)` dictionary to avoid an extra font dependency and
potential FOUT/CSP issues.
## What survived from the original guardrails
- API surface unchanged.
- `escapeHtml()` applied to all interpolated text.
- Chart instances reused where Chart.js is used.
- `saveAllConfig()` posts sequential requests.
- SSE status indicator reflects real connection state.
## What was added beyond the spec
- Atomic config writes with threading.Lock + `.tmp` + `os.replace`
(`admin.py`) plus concurrent-write regression test.
- Four-page navigation shell instead of a single-page dashboard.
- Warnings moved from a card into a navbar bell.
- Quota meter moved from a card into a header chip + modal.
- Verdict mix converted from a Chart.js doughnut to pure HTML/CSS partition
bars (avoids canvas lifecycle issues entirely).
- History split from one chart into five per-metric mini charts.
## Review / process gaps identified from the commit series
1. **Chart-instance assignment bug.** `renderVerdict` initially failed to
assign the created Chart instance back to `verdictChart`, causing repeated
"Canvas is already in use" errors every 30-second poll. The QA evidence
console log contained the error but was missed by the approving pass.
**Lesson: read saved console logs, not just screenshots.**
2. **Commit granularity.** The 8-todo plan specified one commit per todo; the
real work coalesced into fewer larger commits. **Lesson: granular commits
make review and bisection easier.**
3. **Service restart gap.** The atomic-write fix landed on disk but
`llm-router.service` was not restarted, so the running process still used
the old code. **Lesson: after a fix that protects the running service,
restart explicitly.**
4. **Null numeric config fields.** `objective.max_energy_per_request` is
`null` in `config.yaml`; the frontend sends empty string, which
`Optional[float]` rejects. Not data-destructive, but it produces spurious
failure toasts on "Save All".
## Suggested future work
- Extract the repeated cross-page CSS into `admin/frontend/admin.css` to
prevent drift across the four pages.
- Extract the inline `icon()` helper into a shared `admin/frontend/icons.js`
module.
- Null-aware config editor (empty string ↔ null round-trips cleanly).
- Wire `admin_schema.sql` RBAC/audit tables behind authentication.
- Better feedback in `models.html` when the override dropdown is disabled
because a model is already overridden.
---
*Original spec below preserved for context.*

BIN
plans/admin-qa-final.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 214 KiB

BIN
plans/admin-t6-controls.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 86 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 130 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 252 KiB

BIN
plans/admin-t6-index.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 132 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 18 KiB

BIN
plans/admin-t7-controls.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 103 KiB

BIN
plans/admin-t7-fixes.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 198 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 221 KiB

BIN
plans/admin-t7-index.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 222 KiB

BIN
plans/admin-t7-models.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 99 KiB

View File

@@ -1,10 +1,10 @@
# Review: `admin-visual-fixes` implementation
**What it was reviewing:** the 8-todo `admin-visual-fixes` plan
(`code_plans/.omo/plans/admin-visual-fixes.md`), after its orchestration
(`plans/.omo/plans/admin-visual-fixes.md`), after its orchestration
reported 8/8 complete and F1-F4 all APPROVE. Verified against the actual
diff and the plan's own saved QA evidence
(`code_plans/.omo/evidence/task-8-admin-visual-fixes-console.txt`), not the
(`plans/.omo/evidence/task-8-admin-visual-fixes-console.txt`), not the
orchestration summary.
## Verdict: 6 of 8 todos are genuinely fixed; Todo 2 is not, and the plan's own evidence file proves it

View File

@@ -0,0 +1,119 @@
# Admin UI Work Framework
How to produce the class of result the `c22bbae` Controls relayout did — work so
specific it stands alone as a learning input, written as **rules the next page
can apply**, backed by evidence a fresh model can trust.
This framework is about the *craft* of touching the admin portal, not the visual
styles (those live in `admin-design-standards.md`). Read the Standards doc first;
this doc is the method.
## Why this exists
Before `c22bbae`, the admin UI got "fixed" in ways that didn't transfer: a review
category said "card is a single line of buttons," and a future page could easily
re-produce the same dead column. The Controls relayout broke that pattern by
**turning each observed defect into a rule the next page can mechanically apply**
("don't restate the control's own value," "single-line button strip goes full
width"). The commit message is the artifact worth feeding a model: it spells out
the three problems, why each fix is what it is, and the two consistency fixes —
with enough context to stand alone.
## The method (5 moves)
### Move 1 — Diagnose at the surface, in the running browser
Look at the page as a user does: full viewport and at 800px, before and after a
`REFRESH_MS` poll. Name the defect in one concrete sentence before you touch
anything.
Good defect statements (from the real commit):
- "Operational Triggers is a single line of buttons but sat in a col-xl-6 next
to a much taller card, so half that row is permanently empty."
- "Both settings panels used a 5/3/4 Bootstrap grid, which put the control
mid-row at three different widths and spent a whole trailing column restating
the value the control already displays."
- "Rows alternated switches and text inputs, which differ ~10px in natural
height and read as ragged."
These are *specific and falsifiable*. "Fix the controls page layout" is not a
diagnosis — it's a wish.
### Move 2 — Extract the grammar, not the look
Don't record "Controls page uses `grid-template-columns:minmax(0,1fr) auto 176px`."
Record the *transferable principle* behind it, plus the code as evidence:
> A settings list is one CSS-grid row — key | meta | control — not Bootstrap
> `.row/.col-*`. A fixed control track is the point: every switch, input and
> select lands on the same right edge, so the eye scans one column.
Then back it with the exact CSS so a future page can copy it. Grammar lives at
the level of "what column exists and why," not "which pixel value."
Rule templates that transfer:
- **Do not add a column that restates the control's own value.** The meta
track is for facts the control cannot show.
- **A card whose content is a single line of buttons goes full width** with the
buttons in `.card-actions`, not into a `col-xl-6` beside a tall card.
- **Rows with a switch next to a text input need a fixed min-height** or they
read as ragged.
- **Boolean persisted-config controls use switches**, matching runtime knobs —
not bare checkboxes.
### Move 3 — Keep old constraints load-bearing
Every change must be verified against the survival checklist (S1–S5) in
`admin-facelift.md` and the `learnings.md` notepad:
- `escapeHtml()` on every interpolated string
- chart instance reuse (never `.data.labels` mutation on the verdict bar)
- `renderWarnings` inline placeholder rebuild
- `saveAllConfig` stays sequential (`for...of await`, no `Promise.all`)
- SSE status pill reflects real EventSource state
If your change touches any of these, call it out explicitly in the commit.
### Move 4 — Prove it, not just observe it
A single fresh-load screenshot has missed canvas-reuse and selector bugs twice
in this project. Minimum evidence bar for an admin page change:
1. `python -m pytest` — full suite (733 tests) green.
2. Playwright load at 1400px **and** 800px, plus one 30s+ idle past `REFRESH_MS`,
with **zero console errors** (beyond the benign `/favicon.ico 404`).
3. The exact interaction under change exercised by hand (toggle, save, range).
Save the screenshots under `plans/` (e.g. `admin-controls-relayout.jpg`).
### Move 5 — Write the commit as the learning input
The commit message is documentation for the next model. Structure it as:
- **The problem(s)**, each in one concrete sentence (Move 1).
- **Why each fix is what it is** — the reasoning, not just the change.
- **Consistency fixes** called out separately from the main change.
- **Test changes** explained (why the old assertion no longer held).
Follow with a `git show <sha>` that a fresh session can read and immediately
apply.
## Also update the Standards doc in the same commit
When the change teaches a new transferable rule, add it to
`plans/admin-design-standards.md` *in the same commit* so doc and code can
never drift. Write the new section in rule form ("Do not add a column that
restates the control's own value"), not as a description of the one page. That
is what makes it applicable to the *next* page.
## Repository of what this produces
- `plans/admin-design-standards.md` — the accumulated grammar (settings
rows, card-level layout, hard-won rules 1–10, data patterns).
- `.omo/notepads/admin-facelift/learnings.md` — the chronological record of
every diagnosis/fix/QA, including the false-starts that were corrected.
- `plans/*.md` — the post-hoc audits that catch what the QA pass missed.
- Commit messages themselves — the highest-signal, standalone learning input.
## When to use this framework
- Any visible change to `admin/frontend/*.html` or the `/admin` API surface.
- Any downstream admin page you might build (e.g. an Audit/Logs page, wiring in
`admin_schema.sql` RBAC) — apply the grammar first, then the visuals.

View File

Before

Width:  |  Height:  |  Size: 278 KiB

After

Width:  |  Height:  |  Size: 278 KiB

View File

@@ -1,9 +1,9 @@
# Spec: two design decisions deferred out of the context-pruning/framing fix pass
**Origin.** The fix pass for
[`context-pruning-and-framing-review.md`](../code_reviews/context-pruning-and-framing-review.md)
[`context-pruning-and-framing-review.md`](context-pruning-and-framing-review.md)
(reviewed in
[`context-pruning-and-framing-fixes-review.md`](../code_reviews/context-pruning-and-framing-fixes-review.md))
[`context-pruning-and-framing-fixes-review.md`](context-pruning-and-framing-fixes-review.md))
explicitly declined two items as design decisions rather than confirmed
bugs: the `TaskRequest.context` dual-use ambiguity (review finding #6,
partially mitigated but not resolved), and classifying once per session

BIN
plans/layout-800.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 216 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 204 KiB

BIN
plans/layout-controls.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 301 KiB

BIN
plans/layout-load.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 300 KiB

BIN
plans/layout-reload.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 415 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 220 KiB

View File

@@ -8,7 +8,7 @@ request path."* The embedding step was cut on purpose, to keep pinch pure
and offline-testable — not because it was a bad idea. This spec is that step,
scoped to fit in without giving up the purity that cut it in the first
place. It also settles the scoping note in
[`magic-brainstorming-review.md`](../code_reviews/magic-brainstorming-review.md)
[`magic-brainstorming-review.md`](magic-brainstorming-review.md)
(idea #3): pinch's mechanism only ever touches old tool results, and this
spec stays inside that boundary rather than widening it.

View File

@@ -1,9 +1,9 @@
# Review: embedding-based pinch relevance + upstream failover/circuit breaker (commit `03f62e2`)
**What it was reviewing:** opencode/Atlas's implementation of both
`code_plans/pinch-embedding-relevance.md` and
`code_plans/upstream-failover-and-circuit-breaker.md` in one commit, per the
`.omo/plans/pinch-embedding-relevance.md` work plan approved earlier.
`plans/pinch-embedding-relevance.md` and
`plans/upstream-failover-and-circuit-breaker.md` in one commit, per the
`plans/.omo/plans/pinch-embedding-relevance.md` work plan approved earlier.
Verified against the actual diff (`git show 03f62e2`) rather than the commit
message, and ran the full suite directly: 673 passed (up from 562).
@@ -91,7 +91,7 @@ model with a failure history has its cooldown monotonically ratchet toward
`max_cooldown_seconds` on every subsequent failure**, even failures
separated by long stretches of trouble-free service. That is the opposite
of "passive recovery rebuilds trust," which was the explicit point of the
design (`code_plans/upstream-failover-and-circuit-breaker.md`: "a success
design (`plans/upstream-failover-and-circuit-breaker.md`: "a success
clears the entry entirely... the next failure after a success restarts at
`initial_cooldown_seconds`, not wherever the backoff had climbed to").

View File

@@ -1,7 +1,7 @@
# Review: README restructure (commit `14a7653`)
**What it was reviewing:** opencode's implementation of
`documentation_plans/readme-restructure.md` — reshaping `README.md` around a
`plans/readme-restructure.md` — reshaping `README.md` around a
pitch → features → install → per-command usage structure borrowed from
`supabase-plus`'s README shape. Checked the actual diff (`git show
14a7653`) line by line against the spec's explicit constraints, not just the

View File

@@ -0,0 +1,684 @@
# root-reorg-src-config-plans — plan review
Target: `plans/.omo/plans/root-reorg-src-config-plans.md`
Round 1 → v1 (315 lines). Round 2 → v2 (430). Round 3 → v3 (434). Round 4 → **v4** (450).
---
# Round 4 — v4 (post dual-review)
**Verdict: one line to fix, then go.** All three Round-3 findings are applied, and
more carefully than asked — V3-1 is guarded in four separate places, and V3-2's
"don't over-apply" caveat was carried into an explicit *must NOT do*. The four
dual-review edits are all sound. One of them over-corrected.
## V4-1. The concurrency test's `ROOT /` source path does need to change
Plan line 279:
> `tests/test_admin_config.py:142-186` concurrency test stays unchanged — it passes
> explicit `config_yaml` paths directly to `_persist_config_value`, never uses
> `base_dir`; its own `config.yaml.bak.*` glob at 182 is correct as-is.
That is right about the *destination* and the *glob*, and wrong about one line in
between. The block contains:
```python
145: config_yaml = tmp_path / "config.yaml" # destination — STAYS
146: shutil.copyfile(ROOT / "config.yaml", config_yaml) # SOURCE — must move
...
182: backups = list(tmp_path.glob("config.yaml.bak.*")) # glob — STAYS
```
Line 146's `ROOT / "config.yaml"` is the **repo's real config file**, which moves
to `config/config.yaml` in this very wave. Left alone, the test dies with
`FileNotFoundError`.
The rule that actually holds, and is worth stating this way in the plan:
> Inside 142-186, anything anchored on `tmp_path` stays (it is an explicit path
> handed straight to `_persist_config_value`). Anything anchored on `ROOT` moves
> like every other `ROOT /` reference in the suite.
**And the safety net is instructed to dismiss it.** Wave 2's new sanity grep
(plan line 286) reads:
> `git grep -n 'ROOT / "config.yaml"…' tests/` returns **only the expected
> concurrency-test sites at `tests/test_admin_config.py:142-186`** (which pass
> explicit paths) or nothing at all.
Line 146 is precisely the hit that grep would produce, and the acceptance
criterion pre-labels it as expected. That turns a catch into a rubber stamp. The
criterion should be "returns nothing at all" — full stop.
Cosmetic, same area: plan line 275 has an unbalanced quote —
`tmp_path / "config" / "config.yaml` — in a line a worker will copy.
## Verified by running it, not by reading it
- **The new collection gate is not a false red.** `pytest --collect-only -q | tail -1`
emits `733 tests collected in 1.43s` (with ANSI codes, no trailing blank), so
`grep -q "733 tests"` → PASS. Worth having checked; `tail -1` guards break this
way often.
- **The `"status":"success"` gate is achievable and exactly matched.**
`feedback.py --dry-run` exits `rc=0` against the live `router.db`, and FastAPI
renders `{"status":"success"}` with no space after the colon — so the grep
pattern matches literally. Both halves had to be true for that gate to work.
- `deploy/README.md` line references are accurate: 12-14 (the unit table), 34-35
(the sed install loop), 94 (the oneshot-buffering note naming `poller.py` and
`seed_energy.py`).
## On Metis's second gap
"pytest silent zero-collection if `pythonpath=["src"]` is set before `src/` exists"
is not a real failure mode — an unimportable test module produces a **collection
error**, loudly, not a silent zero; and `testpaths = ["tests"]` still resolves. The
added `git mv`-before-pyproject ordering and the count assertion cost nothing and
are fine to keep as belt-and-braces. Flagging only so the reasoning isn't reused
somewhere it matters.
Momus's rollback warning was worth taking, and the pre-flight note now states it
correctly: restoring the backup units *alone*, after modules are already in `src/`,
crashes — a full revert needs `git reset --hard $PREFLIGHT_SHA` **and** the units.
## On the documented tradeoff
Keeping the service up through Waves 1-2 is a reasonable owner's call, and it is
now recorded as one rather than left implicit. Two things make the residual risk
smaller than it reads: both timers are stopped at pre-flight, so the likeliest
trigger is gone; and the exposure is a genuine crash-loop only if something
restarts the process in that window. The one path still open is the admin portal's
own `/admin/api/restart-service` button — worth simply not touching until Wave 3
lands.
## Go / no-go
Fix V4-1 (one line in the plan, plus the grep criterion and the stray quote) and
execute. Nothing else is outstanding.
---
# Round 3 — v3
**Verdict: two real defects and one ordering fix, all small. Fix them before
spending the dual review pass.** All ten claimed v2 fixes are present and correct;
I re-verified each. Both new defects are in the same place — `admin.py`'s config
backup write — and neither is a consequence of the v2→v3 reordering; they have
been carried unnoticed since v1, mine included.
## Must fix
### V3-1. The `admin.py:370` backup-path change is wrong and 500s the Controls save
Wave 2 instructs:
> `src/admin.py:370` backup path `f"config.yaml.bak.{ts}"` → `f"config/config.yaml.bak.{ts}"`
The actual code is:
```python
backup = config_path.with_name(
f"config.yaml.bak.{int(time.time())}"
)
```
`Path.with_name()` replaces only the **filename component and keeps the parent**.
Once `config_path` is `<root>/config/config.yaml`, it already produces
`<root>/config/config.yaml.bak.<ts>`. **The line needs no change at all.**
Applying the instruction gives `with_name("config/config.yaml.bak.123")`, and
`with_name` rejects any argument containing a separator. Verified on this repo's
venv (Python 3.14.6):
```
unchanged -> /repo/config/config.yaml.bak.123
v3 change -> RAISES ValueError Invalid name 'config/config.yaml.bak.123'
```
That raises out of `_persist_config_value`, so **every persisted config edit from
the Controls page returns 500** — the feature that just shipped in the admin
facelift work.
It does fail loudly: `tests/test_admin_config.py:128`
(`test_config_POST_creates_backup_before_write`) and the concurrency test at 142
both exercise the write. So it costs a debug cycle rather than shipping broken.
But it should not be in the plan.
**Fix:** delete that bullet from Wave 2. Add it to the "must NOT do" list instead,
with the `with_name` reason, so nobody re-derives it.
*(This line was in v1 and v2 too. Round 1 flagged `admin.py:370` only for the
gitignore-pattern question and did not check `with_name` — that was my miss.)*
Related, and still correct: `.gitignore`'s `config.yaml.bak.*` has no slash, so it
matches by basename at any depth. `config/config.yaml.bak.*` stays ignored with no
`.gitignore` change — which now matters, since that is genuinely where they land.
### V3-2. `tests/test_admin_config.py` isn't in Wave 2's list and needs real changes
The `client` fixture (lines 45-58) builds the admin router with
`base_dir=str(tmp_path)` and stages the config at the top of tmp_path:
```python
config_yaml = tmp_path / "config.yaml"
shutil.copyfile(ROOT / "config.yaml", config_yaml)
...
router = build_router(None, _db_factory, base_dir=str(tmp_path))
```
Wave 2 changes `admin.py:439` to `Path(base_dir) / "config" / "config.yaml"`, so
the fixture must stage into a `config/` subdirectory instead:
```python
config_yaml = tmp_path / "config" / "config.yaml"
config_yaml.parent.mkdir(parents=True, exist_ok=True)
shutil.copyfile(ROOT / "config" / "config.yaml", config_yaml)
```
and the backup assertion at line 135 becomes:
```python
backups = sorted(tmp_path.glob("config/config.yaml.bak.*"))
```
**Note the asymmetry**, because it is easy to over-apply: the concurrency test at
142-186 calls `_persist_config_value(config_yaml, ...)` with an **explicit** path
and never goes through `base_dir`. Its `tmp_path / "config.yaml"` staging and its
glob at line 182 are correct as they stand and must be left alone.
Wave 2's test list names `test_proficiency.py`, `test_config_endpoints.py`,
`test_poller_parsing.py`, `test_admin_frontend.py` and the generic
`ROOT / "config.yaml"` sweep. `test_admin_config.py` matches the generic sweep for
lines 54 and 146, so a mechanical pass would rewrite the source path and leave the
destination and the glob wrong.
### V3-3. Wave 1 breaks the admin triggers, Wave 3 restarts production into that state, and the Wave 1 smoke cannot detect it
After Wave 1, `_repo_root` is corrected but the spawn argv is still
`[sys.executable, "poller.py"]` with `cwd=<repo root>` — and `poller.py` now lives
in `src/`. Wave 4 is what fixes the argv, three waves later.
Wave 1's acceptance includes:
```bash
curl -sf -X POST localhost:8081/admin/api/apply-feedback?dry_run=true
```
**This cannot fail.** Verified: a missing script exits `rc=2`, and `_run_steps`
turns every failure — spawn error, timeout, nonzero exit — into a `_job(...)` dict
returned as **HTTP 200** with `{"status": "failed"}`. The endpoint has no raise
path, and `curl -sf` only inspects the status code.
That is the third false-green of this exact shape across three revisions (v1's
Wave 1 `import metrics`, v2's Wave 1 `python -m pytest`, now this). The pattern
worth internalising: **a check that passes before the change is applied is not a
check.**
Two fixes, take both:
1. **Assert the job status, not the HTTP code**, at every occurrence of that curl:
```bash
curl -sf -X POST 'localhost:8081/admin/api/apply-feedback?dry_run=true' \
| grep -q '"status":"success"' || { echo "TRIGGER BROKEN"; exit 1; }
```
2. **Fold Wave 4 into Wave 1.** It is four argv lines plus the
`test_admin_triggers.py` assertions, it depends on Wave 1 alone, and it removes
an intermediate state in which production's Controls buttons are silently
dead — a state **Wave 3 currently restarts the live service into**. This is the
same argument v3 already accepted for folding pyproject into Wave 1.
The live service is unaffected during Waves 1-2 (it runs already-loaded code), so
the exposure begins precisely at the Wave 3 restart. Folding closes it.
## Verified, no action
- **Wave 3's reinstall matches the documented install exactly.** `REPO=$(pwd)` and
the sed loop are character-for-character what `deploy/README.md:31-36` does. V2-1
is properly resolved.
- All four `-m` targets (`poller`, `tier`, `seed_energy`, `feedback`) have
`if __name__ == "__main__"` guards.
- No `mypy.ini` / `ruff.toml` / `setup.cfg` / `tox.ini` exists, so no lint config
references root `.py` paths.
- No other invoker of root `.py` scripts in `opencode.json`, `package.json`, or
`.claude/`.
- Root-entry arithmetic checks out: `51 − 29 + 1 − 4 + 1 − 4 + 1 = 17`.
- `tests/test_admin_triggers.py` monkeypatches `asyncio.create_subprocess_exec`,
so Wave 4 spawns nothing real; the cited line numbers (101, 107-108, 140, 149,
161, 168) are accurate.
## Round 2 findings — resolution
| # | Round-2 finding | v3 status |
|---|---|---|
| V2-1 | Wave 7 hardcodes a personal path, installs via `cp` | **Fixed.** `%h/llm-router` kept; sed loop adopted verbatim; "do NOT copy units with `cp`" is an explicit guard. |
| V2-2 | Subprocess probes lose `PYTHONPATH` | **Fixed.** Both call sites listed, with the env dict and the reason. |
| V2-3 | `test_admin_frontend.py` `.parent` assertion | **Fixed.** Line 56 → `.parent.parent`, in Wave 1. |
| V2-4 | Wave 2→7 armed window; wrong dependency | **Fixed.** Timers stopped at pre-flight; reinstall moved to Wave 3; matrix now carries an explicit "production window armed?" column. |
| V2-5 | Wave 1 committed a broken tree | **Fixed.** pythonpath folded into the move commit. |
| V2-6 | `load_config` stays cwd-relative | **Decided, and recorded as a decision** in both Scope-OUT and Wave 2. Correct handling — it is now a choice rather than an omission. |
| V2-7 | Smoke port 9000 vs convention | **Fixed.** 8081, citing CLAUDE.md. |
| V2-8 | Commit count | **Fixed.** "6 waves, 5 git commits." |
| V2-9 | Root target | **Fixed.** Exactly 17. |
| V2-10 | Update the stale comment in the unit files | **Fixed.** Wave 3 updates it for `config/` and PYTHONPATH. |
| V2-11 | `code_plans/` count | **Fixed by removal.** The bad number is gone; the criterion is "all tracked files, no deletions." |
Also correctly carried over: the `config/__init__.py` guard (confirmed — a
`config/` directory does not shadow `src/config.py`, because PEP 420 records it as
a namespace *portion* and a regular module found later on `sys.path` wins; adding
`__init__.py` would flip that to path order).
## Recommended edits before execution
1. **V3-1** — delete the `admin.py:370` bullet; move it to "must NOT do" with the `with_name` reason.
2. **V3-2** — add `tests/test_admin_config.py` to Wave 2: fixture stages into `tmp_path/config/`, line 135 glob gains the `config/` prefix, lines 142-186 left alone.
3. **V3-3** — assert `"status":"success"` in every trigger curl, and fold Wave 4 into Wave 1.
With these three applied v3 is ready to execute, and worth the dual
high-accuracy review pass.
---
# Round 2 — v2
**Verdict: close. Four must-fix items, then it's executable.** Every Round-1
blocker is genuinely resolved, the reference list is now verified rather than
asserted, and the PYTHONPATH decision is the right call and correctly argued. The
remaining problems are all in the two places v1 didn't reach: the systemd install
mechanism, and the tests that spawn real subprocesses.
## Correction to Round 1
My Round-1 item B3 said *"`deploy/*.service` should be corrected to
`%h/Sources/6krrt` (or the drift documented)."* **That was wrong**, and v2's Wave 7
implements it. `%h/llm-router` is not drift — it is a deliberate placeholder,
substituted at install time by the loop already in `deploy/README.md:34-35`:
```bash
for u in deploy/llm-router*.{service,timer}; do
sed "s|%h/llm-router|${REPO}|g" "$u" > ~/.config/systemd/user/"$(basename "$u")"
done
```
The live units are that loop's *output*. See V2-1.
---
## Must fix
### V2-1. Wave 7 hardcodes a personal path and bypasses the documented install
Wave 7 changes `deploy/*.service` to `WorkingDirectory=%h/Sources/6krrt`,
`EnvironmentFile=%h/Sources/6krrt/.env`, `ReadWritePaths=%h/Sources/6krrt`,
`Documentation=file:%h/Sources/6krrt/CLAUDE.md`, and installs with a plain `cp`.
Two consequences:
1. **The repo ships one developer's home directory.** `%h/llm-router` is the
token `deploy/README.md` greps for; replacing it means the sed finds nothing
and every other clone installs units pointing at `~/Sources/6krrt`.
2. **`cp` skips the substitution step entirely**, so the install procedure in
`deploy/README.md` and the one in the plan now disagree.
**Fix:** keep `%h/llm-router` in `deploy/*.service`. Add only the layout changes
there — `Environment=PYTHONPATH=%h/llm-router/src`, `ExecStart=... -m poller`,
etc. Reinstall with the README's existing loop, not `cp`:
```bash
REPO=%h/Sources/6krrt # or: REPO=$(pwd)
systemctl --user stop llm-router.service llm-router-poller.timer llm-router-seed.timer
for u in deploy/llm-router*.{service,timer}; do
sed "s|%h/llm-router|${REPO}|g" "$u" > ~/.config/systemd/user/"$(basename "$u")"
done
systemctl --user daemon-reload
systemctl --user start llm-router.service llm-router-poller.timer llm-router-seed.timer
```
Note the sed must also rewrite the new `PYTHONPATH` line, which it does for free
since that line contains the same token.
### V2-2. Three real-subprocess tests break — pytest's `pythonpath` does not export `PYTHONPATH`
Verified empirically inside a pytest run: **`PYTHONPATH env = None`**. The
`pythonpath` ini option mutates the pytest process's `sys.path`; it sets no
environment variable, so child processes inherit nothing.
Two tests spawn a real interpreter:
```
tests/test_admin_frontend.py:84 subprocess.run([sys.executable, "-c", probe], cwd=str(ROOT))
probe does: import admin
tests/test_admin_health.py:203 subprocess.run([sys.executable, "-c", probe], cwd=str(ROOT))
probe does: import admin; assert 'dispatcher' not in sys.modules
```
With `cwd=ROOT` and the modules now in `src/`, both get
`ModuleNotFoundError: No module named 'admin'` after Wave 2. Neither file appears
in Wave 2's file list.
**Fix:** pass the env explicitly at both call sites:
```python
subprocess.run(
[sys.executable, "-c", probe],
check=True,
cwd=str(ROOT),
env={**os.environ, "PYTHONPATH": str(ROOT / "src")},
capture_output=True,
)
```
This is the price of choosing PYTHONPATH over an editable install — an editable
install would have made these work untouched. The trade is still correct for the
reason v2 gives (a rebuilt `.venv` without `pip install -e .` is a silent
outage); it just has this one bill attached, and the plan should pay it
deliberately rather than discover it.
### V2-3. `tests/test_admin_frontend.py:53-58` hardcodes the old `__file__` depth
```python
def test_admin_index_file_exists_at_module_derived_path():
"""The served file lives at Path(__file__).parent/admin/frontend/index.html."""
expected = (
Path(admin.__file__).resolve().parent / "admin" / "frontend" / "index.html"
)
assert expected.is_file(), f"frontend file missing at {expected}"
```
After Wave 2 this resolves to `src/admin/frontend/index.html` and fails. Needs
`.parent.parent`, and the docstring updated to match.
v2's test list names only `tests/test_admin_frontend.py:71` — which is the
`admin.load_config('config.yaml')` string *inside the subprocess probe*, correctly
caught for Wave 3. The `.parent` assertion eight lines above it was missed. This
is the same test file that documents the `_REPO_ROOT` coupling, so it is the one
place the anchor change is asserted.
### V2-4. The Wave 2 → Wave 7 window leaves production armed, and the timers will fire in it
Quantified from the live units:
```
llm-router-poller.timer OnUnitActiveSec=2h
llm-router-seed.timer OnUnitActiveSec=6h
```
Between Wave 2 (modules move) and Wave 7 (units reinstalled), the installed units
still say `ExecStart=… uvicorn dispatcher:app` with no `PYTHONPATH` and
`WorkingDirectory=%h/Sources/6krrt`. The dispatcher survives only because it is
already loaded. Anything that restarts it — the admin portal's own
`/admin/api/restart-service` button, an OOM, a reboot, `Restart=always` after any
crash — brings it back into a crash loop on the port `opencode.json` uses for the
user's own inference. And a multi-hour reorg **will** cross at least one 2-hour
poller tick, which fails silently; `stale_after_days: 3` then starts the clock on
a catalog that routes nothing.
**Two fixes, both cheap:**
1. **Pre-flight:** `systemctl --user stop llm-router-poller.timer llm-router-seed.timer`
before Wave 1, restart them in Wave 7. Add to the pre-flight block next to
`PREFLIGHT_SHA`.
2. **Move the unit reinstall to immediately after Wave 3.** The dependency matrix
says Wave 7 depends on Wave 4 — it does not. Wave 4 changes `admin.py`'s
*internal* spawn argv; the unit files never reference it. Wave 7 needs Wave 2
(`src/`) and Wave 3 (`config/` paths) and nothing else. Reordering cuts the
exposure window from five waves to two, and leaves Waves 4/5/6 as
doc/test-only and production-safe.
---
## Should fix
### V2-5. Wave 1 commits a tree where the `pytest` console script is broken
Wave 1 sets `pythonpath = ["src"]` while the modules are still at the root and
`src/` does not exist. Collection then works only under `python -m pytest`, which
adds the cwd — and the existing comment in `pyproject.toml` says in as many words
that the setting exists precisely because the bare `pytest` console script does
*not*:
> Running as `python -m pytest` happens to add the cwd and hides this; the
> `pytest` console script does not.
So Wave 1's acceptance criterion (`python -m pytest --collect-only -q` passes) is
a false-green of exactly the kind that made v1's Wave 1 dangerous — it passes for
a reason unrelated to what it claims to check.
**Fix:** move the `pythonpath` line into Wave 2's commit, alongside the move it
describes. Keep `.gitignore` and the pre-flight SHA in Wave 1.
### V2-6. `config.py`'s default stays cwd-relative — worth fixing while you're in there
`load_config(path="config/config.yaml")` still resolves against the current
directory. Every invocation from anywhere but the repo root breaks, which is the
whole reason the live unit carries this comment:
> `# config.yaml, router.db and router.log are all referenced as relative paths,`
> `# so this has to be the repo root.`
You are already introducing `_REPO_ROOT` in `admin.py` and already editing all 11
call sites. Anchoring the default the same way removes the class:
```python
_REPO_ROOT = Path(__file__).resolve().parent.parent
def load_config(path: str | Path | None = None) -> RouterConfig:
path = Path(path) if path is not None else _REPO_ROOT / "config" / "config.yaml"
```
Ten of the eleven call sites then collapse to a bare `load_config()`, and the
9000/8081 smoke stops depending on cwd.
This is a small **logic** change, not a relocation, so it sits just outside the
plan's stated "pure path refactoring" boundary. Flagging it as a decision to make
on purpose — taking it or declining it are both fine; arriving at it by accident
in Wave 3 is not.
---
## Minor
### V2-7. Smoke port 9000 contradicts the project's own convention
CLAUDE.md, in the section written after the `pkill` outage:
> **Convention going forward: 8080 is production, always.** … A throwaway
> instance (manual iteration, Playwright smoke tests against the admin frontend,
> anything that isn't "use the real router") binds **8081** instead.
v2's smoke block uses 9000 throughout. Use 8081, or amend the CLAUDE.md
convention in Wave 5 — but don't leave the repo asserting two different answers,
given what that convention was written to prevent.
### V2-8. Commit count is 6, not 7
Waves 1-5 produce 5 commits; Wave 6 states outright that it generates none
(untracked files); Wave 7 produces 1. The TL;DR says 7.
### V2-9. Root target is exactly 17
`51 − 29 + 1 (src) − 4 + 1 (config) − 4 + 1 (plans) = 17`. v2 says "≈ 16". Make
the success-criteria checkbox an exact number so it can actually be checked.
### V2-10. Update the stale comment in the unit files
`deploy/llm-router.service` carries `# config.yaml, router.db and router.log are
all referenced as relative paths, so this has to be the repo root.` It becomes
`config/config.yaml`. The unit files are already in scope for Wave 7.
### V2-11. `code_plans/` is ~320 files, not "100 glob results"
The glob was truncated. It doesn't matter any more — v2 wisely dropped the
`find plans/ -type f | wc -l = 40` criterion in favour of "all tracked files, no
files deleted", which is both correct and checkable. Noting it only so the number
isn't reused elsewhere. (59 tracked, ~260 ignored inside a nested `.omo/` tree.)
---
## Round 1 findings — resolution
| # | Round-1 finding | v2 status |
|---|---|---|
| B1 | `src.` prefix contradicts flat-module premise | **Fixed.** All `src.` prefixes gone; "must NOT use `src.` prefixes" is now an explicit guard. |
| B2 | `packages=find` installs an empty wheel | **Fixed by removal.** No packaging at all; PYTHONPATH instead, with the rationale stated. `src/__init__.py` explicitly forbidden. |
| B3 | Live systemd units not the repo's; service running | **Partly.** Now a first-class wave with stop/reload/verify and a unit backup — but see V2-1 (install mechanism) and V2-4 (ordering). |
| B4 | `admin.py:445-448` frontend paths | **Fixed** via `_REPO_ROOT`. Test assertion still missed — V2-3. |
| B5 | `admin.py:662` subprocess cwd | **Fixed** via `_REPO_ROOT`. |
| F1 | Fabricated `router_cli.py` section | **Fixed.** Section deleted; `router_cli.py` correctly absent from the config call-site list. |
| F2 | Missing `leaderboard.py:135` | **Fixed**, and flagged as "missed in v1". |
| F3 | Test counts low | **Fixed and exceeded.** v2 found 7 bare sites in `test_proficiency.py` (191, 207, 230, 260, 284, 310, 337) — verified accurate, three more than Round 1 listed. |
| F4 | `git rm` on gitignored backups | **Fixed.** Plain `rm`, no commit, explicitly noted. |
| F5 | Wave 6 file counts wrong | **Adequately fixed** — bad number survives but the criterion no longer depends on it (V2-11). |
| F6 | Three different root-entry numbers | **Fixed** to one number; off by one (V2-9). |
| S1 | Unresolved reasoning left in the document | **Fixed.** Clean throughout; single wave ordering. |
| S2 | 733 tests can't catch B3/B4/B5 | **Fixed.** Real `uvicorn` + `/admin/` + trigger-POST smoke per wave. |
| S3 | No rollback procedure | **Fixed.** Per-wave SHA table, pre-flight SHA, and a unit-file backup/restore for the one non-git wave. |
Also correctly carried over: the `config/__init__.py` guard (confirmed — a
`config/` directory does not shadow `src/config.py`, because PEP 420 records it as
a namespace *portion* and a regular module found later on `sys.path` wins; adding
`__init__.py` would flip that to path order).
---
## Recommended edits before execution
1. **V2-1** — revert `deploy/*.service` to the `%h/llm-router` placeholder; install via the README's sed loop, not `cp`.
2. **V2-2** — add `env={**os.environ, "PYTHONPATH": str(ROOT / "src")}` to the two `subprocess.run` probes; add both files to Wave 2's list.
3. **V2-3** — `tests/test_admin_frontend.py:53-58` → `.parent.parent`; add to Wave 2.
4. **V2-4** — stop the two timers at pre-flight; move the unit reinstall to directly after Wave 3 and correct the dependency matrix.
5. **V2-5** — fold the `pythonpath` change into Wave 2's commit.
6. **V2-6** — decide explicitly on the `config.py` absolute-default change.
7. **V2-7 – V2-10** — port 8081, commit count 6, root target 17, unit comment.
With 1-5 applied this is ready to execute. It would also survive the dual
high-accuracy review now, if you still want that pass — the reference list is
sound, which is what was missing before.
---
# Round 1 — v1
## Blockers — will break
### B1. The `src.` prefix and the flat-module premise are mutually exclusive
The plan asserts both, in the same wave:
- Wave 2 "must NOT do": *"Do NOT change any `.py` module's internal imports (bare
`from config import ...` continues to work because editable install puts
`config.py` on sys.path)."*
- Wave 2 acceptance #4/#5, Wave 3 QA, success criteria: `uvicorn src.dispatcher:app`,
`python -m src.poller`, `from src.config import load_config`.
These cannot both hold. With `package-dir = {"": "src"}` the modules install as
**top-level** names — `config`, `dispatcher`, `metrics`. There is no `src.`
namespace at all. To get `src.dispatcher` you would need a package.
Plan must pick one world: flat modules with bare imports, or a `src/` package
with `src.` prefixes. The repo uses bare imports everywhere, so the correct path
is the former.
**Fix:** drop every `src.` prefix; `uvicorn dispatcher:app`, `python -m poller`,
`from config import load_config`, etc.
### B2. Wave 1 installs an empty wheel
The plan says:
> `pyproject.toml` has a `[project]` or `[tool.setuptools]` section with
> `packages = find` or `packages = ["src"]`, `package-dir = {"": "src"}`.
> ... `src/__init__.py` is created
`find_packages()` will not discover flat `.py` files directly under `src/`. It
finds directories with `__init__.py`. With flat modules and `package_dir={"": "src"}`,
the correct setuptools directive is `py-modules = ["admin", "baseline_report", ...]`
(or use a package). `packages = find` produces an empty distribution.
I verified this in a fresh venv: the resulting wheel installed nothing, so
`import dispatcher` failed. The plan's Wave 1 QA (`import src.dispatcher`) would
also fail because `src.dispatcher` does not exist when flat modules are installed
as top-level names.
**Fix:** either go to a proper package `src/llm_router/` (and rewrite every
import), or skip packaging entirely and use `PYTHONPATH=src` in systemd units +
pytest `pythonpath`. For a service that isn't distributed as a package,
PYTHONPATH is simpler and avoids the editable-install footgun in `CLAUDE.md`'s
venv warning.
### B3. The live systemd units aren't the repo's
`~/.config/systemd/user/llm-router.service` uses `WorkingDirectory=%h/Sources/6krrt`
and is already active. `deploy/llm-router.service` still says `%h/llm-router`.
The plan updates deploy/*.service but says nothing about installing the updated
units on the live machine. The production service is active now and will restart
on `Restart=always` after any failure. If the deploy units are not copied to
`~/.config/systemd/user/`, the next restart will use the old units against the new
`src/` layout and crash-loop on port 8080 (which opencode.json uses).
**Fix:** add a first-class production step: stop timers/service, copy updated
units to `~/.config/systemd/user/`, daemon-reload, start, verify `/health`.
### B4. admin.py serves frontend files from the wrong directory after the move
`admin.py` resolves `_module_dir = Path(__file__).resolve().parent` and uses it
for the four `admin/frontend/*.html` paths. After moving `admin.py` into `src/`,
this resolves to `src/admin/frontend/` instead of `admin/frontend/`, so all admin
pages 404.
The plan mentions the subprocess cwd but not the frontend paths.
### B5. admin.py subprocess cwd is wrong
Line 662 sets `_repo_root = str(Path(__file__).resolve().parent)` which becomes
`src/` after the move. Maintenance scripts are run from `cwd=_repo_root`, but the
config file and `router.db` are at the repo root, so spawned scripts will look for
`config.yaml` in `src/`.
**Fix:** same fix as B4 — compute repo root, not file parent.
## Fabrication / unverified references
- The plan quotes `router_cli.py` as doing `config_py = Path(__file__).resolve().parent / "config.yaml"`.
That code does not exist; `router_cli.py` has no `__file__` path resolution and
no `load_config` call. This error is the tell that the reference list was not
actually checked against the files.
- Missing `leaderboard.py:135` `load_config("config.yaml")`.
- Claimed "9 bare-path test sites, not 2" is correct; the plan says 2.
- Claimed "23 affected test files" is correct; the plan says "~20".
- The 14 `config.yaml.bak.*` are gitignored, so `git rm` cannot work — the plan
says `git rm`.
- `code_plans/` holds ~320 files (59 tracked + ~260 ignored nested artifacts), not
38 / 100 / 40 as variously stated.
## What it got right
- The real problem is the 29 loose `.py` files and 14 backups.
- Moving source to `src/` (flat modules, bare imports) is the right conceptual
answer for this codebase.
- The `config/` directory does not shadow `src/config.py` (PEP 420 namespace vs
regular module resolution), as long as `config/__init__.py` is never added.
- No top-level basename collisions across the four doc dirs.
## Recommended revisions before execution
1. Resolve B1: strike every `src.` prefix; modules stay top-level.
2. Fix B2: `py-modules` not `packages`; **delete** the `src/__init__.py` criterion;
make Wave 1's QA assert the install from outside the repo root. Or take the
`PYTHONPATH` route and drop Wave 1.
3. Add B3 as a first-class wave step with the stop/copy/daemon-reload/start/verify
sequence, and reconcile `deploy/*.service` `WorkingDirectory` with the live units.
4. Fold B4 + B5 + `dispatcher.py:2968` into one `_REPO_ROOT` change in `admin.py`.
5. Delete the fabricated `router_cli.py` section; add `leaderboard.py:135`; re-grep
the test sites for all four path spellings.
6. Rewrite Wave 5 as `rm`, not `git rm`; recount Wave 6; pick one root-entry
number.
7. Strip the in-document reasoning and settle on one wave ordering.
8. Add the 8081 smoke checks and a per-wave `git reset --hard` + unit-reinstall
rollback.
**Not ready for the dual high-accuracy review** — that would burn a review pass on
a document whose reference list is unverified. Fix 1-6 first; 7-8 can land in the
same revision.

View File

@@ -3,8 +3,8 @@
**What it was reviewing:** the shipped admin portal — `admin.py`,
`admin_schema.sql`, `admin/frontend/index.html`, the `dispatcher.py` mount
point, and 9 new test files — against the two documents that set its bar:
`code_plans/admin-portal-gap-analysis.md` (15 pre-implementation findings)
and `code_reviews/router-admin-portal-plan-review.md` (the plan review's
`plans/admin-portal-gap-analysis.md` (15 pre-implementation findings)
and `plans/router-admin-portal-plan-review.md` (the plan review's
4-item "don't approve as-is" list). All four of that list's items are
checked below against the actual code, not the plan's description of it.

View File

@@ -1,13 +1,13 @@
# Review: `router-admin-portal` work plan (pre-implementation)
**What it was reviewing:** `code_plans/.omo/plans/router-admin-portal.md` —
**What it was reviewing:** `plans/.omo/plans/router-admin-portal.md` —
a 10-todo plan to add an integrated `/admin` FastAPI web portal (dashboards
+ operator controls: refresh catalog, reseed energy, restart service,
runtime toggles, allowlisted config writes, model-availability overrides).
Reviewed before any implementation exists, against the actual current code
rather than the plan's own descriptions, and cross-checked against
`code_plans/admin-portal-gap-analysis.md` (a 15-finding critique of an
earlier draft, `code_plans/.omo/drafts/router-admin-portal.md`) to see
`plans/admin-portal-gap-analysis.md` (a 15-finding critique of an
earlier draft, `plans/.omo/drafts/router-admin-portal.md`) to see
whether the final plan actually incorporated what that analysis found.
## Verdict: citations are excellent; one likely functional gap and one unexamined security assumption before this should be approved as-is

View File

@@ -1,9 +1,9 @@
# session-cache-and-baseline-comparator review
Source: manual review of the batch built from
[`session-classification-cache-ttl.md`](../code_plans/session-classification-cache-ttl.md),
[`baseline-routing-comparator.md`](../code_plans/baseline-routing-comparator.md),
and [`classifier-input-scope-check.md`](../code_plans/classifier-input-scope-check.md).
[`session-classification-cache-ttl.md`](session-classification-cache-ttl.md),
[`baseline-routing-comparator.md`](baseline-routing-comparator.md),
and [`classifier-input-scope-check.md`](classifier-input-scope-check.md).
Full suite verified green independently (638 passed). One component is clean
and committed; the other has a bug confirmed live against `router.db`, not
just read off the diff — and is not actually committed yet, despite the

View File

@@ -2,13 +2,13 @@
**Origin.** PR #3 collapsed four independent `messages` → `image_url` part
traversals into one generator, `capabilities.iter_image_url_values`
(`code_reviews/capability-gate-followups.md`, item #6). The plan specified
(`plans/capability-gate-followups.md`, item #6). The plan specified
the two shapes the four originals agreed on (`image_url` as a dict with
`url`, or as a bare string) but not the case where a part is typed
`image_url` and carries neither — well-typed, but nothing to extract. The
merged helper silently skipped that case instead of yielding a placeholder,
which dropped a degenerate part out of two fail-closed checks at once (see
`code_reviews/pr-3-capability-gate-followups-review.md` for the live
`plans/pr-3-capability-gate-followups-review.md` for the live
repro and the fix). The follow-up opencode session that implemented the fix
named the general lesson explicitly, which is what this document is
formalizing: **a plan that proposes collapsing N implementations into one
@@ -94,7 +94,7 @@ This is a discipline to apply going forward, not a code change. Concretely:
### 1. Template for future plans
When a `code_plans/` (or `code_reviews/`) entry proposes extracting a
When a `plans/` (or `plans/`) entry proposes extracting a
shared helper from multiple call sites, include a table shaped like this
alongside the proposed code:
@@ -134,7 +134,7 @@ not a finding.
Keep this document as the reference the next relevant plan cites, rather
than duplicating the rule inline each time. A future plan proposing a
shared-helper extraction should link back to this file
(`code_plans/shared-helper-unparseable-input-contract.md`) and fill in the
(`plans/shared-helper-unparseable-input-contract.md`) and fill in the
contract table from §1 as part of the plan, the same way
`code_reviews/capability-gate-followups.md` gave concrete before/after code
`plans/capability-gate-followups.md` gave concrete before/after code
for each finding rather than describing it in prose only.

View File

Before

Width:  |  Height:  |  Size: 599 KiB

After

Width:  |  Height:  |  Size: 599 KiB

View File

Before

Width:  |  Height:  |  Size: 601 KiB

After

Width:  |  Height:  |  Size: 601 KiB

View File

@@ -0,0 +1,73 @@
# Plan: tui-panel-polish
**What you'll get:** All 5 panels get consistent bordered containers with
visible gaps. Verdict Mix popup goes from 40→50 chars so the table renders
properly.
**Why:** Panels are flat widgets (`<Static title>\n<DataTable>`) with no
containers, no margins. Only the quota panel uses a `Vertical` container.
Boring fix: wrap everything in `Vertical{border, margin-top: 1}` — same
pattern that already works for quota.
**Not doing:** No Grid, no Horizontal, no Footer, no colors/themes, no new
features, no backend changes.
**Effort:** Short | **Risk:** Low | **Files:** `tui.py` + `tui_screens.py` only
## Todos
| # | Scope |
|---|-------|
| 1 | Wrap all 5 panels in bordered `Vertical` containers with `margin-top: 1` on 4 new ones |
| 2 | CSS: `.panel-title` margin 1→2, add `margin-top: 1` selectors for new containers |
| 3 | `VerdictMixScreen` width: 50 (was 40) |
| 4 | Warnings render: target inner `#warnings-content` `Static` instead of `#warnings-panel` |
| 5 | `python -m pytest tests/test_tui.py -v` — fix DOM mismatches |
| 6 | Import smoke test: all 3 modules, exit 0 |
## Success criteria
1. All 5 panels bordered consistently
2. ≥1-line gap between each panel border
3. Warnings panel matches border style
4. `VerdictMixScreen` table renders without cramped wrapping
5. `tests/test_tui.py` — 0 failures
6. No new lint warnings
---
## Quick sanity check against the current code (not a full review — flagging one gap before execution)
Verified the plan's diagnosis directly against `tui.py`/`tui_screens.py`,
since it's cheap and this is about to be executed:
- **Confirmed accurate:** only `quota-panel` is wrapped in a `Vertical`
today (`tui.py:209`); the other four panels really are a bare
`Static(..., classes="panel-title")` immediately followed by a
`DataTable`/`Static` with no wrapping container (`tui.py:212-219`).
`VerdictMixScreen { width: 40; }` is exact (`tui_screens.py:83`).
- **Todo 4 is necessary, not optional:** `#warnings-panel` is currently the
`Static` itself (`tui.py:219`), and two call sites type it as `Static`
directly — `query_one("#warnings-panel", Static).can_focus = True`
(`tui.py:317`) and the render path at `tui.py:527`. Once todo 1 makes
`#warnings-panel` the new *outer* `Vertical` container (to keep the
`_panels`/number-key focus list at `tui.py:194-200` pointing at the right
id), both of those call sites break unless they're repointed at a new
inner `#warnings-content` `Static` — todo 4 catches exactly this, good.
- **One thing the plan doesn't address:** `DataTable` already has its own
border today — `DataTable { border: round $primary; ... }`
(`tui.py:145-149`), a blanket rule hitting all three table panels
(decision/model/category). `#quota-panel` gets its border from the
*wrapping* `Vertical` (`tui.py:150-154`), and its children
(`quota-progress`/`quota-legend`) have no border of their own — so quota
has exactly one border today. If todo 1 wraps the three `DataTable`
panels in a new bordered `Vertical` without also dropping the
`DataTable`'s own `border: round $primary`, those three panels will end
up with two nested rounded borders (the new wrapper's + the table's
existing one), while quota and warnings get one — the opposite of success
criterion 1 ("bordered consistently"). Worth folding into todo 1 or 2:
either remove the border from the blanket `DataTable` rule (letting the
new wrapper supply it uniformly) or give the wrapper no border for those
three and rely on the table's own — either works, but the plan should say
which before execution rather than leaving it to be discovered mid-todo.

View File

@@ -0,0 +1,96 @@
# Review: verdict-bar chart breaks on its first live update
**Scope:** just this one defect, introduced in `906c9e6`
("fix(admin): replace stretched verdict doughnut with partition bar and
split history into per-metric mini charts"), in
`admin/frontend/index.html`. Not a review of the rest of that commit (the
history mini-charts split was checked separately and is sound).
## Verdict: real regression, invisible on first load, breaks on every subsequent poll
`906c9e6` replaced the Verdict Mix doughnut with a horizontal "partition
bar" — one Chart.js `bar` dataset holding all N verdict counts, rendered as
a single row split into colored segments. That design depends on the chart
always having exactly **one** category on its index axis, with all N values
living inside that one dataset's `data` array.
Chart creation gets this right. `renderChart`'s `verdict-bar` branch
(`index.html:844-857`) ignores whatever `labels` it was called with and
hardcodes a single-element array:
```js
return new Chart(canvas, { type: 'bar', data: { labels: [''], datasets }, options: opts });
```
But `updateChart` — the path taken on every render *after* the first,
since `renderVerdict` calls it whenever `verdictChart` already exists
(`index.html:543`) — does not:
```js
function updateChart(name, labels, values, type, colors) { // index.html:898
if (name !== 'verdict' || !verdictChart) return;
verdictChart.data.labels = labels; // index.html:900
verdictChart.data.datasets[0].data = values;
verdictChart.data.datasets[0].backgroundColor = colors;
verdictChart.update();
}
```
`labels` here is `Object.keys(mix)` (`index.html:527`) — e.g.
`['pass', 'fail']`, one entry per verdict category, not the single-element
array the chart was created with. Line 900 overwrites `data.labels` with
that real array.
For a `stacked` bar chart on `indexAxis: 'y'` with a single dataset,
Chart.js uses `data.labels.length` to decide how many category rows to
draw. `stacked: true` (`index.html:854-855`) only merges multiple
*datasets* that share an index position — it does nothing to merge
multiple *values within one dataset* onto a single row. So the moment
`data.labels` goes from `['']` to `['pass', 'fail']`, the same one dataset
that used to render as one bar with two colored segments instead renders
as **two separate bars**, one per label. The "partition bar" concept only
holds together as long as `labels` stays a single blank entry, which is
exactly the invariant `updateChart` breaks.
## Why this passed the round of QA that caught the *last* verdict-chart bug
The previous regression (`verdictChart` never being assigned, causing
"Canvas is already in use") was caught by explicitly waiting 36+ seconds to
span a `REFRESH_MS` poll cycle
(`plans/.omo/evidence/task-8-admin-visual-fixes-v3-verdict.md`). That
discipline wasn't applied to `906c9e6` — it's a pure frontend diff with no
Playwright run, no screenshot, and no unit test attached
(`git show 906c9e6 --stat` touches only `admin/frontend/index.html`). A
single fresh-load screenshot would show a correct-looking single bar, since
first render always goes through `renderChart`, not `updateChart` — the
bug only shows up starting from the second render, i.e. the first 30-second
poll or SSE-triggered refresh after page load. Same blind spot as the last
bug, in a new function.
## Fix
`updateChart` is only ever called for `'verdict'` (the `name !== 'verdict'`
guard at the top makes it single-purpose), so there's no other caller
relying on the `data.labels` assignment. Drop it:
```js
function updateChart(name, labels, values, type, colors) {
if (name !== 'verdict' || !verdictChart) return;
verdictChart.data.datasets[0].data = values;
verdictChart.data.datasets[0].backgroundColor = colors;
verdictChart.update();
}
```
`labels` becomes an unused parameter at that point; either drop it from the
signature and its one call site (`index.html:543`), or leave it for
signature symmetry with `renderChart` — cosmetic, doesn't affect
correctness either way.
## Verification recommendation
A screenshot alone won't catch this class of bug twice. Confirm the fix the
same way the last one was confirmed: load `/admin/`, capture the verdict
bar on first render, wait past one `REFRESH_MS` cycle (30s+), capture again,
and diff — the bar should still be one row with the same segment count
before and after, not N separate rows.

Binary file not shown.

After

Width:  |  Height:  |  Size: 100 KiB

BIN
plans/viewport-1400-mid.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 102 KiB

BIN
plans/viewport-1400-v2.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 129 KiB

BIN
plans/viewport-1400.png Normal file

Binary file not shown.

After

Width:  |  Height:  |  Size: 78 KiB

Some files were not shown because too many files have changed in this diff Show More