diff --git a/.gitignore b/.gitignore index d23c91e..4de9287 100644 --- a/.gitignore +++ b/.gitignore @@ -7,3 +7,11 @@ router.log .venv/ venv/ .omo/ +config.yaml.bak.* +node_modules/ +package.json +package-lock.json +.playwright-mcp/ +.mypy_cache/ +.pytest_cache/ +.ruff_cache/ diff --git a/AGENTS.md b/AGENTS.md index 568b657..2388947 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -18,10 +18,10 @@ code. | Language | Python 3.10+ (3.10 floor tested; 3.14 also verified) | | Framework | FastAPI + uvicorn | | Database | SQLite (`router.db`) | -| Config | `config.yaml` + Pydantic (`config.py`), `extra="forbid"` | +| Config | `config/config.yaml` + Pydantic (`src/config.py`), `extra="forbid"` | | HTTP client | `requests` (pinned in `requirements.txt`) — do NOT add httpx2/aiohttp without a requirements bump | | TUI | `textual==8.2.8` — imported only by `tui*.py` modules, never by the dispatch path | -| Testing | `pytest`, 562 tests, all offline (no provider or local-model calls) | +| Testing | `pytest`, 733 tests, all offline (no provider or local-model calls) | | Dependencies | Pinned. Bump deliberately, never use `>=` | ## Module map and file boundaries @@ -48,6 +48,19 @@ code. | `proficiency.py` | Score blending: leaderboard + self-eval → weighted composite | | `iteration.py` | Retry budget per tier, matching retry to failure kind | +### Admin portal + +| File | Role | +|---|---| +| `admin.py` | FastAPI sub-router mounted by dispatcher at `/admin`; serves API + static HTML | +| `admin/frontend/index.html` | Dashboard: quota chip, per-model usage, verdict mix, category breakdown, history mini-charts, recent decisions | +| `admin/frontend/models.html` | Model availability table with override dropdown | +| `admin/frontend/decisions.html` | Full decision log with filter + search | +| `admin/frontend/controls.html` | Operational triggers, runtime knobs, persisted config editor | +| `admin_schema.sql` | RBAC/audit schema ready for future auth work (config/admin_schema.sql) | + +When editing the admin UI, follow **`plans/admin-design-standards.md`** — design tokens, glass styling, responsive rules, and hard-won gotchas (e.g. navbar z-index, Chart.js canvas reuse, scroll-context bug). Follow **`plans/admin-work-framework.md`** for the *method* — surface-level diagnosis, grammar-first rule extraction, the pytest + Playwright-1400/800 + REFRESH_MS-idle evidence bar, and writing the commit as a standalone learning input. Key transferable rules: a settings list is one CSS-grid row (key | meta | fixed 176px control) — never restate a control's own value in a meta column, and a single-line button card goes full width with buttons in `.card-actions`, not into a `col-xl-6` beside a tall card. + ### TUI modules (import `textual`; never imported by the dispatch path) | File | Role | @@ -77,7 +90,7 @@ code. - No `# type: ignore`, no `as any`, no `@ts-ignore` equivalent. - `except Exception` is acceptable at top-level boundaries with `# noqa: BLE001` comment (see `dispatcher._refresh`, `persist_route_decision`). -- Config is strict: every knob belongs in `config.yaml`, not only in a +- Config is strict: every knob belongs in `config/config.yaml`, not only in a Pydantic default. A default the file never mentions is invisible to a tuner. - Named constants use `Final` in new code (`events.py`); existing code is inconsistent — don't refactor just for this. @@ -95,7 +108,7 @@ code. - SQLite, `PRAGMA foreign_keys = ON`. - `_db()` in `dispatcher.py` returns a `sqlite3.Connection` with `row_factory = sqlite3.Row`. -- Schema in `schema.sql` uses `CREATE TABLE IF NOT EXISTS` — safe to re-run. +- Schema in `config/schema.sql` uses `CREATE TABLE IF NOT EXISTS` — safe to re-run. - Code-side table creation (`ensure_route_decisions`, `proficiency_store.ensure_columns`) mirrors the schema for live DBs that predate a feature. @@ -116,25 +129,36 @@ code. # Setup python -m venv .venv && source .venv/bin/activate pip install -r requirements.txt -sqlite3 router.db < schema.sql -cp .env.example .env # fill in NEURALWATT_API_KEY +sqlite3 router.db < config/schema.sql +cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root) -# Start the service (binds 127.0.0.1:8080) -python -m uvicorn dispatcher:app --reload -# Or via systemd: -systemctl --user start llm-router.service +# Port 8080 is the systemd-managed PRODUCTION instance. opencode's own +# model traffic goes through it (see opencode.json's baseURL) and +# `Restart=always` resurrects it ~5s after any kill — so never +# `pkill`/`kill` anything matching uvicorn/dispatcher/8080 to "free the +# port". That fights the supervisor, and on this repo it can cut off your +# own inference mid-task. See CLAUDE.md's "A bare SIGTERM could hang the +# process forever" section for the incident this note comes from. +# +# To pick up a dispatcher.py change on the real instance: +systemctl --user restart llm-router.service + +# For an ad hoc/manual run — iterating with --reload, a throwaway instance +# for Playwright smoke tests against the admin frontend, anything that +# isn't "use the real router" — bind a different port so it can't collide: +PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081 # Run the TUI (service must be running) -python tui.py +PYTHONPATH=src python -m tui # Run the full test suite python -m pytest # Quick routing probe (no spend) -python router_cli.py "Refactor this Django view" +PYTHONPATH=src python -m router_cli "Refactor this Django view" # Populate the catalog -python poller.py && python tier.py +PYTHONPATH=src python -m poller && PYTHONPATH=src python -m tier ``` ## What's NOT built yet (open items) @@ -149,9 +173,9 @@ See `CLAUDE.md` → "What's NOT built yet — pick up here" for the full list. ## After a code change -- Run `python -m pytest` — 562 tests, ~30s. +- Run `python -m pytest` — 733 tests, ~35s. - If you changed `dispatcher.py`, restart the systemd service: `systemctl --user restart llm-router.service` (it doesn't auto-reload code). -- If you changed the TUI, run `python tui.py` to verify it starts. +- If you changed the TUI, run `PYTHONPATH=src python -m tui` to verify it starts. - Check `lsp_diagnostics` on changed files. - Match the existing commit-message style: `fix:`, `feat:`, `docs:`, `test:`. diff --git a/CLAUDE.md b/CLAUDE.md index fc3b62d..1b61355 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -196,15 +196,15 @@ rather than from months of history. ## What's built and working -- `schema.sql` — `models`, `proficiency`, `energy_observations`. Applies - cleanly (`sqlite3 router.db < schema.sql`). `models` carries the serving +- `config/schema.sql` — `models`, `proficiency`, `energy_observations`. Applies + cleanly (`sqlite3 router.db < config/schema.sql`). `models` carries the serving class columns (below); `energy_observations` carries real carbon/cost. - `poller.py` — fetches Neuralwatt's `/models` endpoint (public, unauthenticated), normalizes, upserts, marks stale rows. **Verified against the live API**: 19 models, and the `metadata.pricing` / `metadata.capabilities` / `metadata.limits` field mappings are confirmed correct. -- `config.yaml` / `config.py` — weights, thresholds, provider settings, +- `config/config.yaml` / `src/config.py` — weights, thresholds, provider settings, Pydantic-validated. - `scoring.py` — one `normalize_inverted` (cost and eco normalize identically; they differ only in what is fed to them) plus the weighted @@ -212,7 +212,7 @@ rather than from months of history. - `seed_energy.py` — runs a fixed reference task N times per routable model and writes `energy_observations` rows tagged `seed_reference`. This is what makes `cost` and `eco` real numbers instead of the neutral 0.5. Re-run it - after the catalog gains models: `python seed_energy.py --samples 5` + after the catalog gains models: `PYTHONPATH=src python -m seed_energy --samples 5` (13 models x 5 = 65 calls, and the whole sweep cost **under a cent**). - `tiering.py` / `tier.py` — pure tier resolver + the DB pass that applies it. - `routing.py` — pure hard filters and ranking. @@ -258,7 +258,7 @@ rather than from months of history. `route_decisions`, per-model aggregates over `energy_observations`, verification verdict mix, and top proficiency by category. - `tui.py` — Textual terminal dashboard over `GET /metrics` and - `GET /events/decisions`. A foreground entrypoint (`python tui.py`), not a + `GET /events/decisions`. A foreground entrypoint (`PYTHONPATH=src python -m tui`), not a service. `textual` is imported only in the TUI modules (`tui.py`, `tui_screens.py`, `tui_sse.py`), so the router's dispatch path has no UI dependency. The dashboard has a live routing-decisions feed (via the SSE @@ -270,7 +270,19 @@ rather than from months of history. - `router_cli.py` — one-shot `/route` probe. Posts a task to the running router and prints the decision tree, or emits raw JSON with `--json`. Spends no quota because it only routes. -- `tests/` — 562 tests across 27 files, all passing, all offline. Verified on +- `admin.py` / `config/admin_schema.sql` / `admin/frontend/*.html` — `/admin` management + portal served by the running router, loopback-only, no auth. Liquid-glass dark UI + (see `plans/admin-design-standards.md`). Pages: dashboard (`index.html`) + with quota chip, per-model usage, verdict-mix/category-breakdown bars, history + mini-charts and recent decisions; models (`models.html`) with availability + overrides; decisions (`decisions.html`) log with filter/search; controls + (`controls.html`) for operational triggers, runtime knobs, and persisted config + edits. SSE status dot + warnings bell are shared chrome. Operational triggers: + refresh catalog, seed energy, apply feedback, restart service. Runtime toggles + reset on restart. Persisted config edits limited to an allowlist. Model + availability overrides feed routing hard filters. Bucketed history endpoint over + `energy_observations` and `route_decisions`. +- `tests/` — 733 tests across 41 files, all passing, all offline. Verified on Python 3.10 and 3.14; nothing declares `requires-python`, so 3.10 is the tested floor rather than a promised one. @@ -289,7 +301,7 @@ failed write is logged and swallowed, because a decision record is worth having but never worth failing or slowing a request for. The table is created from code at module load (`_ensure_route_decisions_table`) and again on every write (`ensure_route_decisions` inside `persist_route_decision`), mirroring -the `proficiency_store.ensure_columns` migration pattern: `schema.sql` is + the `proficiency_store.ensure_columns` migration pattern: `config/schema.sql` is `CREATE TABLE IF NOT EXISTS`, but a live `router.db` predating this table needs the code-side migration. Both the module-load hook and the write-path guarantee are idempotent and leave existing rows intact. @@ -309,7 +321,7 @@ request body. Two new gates live in `routing.rejection_reason` and `json_schema` and `routing.require_json_mode` is true. A model passes only if `supports_json_mode = 1`; `NULL` also fails closed, for the same reason. -Both gates default to **on** in `config.yaml`, because a wrong guess produces a +Both gates default to **on** in `config/config.yaml`, because a wrong guess produces a 400. This is a deliberate asymmetry against the tool-proficiency gate below: capability **flags** fail closed on unknown, while quality **measurements** (tool proficiency, energy) admit on absent evidence ("unproven, not bad"). @@ -336,11 +348,11 @@ so failing early is better than an opaque provider 400. ### Local Ollama vision fallback -`local_vision:` in `config.yaml` configures a fallback path for image requests +`local_vision:` in `config/config.yaml` configures a fallback path for image requests that find no cloud vision candidate — the cloud catalog excludes the cost leader (deepseek) on vision, so without this fallback every image request that would otherwise have routed there 422s instead. It is a **core feature**, -**enabled by default** (`enabled: true`, both in `config.yaml` and in +**enabled by default** (`enabled: true`, both in `config/config.yaml` and in `LocalVisionConfig`'s own default, so a config that omits the section still gets it). Disable it explicitly (`enabled: false`) on a host with no local Ollama, or one that hasn't pulled the vision model. @@ -1043,21 +1055,21 @@ in the same pass — it was declared, never read, and shadowed the ```bash python -m venv .venv && source .venv/bin/activate pip install -r requirements.txt -sqlite3 router.db < schema.sql -cp .env.example .env # fill in NEURALWATT_API_KEY -python poller.py # populate the catalog -python tier.py # resolve tiers -python config.py # sanity-check config loads -python -m uvicorn dispatcher:app --reload +sqlite3 router.db < config/schema.sql +cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root) +PYTHONPATH=src python -m poller # populate the catalog +PYTHONPATH=src python -m tier # resolve tiers +PYTHONPATH=src python -m config # sanity-check config loads +PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081 ``` -Then set what is deployment-specific in `config.yaml`: `classifier.model` and +Then set what is deployment-specific in `config/config.yaml`: `classifier.model` and `classifier.base_url` for your Ollama, and `objective.plan_kwh_per_period` to your own plan's quota (it is reported in `/health` as burn against the allowance; it does not gate anything). Ollama must be reachable with the classifier model pulled — the name must -match `classifier.model` in `config.yaml`, which points at a **Modelfile-tagged +match `classifier.model` in `config/config.yaml`, which points at a **Modelfile-tagged variant**, not the base library tag. Ollama loads a model at its library Modelfile's default context unless told otherwise, and the base tag never was — measured live on a 24GB card, `mistral-nemo:12b` alone came up at @@ -1192,12 +1204,51 @@ are still honored correctly regardless of that policy (systemd tracks deliberate stops separately from the Restart= decision), so this only adds self-healing for the unexplained case. -The signal's actual source is still open — live investigation and audit -trail in `code_plans/router-unreachable-signal-investigation.md`. Confirmed -so far: it is not `systemctl`, not the admin portal's restart trigger, not -suspend/resume, not the OOM killer, and — via `auditd` — not delivered -through the `kill` or `tgkill` syscalls either, which is why the watch was -extended to `pidfd_send_signal` and the `rt_*sigqueueinfo` syscalls. +**Resolved 2026-08-29 — the source was opencode, killing its own supply +line.** Full audit trail in +`plans/router-unreachable-signal-investigation.md`; the syscall-level +extension to `pidfd_send_signal` (past `kill`/`tgkill`, both cleared by +`auditd`) is what finally caught it. `sudo ausearch -k routerkill` matched +five incidents in one afternoon to `pidfd_send_signal(..., SIGTERM)` / +`SIGKILL`-after-escalation calls from a non-interactive `zsh -c "pkill ..."` +(never in `~/.zsh_history`, since it's not a login shell), and +`~/.local/share/opencode/log/opencode.log` matched every one of those +timestamps, to the millisecond, to one opencode run testing the admin +frontend: `pkill -f "uvicorn dispatcher:app"` (or `.*dispatcher`, or plain +`"8080"`), then `python -m uvicorn dispatcher:app --host 127.0.0.1 --port +8080` for its own throwaway instance. When the port came back occupied 5s +later (`Restart=always` resurrecting the real service), the agent read that +as "the kill didn't work" and escalated to `pkill -9` — the one case +(16:01:00) that arrived as a bare `SIGKILL`, `status=9/KILL`, rather than a +caught `SIGTERM`. + +This was self-inflicted in a sharper way than it looks: `opencode.json` +points opencode's *own* model traffic at `http://127.0.0.1:8080/v1` — the +same production instance it was killing to test against. Every kill briefly +cut off the agent's own inference supply. + +The actual bug was in `AGENTS.md`, not in this service: its "How to run +things" section told an agent to bring the router up with a bare +`python -m uvicorn dispatcher:app --reload` (implicitly on 8080, no port +flag) as an equally-valid alternative to `systemctl --user start`, with no +warning that 8080 is normally already held by the supervised instance. An +agent following that instruction and finding the port taken has no way to +know the right move is `systemctl --user restart` (which the doc *does* say +two sections later, for the "changed dispatcher.py" case, but not for "I +want to smoke-test against a running instance"). Fixed there: manual/ad hoc +runs now bind `--port 8081` explicitly, and the section says outright not to +`pkill`/`kill` anything matching `uvicorn`/`dispatcher`/`8080` — that's the +systemd-managed instance, `Restart=always` will fight you, and on this repo +it may be your own model access. + +**Convention going forward: 8080 is production, always.** It's the port +baked into `opencode.json`, every curl example in this file, the systemd +unit, and the admin frontend's own fetches — moving it would touch more +surface than the problem is worth. A throwaway instance (manual iteration, +Playwright smoke tests against the admin frontend, anything that isn't "use +the real router") binds **8081** instead, and nothing should ever send a +kill signal to a process matched by name/port rather than by a PID it +started itself. ## Pointing a coding agent at it diff --git a/README.md b/README.md index 8578ac9..4e22881 100644 --- a/README.md +++ b/README.md @@ -30,9 +30,14 @@ measurement is how you check whether it still holds for you. - **Fall back to local vision when no cloud row supports images.** If no vision-capable catalog candidate survives the hard filters, the router proxies the request to a local Ollama vision model instead of returning 422. -- **Watch decisions arrive live.** `python tui.py` opens a terminal dashboard +- **Watch decisions arrive live.** `PYTHONPATH=src python -m tui` opens a terminal dashboard that follows `/events/decisions` as decisions are recorded, with no polling delay. +- **Manage it from a browser.** `GET /admin/` serves a four-page glass dark-mode + portal from the running router: a dashboard (quota burn, per-model usage, + verdict mix, history), a model-availability table with routing overrides, a + searchable decision log, and a controls page for operational triggers, runtime + toggles, and allowlisted `config/config.yaml` edits. Loopback-only, no auth. - **Verify before learning.** Every routed response is structurally parsed in the background; larger prose answers get an async local-LLM spot-check, and failures fold back into per-model proficiency through `feedback.py`. @@ -54,6 +59,7 @@ measurement is how you check whether it still holds for you. - [Ask an image question](#ask-an-image-question) - [Force JSON output](#force-json-output) - [Watch it live](#watch-it-live) + - [Admin web portal](#admin-web-portal) - [Probe routing without spending](#probe-routing-without-spending) - [At a Glance](#at-a-glance) - [Verification Pipeline](#verification-pipeline) @@ -97,15 +103,15 @@ Nothing else is assumed about the host — routing itself is SQLite and arithmet ```bash python -m venv .venv && source .venv/bin/activate pip install -r requirements.txt -sqlite3 router.db < schema.sql -cp .env.example .env # fill in NEURALWATT_API_KEY -python poller.py # populate the catalog -python tier.py # resolve tiers -python config.py # sanity-check config loads -python -m uvicorn dispatcher:app --reload +sqlite3 router.db < config/schema.sql +cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root) +PYTHONPATH=src python -m poller # populate the catalog +PYTHONPATH=src python -m tier # resolve tiers +PYTHONPATH=src python -m config # sanity-check config loads +PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081 ``` -Then edit `config.yaml` for your own setup — at minimum: +Then edit `config/config.yaml` for your own setup — at minimum: | Key | Why | |---|---| @@ -115,7 +121,7 @@ Then edit `config.yaml` for your own setup — at minimum: | `objective.assumed_cache_rate` | 0.917 was measured from one client's traffic (40.7M tokens). Check yours against the provider's per-session cache-hit figures | | `session_cache.enabled` | off by default; caches category/tier per session for `staleness_minutes` to skip repeat classifier round-trips on long agent sessions | -`python seed_energy.py` is optional. It sweeps a fixed reference workload to +`PYTHONPATH=src python -m seed_energy` is optional. It sweeps a fixed reference workload to populate `eco`, which is logged but is not an objective — routing works without it. It costs real money and quota, so it is not in the path above. @@ -127,9 +133,9 @@ Five user units cover continuous dispatch, catalog polling, and periodic energy |---|---|---| | `llm-router.service` | Continuous | FastAPI dispatcher | | `llm-router-poller.timer` | 2 min after boot, then every 2 h | triggers the poller unit | -| `llm-router-poller.service` | oneshot | `poller.py` → `tier.py` | +| `llm-router-poller.service` | oneshot | `python -m poller` → `python -m tier` | | `llm-router-seed.timer` | Every 6 h | triggers the seed unit | -| `llm-router-seed.service` | oneshot | small `seed_energy.py` sweep | +| `llm-router-seed.service` | oneshot | small `python -m seed_energy` sweep | **The poller timer is load-bearing, not optional.** `freshness.stale_after_days` is 3 with `exclude_stale: true` — an unpolled catalog marks every row stale @@ -250,7 +256,7 @@ curl -s -X POST localhost:8080/v1/chat/completions -H 'content-type: application - Only inline `data:` URIs are accepted; remote `http(s)` image URLs are declined to avoid SSRF. - Image count and total payload size are bounded before the local call is made. -- Configure the fallback in `config.yaml` under `local_vision:`. +- Configure the fallback in `config/config.yaml` under `local_vision:`. ### Force JSON output @@ -277,23 +283,100 @@ curl -s localhost:8080/metrics | python -m json.tool curl -s localhost:8080/events/decisions # Terminal dashboard with live routing feed -python tui.py +PYTHONPATH=src python -m tui ``` - **`GET /metrics`** returns quota burn, coverage, recent decisions, per-model totals, verdict mix, and top proficiency. Loopback-only, no auth. - **`GET /events/decisions`** is a Server-Sent Events stream of routing decisions. It replays recent decisions, then streams new ones as they happen; `:heartbeat` keepalive comments keep the connection alive between events. -- **`python tui.py`** is a Textual dashboard with a live routing feed, a category → model breakdown panel, a detail popup (press Enter or `e`), and quota/per-model/verdict/warnings panels. It is a separate entrypoint, not a systemd unit. -- **`baseline_report.py --since 2026-08-01`** is a read-only retrospective that replays recent decisions against two trivial counterfactuals — always pick the cheapest eligible candidate and always pick the best-proficiency eligible candidate — and reports total/mean cost and proficiency plus a dominance share. Add `--category coding_refactor` to scope it, or `--csv` for parseable output. +- **`PYTHONPATH=src python -m tui`** is a Textual dashboard with a live routing feed, a category → model breakdown panel, a detail popup (press Enter or `e`), and quota/per-model/verdict/warnings panels. It is a separate entrypoint, not a systemd unit. +- **`baseline_report.py --since 2026-08-01`** is a read-only retrospective (`PYTHONPATH=src python -m baseline_report`) that replays recent decisions against two trivial counterfactuals — always pick the cheapest eligible candidate and always pick the best-proficiency eligible candidate — and reports total/mean cost and proficiency plus a dominance share. Add `--category coding_refactor` to scope it, or `--csv` for parseable output. `textual` is pinned in `requirements.txt` solely for the TUI modules (`tui.py`, `tui_screens.py`, `tui_sse.py`). It is imported only by these modules; the FastAPI service dispatch path never touches it, so the router itself has no UI dependency. +### Admin web portal + +`GET /admin/` serves a management portal from the running router on the same +loopback-only bind as the API. It has no auth layer yet, so like the other +endpoints it is reachable only from `127.0.0.1`. + +The portal is a **four-page, glass dark-mode** dashboard built on Tabler/Bootstrap +with a few lighter inline SVG icons (see `plans/admin-design-standards.md` +for the visual language and `plans/admin-work-framework.md` for the method +used to evolve it). It is implemented in `admin.py`, with `config/admin_schema.sql` for +its future tables and `admin/frontend/` for the browser UI. + +**Dashboard** — `GET /admin/` aggregates the router at a glance: a quota chip +(burn against `objective.plan_kwh_per_period`), per-model usage bars, verdict +mix and category breakdown, five history mini-charts, and a recent-decisions +table. + +
+ Admin dashboard +
+ +**Models** — `GET /admin/models` is the model-availability table. Each row shows +the serving class, tier, status, and an override dropdown (`active` / +`deprecated` / `stale`) that writes through to the routing hard filters. + +
+ Admin models page +
+ +**Controls** — `GET /admin/controls` holds the operational triggers, runtime +knobs, and the persisted `config/config.yaml` editor. Changes are marked as dirty and +written only on save. + +
+ Admin controls page +
+ +**Decisions** — `GET /admin/decisions` is the full decision log with +kind/category/tier filters and free-text search. + +
+ Admin decisions log +
+ +**Read-only dashboards** — `GET /admin/api/snapshot` exposes data for: + +- quota burn against `objective.plan_kwh_per_period` +- per-model usage from `energy_observations` +- live routing decisions from `route_decisions` +- verdict mix and scoring coverage +- history: `GET /admin/api/history?range=6h` (also `1h`, `24h`, `7d`, `30d`) + +**Operational triggers** — async, fire-and-forget maintenance jobs: + +- `POST /admin/api/refresh-catalog` runs the poller and tier pass +- `POST /admin/api/seed-energy?samples=N` starts a reference sweep +- `POST /admin/api/apply-feedback?dry_run=true` runs `feedback.py` +- `POST /admin/api/restart-service` restarts the running systemd unit + +**Runtime toggles** + +`GET /admin/api/runtime` shows persisted-vs-runtime values. `POST /admin/api/runtime/{knob}` +flips in-memory settings such as `log_route_decisions` and `local_llm_enabled`. +Changes take effect immediately but reset on restart. + +**Persisted config edits** + +`GET /admin/api/config` lists allowlisted keys. `POST /admin/api/config/{key}` +writes one allowlisted key back to `config/config.yaml` with a timestamped backup +and whole-config validation. Arbitrary keys are rejected. + +**Model availability overrides** + +`POST /admin/api/models/{model_id}/{provider}/availability` marks a model as +`active`, `deprecated`, or `stale`. `DELETE` on the same path removes the override. +Deprecation feeds into the routing hard filters. + ### Probe routing without spending `router_cli.py` is a one-shot shell probe that POSTs to `/route` once and prints the full decision tree. ```bash -python router_cli.py "Refactor this Django view into service objects" -python router_cli.py "Summarize this diff" --category summarization --tier 2 +PYTHONPATH=src python -m router_cli "Refactor this Django view into service objects" +PYTHONPATH=src python -m router_cli "Summarize this diff" --category summarization --tier 2 ``` - Prints the selected model, candidates, estimated cost, and estimated proficiency. @@ -377,8 +460,8 @@ Observation → learning. `feedback.py` folds verification failures into 23-task benchmark: ```bash -python feedback.py --dry-run # preview what would change -python feedback.py # apply +PYTHONPATH=src python -m feedback --dry-run # preview what would change +PYTHONPATH=src python -m feedback # apply ``` Key behaviors: @@ -458,11 +541,11 @@ Key behaviors: | **Local Classification** | Ollama, OpenAI-compatible — `localhost:11434/v1` or an Ollama across your VPN | | **Local Model** | `classifier.model` — `mistral-nemo:12b` by default; any Ollama model works | | **Cloud Provider** | Neuralwatt only | -| **Config** | `config.yaml` loaded & validated by Pydantic (`config.py`) | +| **Config** | `config/config.yaml` loaded & validated by Pydantic (`src/config.py`) | | **OpenAI Client** | `openai==3.0.0` (official SDK) | | **HTTP** | `requests` for poller, `httpx` (via openai/uvicorn) | -| **Testing** | `pytest` — 562 tests across 27 files, all offline | -| **Config Files** | `config.yaml`, `leaderboards.yaml`, `evals/tasks.yaml` | +| **Testing** | `pytest` — 733 tests across 41 files, all offline | +| **Config Files** | `config/config.yaml`, `config/leaderboards.yaml`, `evals/tasks.yaml` | | **Deployment** | systemd user units (`.service` + `.timer` files in `deploy/`) | | **Integration** | `opencode.json` in the repo routes through it by default; any OpenAI-compatible client works | @@ -505,6 +588,7 @@ restarts on boot shouldn't change its dependency tree underneath itself. | **`config.py`** | YAML loader + Pydantic validators (blend weights sum to 1, valid tiers, endpoints separately addressable) | Yes (file) | | **`metrics.py`** | Read-only aggregations for `/health` and `GET /metrics`: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency | Yes (DB) | | **`events.py`** | In-memory decision-event broker for the TUI's live feed: bounded ring buffer + thread-safe fan-out to SSE subscribers | Pure | +| **`admin.py`** | `/admin` management portal: serves the four frontend pages + read/write API (snapshot, history, models/availability, runtime knobs, allowlisted config edits, operational triggers) | Yes (DB, network, file) | | **`tui.py`** | Textual terminal dashboard over `GET /metrics` and `GET /events/decisions`; live routing feed, detail popup, category breakdown. Foreground tool, not a service | Yes (network) | | **`tui_model.py`** | Pure data layer for the TUI: `build_model`, `build_category_breakdown`, `decision_row` — no Textual import, testable without a terminal | Pure | | **`tui_sse.py`** | Background-thread SSE consumer for the TUI: reconnects on failure, marshals live decisions onto the UI thread | Yes (network) | @@ -570,7 +654,7 @@ parses them into `access_level` and `routing.allowed_access_levels` (default | `inherited_from` | TEXT | Model this row was copied from, NULL if measured directly | | `last_updated` | TEXT | ISO8601 | -**Category set** (9 categories, defined in `config.yaml`): +**Category set** (9 categories, defined in `config/config.yaml`): | Category | Example | Scoring type | |---|---|---| @@ -741,7 +825,7 @@ rows with `supports_vision = 1` survive the hard filters. If **no** cloud candidate survives, the router can fall back to a local vision model instead of returning 422. -`local_vision:` in `config.yaml` controls this path: +`local_vision:` in `config/config.yaml` controls this path: | Key | Default | Purpose | |---|---|---| @@ -753,7 +837,7 @@ of returning 422. | `max_images` | `4` | Refuse requests with more image parts | | `max_image_bytes` | `9437184` (9 MiB) | Refuse requests whose image payload exceeds this | -The fallback is **enabled by default** both in `config.yaml` and in +The fallback is **enabled by default** both in `config/config.yaml` and in `LocalVisionConfig`, so omitting the section still turns it on. Disable it explicitly (`enabled: false`) on a host with no local Ollama or one that has not pulled the vision model. @@ -883,6 +967,7 @@ allowance. | `POST` | `/dispatch` | Same as `/route`, plus complete the provider call, stream response, log observation | | `GET` | `/v1/models` | OpenAI-compatible model list (router virtual models + catalog) | | `POST` | `/v1/chat/completions` | OpenAI-compatible completions — routes then proxies, **streaming supported** | +| `GET` | `/admin` | Loopback-only web management portal (read-only dashboards, operational triggers, runtime toggles, allowlisted config edits) | `/metrics` returns a single JSON object with these top-level keys: @@ -951,7 +1036,7 @@ rank id=r9116d9 pos=0 model=deepseek-v4-flash prof=1 est_usd=0.00016296 **No conversation text is logged at any level**, prompt or answer — prompts here run 60k–150k tokens and the journal is on disk. A test enforces it. -Set the level in `config.yaml` (`logging.level`), or override it without +Set the level in `config/config.yaml` (`logging.level`), or override it without touching a tracked file: ```bash @@ -976,10 +1061,10 @@ out clean. ## Self-Eval Harness (`eval_proficiency.py`) ```bash -python eval_proficiency.py # every routable model × every task -python eval_proficiency.py --models kimi-k3 # subset of models -python eval_proficiency.py --categories coding_general -python eval_proficiency.py --dry-run # plan only +PYTHONPATH=src python -m eval_proficiency # every routable model × every task +PYTHONPATH=src python -m eval_proficiency --models kimi-k3 # subset of models +PYTHONPATH=src python -m eval_proficiency --categories coding_general +PYTHONPATH=src python -m eval_proficiency --dry-run # plan only ``` - Runs every task through the target provider, scores it, writes to `proficiency`. @@ -1037,7 +1122,7 @@ Several settings keep it from cascading failures: unparseable output. A coding agent would rather have a mid-tier answer than an error. Escalation deliberately skips fallbacks so an unavailable local model doesn't silently promote every request to the frontier tier. -- **Session classification cache** (`session_cache:` in `config.yaml`, +- **Session classification cache** (`session_cache:` in `config/config.yaml`, off by default) remembers the last `task_category`/`task_tier` decision per session for `staleness_minutes` (default 20), so a long agent session skips the classifier round-trip on every turn. It only ever short-circuits the @@ -1068,7 +1153,7 @@ carries the same modality block. ## Testing ```bash -python -m pytest # 562 tests +python -m pytest # 733 tests python -m pytest --cov # with coverage ``` @@ -1093,6 +1178,7 @@ config as arguments, so the suite runs offline on a clean checkout. | `test_session_identity.py` | Outcome attribution: session matching, ambiguity refusal | | `test_config_endpoints.py` | Classifier and verifier are separately addressable; guards on the split | | `test_metrics_endpoint.py` | `/metrics` endpoint, SSE `/events/decisions` headers + replay/stream behavior | +| `test_admin_frontend.py` | Admin portal serves the four HTML pages with their expected markers and Chart.js asset | | `test_events.py` | Decision-event broker: publish, subscribe/replay, unsubscribe, full-subscriber eviction | | `test_tui.py` | TUI data model, category breakdown, detail popup, live SSE decision handling, keyboard controls | | `test_context_prune.py` | Context pruning: image_url handling, structured content, recency guards, stats accuracy | @@ -1119,7 +1205,7 @@ ollama create qwen3-vl-router:4b -f Modelfile.vision With these tags the combined resident footprint was measured at roughly **15.9GB** in the worst case: classifier and verifier share one `mistral-nemo-router:12b` instance at ~8.6GB, with the `qwen3-vl-router:4b` vision fallback loaded alongside it. That leaves real headroom on a 24GB card. -`config.yaml` already points at these tags by default: +`config/config.yaml` already points at these tags by default: ```yaml classifier: diff --git a/admin/frontend/controls.html b/admin/frontend/controls.html new file mode 100644 index 0000000..9746cb6 --- /dev/null +++ b/admin/frontend/controls.html @@ -0,0 +1,652 @@ + + + + + +Controls · LLM Router Admin + + + + + + + + + + + + + + +
+
+ +
+
+
+ + +
+
+
+

Operational Triggers

+
+ + + + +
+
+
+ Fire-and-forget maintenance jobs. + +
+
+
+ + +
+
+

Runtime Knobs

+
+

In-memory only — reverts on restart, no config.yaml write.

+
+
+
+
+ +
+
+
+

Persisted Config

+
+ + +
+
+
+

Allowlisted keys, written to config.yaml.

+
+
Loading config…
+
+
+
+
+ +
+
+
+ +
+
+ + + + + + + diff --git a/admin/frontend/decisions.html b/admin/frontend/decisions.html new file mode 100644 index 0000000..5f0c210 --- /dev/null +++ b/admin/frontend/decisions.html @@ -0,0 +1,553 @@ + + + + + +Decisions · LLM Router Admin + + + + + + + + + + + + + + +
+
+ +
+
+
+
+
+
+
+ + + + + + +
+ +
+
+ + + + + + + + + + +
TimeKindCategoryTierSourceModelFlagsCostProfRejected
Loading…
+
+ +
+
+
+
+
+ +
+
+ + + + diff --git a/admin/frontend/index.html b/admin/frontend/index.html index fd91224..b5e6c6d 100644 --- a/admin/frontend/index.html +++ b/admin/frontend/index.html @@ -1,292 +1,461 @@ - + -Admin Dashboard — LLM Router +Dashboard · LLM Router Admin + + + + + + -
-
-

admin@router ▸ dashboard

- - connecting… -
-
—
-
-
- -
-
-

⚡ Quota Meter

-
Loading…
-
-
-

◈ Model Availability

-
- - -
modelprovidertierstatusoverride
Loading…
-
-
-
- - -
-
-

⟁ Recent Decisions

-
- - -
timekindcategorytiermodelcostprof
-
-
-
-

▐ Per-Model Usage

-
-
Loading…
-
-
-
- - -
-
-

◉ Verdict Mix

-
-
-
-
-

◆ Category Breakdown

-
-
Loading…
-
-
-
-

⚠ Warnings

-
- -
-
-
- - -
-
-

- ◈ History - - - - - - + +

+ +
+ - -
-

⚙ Controls

- - -
-
Operational Triggers — fire-and-forget maintenance jobs
-
- - - - + +
+
+ +
+
+
- -
-
Runtime Knobs — toggle in-memory; no config.yaml write
-
-
+ +
+
+
+

Per-Model Usage

+
+ + + +
+
+
+
+
Loading…
+
+
+
+
- -
-
Persisted Config — allowlisted keys only; persisted to config.yaml
- - -
Loading config…
- - + +
+
+

Verdict Mix

+
+
+
+
+
+

Category Breakdown

+
+
+
+ + +
+
+
+

History

+
+ + + + + +
+
+
+
+
+
Decisions
+
+
+
+
Requests
+
+
+
+
Cost (USD)
+
+
+
+
Energy (kWh)
+
+
+
+
Carbon (g)
+
+
+
+
+
+
+ + +
+
+
+

Recent Decisions

+ View all → +
+
+
+ + + +
timekindcategorytiermodelcostprof
+
+
+
+
+ +
+
+
+
+
+ © 2026 adLee · 6krrt, local LLM model router · admin +
+
+
+
+ + + - -
+ + + + + + + + + + + + + + + + + +
+
+ +
+
+
+
+
+
+
+ +

Model Availability

+
+
Change the override dropdown to mark a model active, deprecated, or stale
+
+
+ + + + + + + + + + + + + +
ModelProviderTierStatusOverride
Loading…
+
+
+
+
+
+
+
+
+ © 2026 adLee · 6krrt, local LLM model router · admin +
+
+
+
+ + + + + + + diff --git a/assets/6krrt-logo-outlined.svg b/assets/6krrt-logo-outlined.svg new file mode 100644 index 0000000..6f7645f --- /dev/null +++ b/assets/6krrt-logo-outlined.svg @@ -0,0 +1,18 @@ + + \ No newline at end of file diff --git a/assets/screenshots/controls.png b/assets/screenshots/controls.png new file mode 100644 index 0000000..47533cf Binary files /dev/null and b/assets/screenshots/controls.png differ diff --git a/assets/screenshots/dashboard.png b/assets/screenshots/dashboard.png new file mode 100644 index 0000000..7408a51 Binary files /dev/null and b/assets/screenshots/dashboard.png differ diff --git a/assets/screenshots/decisions.png b/assets/screenshots/decisions.png new file mode 100644 index 0000000..0387595 Binary files /dev/null and b/assets/screenshots/decisions.png differ diff --git a/assets/screenshots/models.png b/assets/screenshots/models.png new file mode 100644 index 0000000..d50f76c Binary files /dev/null and b/assets/screenshots/models.png differ diff --git a/admin_schema.sql b/config/admin_schema.sql similarity index 100% rename from admin_schema.sql rename to config/admin_schema.sql diff --git a/config.yaml b/config/config.yaml similarity index 98% rename from config.yaml rename to config/config.yaml index c536550..b0559ce 100644 --- a/config.yaml +++ b/config/config.yaml @@ -18,7 +18,7 @@ objective: # currently rest on 2-3 samples per category, so a 0.05 gap is # indistinguishable from sampling variation and paying for it buys noise. # Narrow it as samples accumulate. - quality_tolerance: 0.10 + quality_tolerance: 0.1 # Cost is priced per-request from catalog prices, NOT from a benchmark # sweep. A fixed 400-token reference task ranked glm-5.2-fast 3.2x cheaper @@ -52,7 +52,7 @@ objective: # wall you hit mid-task. So this is the cost mandate stated as a guarantee. # For scale: the reference task runs ~5e-06 kWh on the cheapest model and # ~2.2e-04 on the most expensive. - max_energy_per_request: null + max_energy_per_request: # The subscription's kWh allowance per billing period, for reporting burn in # /health. Set to match your plan; null disables the report. NeuralWatt also @@ -115,15 +115,15 @@ proficiency: leaderboard_weight: 0.3 self_eval_weight: 0.7 categories: - - coding_general - - coding_refactor - - debugging - - docs_writing - - summarization - - translation - - reasoning_math - - tool_use_agentic - - general_chat + - coding_general + - coding_refactor + - debugging + - docs_writing + - summarization + - translation + - reasoning_math + - tool_use_agentic + - general_chat escalation: enabled: true @@ -203,7 +203,7 @@ session_cache: # Off by default, matching every other new-and-unproven knob in this # project: ship it, watch route_decisions.source="cached" on real traffic, # then decide the right default. - enabled: false + enabled: true staleness_minutes: 20 circuit_breaker: @@ -224,7 +224,7 @@ routing: # routing excludes anything not listed here. Add 'preview'/'canary' only if # the account actually holds the grant — otherwise dispatch earns a 403. allowed_access_levels: - - public + - public # '-flex' rows are held server-side during peak until a capacity gap opens. # That's correct for overnight/batch agent work and wrong for anything @@ -274,7 +274,7 @@ routing: # off, let POST /outcome report real pass/fail, and compare # tool_use_agentic proficiency for deepseek before and after. That is the # one signal here that knows whether the work actually worked. - min_tool_proficiency: null + min_tool_proficiency: tool_use_category: tool_use_agentic # Request-side capability gates. These read the request body (image parts, @@ -299,7 +299,7 @@ local_vision: # and the model must be pulled (`ollama pull qwen3-vl:4b`) on that host. enabled: true base_url: "http://localhost:11434/v1" - api_key_env: null + api_key_env: # A Modelfile-tagged variant of qwen3-vl:4b, not the base library tag. # Measured live: the base tag comes up at Ollama's own default num_ctx # (32768) and costs 9.4GB loaded — resident alongside the classifier's @@ -398,7 +398,7 @@ classifier: base_url: "http://localhost:11434/v1" # Unset means unauthenticated, which is the Ollama case. Name the env var # holding the key when the endpoint actually checks one. - api_key_env: null + api_key_env: # A Modelfile-tagged variant of mistral-nemo:12b, not the base library tag # — must match a model `ollama list` reports. Ollama loads a model at its # library Modelfile's default context unless told otherwise, and the base @@ -517,3 +517,4 @@ logging: # systemctl --user edit llm-router # Environment="LLM_ROUTER_LOG_LEVEL=debug" # systemctl --user restart llm-router level: info + diff --git a/leaderboards.yaml b/config/leaderboards.yaml similarity index 100% rename from leaderboards.yaml rename to config/leaderboards.yaml diff --git a/schema.sql b/config/schema.sql similarity index 100% rename from schema.sql rename to config/schema.sql diff --git a/deploy/README.md b/deploy/README.md index 8bea6d8..0ac26c1 100644 --- a/deploy/README.md +++ b/deploy/README.md @@ -9,9 +9,9 @@ all. | file | what it does | |---|---| | `llm-router.service` | the FastAPI dispatcher, on `127.0.0.1:8080` | -| `llm-router-poller.service` | one-shot: `poller.py` then `tier.py` | +| `llm-router-poller.service` | one-shot: `PYTHONPATH=src python -m poller` then `PYTHONPATH=src python -m tier` | | `llm-router-poller.timer` | fires the poller 2 min after boot, then every 2 h | -| `llm-router-seed.service` | one-shot: a small `seed_energy.py` reference sweep | +| `llm-router-seed.service` | one-shot: a small `PYTHONPATH=src python -m seed_energy` reference sweep | | `llm-router-seed.timer` | every 6 h — energy attribution drifts with pool load across hours, so the median has to span time rather than one sweep | These are **user** units — no root, and they run as you with your own @@ -37,6 +37,10 @@ done systemctl --user daemon-reload systemctl --user enable --now llm-router.service llm-router-poller.timer llm-router-seed.timer +# Note: the units rely on `Environment=PYTHONPATH=%h/llm-router/src` (rewritten +# by the same sed to your REPO) so the modules under `src/` are importable +# without an editable install. .env is read from the repo root and stays there. + # 3. Survive logout/reboot (user units stop with your session otherwise) loginctl enable-linger "$USER" @@ -54,17 +58,17 @@ journalctl --user -u 'llm-router*' -f # everything, live (quote the g journalctl --user -u llm-router -f -o cat # the request log, message only journalctl --user -u llm-router -p warning # fallbacks, retries, refusals journalctl --user -u llm-router-poller.service # catalog refreshes -systemctl --user restart llm-router.service # after editing config.yaml +systemctl --user restart llm-router.service # after editing config/config.yaml systemctl --user start llm-router-poller.service # force a refresh now ``` -`config.yaml` is read once at startup, so weight and threshold changes need a +`config/config.yaml` is read once at startup, so weight and threshold changes need a restart. The catalog is read per-request, so a poller run takes effect immediately. ### Turning up the logs -`logging.level` in config.yaml is the documented setting, but flipping it means +`logging.level` in config/config.yaml is the documented setting, but flipping it means editing a tracked file. For a running service use a drop-in instead: ```bash @@ -91,7 +95,7 @@ priorities when systemd owns its stderr (`SyslogLevelPrefix` is on by default). A foreground `uvicorn` prints them clean, so the same binary is readable either way. -**The oneshot units buffer.** `poller.py` and `seed_energy.py` print progress +**The oneshot units buffer.** `python -m poller` and `python -m seed_energy` print progress with plain `print()`, and Python block-buffers stdout when it is not a terminal, so their output arrives in one dump at exit rather than progressively. Add `Environment="PYTHONUNBUFFERED=1"` to those units if you @@ -123,7 +127,7 @@ sudo nano /etc/systemd/system/ollama.service.d/override.conf sudo systemctl daemon-reload && sudo systemctl restart ollama ``` -On the **client** host, in `config.yaml`: +On the **client** host, in `config/config.yaml`: ```yaml classifier: diff --git a/deploy/llm-router-poller.service b/deploy/llm-router-poller.service index 39dd4f7..3a09f30 100644 --- a/deploy/llm-router-poller.service +++ b/deploy/llm-router-poller.service @@ -7,12 +7,13 @@ Wants=network-online.target [Service] Type=oneshot WorkingDirectory=%h/llm-router +Environment=PYTHONPATH=%h/llm-router/src EnvironmentFile=%h/llm-router/.env -# poller.py refreshes the catalog; tier.py re-resolves tiers from it. Tiers +# poller refreshes the catalog; tier re-resolves tiers from it. Tiers # are derived from cost and reasoning fields the poll may have changed, so -# they always run as a pair. -ExecStart=%h/llm-router/.venv/bin/python poller.py -ExecStart=%h/llm-router/.venv/bin/python tier.py +# they always run as a pair. PYTHONPATH above resolves src/ for `-m` imports. +ExecStart=%h/llm-router/.venv/bin/python -m poller +ExecStart=%h/llm-router/.venv/bin/python -m tier NoNewPrivileges=true PrivateTmp=true diff --git a/deploy/llm-router-seed.service b/deploy/llm-router-seed.service index db7e1a7..b78ea26 100644 --- a/deploy/llm-router-seed.service +++ b/deploy/llm-router-seed.service @@ -7,10 +7,12 @@ Wants=network-online.target [Service] Type=oneshot WorkingDirectory=%h/llm-router +Environment=PYTHONPATH=%h/llm-router/src EnvironmentFile=%h/llm-router/.env # Fewer samples per run than a manual sweep, because the point is coverage -# across TIME rather than depth at one moment — see the timer. -ExecStart=%h/llm-router/.venv/bin/python seed_energy.py --samples 3 +# across TIME rather than depth at one moment — see the timer. PYTHONPATH +# above resolves src/ for the `-m` import. +ExecStart=%h/llm-router/.venv/bin/python -m seed_energy --samples 3 NoNewPrivileges=true PrivateTmp=true diff --git a/deploy/llm-router.service b/deploy/llm-router.service index 54aa58f..8d76add 100644 --- a/deploy/llm-router.service +++ b/deploy/llm-router.service @@ -10,9 +10,12 @@ Wants=network-online.target [Service] Type=exec -# config.yaml, router.db and router.log are all referenced as relative paths, -# so this has to be the repo root. +# router.db and router.log are referenced as relative paths, so this has to be +# the repo root. config.yaml now lives under config/ and the Python modules +# under src/, so PYTHONPATH points at src/ for the dispatcher import to resolve. WorkingDirectory=%h/llm-router +# Resolves `dispatcher` (and any other src/ module) for the ExecStart below. +Environment=PYTHONPATH=%h/llm-router/src # Holds NEURALWATT_API_KEY. Create it with: # echo "NEURALWATT_API_KEY=$NEURALWATT_API_KEY" > .env && chmod 600 .env EnvironmentFile=%h/llm-router/.env diff --git a/code_plans/.playwright-mcp/console-2026-08-29T06-11-18-590Z.log b/plans/.playwright-mcp/console-2026-08-29T06-11-18-590Z.log similarity index 100% rename from code_plans/.playwright-mcp/console-2026-08-29T06-11-18-590Z.log rename to plans/.playwright-mcp/console-2026-08-29T06-11-18-590Z.log diff --git a/code_plans/.playwright-mcp/console-2026-08-29T07-07-39-971Z.log b/plans/.playwright-mcp/console-2026-08-29T07-07-39-971Z.log similarity index 100% rename from code_plans/.playwright-mcp/console-2026-08-29T07-07-39-971Z.log rename to plans/.playwright-mcp/console-2026-08-29T07-07-39-971Z.log diff --git a/code_plans/.playwright-mcp/console-2026-08-29T07-17-18-482Z.log b/plans/.playwright-mcp/console-2026-08-29T07-17-18-482Z.log similarity index 100% rename from code_plans/.playwright-mcp/console-2026-08-29T07-17-18-482Z.log rename to plans/.playwright-mcp/console-2026-08-29T07-17-18-482Z.log diff --git a/code_plans/.playwright-mcp/console-2026-08-29T14-42-18-565Z.log b/plans/.playwright-mcp/console-2026-08-29T14-42-18-565Z.log similarity index 100% rename from code_plans/.playwright-mcp/console-2026-08-29T14-42-18-565Z.log rename to plans/.playwright-mcp/console-2026-08-29T14-42-18-565Z.log diff --git a/code_plans/.playwright-mcp/console-2026-08-29T14-47-56-594Z.log b/plans/.playwright-mcp/console-2026-08-29T14-47-56-594Z.log similarity index 100% rename from code_plans/.playwright-mcp/console-2026-08-29T14-47-56-594Z.log rename to plans/.playwright-mcp/console-2026-08-29T14-47-56-594Z.log diff --git a/code_plans/.playwright-mcp/console-2026-08-29T15-00-11-582Z.log b/plans/.playwright-mcp/console-2026-08-29T15-00-11-582Z.log similarity index 100% rename from code_plans/.playwright-mcp/console-2026-08-29T15-00-11-582Z.log rename to plans/.playwright-mcp/console-2026-08-29T15-00-11-582Z.log diff --git a/code_plans/.playwright-mcp/console-2026-08-29T15-01-33-174Z.log b/plans/.playwright-mcp/console-2026-08-29T15-01-33-174Z.log similarity index 100% rename from code_plans/.playwright-mcp/console-2026-08-29T15-01-33-174Z.log rename to plans/.playwright-mcp/console-2026-08-29T15-01-33-174Z.log diff --git a/code_plans/.playwright-mcp/page-2026-08-29T06-11-19-085Z.yml b/plans/.playwright-mcp/page-2026-08-29T06-11-19-085Z.yml similarity index 100% rename from code_plans/.playwright-mcp/page-2026-08-29T06-11-19-085Z.yml rename to plans/.playwright-mcp/page-2026-08-29T06-11-19-085Z.yml diff --git a/code_plans/.playwright-mcp/page-2026-08-29T06-11-46-660Z.yml b/plans/.playwright-mcp/page-2026-08-29T06-11-46-660Z.yml similarity index 100% rename from code_plans/.playwright-mcp/page-2026-08-29T06-11-46-660Z.yml rename to plans/.playwright-mcp/page-2026-08-29T06-11-46-660Z.yml diff --git a/code_plans/.playwright-mcp/page-2026-08-29T07-07-40-410Z.yml b/plans/.playwright-mcp/page-2026-08-29T07-07-40-410Z.yml similarity index 100% rename from code_plans/.playwright-mcp/page-2026-08-29T07-07-40-410Z.yml rename to plans/.playwright-mcp/page-2026-08-29T07-07-40-410Z.yml diff --git a/code_plans/.playwright-mcp/page-2026-08-29T07-17-18-936Z.yml b/plans/.playwright-mcp/page-2026-08-29T07-17-18-936Z.yml similarity index 100% rename from code_plans/.playwright-mcp/page-2026-08-29T07-17-18-936Z.yml rename to plans/.playwright-mcp/page-2026-08-29T07-17-18-936Z.yml diff --git a/code_plans/.playwright-mcp/page-2026-08-29T14-42-18-860Z.yml b/plans/.playwright-mcp/page-2026-08-29T14-42-18-860Z.yml similarity index 100% rename from code_plans/.playwright-mcp/page-2026-08-29T14-42-18-860Z.yml rename to plans/.playwright-mcp/page-2026-08-29T14-42-18-860Z.yml diff --git a/code_plans/.playwright-mcp/page-2026-08-29T14-42-58-029Z.yml b/plans/.playwright-mcp/page-2026-08-29T14-42-58-029Z.yml similarity index 100% rename from code_plans/.playwright-mcp/page-2026-08-29T14-42-58-029Z.yml rename to plans/.playwright-mcp/page-2026-08-29T14-42-58-029Z.yml diff --git a/code_plans/.playwright-mcp/page-2026-08-29T14-44-04-172Z.yml b/plans/.playwright-mcp/page-2026-08-29T14-44-04-172Z.yml similarity index 100% rename from code_plans/.playwright-mcp/page-2026-08-29T14-44-04-172Z.yml rename to plans/.playwright-mcp/page-2026-08-29T14-44-04-172Z.yml diff --git a/code_plans/.playwright-mcp/page-2026-08-29T14-47-57-000Z.yml b/plans/.playwright-mcp/page-2026-08-29T14-47-57-000Z.yml similarity index 100% rename from code_plans/.playwright-mcp/page-2026-08-29T14-47-57-000Z.yml rename to plans/.playwright-mcp/page-2026-08-29T14-47-57-000Z.yml diff --git a/code_plans/.playwright-mcp/page-2026-08-29T14-48-31-579Z.yml b/plans/.playwright-mcp/page-2026-08-29T14-48-31-579Z.yml similarity index 100% rename from code_plans/.playwright-mcp/page-2026-08-29T14-48-31-579Z.yml rename to plans/.playwright-mcp/page-2026-08-29T14-48-31-579Z.yml diff --git a/code_plans/.playwright-mcp/page-2026-08-29T15-00-11-987Z.yml b/plans/.playwright-mcp/page-2026-08-29T15-00-11-987Z.yml similarity index 100% rename from code_plans/.playwright-mcp/page-2026-08-29T15-00-11-987Z.yml rename to plans/.playwright-mcp/page-2026-08-29T15-00-11-987Z.yml diff --git a/code_plans/.playwright-mcp/page-2026-08-29T15-01-33-298Z.yml b/plans/.playwright-mcp/page-2026-08-29T15-01-33-298Z.yml similarity index 100% rename from code_plans/.playwright-mcp/page-2026-08-29T15-01-33-298Z.yml rename to plans/.playwright-mcp/page-2026-08-29T15-01-33-298Z.yml diff --git a/code_plans/.playwright-mcp/task-8-admin-visual-fixes-v2-full.png b/plans/.playwright-mcp/task-8-admin-visual-fixes-v2-full.png similarity index 100% rename from code_plans/.playwright-mcp/task-8-admin-visual-fixes-v2-full.png rename to plans/.playwright-mcp/task-8-admin-visual-fixes-v2-full.png diff --git a/code_plans/.playwright-mcp/task-8-admin-visual-fixes-v2.png b/plans/.playwright-mcp/task-8-admin-visual-fixes-v2.png similarity index 100% rename from code_plans/.playwright-mcp/task-8-admin-visual-fixes-v2.png rename to plans/.playwright-mcp/task-8-admin-visual-fixes-v2.png diff --git a/code_plans/.playwright-mcp/task-8-v2-console-clean.log b/plans/.playwright-mcp/task-8-v2-console-clean.log similarity index 100% rename from code_plans/.playwright-mcp/task-8-v2-console-clean.log rename to plans/.playwright-mcp/task-8-v2-console-clean.log diff --git a/code_plans/.playwright-mcp/task-8-v2-console-final.log b/plans/.playwright-mcp/task-8-v2-console-final.log similarity index 100% rename from code_plans/.playwright-mcp/task-8-v2-console-final.log rename to plans/.playwright-mcp/task-8-v2-console-final.log diff --git a/code_plans/.playwright-mcp/task-8-v2-console-full.log b/plans/.playwright-mcp/task-8-v2-console-full.log similarity index 100% rename from code_plans/.playwright-mcp/task-8-v2-console-full.log rename to plans/.playwright-mcp/task-8-v2-console-full.log diff --git a/plans/admin-controls-relayout.jpg b/plans/admin-controls-relayout.jpg new file mode 100644 index 0000000..6448c02 Binary files /dev/null and b/plans/admin-controls-relayout.jpg differ diff --git a/plans/admin-design-standards.md b/plans/admin-design-standards.md new file mode 100644 index 0000000..ff14e92 --- /dev/null +++ b/plans/admin-design-standards.md @@ -0,0 +1,364 @@ +# Admin Portal Design Standards + +This document captures the design system used by the Claude-built admin portal uplift (commits around `23b74a3`–`906c9e6`) so future edits to `admin/frontend/` stay visually consistent. + +Intended readers: any agent or human touching `admin/frontend/*.html`, `admin.py`, or the admin API surface. + +**Supplementary sources** (read these too before touching the admin UI): +- `.omo/notepads/admin-facelift/learnings.md` — the running log of every visual QA finding, root cause, and fix across the whole 10-wave facelift. This is the single richest record of what went wrong and how it was fixed. +- `plans/admin-facelift.md` — the original spec (survival guardrails S1–S5, device contracts). +- `plans/admin-visual-fixes-review.md` and `plans/router-admin-portal-implementation-review.md` — post-hoc audits that caught regressions the QA pass missed. + +## Philosophy + +- **Liquid-glass dashboard, not a marketing page.** Background is a subtle gradient wash; cards are translucent, blurred, and softly lit from the top. The goal is information density with low visual fatigue. +- **One visual language everywhere.** Use the same frosted-card treatment, pill buttons, and glass badges on every page so the portal feels like a single app. +- **Subtle cues over neon alerts.** Status colors are desaturated translucent tints. Warnings live in the navbar bell, not in a full card. + +## Theme + +- Always dark mode. Root is ``. +- **Base background** (body): + - Three overlapping `radial-gradient` blobs: + - amber (`rgba(245,158,11,.10)`) top-left + - blue (`rgba(59,130,246,.10)`) top-right + - purple (`rgba(139,92,246,.07)`) bottom + - Base color `#111827`. + - `background-attachment: fixed` so it stays put during scroll. + +## Layout shell (every page) + +1. Bootstrap 5 / Tabler page skeleton: + ```html + +
+
+ +
…
+
…
+
+
+ ``` +2. Scrolling context: **`` scrolls, `` does not.** This works around a Chrome/Linux phantom-margin bug caused by the HTML element also being the scrollbar container. + ```css + html { overflow:hidden; height:100%; margin:0!important; padding:0!important } + body { overflow-y:auto; height:100%; margin:0!important; padding:0!important } + ``` +3. Container width: `container-xl` inside `.page-body`. +4. Cards grid: `.row.row-cards > .col-xl-6` or `.col-12`. + +## Navbar + +- Glass background: `rgba(31,41,55,.6)` with `backdrop-filter: blur(16px) saturate(160%)`. +- Bottom border: `1px solid rgba(255,255,255,.06)`. +- Explicit `position: relative; z-index: 1030;` on `header.navbar` so its dropdowns paint above later backdrop-filter stacking contexts (e.g. the quota chip). +- Container padding pinned to exactly `20px` left/right so the brand edge stays aligned across breakpoints: + ```css + header.navbar > .container-fluid { padding-left:20px!important; padding-right:20px!important; } + ``` +- Brand wordmark: **6krrt LLM Router**, in `'Quicksand', var(--tblr-font-sans-serif)` at `font-weight: 700; letter-spacing: 0.01em`. The leading `6` is `.brand-six` — brand amber, `1.14em` — because the mascot is itself a 6. Keep the wordmark inside a single `.brand-word` span: the anchor is `inline-flex` with a `gap`, so a bare span around the `6` becomes its own flex item and that gap opens up between `6` and `krrt`. +- Logo SVG inline; force `filter:none!important` and size `64x64px` to defeat Tabler's autodark invert filter. The size lives in three places that must agree — the CSS rule, and the inline `width`/`height` attributes **and** `style="width:64px;height:64px"` on the `` (that inline style outranks the stylesheet, so changing only the CSS silently does nothing). +- **64px brand, 69px bar.** The mascot's face is ~40% of his height, so at 45px his eyes were ~4px. The extra 19px comes out of padding rather than out of the bar: `.navbar-brand` gives up its 8px top/bottom, and `header.navbar` drops 4px to 2px. Net bar height 69px, 1px *shorter* than the 45px version. Re-measure `header.navbar.offsetHeight` before touching either number. + +### Logo + +Two variants in `assets/`: + +| File | Use | +|---|---| +| `6krrt-logo.svg` | plain mascot (Inkscape source) | +| `6krrt-logo-outlined.svg` | + white keyline; **this is what the admin portal uses** | + +- The keyline is a 3.6-unit white outline around the head and tail, so the amber + silhouette separates from `#111827` and from darker browser tab bars. +- It is **not** a stroke on the amber shapes — a stroke centers on the path and + would eat half its width into the body. It is the same two shapes drawn first, + larger, in white: `circle r 42 -> 45.6`, tail `stroke-width 32 -> 39.2`. Draw + order is the whole mechanism; the white pair must stay above the amber pair. +- Eyes, smile, antenna and wheels already carry their own white rings and need + nothing added. +- Keep the keyline well under 9 units: by then the white has grown across the + counter of the "6" and closed the hole that makes the silhouette a digit. +- The favicon is the same artwork as a percent-encoded `data:image/svg+xml` URI + in each page's `` + and let `renderStaticIcons()` fill it on init. Using the literal template in + static markup renders the raw `${icon('x')}` text into the page — a real + regression that shipped once. +- **Tabler exposes Bootstrap components under the `tabler.*` global, NOT + `bootstrap.*`.** `tabler.min.js` does not expose a `bootstrap` global. + Always `new tabler.Toast(...)`, never `bootstrap.Toast(...)` (would throw + `ReferenceError`). +- **Keep a Chart.js `