- Bump test count from 923 to 1923 across ~90 files (README + AGENTS.md)
- Add links to docs/config-local-overlay.md and docs/pinch.md
- Rewrite single-provider ('NeuralWatt only') framing to multi-provider;
note NeuralWatt's unique power telemetry, any OpenAI-compatible carrier works
- Reconcile AGENTS.md '10 reference files' claim with 13 actual docs files
- Standardize NeuralWatt casing throughout
- Update deploy units line to mention backup/offsite snapshot timers
17 KiB
AGENTS.md — working guide for AI agents on this repo
The three-tier docs:
README.md— concise front door. Deep-dive module-by-module walkthroughs are indocs/(13 reference files:admin-portal.md,api.md,architecture.md,clients.md,config-local-overlay.md,data-model.md,evaluation.md,incidents.md,local-models.md,operations.md,pinch.md,routing.md,verification.md). The README defers todocs/for module-level detail.CLAUDE.md— working state + immediate next steps + design rationale (the one to trust on what is currently true).design/local-llm-model-router.md— architecture and rationale, including parts still unbuilt.
This file is the quick-reference for an agent starting work: where things live, what's safe to touch, and the conventions that aren't obvious from the code.
Stack snapshot
| Dimension | Value |
|---|---|
| Language | Python 3.10+ (3.10 floor tested; 3.14 also verified) |
| Framework | FastAPI + uvicorn |
| Database | SQLite (router.db) |
| Config | config/config.yaml + Pydantic (src/config.py), extra="forbid" |
| HTTP client | requests (pinned in requirements.txt) — do NOT add httpx2/aiohttp without a requirements bump |
| TUI | textual==8.2.8 — imported only by tui*.py modules, never by the dispatch path |
| Testing | pytest, 1923 tests, all offline (no provider or local-model calls) |
| Dependencies | Pinned. Bump deliberately, never use >= |
Module map and file boundaries
Dispatcher / service (I/O: DB + network)
| File | Role | Agent notes |
|---|---|---|
dispatcher.py |
FastAPI service: routes, calls providers, logs, streams | ~2500 LOC; only edit the specific function you need. Must not import metrics (circular). Imports events for SSE fan-out. |
metrics.py |
Read-only aggregations for /health and /metrics |
Takes (conn, cfg) args — must not import dispatcher to avoid circular import. |
events.py |
In-memory decision-event broker | Pure stdlib (queue, collections.deque). Thread-safe. No Textual import. The SSE endpoint lives in dispatcher.py and calls events.subscribe()/publish_decision(). |
config.py |
Pydantic models + YAML loader | All config models inherit StrictModel (extra="forbid"). Unknown keys fail at load. |
capabilities.py |
Request-side capability detection (tools, images, json_mode, reasoning) |
Reads from the OpenAI-format request body, not from the classifier. |
context_prune.py |
Relevance-based context pruning (pinch) |
Ships enabled by default (pinch.enabled: true; pinch.relevance.enabled: true; budget-gated: the embed only fires when the conversation exceeds budget_tokens). Imports PinchConfig from config.py (no circular import: config.py doesn't import context_prune). |
logs.py |
Structured logging (logfmt, journald, ContextVar trace ids) | logs.bind() exists for StreamingResponse generators that lose the ContextVar. |
Pure scoring/routing modules (no I/O)
| File | Role |
|---|---|
scoring.py |
normalize_inverted + weighted composite (quality-first, cost as tiebreak) |
routing.py |
Hard filters (select_candidates) + ranking (rank_candidates) |
tiering.py |
Pure tier resolver: 1=cheap+small, 2=mid, 3=frontier |
proficiency.py |
Score blending: leaderboard + self-eval → weighted composite |
iteration.py |
Retry budget per tier, matching retry to failure kind |
Admin portal
| File | Role |
|---|---|
admin.py |
FastAPI sub-router mounted by dispatcher at /admin; serves API + static HTML |
admin/frontend/index.html |
Dashboard: quota chip, per-model usage, verdict mix, category breakdown, history mini-charts, recent decisions |
admin/frontend/models.html |
Model availability table with override dropdown |
admin/frontend/decisions.html |
Full decision log with filter + search |
admin/frontend/controls.html |
Operational triggers, runtime knobs, persisted config editor |
admin_schema.sql |
RBAC/audit schema ready for future auth work (config/admin_schema.sql) |
When editing the admin UI, follow plans/admin-design-standards.md — design tokens, glass styling, responsive rules, and hard-won gotchas (e.g. navbar z-index, Chart.js canvas reuse, scroll-context bug). Follow plans/admin-work-framework.md for the method — surface-level diagnosis, grammar-first rule extraction, the pytest + Playwright-1400/800 + REFRESH_MS-idle evidence bar, and writing the commit as a standalone learning input. Key transferable rules: a settings list is one CSS-grid row (key | meta | fixed 176px control) — never restate a control's own value in a meta column, and a single-line button card goes full width with buttons in .card-actions, not into a col-xl-6 beside a tall card.
TUI modules (import textual; never imported by the dispatch path)
| File | Role |
|---|---|
tui.py |
DashboardApp — main Textual app with CSS, bindings, compose, render |
tui_model.py |
Pure data layer: build_model, build_category_breakdown, decision_row — no Textual import, testable without a terminal |
tui_sse.py |
Background-thread SSE consumer (DecisionStream): reconnects on failure, marshals decisions to UI thread via call_from_thread |
tui_screens.py |
DecisionDetailScreen — ModalScreen showing full decision JSON (Enter or e key) |
Other entrypoints
| File | Role |
|---|---|
poller.py |
Fetches Neuralwatt catalog, upserts models table; also refreshes provider='ollama-local' rows each poll via upsert_local_dispatch_models |
seed_energy.py |
Reference workload sweep → energy_observations |
seed_local_dispatch_energy.py |
Local reference-shape sweep → per-token rates for ollama-local rows |
eval_proficiency.py |
Self-eval harness → proficiency table |
feedback.py |
Folds verification failures into proficiency |
tier.py |
DB tiering pass |
router_cli.py |
One-shot /route probe |
proficiency_store.py |
DB write path for proficiency |
Conventions
Code style
- The codebase uses
Optional[X](notX | None) anddictreturn types throughout — match the existing style in the file you're editing. - No
# type: ignore, noas any, no@ts-ignoreequivalent. except Exceptionis acceptable at top-level boundaries with# noqa: BLE001comment (seedispatcher._refresh,persist_route_decision).- Config is strict: every knob belongs in
config/config.yaml, not only in a Pydantic default. A default the file never mentions is invisible to a tuner. - Named constants use
Finalin new code (events.py); existing code is inconsistent — don't refactor just for this.
Import discipline
metrics.pymust not importdispatcher(circular import).events.pyimports nothing from the project (pure stdlib).tui_model.pyimportsrequestsbut nottextual— it's the testable data layer.tui_sse.pyimportsrequestsbut nottextual— it's a thread worker.tui.pyandtui_screens.pyimporttextual— that's fine, they're TUI.- The service dispatch path (
dispatcher.py→ provider call) never touchestextual.
Database
- SQLite,
PRAGMA foreign_keys = ON. _db()indispatcher.pyreturns asqlite3.Connectionwithrow_factory = sqlite3.Row.- Schema in
config/schema.sqlusesCREATE TABLE IF NOT EXISTS— safe to re-run. - Code-side table creation (
ensure_route_decisions,proficiency_store.ensure_columns) mirrors the schema for live DBs that predate a feature.
Testing
- All tests are offline. No test calls a provider or local model.
- Pure modules take rows and config as arguments, so they're testable without DB/network.
- Test files that exercise FastAPI use
starlette.testclient.TestClientwith a temp SQLite DB (tests/test_metrics_endpoint.py). - TUI tests use
App.run_test()with a stubbed fetcher and a_no_real_networksafety-net fixture that guards bothrequests.getand the SSE consumer. - SSE tests must not use
TestClient.stream()on the infinite endpoint — it hangs. Test the generator directly (dispatcher._decision_event_stream()) or stub the generator to a bounded one for header/route checks.
Shared scripts (scripts/)
Mechanics every agent on this repo ends up doing by hand, wrapped once so
they are done the same way each time. Agent-agnostic -- plain scripts, plain
text output. See scripts/README.md for the reasoning behind each.
| script | what it does |
|---|---|
scripts/preflight.sh |
run this before starting any multi-step work. Prints the shared stash stack, then refuses a shared-checkout / wrong-branch / dirty / stale start |
scripts/sandbox.sh start|stop|status |
throwaway router on 8081 for exercising the admin portal against real data; never binds 8080, only kills a pid it started itself |
scripts/mkpr.sh <body.md> "<title>" |
cut a gitea PR with a long markdown body (tea has no --description-file), refusing an unpushed or non-ahead branch |
Prefer these over hand-rolling the equivalent command: each encodes a constraint that is easy to get wrong once and expensive to get wrong twice (production port, kill-by-name, WAL-safe DB copies, overlay protection).
The rule that matters most
When a tracker, hook, test, or the working tree disagrees with your own account of what you did, assume YOUR ACCOUNT is wrong first. Check the artifact before concluding the tool is broken, and never modify tracking state to stop a signal you have not disproven.
Both of the expensive failures on 2026-09-10 had one shape:
external signal contradicts internal belief
-> conclude the tool is broken
-> modify state to silence it
- A clean working tree contradicted "I made edits", so the conclusion was
"subagent edits are not persisting" and subagents were re-fired twice. The
work was in
stash@{0}the whole time. The tree was telling the truth. - A continuation hook contradicted "the work is complete", so the conclusion was "the plan tracker cannot count" and the proposed fix was to clear the field the hook reads. Three of seven todos were marked done and had never been started -- two of them were not in the PR's diff at all. The hook was telling the truth.
The tracker being genuinely buggy does not make the signal wrong. Fix the mechanism AND correct the state it was misreporting; a fixed parser reading 7/7 against work that is 4/7 done is worse than a broken one reading 0/7, because it turns a loud wrong signal into a quiet one.
Git hygiene
If the working tree looks unexpectedly clean, run git stash list BEFORE
concluding anything about your tooling.
That single line is here because of a real run on 2026-09-10. An executor
stashed its own partial work with a descriptive message, came back to a clean
tree, concluded "subagent edits are not persisting", and re-fired subagents
twice against that theory. The work was sitting in stash@{0} the whole time.
An unexpectedly clean tree is far more often a stash than a broken editor.
scripts/preflight.sh prints the stack unconditionally for this reason.
Never git stash or git stash pop bare. This repo has ~20 worktrees on
this machine and they all share ONE stash stack. Another session pushing an
entry shifts every index, and pop deletes on success -- so a wrong index is
destructive and unrecoverable. Always:
git stash push -u -m '<unique-tag>' # push with a tag you can find
git stash list --format='%gd %H %gs' # capture the SHA
git stash apply <SHA> # apply by SHA, never pop by index
Work in a worktree, not the shared checkout. /home/alee/Sources/6krrt is
what the operator uses from their own terminal. Mutating it collides with
their git pull mid-run, and a leftover branch there sends your commits
somewhere nobody is looking -- in the same 2026-09-10 run, execution happened
on docs/admin-screenshots, a merged PR's branch, against a tree missing two
already-merged PRs.
The plan is the contract. If you believe an approved plan item is wrong, STOP and say so. Do not implement the alternative. The same run re-derived a design that had been explicitly reviewed and rejected, because the rejected version sounded more reasonable in isolation -- the reasoning against it was in the plan's own "Must NOT" list.
How to run things
# Setup
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
sqlite3 router.db < config/schema.sql
cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
# Port 8080 is the systemd-managed PRODUCTION instance. opencode's own
# model traffic goes through it (see opencode.json's baseURL) and
# `Restart=always` resurrects it ~5s after any kill — so never
# `pkill`/`kill` anything matching uvicorn/dispatcher/8080 to "free the
# port". That fights the supervisor, and on this repo it can cut off your
# own inference mid-task. See CLAUDE.md's "A bare SIGTERM could hang the
# process forever" section for the incident this note comes from.
#
# To pick up a dispatcher.py change on the real instance:
systemctl --user restart llm-router.service
# For an ad hoc/manual run — iterating with --reload, a throwaway instance
# for Playwright smoke tests against the admin frontend, anything that
# isn't "use the real router" — bind a different port so it can't collide:
PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
# Run the TUI (service must be running)
PYTHONPATH=src python -m tui
# Run the full test suite
python -m pytest
# Quick routing probe (no spend)
PYTHONPATH=src python -m router_cli "Refactor this Django view"
# Populate the catalog
PYTHONPATH=src python -m poller && PYTHONPATH=src python -m tier
Gitea / tea CLI (not gh)
This repo is hosted on self-hosted Gitea (git.adlee.work/alee/6krrt).
The GitHub CLI gh is not installed and will fail — use tea instead.
-
Login:
tea loginasalee. Thealeelogin is NOT the default tea login — always pass--remote originso tea resolves the right context. -
Remote:
originresolves togit.adlee.work/alee/6krrt. -
List open PRs:
tea pulls --remote origin -
PR metadata (JSON):
tea pulls <PR_INDEX> --remote origin --fields index,state,draft,title,mergeable,base,head,body -o jsonUse the
baseandheadfrom this output for local diff commands. -
PR creation / update / maintenance: use the
tea prfamily (e.g.tea pr create --base main --title ... --description ...). Confirm againsttea pr --helpfor the current flag set. -
Post a comment to a PR (uses the
tea apiproxy to Gitea's REST API):tea api --remote origin "/repos/{owner}/{repo}/issues/<PR_INDEX>/comments" -F body=@-- CRITICAL: use
-F(typed field), not-f.-f key=@filewrites the literal string@file; only-Freads stdin or the contents of a file. - The endpoint must include the
/repos/prefix — omitting it returns a 404. tea apialready prints JSON to stdout — do not pass-o jsonon top of that (you would write the body to a file literally namedjson).- Fallback if stdin is awkward:
-F body=@/tmp/review-body.md.
- CRITICAL: use
-
Getting a PR diff: prefer local
git diff origin/<base>...origin/<head>(base and head from the PR JSON). In tea v0.14.1 thediff/patchfields return empty even when requested — do not rely on them. -
Gitea source links (not GitHub
bloblinks):https://git.adlee.work/alee/6krrt/src/commit/<FULL-SHA>/<path>#L<start>-L<end>(segment issrc/commit, notblob). -
Code-review plugin: this repo ships a project-local fork of the code-review plugin (
code-review-tea) that already wires theseteacalls. Prefer invoking it (e.g./code-review-tea:code-review) over hand-rolling.
What's NOT built yet (open items)
- Leaderboard priors unfilled —
leaderboards.yamlships empty. - Three models unsettled on eco stability (attribution noise).
- Retry does not reach streaming —
POST /outcomeis the answer for streamed traffic. - Local energy not on the ledger — local classifier/verifier electricity is unmeasured.
- Session-directory attribution picks the wrong directory — heuristic resolves to a dependency's source dir instead of the project being edited.
See CLAUDE.md → "What's NOT built yet — pick up here" for the full list.
After a code change
- Run
python -m pytest— 1923 tests, ~35s. - If you changed
dispatcher.py, restart the systemd service:systemctl --user restart llm-router.service(it doesn't auto-reload code). - If you changed the TUI, run
PYTHONPATH=src python -m tuito verify it starts. - Check
lsp_diagnosticson changed files. - Match the existing commit-message style:
fix:,feat:,docs:,test:.