# AGENTS.md — working guide for AI agents on this repo The three-tier docs: - **`README.md`** — concise front door. Deep-dive module-by-module walkthroughs are in **`docs/`** (13 reference files: `admin-portal.md`, `api.md`, `architecture.md`, `clients.md`, `config-local-overlay.md`, `data-model.md`, `evaluation.md`, `incidents.md`, `local-models.md`, `operations.md`, `pinch.md`, `routing.md`, `verification.md`). The README defers to `docs/` for module-level detail. - **`CLAUDE.md`** — working state + immediate next steps + design rationale (the one to trust on what is currently true). - **`design/local-llm-model-router.md`** — architecture and rationale, including parts still unbuilt. This file is the quick-reference for an agent starting work: where things live, what's safe to touch, and the conventions that aren't obvious from the code. ## Worktrees: where dispatched work happens (read before editing anything) **`/home/alee/Sources/6krrt` is the main checkout, and it is production.** The live router restarts from its `config/config.yaml`, and the live `router.db` and `config/config.local.yaml` live there. Its uncommitted files are the owner's. Your session opens in that directory by default, even when your task belongs to a worktree. So a task that names no worktree edits production files. That happened on 2026-09-28: a worker rewrote the main checkout's live `config.yaml`. **If you are a worker:** - Your task names a worktree under `/home/alee/Sources/6krrt-worktrees/`. `cd` there first, and read and edit only paths under it. - Check where you are before your first edit: `git -C rev-parse --show-toplevel` must print the worktree, not `/home/alee/Sources/6krrt`. - If your task names no worktree, stop and ask the orchestrator. Do not default to the main checkout. - Never edit, commit, stash, reset or check out anything in the main checkout unless your task explicitly says "main checkout". **If you dispatch work (Atlas, any orchestrator):** - Begin every `task()` prompt with this line: `WORKTREE: . cd there first; never edit under /home/alee/Sources/6krrt/.` Include it even for a small or read-only task. - Give every `task()` a `category` or `subagent_type`. A call with neither creates no worker, yet still reports "completed". - Never dispatch to `oh-my-claudecode:*` agents. They are pinned to Anthropic models this opencode cannot reach, and return "completed" with zero work. - Run one worker at a time per worktree, unless the plan says the files are disjoint. - Before ticking a todo, confirm its commit exists: `git -C log --oneline`. Additionally: - A worker, after committing, runs `python3 scripts/verify_commit.py --repo HEAD` and pastes its output before reporting done. - An orchestrator, before reporting a wave done, runs `scripts/oc_dispatch_audit.py --worktree ` and pastes its output. - A guardrail block is a rule, not an error to route around: fix the call, do not work around the plugin. - A test of code that calls an external service (Ollama, a provider, the opencode API) fakes it at the HTTP boundary (`requests.post`, `fetch`), never by replacing the project's own wrapper. On 2026-09-29 the local_decision dispatcher bug (category names passed where option letters were expected) passed tests that faked `classify_choice` and failed every live call. The Git hygiene section below says why the main checkout is shared. ## Stack snapshot | Dimension | Value | |---|---| | Language | Python 3.10+ (3.10 floor tested; 3.14 also verified) | | Framework | FastAPI + uvicorn | | Database | SQLite (`router.db`) | | Config | `config/config.yaml` + Pydantic (`src/config.py`), `extra="forbid"` | | HTTP client | `requests` (pinned in `requirements.txt`) — do NOT add httpx2/aiohttp without a requirements bump | | TUI | `textual==8.2.8` — imported only by `tui*.py` modules, never by the dispatch path | | Testing | `pytest`, 1923 tests, all offline (no provider or local-model calls) | | Dependencies | Pinned. Bump deliberately, never use `>=` | ## Module map and file boundaries ### Dispatcher / service (I/O: DB + network) | File | Role | Agent notes | |---|---|---| | `dispatcher.py` | FastAPI service: routes, calls providers, logs, streams | ~2500 LOC; only edit the specific function you need. **Must not import `metrics`** (circular). Imports `events` for SSE fan-out. | | `metrics.py` | Read-only aggregations for `/health` and `/metrics` | Takes `(conn, cfg)` args — **must not import `dispatcher`** to avoid circular import. | | `events.py` | In-memory decision-event broker | Pure stdlib (`queue`, `collections.deque`). Thread-safe. No Textual import. The SSE endpoint lives in `dispatcher.py` and calls `events.subscribe()`/`publish_decision()`. | | `config.py` | Pydantic models + YAML loader | All config models inherit `StrictModel` (`extra="forbid"`). Unknown keys fail at load. | | `capabilities.py` | Request-side capability detection (`tools`, `images`, `json_mode`, `reasoning`) | Reads from the OpenAI-format request body, not from the classifier. | | `context_prune.py` | Relevance-based context pruning (`pinch`) | Ships **enabled** by default (`pinch.enabled: true`; `pinch.relevance.enabled: true`; budget-gated: the embed only fires when the conversation exceeds `budget_tokens`). Imports `PinchConfig` from `config.py` (no circular import: `config.py` doesn't import `context_prune`). | | `logs.py` | Structured logging (logfmt, journald, ContextVar trace ids) | `logs.bind()` exists for StreamingResponse generators that lose the ContextVar. | ### Pure scoring/routing modules (no I/O) | File | Role | |---|---| | `scoring.py` | `normalize_inverted` + weighted composite (quality-first, cost as tiebreak) | | `routing.py` | Hard filters (`select_candidates`) + ranking (`rank_candidates`) | | `tiering.py` | Pure tier resolver: 1=cheap+small, 2=mid, 3=frontier | | `proficiency.py` | Score blending: leaderboard + self-eval → weighted composite | | `iteration.py` | Retry budget per tier, matching retry to failure kind | ### Admin portal | File | Role | |---|---| | `admin.py` | FastAPI sub-router mounted by dispatcher at `/admin`; serves API + static HTML | | `admin/frontend/index.html` | Dashboard: quota chip, per-model usage, verdict mix, category breakdown, history mini-charts, recent decisions | | `admin/frontend/models.html` | Model availability table with override dropdown | | `admin/frontend/decisions.html` | Full decision log with filter + search | | `admin/frontend/controls.html` | Operational triggers, runtime knobs, persisted config editor | | `admin_schema.sql` | RBAC/audit schema ready for future auth work (config/admin_schema.sql) | When editing the admin UI, follow **`plans/admin-design-standards.md`** — design tokens, glass styling, responsive rules, and hard-won gotchas (e.g. navbar z-index, Chart.js canvas reuse, scroll-context bug). Follow **`plans/admin-work-framework.md`** for the *method* — surface-level diagnosis, grammar-first rule extraction, the pytest + Playwright-1400/800 + REFRESH_MS-idle evidence bar, and writing the commit as a standalone learning input. Key transferable rules: a settings list is one CSS-grid row (key | meta | fixed 176px control) — never restate a control's own value in a meta column, and a single-line button card goes full width with buttons in `.card-actions`, not into a `col-xl-6` beside a tall card. ### TUI modules (import `textual`; never imported by the dispatch path) | File | Role | |---|---| | `tui.py` | `DashboardApp` — main Textual app with CSS, bindings, compose, render | | `tui_model.py` | Pure data layer: `build_model`, `build_category_breakdown`, `decision_row` — no Textual import, testable without a terminal | | `tui_sse.py` | Background-thread SSE consumer (`DecisionStream`): reconnects on failure, marshals decisions to UI thread via `call_from_thread` | | `tui_screens.py` | `DecisionDetailScreen` — `ModalScreen` showing full decision JSON (Enter or `e` key) | ### Other entrypoints | File | Role | |---|---| | `poller.py` | Fetches Neuralwatt catalog, upserts `models` table; also refreshes `provider='ollama-local'` rows each poll via `upsert_local_dispatch_models` | | `seed_energy.py` | Reference workload sweep → `energy_observations` | | `seed_local_dispatch_energy.py` | Local reference-shape sweep → per-token rates for `ollama-local` rows | | `eval_proficiency.py` | Self-eval harness → `proficiency` table | | `feedback.py` | Folds verification failures into `proficiency` | | `tier.py` | DB tiering pass | | `router_cli.py` | One-shot `/route` probe | | `proficiency_store.py` | DB write path for `proficiency` | ## Conventions ### Code style - The codebase uses `Optional[X]` (not `X | None`) and `dict` return types throughout — match the existing style in the file you're editing. - No `# type: ignore`, no `as any`, no `@ts-ignore` equivalent. - `except Exception` is acceptable at top-level boundaries with `# noqa: BLE001` comment (see `dispatcher._refresh`, `persist_route_decision`). - Config is strict: every knob belongs in `config/config.yaml`, not only in a Pydantic default. A default the file never mentions is invisible to a tuner. - Named constants use `Final` in new code (`events.py`); existing code is inconsistent — don't refactor just for this. ### Import discipline - `metrics.py` **must not import `dispatcher`** (circular import). - `events.py` imports nothing from the project (pure stdlib). - `tui_model.py` imports `requests` but **not `textual`** — it's the testable data layer. - `tui_sse.py` imports `requests` but **not `textual`** — it's a thread worker. - `tui.py` and `tui_screens.py` import `textual` — that's fine, they're TUI. - The service dispatch path (`dispatcher.py` → provider call) **never touches `textual`**. ### Database - SQLite, `PRAGMA foreign_keys = ON`. - `_db()` in `dispatcher.py` returns a `sqlite3.Connection` with `row_factory = sqlite3.Row`. - Schema in `config/schema.sql` uses `CREATE TABLE IF NOT EXISTS` — safe to re-run. - Code-side table creation (`ensure_route_decisions`, `proficiency_store.ensure_columns`) mirrors the schema for live DBs that predate a feature. ### Testing - All tests are offline. No test calls a provider or local model. - Pure modules take rows and config as arguments, so they're testable without DB/network. - Test files that exercise FastAPI use `starlette.testclient.TestClient` with a temp SQLite DB (`tests/test_metrics_endpoint.py`). - TUI tests use `App.run_test()` with a stubbed fetcher and a `_no_real_network` safety-net fixture that guards both `requests.get` and the SSE consumer. - **SSE tests must not use `TestClient.stream()` on the infinite endpoint** — it hangs. Test the generator directly (`dispatcher._decision_event_stream()`) or stub the generator to a bounded one for header/route checks. ## Shared scripts (`scripts/`) Mechanics every agent on this repo ends up doing by hand, wrapped once so they are done the same way each time. Agent-agnostic -- plain scripts, plain text output. See `scripts/README.md` for the reasoning behind each. | script | what it does | |---|---| | `scripts/preflight.sh` | **run this before starting any multi-step work.** Prints the shared stash stack, then refuses a shared-checkout / wrong-branch / dirty / stale start | | `scripts/sandbox.sh start\|stop\|status` | throwaway router on **8081** for exercising the admin portal against real data; never binds 8080, only kills a pid it started itself | | `scripts/mkpr.sh ""` | cut a gitea PR with a long markdown body (`tea` has no `--description-file`), refusing an unpushed or non-ahead branch | Prefer these over hand-rolling the equivalent command: each encodes a constraint that is easy to get wrong once and expensive to get wrong twice (production port, kill-by-name, WAL-safe DB copies, overlay protection). ## The rule that matters most **When a tracker, hook, test, or the working tree disagrees with your own account of what you did, assume YOUR ACCOUNT is wrong first.** Check the artifact before concluding the tool is broken, and never modify tracking state to stop a signal you have not disproven. Both of the expensive failures on 2026-09-10 had one shape: ``` external signal contradicts internal belief -> conclude the tool is broken -> modify state to silence it ``` - A clean working tree contradicted "I made edits", so the conclusion was "subagent edits are not persisting" and subagents were re-fired twice. The work was in `stash@{0}` the whole time. The tree was telling the truth. - A continuation hook contradicted "the work is complete", so the conclusion was "the plan tracker cannot count" and the proposed fix was to clear the field the hook reads. Three of seven todos were marked done and had never been started -- two of them were not in the PR's diff at all. The hook was telling the truth. The tracker being genuinely buggy does not make the signal wrong. Fix the mechanism AND correct the state it was misreporting; a fixed parser reading 7/7 against work that is 4/7 done is worse than a broken one reading 0/7, because it turns a loud wrong signal into a quiet one. ## Git hygiene **If the working tree looks unexpectedly clean, run `git stash list` BEFORE concluding anything about your tooling.** That single line is here because of a real run on 2026-09-10. An executor stashed its own partial work with a descriptive message, came back to a clean tree, concluded "subagent edits are not persisting", and re-fired subagents twice against that theory. The work was sitting in `stash@{0}` the whole time. An unexpectedly clean tree is far more often a stash than a broken editor. `scripts/preflight.sh` prints the stack unconditionally for this reason. **Never `git stash` or `git stash pop` bare.** This repo has ~20 worktrees on this machine and they all share ONE stash stack. Another session pushing an entry shifts every index, and `pop` deletes on success -- so a wrong index is destructive and unrecoverable. Always: ```bash git stash push -u -m '<unique-tag>' # push with a tag you can find git stash list --format='%gd %H %gs' # capture the SHA git stash apply <SHA> # apply by SHA, never pop by index ``` **Work in a worktree, not the shared checkout.** `/home/alee/Sources/6krrt` is what the operator uses from their own terminal. Mutating it collides with their `git pull` mid-run, and a leftover branch there sends your commits somewhere nobody is looking -- in the same 2026-09-10 run, execution happened on `docs/admin-screenshots`, a merged PR's branch, against a tree missing two already-merged PRs. **The plan is the contract.** If you believe an approved plan item is wrong, STOP and say so. Do not implement the alternative. The same run re-derived a design that had been explicitly reviewed and rejected, because the rejected version sounded more reasonable in isolation -- the reasoning against it was in the plan's own "Must NOT" list. ## How to run things ```bash # Setup python -m venv .venv && source .venv/bin/activate pip install -r requirements.txt sqlite3 router.db < config/schema.sql cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root) # Port 8080 is the systemd-managed PRODUCTION instance. opencode's own # model traffic goes through it (see opencode.json's baseURL) and # `Restart=always` resurrects it ~5s after any kill — so never # `pkill`/`kill` anything matching uvicorn/dispatcher/8080 to "free the # port". That fights the supervisor, and on this repo it can cut off your # own inference mid-task. See CLAUDE.md's "A bare SIGTERM could hang the # process forever" section for the incident this note comes from. # # To pick up a dispatcher.py change on the real instance: systemctl --user restart llm-router.service # For an ad hoc/manual run — iterating with --reload, a throwaway instance # for Playwright smoke tests against the admin frontend, anything that # isn't "use the real router" — bind a different port so it can't collide: PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081 # Run the TUI (service must be running) PYTHONPATH=src python -m tui # Run the full test suite python -m pytest # Quick routing probe (no spend) PYTHONPATH=src python -m router_cli "Refactor this Django view" # Populate the catalog PYTHONPATH=src python -m poller && PYTHONPATH=src python -m tier ``` ## Gitea / `tea` CLI (not `gh`) This repo is hosted on **self-hosted Gitea** (`git.adlee.work/alee/6krrt`). The GitHub CLI `gh` is **not installed** and will fail — use `tea` instead. - **Login:** `tea login` as `alee`. The `alee` login is NOT the default tea login — always pass `--remote origin` so tea resolves the right context. - **Remote:** `origin` resolves to `git.adlee.work/alee/6krrt`. - **List open PRs:** `tea pulls --remote origin` - **PR metadata (JSON):** ``` tea pulls <PR_INDEX> --remote origin --fields index,state,draft,title,mergeable,base,head,body -o json ``` Use the `base` and `head` from this output for local diff commands. - **PR creation / update / maintenance:** use the `tea pr` family (e.g. `tea pr create --base main --title ... --description ...`). Confirm against `tea pr --help` for the current flag set. - **Post a comment to a PR** (uses the `tea api` proxy to Gitea's REST API): ``` tea api --remote origin "/repos/{owner}/{repo}/issues/<PR_INDEX>/comments" -F body=@- ``` - **CRITICAL:** use `-F` (typed field), **not** `-f`. `-f key=@file` writes the literal string `@file`; only `-F` reads stdin or the contents of a file. - The endpoint **must** include the `/repos/` prefix — omitting it returns a 404. - `tea api` already prints JSON to stdout — do **not** pass `-o json` on top of that (you would write the body to a file literally named `json`). - Fallback if stdin is awkward: `-F body=@/tmp/review-body.md`. - **Getting a PR diff:** prefer local `git diff origin/<base>...origin/<head>` (base and head from the PR JSON). In tea v0.14.1 the `diff`/`patch` fields return empty even when requested — do not rely on them. - **Gitea source links** (not GitHub `blob` links): `https://git.adlee.work/alee/6krrt/src/commit/<FULL-SHA>/<path>#L<start>-L<end>` (segment is `src/commit`, not `blob`). - **Code-review plugin:** this repo ships a project-local fork of the code-review plugin (`code-review-tea`) that already wires these `tea` calls. Prefer invoking it (e.g. `/code-review-tea:code-review`) over hand-rolling. ## What's NOT built yet (open items) 1. **Leaderboard priors unfilled** — `leaderboards.yaml` ships empty. 2. **Three models unsettled** on eco stability (attribution noise). 3. **Retry does not reach streaming** — `POST /outcome` is the answer for streamed traffic. 4. **Local energy not on the ledger** — local classifier/verifier electricity is unmeasured. 5. **Session-directory attribution picks the wrong directory** — heuristic resolves to a dependency's source dir instead of the project being edited. See `CLAUDE.md` → "What's NOT built yet — pick up here" for the full list. ## After a code change - Run `python -m pytest` — 1923 tests, ~35s. - If you changed `dispatcher.py`, restart the systemd service: `systemctl --user restart llm-router.service` (it doesn't auto-reload code). - If you changed the TUI, run `PYTHONPATH=src python -m tui` to verify it starts. - Check `lsp_diagnostics` on changed files. - Match the existing commit-message style: `fix:`, `feat:`, `docs:`, `test:`.