363 lines
19 KiB
Markdown
363 lines
19 KiB
Markdown
# AGENTS.md — working guide for AI agents on this repo
|
|
|
|
The three-tier docs:
|
|
- **`README.md`** — concise front door. Deep-dive module-by-module walkthroughs
|
|
are in **`docs/`** (13 reference files: `admin-portal.md`, `api.md`,
|
|
`architecture.md`, `clients.md`, `config-local-overlay.md`, `data-model.md`,
|
|
`evaluation.md`, `incidents.md`, `local-models.md`, `operations.md`,
|
|
`pinch.md`, `routing.md`, `verification.md`).
|
|
The README defers to `docs/` for module-level detail.
|
|
- **`CLAUDE.md`** — working state + immediate next steps + design rationale
|
|
(the one to trust on what is currently true).
|
|
- **`design/local-llm-model-router.md`** — architecture and rationale,
|
|
including parts still unbuilt.
|
|
|
|
This file is the quick-reference for an agent starting work: where things
|
|
live, what's safe to touch, and the conventions that aren't obvious from the
|
|
code.
|
|
|
|
## Worktrees: where dispatched work happens (read before editing anything)
|
|
|
|
**`/home/alee/Sources/6krrt` is the main checkout, and it is production.** The
|
|
live router restarts from its `config/config.yaml`, and the live `router.db` and
|
|
`config/config.local.yaml` live there. Its uncommitted files are the owner's.
|
|
|
|
Your session opens in that directory by default, even when your task belongs to
|
|
a worktree. So a task that names no worktree edits production files. That
|
|
happened on 2026-09-28: a worker rewrote the main checkout's live
|
|
`config.yaml`.
|
|
|
|
**If you are a worker:**
|
|
- Your task names a worktree under `/home/alee/Sources/6krrt-worktrees/<name>`.
|
|
`cd` there first, and read and edit only paths under it.
|
|
- Check where you are before your first edit:
|
|
`git -C <worktree> rev-parse --show-toplevel` must print the worktree, not
|
|
`/home/alee/Sources/6krrt`.
|
|
- If your task names no worktree, stop and ask the orchestrator. Do not default
|
|
to the main checkout.
|
|
- Never edit, commit, stash, reset or check out anything in the main checkout
|
|
unless your task explicitly says "main checkout".
|
|
|
|
**If you dispatch work (Atlas, any orchestrator):**
|
|
- Begin every `task()` prompt with this line:
|
|
`WORKTREE: <absolute worktree path>. cd there first; never edit under /home/alee/Sources/6krrt/.`
|
|
Include it even for a small or read-only task.
|
|
- Give every `task()` a `category` or `subagent_type`. A call with neither
|
|
creates no worker, yet still reports "completed".
|
|
- Never dispatch to `oh-my-claudecode:*` agents. They are pinned to Anthropic
|
|
models this opencode cannot reach, and return "completed" with zero work.
|
|
- Run one worker at a time per worktree, unless the plan says the files are
|
|
disjoint.
|
|
- Before ticking a todo, confirm its commit exists:
|
|
`git -C <worktree> log --oneline`.
|
|
|
|
Additionally:
|
|
- A worker, after committing, runs
|
|
`python3 scripts/verify_commit.py --repo <worktree> HEAD` and pastes its
|
|
output before reporting done.
|
|
- An orchestrator, before reporting a wave done, runs
|
|
`scripts/oc_dispatch_audit.py <session> --worktree <path>` and pastes its
|
|
output.
|
|
- A guardrail block is a rule, not an error to route around: fix the call,
|
|
do not work around the plugin.
|
|
- A test of code that calls an external service (Ollama, a provider, the
|
|
opencode API) fakes it at the HTTP boundary (`requests.post`, `fetch`),
|
|
never by replacing the project's own wrapper. On 2026-09-29 the
|
|
local_decision dispatcher bug (category names passed where option letters
|
|
were expected) passed tests that faked `classify_choice` and failed every
|
|
live call.
|
|
|
|
The Git hygiene section below says why the main checkout is shared.
|
|
|
|
## Stack snapshot
|
|
|
|
| Dimension | Value |
|
|
|---|---|
|
|
| Language | Python 3.10+ (3.10 floor tested; 3.14 also verified) |
|
|
| Framework | FastAPI + uvicorn |
|
|
| Database | SQLite (`router.db`) |
|
|
| Config | `config/config.yaml` + Pydantic (`src/config.py`), `extra="forbid"` |
|
|
| HTTP client | `requests` (pinned in `requirements.txt`) — do NOT add httpx2/aiohttp without a requirements bump |
|
|
| TUI | `textual==8.2.8` — imported only by `tui*.py` modules, never by the dispatch path |
|
|
| Testing | `pytest`, 1923 tests, all offline (no provider or local-model calls) |
|
|
| Dependencies | Pinned. Bump deliberately, never use `>=` |
|
|
|
|
## Module map and file boundaries
|
|
|
|
### Dispatcher / service (I/O: DB + network)
|
|
|
|
| File | Role | Agent notes |
|
|
|---|---|---|
|
|
| `dispatcher.py` | FastAPI service: routes, calls providers, logs, streams | ~2500 LOC; only edit the specific function you need. **Must not import `metrics`** (circular). Imports `events` for SSE fan-out. |
|
|
| `metrics.py` | Read-only aggregations for `/health` and `/metrics` | Takes `(conn, cfg)` args — **must not import `dispatcher`** to avoid circular import. |
|
|
| `events.py` | In-memory decision-event broker | Pure stdlib (`queue`, `collections.deque`). Thread-safe. No Textual import. The SSE endpoint lives in `dispatcher.py` and calls `events.subscribe()`/`publish_decision()`. |
|
|
| `config.py` | Pydantic models + YAML loader | All config models inherit `StrictModel` (`extra="forbid"`). Unknown keys fail at load. |
|
|
| `capabilities.py` | Request-side capability detection (`tools`, `images`, `json_mode`, `reasoning`) | Reads from the OpenAI-format request body, not from the classifier. |
|
|
| `context_prune.py` | Relevance-based context pruning (`pinch`) | Ships **enabled** by default (`pinch.enabled: true`; `pinch.relevance.enabled: true`; budget-gated: the embed only fires when the conversation exceeds `budget_tokens`). Imports `PinchConfig` from `config.py` (no circular import: `config.py` doesn't import `context_prune`). |
|
|
| `logs.py` | Structured logging (logfmt, journald, ContextVar trace ids) | `logs.bind()` exists for StreamingResponse generators that lose the ContextVar. |
|
|
|
|
### Pure scoring/routing modules (no I/O)
|
|
|
|
| File | Role |
|
|
|---|---|
|
|
| `scoring.py` | `normalize_inverted` + weighted composite (quality-first, cost as tiebreak) |
|
|
| `routing.py` | Hard filters (`select_candidates`) + ranking (`rank_candidates`) |
|
|
| `tiering.py` | Pure tier resolver: 1=cheap+small, 2=mid, 3=frontier |
|
|
| `proficiency.py` | Score blending: leaderboard + self-eval → weighted composite |
|
|
| `iteration.py` | Retry budget per tier, matching retry to failure kind |
|
|
|
|
### Admin portal
|
|
|
|
| File | Role |
|
|
|---|---|
|
|
| `admin.py` | FastAPI sub-router mounted by dispatcher at `/admin`; serves API + static HTML |
|
|
| `admin/frontend/index.html` | Dashboard: quota chip, per-model usage, verdict mix, category breakdown, history mini-charts, recent decisions |
|
|
| `admin/frontend/models.html` | Model availability table with override dropdown |
|
|
| `admin/frontend/decisions.html` | Full decision log with filter + search |
|
|
| `admin/frontend/controls.html` | Operational triggers, runtime knobs, persisted config editor |
|
|
| `admin_schema.sql` | RBAC/audit schema ready for future auth work (config/admin_schema.sql) |
|
|
|
|
When editing the admin UI, follow **`plans/admin-design-standards.md`** — design tokens, glass styling, responsive rules, and hard-won gotchas (e.g. navbar z-index, Chart.js canvas reuse, scroll-context bug). Follow **`plans/admin-work-framework.md`** for the *method* — surface-level diagnosis, grammar-first rule extraction, the pytest + Playwright-1400/800 + REFRESH_MS-idle evidence bar, and writing the commit as a standalone learning input. Key transferable rules: a settings list is one CSS-grid row (key | meta | fixed 176px control) — never restate a control's own value in a meta column, and a single-line button card goes full width with buttons in `.card-actions`, not into a `col-xl-6` beside a tall card.
|
|
|
|
### TUI modules (import `textual`; never imported by the dispatch path)
|
|
|
|
| File | Role |
|
|
|---|---|
|
|
| `tui.py` | `DashboardApp` — main Textual app with CSS, bindings, compose, render |
|
|
| `tui_model.py` | Pure data layer: `build_model`, `build_category_breakdown`, `decision_row` — no Textual import, testable without a terminal |
|
|
| `tui_sse.py` | Background-thread SSE consumer (`DecisionStream`): reconnects on failure, marshals decisions to UI thread via `call_from_thread` |
|
|
| `tui_screens.py` | `DecisionDetailScreen` — `ModalScreen` showing full decision JSON (Enter or `e` key) |
|
|
|
|
### Other entrypoints
|
|
|
|
| File | Role |
|
|
|---|---|
|
|
| `poller.py` | Fetches Neuralwatt catalog, upserts `models` table; also refreshes `provider='ollama-local'` rows each poll via `upsert_local_dispatch_models` |
|
|
| `seed_energy.py` | Reference workload sweep → `energy_observations` |
|
|
| `seed_local_dispatch_energy.py` | Local reference-shape sweep → per-token rates for `ollama-local` rows |
|
|
| `eval_proficiency.py` | Self-eval harness → `proficiency` table |
|
|
| `feedback.py` | Folds verification failures into `proficiency` |
|
|
| `tier.py` | DB tiering pass |
|
|
| `router_cli.py` | One-shot `/route` probe |
|
|
| `proficiency_store.py` | DB write path for `proficiency` |
|
|
|
|
## Conventions
|
|
|
|
### Code style
|
|
- The codebase uses `Optional[X]` (not `X | None`) and `dict` return types
|
|
throughout — match the existing style in the file you're editing.
|
|
- No `# type: ignore`, no `as any`, no `@ts-ignore` equivalent.
|
|
- `except Exception` is acceptable at top-level boundaries with
|
|
`# noqa: BLE001` comment (see `dispatcher._refresh`, `persist_route_decision`).
|
|
- Config is strict: every knob belongs in `config/config.yaml`, not only in a
|
|
Pydantic default. A default the file never mentions is invisible to a tuner.
|
|
- Named constants use `Final` in new code (`events.py`); existing code is
|
|
inconsistent — don't refactor just for this.
|
|
|
|
### Import discipline
|
|
- `metrics.py` **must not import `dispatcher`** (circular import).
|
|
- `events.py` imports nothing from the project (pure stdlib).
|
|
- `tui_model.py` imports `requests` but **not `textual`** — it's the testable data layer.
|
|
- `tui_sse.py` imports `requests` but **not `textual`** — it's a thread worker.
|
|
- `tui.py` and `tui_screens.py` import `textual` — that's fine, they're TUI.
|
|
- The service dispatch path (`dispatcher.py` → provider call) **never touches
|
|
`textual`**.
|
|
|
|
### Database
|
|
- SQLite, `PRAGMA foreign_keys = ON`.
|
|
- `_db()` in `dispatcher.py` returns a `sqlite3.Connection` with
|
|
`row_factory = sqlite3.Row`.
|
|
- Schema in `config/schema.sql` uses `CREATE TABLE IF NOT EXISTS` — safe to re-run.
|
|
- Code-side table creation (`ensure_route_decisions`, `proficiency_store.ensure_columns`)
|
|
mirrors the schema for live DBs that predate a feature.
|
|
|
|
### Testing
|
|
- All tests are offline. No test calls a provider or local model.
|
|
- Pure modules take rows and config as arguments, so they're testable without DB/network.
|
|
- Test files that exercise FastAPI use `starlette.testclient.TestClient` with a
|
|
temp SQLite DB (`tests/test_metrics_endpoint.py`).
|
|
- TUI tests use `App.run_test()` with a stubbed fetcher and a `_no_real_network`
|
|
safety-net fixture that guards both `requests.get` and the SSE consumer.
|
|
- **SSE tests must not use `TestClient.stream()` on the infinite endpoint** — it
|
|
hangs. Test the generator directly (`dispatcher._decision_event_stream()`)
|
|
or stub the generator to a bounded one for header/route checks.
|
|
|
|
## Shared scripts (`scripts/`)
|
|
|
|
Mechanics every agent on this repo ends up doing by hand, wrapped once so
|
|
they are done the same way each time. Agent-agnostic -- plain scripts, plain
|
|
text output. See `scripts/README.md` for the reasoning behind each.
|
|
|
|
| script | what it does |
|
|
|---|---|
|
|
| `scripts/preflight.sh` | **run this before starting any multi-step work.** Prints the shared stash stack, then refuses a shared-checkout / wrong-branch / dirty / stale start |
|
|
| `scripts/sandbox.sh start\|stop\|status` | throwaway router on **8081** for exercising the admin portal against real data; never binds 8080, only kills a pid it started itself |
|
|
| `scripts/mkpr.sh <body.md> "<title>"` | cut a gitea PR with a long markdown body (`tea` has no `--description-file`), refusing an unpushed or non-ahead branch |
|
|
|
|
Prefer these over hand-rolling the equivalent command: each encodes a
|
|
constraint that is easy to get wrong once and expensive to get wrong twice
|
|
(production port, kill-by-name, WAL-safe DB copies, overlay protection).
|
|
|
|
## The rule that matters most
|
|
|
|
**When a tracker, hook, test, or the working tree disagrees with your own
|
|
account of what you did, assume YOUR ACCOUNT is wrong first.** Check the
|
|
artifact before concluding the tool is broken, and never modify tracking state
|
|
to stop a signal you have not disproven.
|
|
|
|
Both of the expensive failures on 2026-09-10 had one shape:
|
|
|
|
```
|
|
external signal contradicts internal belief
|
|
-> conclude the tool is broken
|
|
-> modify state to silence it
|
|
```
|
|
|
|
- A clean working tree contradicted "I made edits", so the conclusion was
|
|
"subagent edits are not persisting" and subagents were re-fired twice. The
|
|
work was in `stash@{0}` the whole time. The tree was telling the truth.
|
|
- A continuation hook contradicted "the work is complete", so the conclusion
|
|
was "the plan tracker cannot count" and the proposed fix was to clear the
|
|
field the hook reads. Three of seven todos were marked done and had never
|
|
been started -- two of them were not in the PR's diff at all. The hook was
|
|
telling the truth.
|
|
|
|
The tracker being genuinely buggy does not make the signal wrong. Fix the
|
|
mechanism AND correct the state it was misreporting; a fixed parser reading
|
|
7/7 against work that is 4/7 done is worse than a broken one reading 0/7,
|
|
because it turns a loud wrong signal into a quiet one.
|
|
|
|
## Git hygiene
|
|
|
|
**If the working tree looks unexpectedly clean, run `git stash list` BEFORE
|
|
concluding anything about your tooling.**
|
|
|
|
That single line is here because of a real run on 2026-09-10. An executor
|
|
stashed its own partial work with a descriptive message, came back to a clean
|
|
tree, concluded "subagent edits are not persisting", and re-fired subagents
|
|
twice against that theory. The work was sitting in `stash@{0}` the whole time.
|
|
An unexpectedly clean tree is far more often a stash than a broken editor.
|
|
`scripts/preflight.sh` prints the stack unconditionally for this reason.
|
|
|
|
**Never `git stash` or `git stash pop` bare.** This repo has ~20 worktrees on
|
|
this machine and they all share ONE stash stack. Another session pushing an
|
|
entry shifts every index, and `pop` deletes on success -- so a wrong index is
|
|
destructive and unrecoverable. Always:
|
|
|
|
```bash
|
|
git stash push -u -m '<unique-tag>' # push with a tag you can find
|
|
git stash list --format='%gd %H %gs' # capture the SHA
|
|
git stash apply <SHA> # apply by SHA, never pop by index
|
|
```
|
|
|
|
**Work in a worktree, not the shared checkout.** `/home/alee/Sources/6krrt` is
|
|
what the operator uses from their own terminal. Mutating it collides with
|
|
their `git pull` mid-run, and a leftover branch there sends your commits
|
|
somewhere nobody is looking -- in the same 2026-09-10 run, execution happened
|
|
on `docs/admin-screenshots`, a merged PR's branch, against a tree missing two
|
|
already-merged PRs.
|
|
|
|
**The plan is the contract.** If you believe an approved plan item is wrong,
|
|
STOP and say so. Do not implement the alternative. The same run re-derived a
|
|
design that had been explicitly reviewed and rejected, because the rejected
|
|
version sounded more reasonable in isolation -- the reasoning against it was
|
|
in the plan's own "Must NOT" list.
|
|
|
|
## How to run things
|
|
|
|
```bash
|
|
# Setup
|
|
python -m venv .venv && source .venv/bin/activate
|
|
pip install -r requirements.txt
|
|
sqlite3 router.db < config/schema.sql
|
|
cp .env.example .env # fill in NEURALWATT_API_KEY (.env stays at repo root)
|
|
|
|
# Port 8080 is the systemd-managed PRODUCTION instance. opencode's own
|
|
# model traffic goes through it (see opencode.json's baseURL) and
|
|
# `Restart=always` resurrects it ~5s after any kill — so never
|
|
# `pkill`/`kill` anything matching uvicorn/dispatcher/8080 to "free the
|
|
# port". That fights the supervisor, and on this repo it can cut off your
|
|
# own inference mid-task. See CLAUDE.md's "A bare SIGTERM could hang the
|
|
# process forever" section for the incident this note comes from.
|
|
#
|
|
# To pick up a dispatcher.py change on the real instance:
|
|
systemctl --user restart llm-router.service
|
|
|
|
# For an ad hoc/manual run — iterating with --reload, a throwaway instance
|
|
# for Playwright smoke tests against the admin frontend, anything that
|
|
# isn't "use the real router" — bind a different port so it can't collide:
|
|
PYTHONPATH=src python -m uvicorn dispatcher:app --reload --port 8081
|
|
|
|
# Run the TUI (service must be running)
|
|
PYTHONPATH=src python -m tui
|
|
|
|
# Run the full test suite
|
|
python -m pytest
|
|
|
|
# Quick routing probe (no spend)
|
|
PYTHONPATH=src python -m router_cli "Refactor this Django view"
|
|
|
|
# Populate the catalog
|
|
PYTHONPATH=src python -m poller && PYTHONPATH=src python -m tier
|
|
```
|
|
|
|
## Gitea / `tea` CLI (not `gh`)
|
|
|
|
This repo is hosted on **self-hosted Gitea** (`git.adlee.work/alee/6krrt`).
|
|
The GitHub CLI `gh` is **not installed** and will fail — use `tea` instead.
|
|
|
|
- **Login:** `tea login` as `alee`. The `alee` login is NOT the default tea
|
|
login — always pass `--remote origin` so tea resolves the right context.
|
|
|
|
- **Remote:** `origin` resolves to `git.adlee.work/alee/6krrt`.
|
|
|
|
- **List open PRs:** `tea pulls --remote origin`
|
|
|
|
- **PR metadata (JSON):**
|
|
```
|
|
tea pulls <PR_INDEX> --remote origin --fields index,state,draft,title,mergeable,base,head,body -o json
|
|
```
|
|
Use the `base` and `head` from this output for local diff commands.
|
|
|
|
- **PR creation / update / maintenance:** use the `tea pr` family (e.g. `tea pr create --base main --title ... --description ...`). Confirm against `tea pr --help` for the current flag set.
|
|
|
|
- **Post a comment to a PR** (uses the `tea api` proxy to Gitea's REST API):
|
|
```
|
|
tea api --remote origin "/repos/{owner}/{repo}/issues/<PR_INDEX>/comments" -F body=@-
|
|
```
|
|
- **CRITICAL:** use `-F` (typed field), **not** `-f`. `-f key=@file` writes the literal string `@file`; only `-F` reads stdin or the contents of a file.
|
|
- The endpoint **must** include the `/repos/` prefix — omitting it returns a 404.
|
|
- `tea api` already prints JSON to stdout — do **not** pass `-o json` on top of that (you would write the body to a file literally named `json`).
|
|
- Fallback if stdin is awkward: `-F body=@/tmp/review-body.md`.
|
|
|
|
- **Getting a PR diff:** prefer local `git diff origin/<base>...origin/<head>`
|
|
(base and head from the PR JSON). In tea v0.14.1 the `diff`/`patch` fields
|
|
return empty even when requested — do not rely on them.
|
|
|
|
- **Gitea source links** (not GitHub `blob` links):
|
|
`https://git.adlee.work/alee/6krrt/src/commit/<FULL-SHA>/<path>#L<start>-L<end>`
|
|
(segment is `src/commit`, not `blob`).
|
|
|
|
- **Code-review plugin:** this repo ships a project-local fork of the code-review
|
|
plugin (`code-review-tea`) that already wires these `tea` calls. Prefer
|
|
invoking it (e.g. `/code-review-tea:code-review`) over hand-rolling.
|
|
|
|
## What's NOT built yet (open items)
|
|
|
|
1. **Leaderboard priors unfilled** — `leaderboards.yaml` ships empty.
|
|
2. **Three models unsettled** on eco stability (attribution noise).
|
|
3. **Retry does not reach streaming** — `POST /outcome` is the answer for streamed traffic.
|
|
4. **Local energy not on the ledger** — local classifier/verifier electricity is unmeasured.
|
|
5. **Session-directory attribution picks the wrong directory** — heuristic resolves to a dependency's source dir instead of the project being edited.
|
|
|
|
See `CLAUDE.md` → "What's NOT built yet — pick up here" for the full list.
|
|
|
|
## After a code change
|
|
|
|
- Run `python -m pytest` — 1923 tests, ~35s.
|
|
- If you changed `dispatcher.py`, restart the systemd service:
|
|
`systemctl --user restart llm-router.service` (it doesn't auto-reload code).
|
|
- If you changed the TUI, run `PYTHONPATH=src python -m tui` to verify it starts.
|
|
- Check `lsp_diagnostics` on changed files.
|
|
- Match the existing commit-message style: `fix:`, `feat:`, `docs:`, `test:`.
|