Replaces stale nemotron-mini-router:4b references with the current local-dispatch model tag across README, API/client/evaluation docs, data-model notes, and CLAUDE.md. Historical measurement mentions and the benchmark table in config.yaml are preserved. Spec: plans/local-dispatch-inert-and-test-coupling.md §4
40 lines
2.0 KiB
Markdown
40 lines
2.0 KiB
Markdown
> Pointing coding agents and OpenAI-compatible clients at the router. Back to [README](../README.md).
|
||
|
||
## Pointing a Coding Agent at It
|
||
|
||
The `/v1` endpoints are OpenAI-compatible, so any normal client works —
|
||
opencode, an SDK, plain curl. A repo-local `opencode.json` is included, so
|
||
running `opencode` from a clone of this repo routes through the router by
|
||
default. For global use, merge its `provider.llm-router` block into
|
||
`~/.config/opencode/opencode.json`.
|
||
|
||
Virtual model names:
|
||
- `auto` → router picks, interactive mode (flex rows excluded)
|
||
- `auto:batch` → router picks, admits flex rows for async work
|
||
|
||
For `model: "auto"`, eligible `file_summarization` and `diff_checking` tasks may
|
||
route to the configured local model (`qwen2.5-coder-router:14b`) instead of a
|
||
cloud model. Pin the local `model_id` to force that tag. Local rows appear in
|
||
`/v1/models` with `owned_by: "ollama-local"`.
|
||
|
||
**opencode image support**: The repo's `opencode.json` declares
|
||
`"modalities": {"input": ["text", "image"]}` for every `llm-router` model.
|
||
This is required: opencode strips image parts client-side unless the provider
|
||
model declares image input. Without it, images never reach the router at all.
|
||
For global use, make sure the merged `~/.config/opencode/opencode.json` entry
|
||
carries the same modality block.
|
||
|
||
## Session-directory attribution (opencode plugin)
|
||
|
||
The opencode plugin's "most frequent path" heuristic decides which working
|
||
directory a session's traffic is attributed to. It weights **write-like**
|
||
`tool_calls` (`edit`, `write`, `patch`, `apply_patch`, `str_replace`,
|
||
`create`) at **3×**, while read-only mentions (tool results and plain text)
|
||
stay at weight 1. It **tie-breaks on the deepest directory** and **refuses
|
||
ambiguous sessions** rather than guessing.
|
||
|
||
This resolved a real incident where reading a dependency's source 24 times
|
||
outweighed writing to the project dir 22 times: the write-like 3× weighting
|
||
puts the project directory that actually received the edits ahead of a
|
||
dependency's source that was merely read.
|