Covers per-provider quota measurement + credit-aware attenuation (#40), the :path routing fix (#42), quota UI billing-shape clarity (#43), the pinch protected-window size cap (#44, docs/pinch.md verified against what actually shipped), the OpenRouter opt-in allowlist (#45, with the firm x-ai/openai/anthropic exclusions stated as settled, not pending), and the models-page deprecated/stale toggle (#46). Also extends the classifier.mode section with the #48-#50 incident chain (confidence_threshold validation, the gated-model crash, and the natural-language-labels accuracy fix) since it's directly continuous with the local_encoder-is-live-but-not-yet-default note already there.
5.1 KiB
5.1 KiB
Module-by-module reference. Back to README.
Modules
| Module | Role | I/O? |
|---|---|---|
dispatcher.py |
FastAPI service: routes, calls providers, logs, streams | Yes (DB, network) |
poller.py |
Fetches Neuralwatt catalog, normalizes, upserts models table |
Yes (network, DB) |
scoring.py |
normalize_inverted + composite_score weighted formula |
Pure |
routing.py |
Hard filters (select_candidates) + ranking (rank_candidates) |
Pure |
capabilities.py |
Reads tools/image_url/response_format off the request body into RequestCapabilities — the hard filters read this, not a classifier guess |
Pure |
circuit_breaker.py |
Passive availability skip: cooldown + backoff on a 5xx, clears on next success, no poller | Pure (in-memory) |
context_prune.py |
Bounds the message list sent to the local classifier/vision model — image handling, structured content, recency guards | Pure |
tiering.py |
Pure tier resolver: 1=cheap, 2=mid, 3=frontier | Pure |
tier.py |
DB tiering pass: reads models, resolves, writes tier column |
Yes (DB) |
proficiency.py / proficiency_store.py / proficiency_outcome.py |
Blend leaderboard + self-eval into a benchmark prior, accumulate client outcomes, and recompute expected pass rates; the only write paths to proficiency, so blended_score/source never drift |
Yes (DB) |
proficiency_store_core.py |
Core DB primitives for the proficiency table, split out so the empirical-Bayes recompute path shares one write path without a circular import |
Yes (DB) |
exploration.py |
Epsilon-greedy exploration chooser; injected RNG, no mutable state | Pure |
seed_energy.py |
Reference workload sweep: fixed prompt × N runs per model | Yes (network, DB) |
seed_local_dispatch_energy.py |
Standalone reference-shape sweep for ollama-local rows; derives per-token USD rates from measured GPU draw and your tariff |
Yes (network, DB, subprocess) |
seed_local_dispatch_core.py |
Pure derivation logic (through-origin OLS) for seed_local_dispatch_energy.py, split out to stay under the module LOC ceiling |
Pure |
local_energy.py |
Meters the router's OWN local-hardware calls (classifier/verifier), which NeuralWatt never bills and energy_observations never sees; samples nvidia-smi in a background thread into local_energy_observations |
Yes (subprocess, DB) |
eval_proficiency.py |
Self-eval harness: 4 scoring kinds against N models | Yes (network, DB, subprocess) |
verification.py |
Two-layer check: structural parse (always) + local LLM spot-check (async, size-gated). Never executes model output | Pure |
feedback.py |
Folds observed verification failures into proficiency; routing learns from real traffic |
Yes (DB) |
leaderboard.py |
Imports leaderboards.yaml priors into proficiency; --check reports gaps |
Yes (DB) |
iteration.py |
Retry budget per tier, and matching the retry to the failure kind | Pure |
config.py |
YAML loader + Pydantic validators (blend weights sum to 1, valid tiers, endpoints separately addressable) | Yes (file) |
logs.py |
Structured logfmt logging: per-request trace id (ContextVar), journald priority prefixes when systemd owns stderr | Yes (stderr) |
session_cache.py |
Off-by-default in-memory cache of the last task_category/task_tier per session, so a long agent session can skip repeat classifier round-trips |
Pure (in-memory) |
baseline_report.py (repo root, not src/) |
Read-only retrospective: replays route_decisions against two trivial baselines (always-cheapest, always-highest-proficiency) to check whether scoring earns its complexity |
Yes (DB) |
metrics.py |
Read-only aggregations for /health and GET /metrics: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency. Never imports dispatcher |
Yes (DB) |
events.py |
In-memory decision-event broker for the TUI's live feed: bounded ring buffer + thread-safe fan-out to SSE subscribers | Pure |
admin.py |
/admin management portal: serves seven frontend pages (dashboard, models, profiles, proficiency, decisions, controls, providers) + read/write API (snapshot, history, models/availability, profile CRUD, runtime knobs, allowlisted config edits, operational triggers, provider-model allowlist) |
Yes (DB, network, file) |
tui.py |
Textual terminal dashboard over GET /metrics and GET /events/decisions; live routing feed, detail popup, category breakdown. Foreground tool, not a service |
Yes (network) |
tui_model.py |
Pure data layer for the TUI: build_model, build_category_breakdown, decision_row — no Textual import, testable without a terminal |
Pure |
tui_sse.py |
Background-thread SSE consumer for the TUI: reconnects on failure, marshals live decisions onto the UI thread | Yes (network) |
tui_screens.py |
Modal screen for the TUI: DecisionDetailScreen shows the full decision JSON when the user presses Enter |
Pure |
router_cli.py |
One-shot routing probe: POSTs to /route and prints the decision tree |
Yes (network) |