Files
6krrt/docs/architecture.md
adlee-was-taken d4d69a41fd docs: refresh CLAUDE.md and companion docs for PRs #40-#47
Covers per-provider quota measurement + credit-aware attenuation (#40),
the :path routing fix (#42), quota UI billing-shape clarity (#43), the
pinch protected-window size cap (#44, docs/pinch.md verified against
what actually shipped), the OpenRouter opt-in allowlist (#45, with the
firm x-ai/openai/anthropic exclusions stated as settled, not pending),
and the models-page deprecated/stale toggle (#46).

Also extends the classifier.mode section with the #48-#50 incident
chain (confidence_threshold validation, the gated-model crash, and the
natural-language-labels accuracy fix) since it's directly continuous
with the local_encoder-is-live-but-not-yet-default note already there.
2026-09-06 22:10:11 -04:00

5.1 KiB
Raw Permalink Blame History

Module-by-module reference. Back to README.

Modules

Module Role I/O?
dispatcher.py FastAPI service: routes, calls providers, logs, streams Yes (DB, network)
poller.py Fetches Neuralwatt catalog, normalizes, upserts models table Yes (network, DB)
scoring.py normalize_inverted + composite_score weighted formula Pure
routing.py Hard filters (select_candidates) + ranking (rank_candidates) Pure
capabilities.py Reads tools/image_url/response_format off the request body into RequestCapabilities — the hard filters read this, not a classifier guess Pure
circuit_breaker.py Passive availability skip: cooldown + backoff on a 5xx, clears on next success, no poller Pure (in-memory)
context_prune.py Bounds the message list sent to the local classifier/vision model — image handling, structured content, recency guards Pure
tiering.py Pure tier resolver: 1=cheap, 2=mid, 3=frontier Pure
tier.py DB tiering pass: reads models, resolves, writes tier column Yes (DB)
proficiency.py / proficiency_store.py / proficiency_outcome.py Blend leaderboard + self-eval into a benchmark prior, accumulate client outcomes, and recompute expected pass rates; the only write paths to proficiency, so blended_score/source never drift Yes (DB)
proficiency_store_core.py Core DB primitives for the proficiency table, split out so the empirical-Bayes recompute path shares one write path without a circular import Yes (DB)
exploration.py Epsilon-greedy exploration chooser; injected RNG, no mutable state Pure
seed_energy.py Reference workload sweep: fixed prompt × N runs per model Yes (network, DB)
seed_local_dispatch_energy.py Standalone reference-shape sweep for ollama-local rows; derives per-token USD rates from measured GPU draw and your tariff Yes (network, DB, subprocess)
seed_local_dispatch_core.py Pure derivation logic (through-origin OLS) for seed_local_dispatch_energy.py, split out to stay under the module LOC ceiling Pure
local_energy.py Meters the router's OWN local-hardware calls (classifier/verifier), which NeuralWatt never bills and energy_observations never sees; samples nvidia-smi in a background thread into local_energy_observations Yes (subprocess, DB)
eval_proficiency.py Self-eval harness: 4 scoring kinds against N models Yes (network, DB, subprocess)
verification.py Two-layer check: structural parse (always) + local LLM spot-check (async, size-gated). Never executes model output Pure
feedback.py Folds observed verification failures into proficiency; routing learns from real traffic Yes (DB)
leaderboard.py Imports leaderboards.yaml priors into proficiency; --check reports gaps Yes (DB)
iteration.py Retry budget per tier, and matching the retry to the failure kind Pure
config.py YAML loader + Pydantic validators (blend weights sum to 1, valid tiers, endpoints separately addressable) Yes (file)
logs.py Structured logfmt logging: per-request trace id (ContextVar), journald priority prefixes when systemd owns stderr Yes (stderr)
session_cache.py Off-by-default in-memory cache of the last task_category/task_tier per session, so a long agent session can skip repeat classifier round-trips Pure (in-memory)
baseline_report.py (repo root, not src/) Read-only retrospective: replays route_decisions against two trivial baselines (always-cheapest, always-highest-proficiency) to check whether scoring earns its complexity Yes (DB)
metrics.py Read-only aggregations for /health and GET /metrics: quota burn, coverage, recent decisions, per-model totals, verdict mix, top proficiency. Never imports dispatcher Yes (DB)
events.py In-memory decision-event broker for the TUI's live feed: bounded ring buffer + thread-safe fan-out to SSE subscribers Pure
admin.py /admin management portal: serves seven frontend pages (dashboard, models, profiles, proficiency, decisions, controls, providers) + read/write API (snapshot, history, models/availability, profile CRUD, runtime knobs, allowlisted config edits, operational triggers, provider-model allowlist) Yes (DB, network, file)
tui.py Textual terminal dashboard over GET /metrics and GET /events/decisions; live routing feed, detail popup, category breakdown. Foreground tool, not a service Yes (network)
tui_model.py Pure data layer for the TUI: build_model, build_category_breakdown, decision_row — no Textual import, testable without a terminal Pure
tui_sse.py Background-thread SSE consumer for the TUI: reconnects on failure, marshals live decisions onto the UI thread Yes (network)
tui_screens.py Modal screen for the TUI: DecisionDetailScreen shows the full decision JSON when the user presses Enter Pure
router_cli.py One-shot routing probe: POSTs to /route and prints the decision tree Yes (network)