# Spec: retrospective baseline comparator (LLMRouter smallest_llm/largest_llm) Status: done -- baseline_report.py **Origin.** LLMRouter ships `smallest_llm`/`largest_llm` — routers with no scoring logic at all, used purely as a floor to prove a real router beats "always pick cheap" or "always pick strong." This project has no size axis (no parameter counts in the catalog), so the direct port doesn't map cleanly — the useful part isn't "smallest/largest," it's *having a trivial counterfactual at all*. This is also the automated form of something the README already asks a human to do by eye: "If you see one model win everything again, check for dominance before reaching for config." A report script makes that a query instead of a memory. No code changes accompany this document — this is the spec opencode builds from. --- ## Design A new **read-only, standalone script**, `baseline_report.py`, modeled on `metrics.py`'s existing pattern: takes `(conn, cfg)`, never imports `dispatcher`, so it stays outside the import-cycle constraint `metrics.py` was already built to respect. Runnable directly (`python baseline_report.py --since 2026-08-01`), same shape as `leaderboard.py --check` or `eval_proficiency.py` — no new schema, no new table, no write path. ### What it reads `route_decisions` for the historical record of *what was actually decided*: `task_category`, `task_tier`, `required_context_tokens`, `latency_tolerance`, `tools`/`images`/`json_mode` flags, `selected_model`, `est_cost_usd`, `est_proficiency`. Every field the two baselines below need to reconstruct an equivalent decision is already logged here — no join to `energy_observations` required, and deliberately so (see the limitation below). ### The two baselines For each historical decision row, re-derive the **same eligible candidate set** the real decision would have seen — by re-running `routing.py`'s hard filters (`select_candidates`-equivalent: tier, context, latency_tolerance, tools/images/json_mode gates) against the **current** catalog, using the row's own stored `task_tier`/`required_context_tokens`/`latency_tolerance`/ capability flags as the filter inputs. This matters: a baseline of "always pick the globally cheapest model in the catalog" is not an interesting comparison if that model couldn't have legally served half the requests (wrong tier, no vision, wrong latency class). The comparison that's actually useful is "given the same hard constraints the real router respected, would a dumber tie-break have done just as well?" Within that reconstructed eligible set: - **`always_cheapest`** (LLMRouter's `smallest_llm` analogue) — the candidate with the lowest `routing.estimated_cost(row, prompt_tokens= required_context_tokens, completion_tokens=, cache_rate= )`, using the same shape assumptions the real decision's own `est_cost_usd` was computed with. - **`always_best_proficiency`** (LLMRouter's `largest_llm` analogue, adapted — this catalog has no size axis, but proficiency is the actual axis routing optimizes quality on) — the candidate with the highest `proficiency.blended_score` for that row's `task_category`, ties broken by lowest cost. Both are pure functions of already-existing code (`routing.estimated_cost`, a live `proficiency` table read) — no new scoring logic to build or validate. ### Known limitation, stated plainly Both baselines are recomputed against the **current** catalog and **current** proficiency table, not a historical snapshot from when each decision was actually made. Prices, tiers, and proficiency scores drift over time (the README documents several such shifts), so a decision from three weeks ago is compared against today's catalog, not the one it actually saw. This is the same trade-off LLMRouter's own benchmark pipeline makes in the other direction — it "replays pre-recorded model executions" against a frozen dataset rather than live catalogs. Neither approach is wrong; this project doesn't snapshot the catalog per-decision, so "current catalog" is the only version buildable without a new logging table, and it's directionally fine for an aggregate report over a recent window (`--since`) where the catalog hasn't moved much. Flagging it so nobody mistakes this for a historically-exact replay. `runner_up_models` (already logged, top 3 candidates by id/provider only) is **not** sufficient for this on its own — it carries no cost or proficiency values and isn't guaranteed to include the two baseline winners identified above. Re-deriving the eligible set from the catalog, as described, is required either way. ### Output ``` python baseline_report.py --since 2026-08-01 [--category coding_general] [--csv] ``` Aggregate summary: - Total actual cost vs. total `always_cheapest` cost vs. total `always_best_proficiency` cost, over the window. - Mean `est_proficiency` (actual) vs. mean proficiency of each baseline. - **Dominance check**: the count/percentage of decisions where `selected_model == always_cheapest` — this is the automated form of the README's "check for dominance" instruction. A high percentage with near- zero proficiency delta says the real scoring isn't earning its complexity for that slice of traffic; a low percentage, or a percentage that's high only where quality is genuinely tied, says it is. - Break out by `task_category` (not just an overall number) — the README's own finding was that dominance is categorical (`qwen3.6-35b` dominated *before* cost-per-request and tier-from-price were fixed; different categories now have different winners), so an aggregate-only number would hide exactly the thing this report exists to surface. `--csv` mirrors LLMRouter's `aggregate_results.py --csv` for the same reason it works there: a report like this is as much for pasting into a follow-up conversation as for reading in a terminal. ## Recommendation Build it as specified. It's pure read-over-existing-data — no new table, no new write path, no scoring logic beyond calling `estimated_cost` and reading `proficiency` for models the router already knows about. Run it after the next `route_decisions` accumulates a reasonable window (a few hundred rows across categories, going by current volume) and treat a lopsided dominance result as the same kind of signal the README already tells you to distrust if it isn't backed by a quality-tolerance argument.