Files
6krrt/plans/baseline-routing-comparator.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

128 lines
6.3 KiB
Markdown

# Spec: retrospective baseline comparator (LLMRouter smallest_llm/largest_llm)
Status: done -- baseline_report.py
**Origin.** LLMRouter ships `smallest_llm`/`largest_llm` — routers with no
scoring logic at all, used purely as a floor to prove a real router beats
"always pick cheap" or "always pick strong." This project has no size axis
(no parameter counts in the catalog), so the direct port doesn't map
cleanly — the useful part isn't "smallest/largest," it's *having a trivial
counterfactual at all*. This is also the automated form of something the
README already asks a human to do by eye: "If you see one model win
everything again, check for dominance before reaching for config." A
report script makes that a query instead of a memory.
No code changes accompany this document — this is the spec opencode builds
from.
---
## Design
A new **read-only, standalone script**, `baseline_report.py`, modeled on
`metrics.py`'s existing pattern: takes `(conn, cfg)`, never imports
`dispatcher`, so it stays outside the import-cycle constraint `metrics.py`
was already built to respect. Runnable directly (`python baseline_report.py
--since 2026-08-01`), same shape as `leaderboard.py --check` or
`eval_proficiency.py` — no new schema, no new table, no write path.
### What it reads
`route_decisions` for the historical record of *what was actually decided*:
`task_category`, `task_tier`, `required_context_tokens`,
`latency_tolerance`, `tools`/`images`/`json_mode` flags, `selected_model`,
`est_cost_usd`, `est_proficiency`. Every field the two baselines below need
to reconstruct an equivalent decision is already logged here — no join to
`energy_observations` required, and deliberately so (see the limitation
below).
### The two baselines
For each historical decision row, re-derive the **same eligible candidate
set** the real decision would have seen — by re-running `routing.py`'s hard
filters (`select_candidates`-equivalent: tier, context, latency_tolerance,
tools/images/json_mode gates) against the **current** catalog, using the
row's own stored `task_tier`/`required_context_tokens`/`latency_tolerance`/
capability flags as the filter inputs. This matters: a baseline of "always
pick the globally cheapest model in the catalog" is not an interesting
comparison if that model couldn't have legally served half the requests
(wrong tier, no vision, wrong latency class). The comparison that's actually
useful is "given the same hard constraints the real router respected, would
a dumber tie-break have done just as well?"
Within that reconstructed eligible set:
- **`always_cheapest`** (LLMRouter's `smallest_llm` analogue) — the
candidate with the lowest `routing.estimated_cost(row, prompt_tokens=
required_context_tokens, completion_tokens=<config default>, cache_rate=
<objective.assumed_cache_rate>)`, using the same shape assumptions the
real decision's own `est_cost_usd` was computed with.
- **`always_best_proficiency`** (LLMRouter's `largest_llm` analogue,
adapted — this catalog has no size axis, but proficiency is the actual
axis routing optimizes quality on) — the candidate with the highest
`proficiency.blended_score` for that row's `task_category`, ties broken
by lowest cost.
Both are pure functions of already-existing code
(`routing.estimated_cost`, a live `proficiency` table read) — no new scoring
logic to build or validate.
### Known limitation, stated plainly
Both baselines are recomputed against the **current** catalog and **current**
proficiency table, not a historical snapshot from when each decision was
actually made. Prices, tiers, and proficiency scores drift over time (the
README documents several such shifts), so a decision from three weeks ago
is compared against today's catalog, not the one it actually saw. This is
the same trade-off LLMRouter's own benchmark pipeline makes in the other
direction — it "replays pre-recorded model executions" against a frozen
dataset rather than live catalogs. Neither approach is wrong; this project
doesn't snapshot the catalog per-decision, so "current catalog" is the
only version buildable without a new logging table, and it's directionally
fine for an aggregate report over a recent window (`--since`) where the
catalog hasn't moved much. Flagging it so nobody mistakes this for a
historically-exact replay.
`runner_up_models` (already logged, top 3 candidates by id/provider only)
is **not** sufficient for this on its own — it carries no cost or
proficiency values and isn't guaranteed to include the two baseline winners
identified above. Re-deriving the eligible set from the catalog, as
described, is required either way.
### Output
```
python baseline_report.py --since 2026-08-01 [--category coding_general] [--csv]
```
Aggregate summary:
- Total actual cost vs. total `always_cheapest` cost vs. total
`always_best_proficiency` cost, over the window.
- Mean `est_proficiency` (actual) vs. mean proficiency of each baseline.
- **Dominance check**: the count/percentage of decisions where
`selected_model == always_cheapest` — this is the automated form of the
README's "check for dominance" instruction. A high percentage with near-
zero proficiency delta says the real scoring isn't earning its complexity
for that slice of traffic; a low percentage, or a percentage that's high
only where quality is genuinely tied, says it is.
- Break out by `task_category` (not just an overall number) — the README's
own finding was that dominance is categorical (`qwen3.6-35b` dominated
*before* cost-per-request and tier-from-price were fixed; different
categories now have different winners), so an aggregate-only number would
hide exactly the thing this report exists to surface.
`--csv` mirrors LLMRouter's `aggregate_results.py --csv` for the same
reason it works there: a report like this is as much for pasting into a
follow-up conversation as for reading in a terminal.
## Recommendation
Build it as specified. It's pure read-over-existing-data — no new table, no
new write path, no scoring logic beyond calling `estimated_cost` and reading
`proficiency` for models the router already knows about. Run it after the
next `route_decisions` accumulates a reasonable window (a few hundred rows
across categories, going by current volume) and treat a lopsided dominance
result as the same kind of signal the README already tells you to distrust
if it isn't backed by a quality-tolerance argument.