Files
6krrt/plans/baseline-routing-comparator.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

6.3 KiB

Spec: retrospective baseline comparator (LLMRouter smallest_llm/largest_llm)

Status: done -- baseline_report.py

Origin. LLMRouter ships smallest_llm/largest_llm — routers with no scoring logic at all, used purely as a floor to prove a real router beats "always pick cheap" or "always pick strong." This project has no size axis (no parameter counts in the catalog), so the direct port doesn't map cleanly — the useful part isn't "smallest/largest," it's having a trivial counterfactual at all. This is also the automated form of something the README already asks a human to do by eye: "If you see one model win everything again, check for dominance before reaching for config." A report script makes that a query instead of a memory.

No code changes accompany this document — this is the spec opencode builds from.


Design

A new read-only, standalone script, baseline_report.py, modeled on metrics.py's existing pattern: takes (conn, cfg), never imports dispatcher, so it stays outside the import-cycle constraint metrics.py was already built to respect. Runnable directly (python baseline_report.py --since 2026-08-01), same shape as leaderboard.py --check or eval_proficiency.py — no new schema, no new table, no write path.

What it reads

route_decisions for the historical record of what was actually decided: task_category, task_tier, required_context_tokens, latency_tolerance, tools/images/json_mode flags, selected_model, est_cost_usd, est_proficiency. Every field the two baselines below need to reconstruct an equivalent decision is already logged here — no join to energy_observations required, and deliberately so (see the limitation below).

The two baselines

For each historical decision row, re-derive the same eligible candidate set the real decision would have seen — by re-running routing.py's hard filters (select_candidates-equivalent: tier, context, latency_tolerance, tools/images/json_mode gates) against the current catalog, using the row's own stored task_tier/required_context_tokens/latency_tolerance/ capability flags as the filter inputs. This matters: a baseline of "always pick the globally cheapest model in the catalog" is not an interesting comparison if that model couldn't have legally served half the requests (wrong tier, no vision, wrong latency class). The comparison that's actually useful is "given the same hard constraints the real router respected, would a dumber tie-break have done just as well?"

Within that reconstructed eligible set:

  • always_cheapest (LLMRouter's smallest_llm analogue) — the candidate with the lowest routing.estimated_cost(row, prompt_tokens= required_context_tokens, completion_tokens=<config default>, cache_rate= <objective.assumed_cache_rate>), using the same shape assumptions the real decision's own est_cost_usd was computed with.
  • always_best_proficiency (LLMRouter's largest_llm analogue, adapted — this catalog has no size axis, but proficiency is the actual axis routing optimizes quality on) — the candidate with the highest proficiency.blended_score for that row's task_category, ties broken by lowest cost.

Both are pure functions of already-existing code (routing.estimated_cost, a live proficiency table read) — no new scoring logic to build or validate.

Known limitation, stated plainly

Both baselines are recomputed against the current catalog and current proficiency table, not a historical snapshot from when each decision was actually made. Prices, tiers, and proficiency scores drift over time (the README documents several such shifts), so a decision from three weeks ago is compared against today's catalog, not the one it actually saw. This is the same trade-off LLMRouter's own benchmark pipeline makes in the other direction — it "replays pre-recorded model executions" against a frozen dataset rather than live catalogs. Neither approach is wrong; this project doesn't snapshot the catalog per-decision, so "current catalog" is the only version buildable without a new logging table, and it's directionally fine for an aggregate report over a recent window (--since) where the catalog hasn't moved much. Flagging it so nobody mistakes this for a historically-exact replay.

runner_up_models (already logged, top 3 candidates by id/provider only) is not sufficient for this on its own — it carries no cost or proficiency values and isn't guaranteed to include the two baseline winners identified above. Re-deriving the eligible set from the catalog, as described, is required either way.

Output

python baseline_report.py --since 2026-08-01 [--category coding_general] [--csv]

Aggregate summary:

  • Total actual cost vs. total always_cheapest cost vs. total always_best_proficiency cost, over the window.
  • Mean est_proficiency (actual) vs. mean proficiency of each baseline.
  • Dominance check: the count/percentage of decisions where selected_model == always_cheapest — this is the automated form of the README's "check for dominance" instruction. A high percentage with near- zero proficiency delta says the real scoring isn't earning its complexity for that slice of traffic; a low percentage, or a percentage that's high only where quality is genuinely tied, says it is.
  • Break out by task_category (not just an overall number) — the README's own finding was that dominance is categorical (qwen3.6-35b dominated before cost-per-request and tier-from-price were fixed; different categories now have different winners), so an aggregate-only number would hide exactly the thing this report exists to surface.

--csv mirrors LLMRouter's aggregate_results.py --csv for the same reason it works there: a report like this is as much for pasting into a follow-up conversation as for reading in a terminal.

Recommendation

Build it as specified. It's pure read-over-existing-data — no new table, no new write path, no scoring logic beyond calling estimated_cost and reading proficiency for models the router already knows about. Run it after the next route_decisions accumulates a reasonable window (a few hundred rows across categories, going by current volume) and treat a lopsided dominance result as the same kind of signal the README already tells you to distrust if it isn't backed by a quality-tolerance argument.