plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
6.3 KiB
Spec: retrospective baseline comparator (LLMRouter smallest_llm/largest_llm)
Status: done -- baseline_report.py
Origin. LLMRouter ships smallest_llm/largest_llm — routers with no
scoring logic at all, used purely as a floor to prove a real router beats
"always pick cheap" or "always pick strong." This project has no size axis
(no parameter counts in the catalog), so the direct port doesn't map
cleanly — the useful part isn't "smallest/largest," it's having a trivial
counterfactual at all. This is also the automated form of something the
README already asks a human to do by eye: "If you see one model win
everything again, check for dominance before reaching for config." A
report script makes that a query instead of a memory.
No code changes accompany this document — this is the spec opencode builds from.
Design
A new read-only, standalone script, baseline_report.py, modeled on
metrics.py's existing pattern: takes (conn, cfg), never imports
dispatcher, so it stays outside the import-cycle constraint metrics.py
was already built to respect. Runnable directly (python baseline_report.py --since 2026-08-01), same shape as leaderboard.py --check or
eval_proficiency.py — no new schema, no new table, no write path.
What it reads
route_decisions for the historical record of what was actually decided:
task_category, task_tier, required_context_tokens,
latency_tolerance, tools/images/json_mode flags, selected_model,
est_cost_usd, est_proficiency. Every field the two baselines below need
to reconstruct an equivalent decision is already logged here — no join to
energy_observations required, and deliberately so (see the limitation
below).
The two baselines
For each historical decision row, re-derive the same eligible candidate
set the real decision would have seen — by re-running routing.py's hard
filters (select_candidates-equivalent: tier, context, latency_tolerance,
tools/images/json_mode gates) against the current catalog, using the
row's own stored task_tier/required_context_tokens/latency_tolerance/
capability flags as the filter inputs. This matters: a baseline of "always
pick the globally cheapest model in the catalog" is not an interesting
comparison if that model couldn't have legally served half the requests
(wrong tier, no vision, wrong latency class). The comparison that's actually
useful is "given the same hard constraints the real router respected, would
a dumber tie-break have done just as well?"
Within that reconstructed eligible set:
always_cheapest(LLMRouter'ssmallest_llmanalogue) — the candidate with the lowestrouting.estimated_cost(row, prompt_tokens= required_context_tokens, completion_tokens=<config default>, cache_rate= <objective.assumed_cache_rate>), using the same shape assumptions the real decision's ownest_cost_usdwas computed with.always_best_proficiency(LLMRouter'slargest_llmanalogue, adapted — this catalog has no size axis, but proficiency is the actual axis routing optimizes quality on) — the candidate with the highestproficiency.blended_scorefor that row'stask_category, ties broken by lowest cost.
Both are pure functions of already-existing code
(routing.estimated_cost, a live proficiency table read) — no new scoring
logic to build or validate.
Known limitation, stated plainly
Both baselines are recomputed against the current catalog and current
proficiency table, not a historical snapshot from when each decision was
actually made. Prices, tiers, and proficiency scores drift over time (the
README documents several such shifts), so a decision from three weeks ago
is compared against today's catalog, not the one it actually saw. This is
the same trade-off LLMRouter's own benchmark pipeline makes in the other
direction — it "replays pre-recorded model executions" against a frozen
dataset rather than live catalogs. Neither approach is wrong; this project
doesn't snapshot the catalog per-decision, so "current catalog" is the
only version buildable without a new logging table, and it's directionally
fine for an aggregate report over a recent window (--since) where the
catalog hasn't moved much. Flagging it so nobody mistakes this for a
historically-exact replay.
runner_up_models (already logged, top 3 candidates by id/provider only)
is not sufficient for this on its own — it carries no cost or
proficiency values and isn't guaranteed to include the two baseline winners
identified above. Re-deriving the eligible set from the catalog, as
described, is required either way.
Output
python baseline_report.py --since 2026-08-01 [--category coding_general] [--csv]
Aggregate summary:
- Total actual cost vs. total
always_cheapestcost vs. totalalways_best_proficiencycost, over the window. - Mean
est_proficiency(actual) vs. mean proficiency of each baseline. - Dominance check: the count/percentage of decisions where
selected_model == always_cheapest— this is the automated form of the README's "check for dominance" instruction. A high percentage with near- zero proficiency delta says the real scoring isn't earning its complexity for that slice of traffic; a low percentage, or a percentage that's high only where quality is genuinely tied, says it is. - Break out by
task_category(not just an overall number) — the README's own finding was that dominance is categorical (qwen3.6-35bdominated before cost-per-request and tier-from-price were fixed; different categories now have different winners), so an aggregate-only number would hide exactly the thing this report exists to surface.
--csv mirrors LLMRouter's aggregate_results.py --csv for the same
reason it works there: a report like this is as much for pasting into a
follow-up conversation as for reading in a terminal.
Recommendation
Build it as specified. It's pure read-over-existing-data — no new table, no
new write path, no scoring logic beyond calling estimated_cost and reading
proficiency for models the router already knows about. Run it after the
next route_decisions accumulates a reasonable window (a few hundred rows
across categories, going by current volume) and treat a lopsided dominance
result as the same kind of signal the README already tells you to distrust
if it isn't backed by a quality-tolerance argument.