plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
128 lines
6.3 KiB
Markdown
128 lines
6.3 KiB
Markdown
# Spec: retrospective baseline comparator (LLMRouter smallest_llm/largest_llm)
|
|
|
|
Status: done -- baseline_report.py
|
|
|
|
**Origin.** LLMRouter ships `smallest_llm`/`largest_llm` — routers with no
|
|
scoring logic at all, used purely as a floor to prove a real router beats
|
|
"always pick cheap" or "always pick strong." This project has no size axis
|
|
(no parameter counts in the catalog), so the direct port doesn't map
|
|
cleanly — the useful part isn't "smallest/largest," it's *having a trivial
|
|
counterfactual at all*. This is also the automated form of something the
|
|
README already asks a human to do by eye: "If you see one model win
|
|
everything again, check for dominance before reaching for config." A
|
|
report script makes that a query instead of a memory.
|
|
|
|
No code changes accompany this document — this is the spec opencode builds
|
|
from.
|
|
|
|
---
|
|
|
|
## Design
|
|
|
|
A new **read-only, standalone script**, `baseline_report.py`, modeled on
|
|
`metrics.py`'s existing pattern: takes `(conn, cfg)`, never imports
|
|
`dispatcher`, so it stays outside the import-cycle constraint `metrics.py`
|
|
was already built to respect. Runnable directly (`python baseline_report.py
|
|
--since 2026-08-01`), same shape as `leaderboard.py --check` or
|
|
`eval_proficiency.py` — no new schema, no new table, no write path.
|
|
|
|
### What it reads
|
|
|
|
`route_decisions` for the historical record of *what was actually decided*:
|
|
`task_category`, `task_tier`, `required_context_tokens`,
|
|
`latency_tolerance`, `tools`/`images`/`json_mode` flags, `selected_model`,
|
|
`est_cost_usd`, `est_proficiency`. Every field the two baselines below need
|
|
to reconstruct an equivalent decision is already logged here — no join to
|
|
`energy_observations` required, and deliberately so (see the limitation
|
|
below).
|
|
|
|
### The two baselines
|
|
|
|
For each historical decision row, re-derive the **same eligible candidate
|
|
set** the real decision would have seen — by re-running `routing.py`'s hard
|
|
filters (`select_candidates`-equivalent: tier, context, latency_tolerance,
|
|
tools/images/json_mode gates) against the **current** catalog, using the
|
|
row's own stored `task_tier`/`required_context_tokens`/`latency_tolerance`/
|
|
capability flags as the filter inputs. This matters: a baseline of "always
|
|
pick the globally cheapest model in the catalog" is not an interesting
|
|
comparison if that model couldn't have legally served half the requests
|
|
(wrong tier, no vision, wrong latency class). The comparison that's actually
|
|
useful is "given the same hard constraints the real router respected, would
|
|
a dumber tie-break have done just as well?"
|
|
|
|
Within that reconstructed eligible set:
|
|
|
|
- **`always_cheapest`** (LLMRouter's `smallest_llm` analogue) — the
|
|
candidate with the lowest `routing.estimated_cost(row, prompt_tokens=
|
|
required_context_tokens, completion_tokens=<config default>, cache_rate=
|
|
<objective.assumed_cache_rate>)`, using the same shape assumptions the
|
|
real decision's own `est_cost_usd` was computed with.
|
|
- **`always_best_proficiency`** (LLMRouter's `largest_llm` analogue,
|
|
adapted — this catalog has no size axis, but proficiency is the actual
|
|
axis routing optimizes quality on) — the candidate with the highest
|
|
`proficiency.blended_score` for that row's `task_category`, ties broken
|
|
by lowest cost.
|
|
|
|
Both are pure functions of already-existing code
|
|
(`routing.estimated_cost`, a live `proficiency` table read) — no new scoring
|
|
logic to build or validate.
|
|
|
|
### Known limitation, stated plainly
|
|
|
|
Both baselines are recomputed against the **current** catalog and **current**
|
|
proficiency table, not a historical snapshot from when each decision was
|
|
actually made. Prices, tiers, and proficiency scores drift over time (the
|
|
README documents several such shifts), so a decision from three weeks ago
|
|
is compared against today's catalog, not the one it actually saw. This is
|
|
the same trade-off LLMRouter's own benchmark pipeline makes in the other
|
|
direction — it "replays pre-recorded model executions" against a frozen
|
|
dataset rather than live catalogs. Neither approach is wrong; this project
|
|
doesn't snapshot the catalog per-decision, so "current catalog" is the
|
|
only version buildable without a new logging table, and it's directionally
|
|
fine for an aggregate report over a recent window (`--since`) where the
|
|
catalog hasn't moved much. Flagging it so nobody mistakes this for a
|
|
historically-exact replay.
|
|
|
|
`runner_up_models` (already logged, top 3 candidates by id/provider only)
|
|
is **not** sufficient for this on its own — it carries no cost or
|
|
proficiency values and isn't guaranteed to include the two baseline winners
|
|
identified above. Re-deriving the eligible set from the catalog, as
|
|
described, is required either way.
|
|
|
|
### Output
|
|
|
|
```
|
|
python baseline_report.py --since 2026-08-01 [--category coding_general] [--csv]
|
|
```
|
|
|
|
Aggregate summary:
|
|
|
|
- Total actual cost vs. total `always_cheapest` cost vs. total
|
|
`always_best_proficiency` cost, over the window.
|
|
- Mean `est_proficiency` (actual) vs. mean proficiency of each baseline.
|
|
- **Dominance check**: the count/percentage of decisions where
|
|
`selected_model == always_cheapest` — this is the automated form of the
|
|
README's "check for dominance" instruction. A high percentage with near-
|
|
zero proficiency delta says the real scoring isn't earning its complexity
|
|
for that slice of traffic; a low percentage, or a percentage that's high
|
|
only where quality is genuinely tied, says it is.
|
|
- Break out by `task_category` (not just an overall number) — the README's
|
|
own finding was that dominance is categorical (`qwen3.6-35b` dominated
|
|
*before* cost-per-request and tier-from-price were fixed; different
|
|
categories now have different winners), so an aggregate-only number would
|
|
hide exactly the thing this report exists to surface.
|
|
|
|
`--csv` mirrors LLMRouter's `aggregate_results.py --csv` for the same
|
|
reason it works there: a report like this is as much for pasting into a
|
|
follow-up conversation as for reading in a terminal.
|
|
|
|
## Recommendation
|
|
|
|
Build it as specified. It's pure read-over-existing-data — no new table, no
|
|
new write path, no scoring logic beyond calling `estimated_cost` and reading
|
|
`proficiency` for models the router already knows about. Run it after the
|
|
next `route_decisions` accumulates a reasonable window (a few hundred rows
|
|
across categories, going by current volume) and treat a lopsided dominance
|
|
result as the same kind of signal the README already tells you to distrust
|
|
if it isn't backed by a quality-tolerance argument.
|