plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
4.6 KiB
Review: benchmark-sourced-eval-tasks plan
Status: done -- review of shipped work
What it was reviewing: the 11-todo execution plan at
code_plans/.omo/plans/benchmark-sourced-eval-tasks.md, built from the
source brief at plans/benchmark-sourced-eval-tasks.md, after its own
"Metis" gap-analysis pass reported the plan decision-complete and asked for
sign-off on 4 open sanity-check items before /start-work. Verified against
the plan text and, where it made falsifiable claims about external data,
against the actual live sources — not against the plan's own "verified"
labels.
Verdict: approved
Independently re-checked the load-bearing external claims rather than trusting the plan's own "verified" labels, and they hold.
Confirmed correct, verified directly rather than trusted
- CRUXEval fetch (todo 4). Pulled
cruxeval.jsonlmyself: 800 rows, schema{code, input, output, id}exactly as the plan claims. The specificsample_0row it cites as a "verified fact" — input[1, 1, 3, 1, 3, 1], output[(4, 1), (4, 1), (4, 1), (4, 1), (2, 3), (2, 3)]— matches the live file verbatim. eval("f(" + input + ")")vsliteral_eval(input)(todo 4). Correct as specified. CRUXEval'sinputfield is a raw comma-separated argument list (e.g.(1, ), (1, ), (1, 2)), not a single literal —literal_evalwould mis-parse or SyntaxError on multi-arg rows. The plan's explicit warning against usingliteral_evalhere is right, not just cautious.score_toolsemantics (todo 2), read directly atsrc/eval_proficiency.py:197-239: 0.5 floor for the right tool name, string args matched by substring (want.lower() in got.lower()), everything else bystr()equality, andexpect_argsomission falls back to a flat 1.0 for tool-name-only. Matches what todo 2's translation logic assumes exactly, including the array/nested-object-arg omission path.- Exercism exercises exist —
bowling,dominoes,affine-cipherconfirmed present inexercism/python'sexercises/practice/via the GitHub API directly (not just in Aider's derived repo).
Sanity-check items — verdict on each
- Tool-name-only scoring when BFCL ground-truth args are arrays/nested-objects: approved. Still discriminates on the thing that matters most for this category (right tool vs. wrong tool vs. no tool); losing arg-precision credit on those specific rows is an acceptable, well -understood loss.
eval("f(" + input + ")")as the CRUXEval-O recomputation harness: approved — verified correct against the real fetched dataset, see above.- Snapshot
router.dbbefore the first live run: approved, and I'd make this non-negotiable rather than optional. Cheap insurance (cp router.db router.db.pre-benchmark-tasks.bak) against 8 live passes mutating data that feeds production routing. - PR base
neuralwatt-router-service: resolved outside this plan — the user approved the existing PR #10 (head neuralwatt-router-service→base main) directly, so the branch-stacking question this item raised no longer applies. Proceed with the plan's branch strategy as written; no change needed. (Flagging for the record: the plan hedged on this item without actually checkingtea pulls --remote originfor an existing open PR at planning time — worth doing that check up front next time rather than treating base-branch choice as a coin flip deferred to the human.)
One minor, non-blocking inconsistency
Todo 8 (affine-cipher) explicitly says "port check values from the fetched
canonical suite, don't trust memory." Todos 6 and 7 (bowling, dominoes)
state specific numeric canonical values (e.g. the 10th-frame bonus scores)
as flat "verified facts" without that same hedge. This doesn't block
anything structurally — every check gets executed against the real fetched
reference solution by the offline pytest gate before it ships (tests/ test_task_set.py's reference-passes-its-own-checks requirement), so a wrong
recalled number just fails pytest and forces a fix at authoring time rather
than shipping silently wrong. But it's the pytest gate doing that safety
work, not the plan text's own accuracy — worth adding the same "verify
against the fetch, don't trust memory" phrasing to todos 6 and 7 for
consistency, since nothing currently stops a confidently-wrong "verified
fact" from costing an extra authoring cycle before the gate catches it.
Bottom line
Plan is unusually well-grounded for something this size (43-task target,
touches a scoring function, spends real API quota across 8 live passes).
Nothing found here should block /start-work.