plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
14 KiB
Spec: sourcing harder self-eval tasks from external benchmarks
Status: done -- evals/tasks.yaml, 43 tasks
Origin. PYTHONPATH=src python -m eval_proficiency was re-run in full this
session (11 models x 23 tasks, 216 billed calls, $0.79 / 0.098 kWh, 0 call
failures) specifically to see whether more samples of the existing task set
would break the ties CLAUDE.md already flagged as suspicious. Result, doubling
n from 2-3 to 4-6 per model per category:
tool_use_agentic:deepseek-v4-flashheld at exactly 2/6, the same ratio as the earlier 1/3 — real signal, not n=3 noise, but still short ofself_eval_min_samples: 10.coding_general,coding_refactor,debugging: still perfectly flat at 1.00 for every model. This is no longer "not enough samples" — it's "these 3 tasks per category don't discriminate anything, at any sample count." More passes over the same 23 tasks won't fix that; the task set itself needs to get harder.
This spec is the result of "where do we get harder tasks without inventing
plausible-looking difficulty from nothing" — the same objection this project
already raised against leaderboards.yaml's empty priors. No code changes
accompany this document — this is the spec opencode builds from.
Non-goals, stated up front
- No new categories.
proficiency.categoriesinconfig.pyis read throughout scoring/tiering/routing; adding one is a real migration, not a task-authoring change, and nothing here needs it. Every source below gets mapped onto one of the 9 existing categories. - No new scoring kinds. Everything maps onto
code,exact, ortool. (judgecategories —docs_writing,general_chat,summarization,translation,general_chat— aren't the problem this session found; skip them for this pass.) - No repo-context benchmarks. This is the reason SWE-bench (and anything
else shaped like "here's a git checkout, go fix an issue") is not on the
list below despite being the best-known real-bug-fixing benchmark. Every
task in
evals/tasks.yamlis one self-contained prompt scored by one subprocess call or one structural check — the harness has no notion of a repo checkout, and building one is a much bigger lift than the actual problem (three flat categories) justifies. - No verbatim copying. Every source below gets adapted — either
translated into this project's YAML+checks shape, or (for
coding_refactor/debugging) used only to identify which underlying problems are hard enough to be worth hand-authoring a variant of. Nothing here is "download file, paste into tasks.yaml."
A code change this needs, before any task uses it
score_exact (eval_proficiency.py:191) normalizes by regexing out the
last number in the reply (normalize_answer). That's correct for this
project's current exact tasks — they're all arithmetic word problems with a
scalar numeric answer (math_percent_trap, math_rate_trap,
math_counting). It is not correct for CRUXEval-style tasks (below),
where the expected value is a Python literal that can be a list, dict, tuple,
string, or None — e.g. an answer of [1, 2, 3] would get reduced to 3,
silently discarding the structure and scoring against the wrong thing.
Fix, scoped and small: in normalize_answer, try ast.literal_eval on the
stripped text first (covers numbers, strings, lists, dicts, tuples, bools,
None uniformly); fall back to the existing regex-last-number heuristic only
on a ValueError/SyntaxError, which is exactly what a numeric word-problem
answer with prose around it produces. This is additive — every existing
exact task still gets scored the same way, since literal_eval will fail on
"3 minutes" and fall through to the current path unchanged. Needs its own
test alongside the existing test_counting_answer_matches_brute_force-style
ones (tests/test_task_set.py) before any CRUXEval-derived task is added, per
the project's own rule that a check the reference solution can't pass "scores
the task set rather than the model."
tool_use_agentic ← BFCL live_irrelevance
Source: BFCL_v4_live_irrelevance.json in
ShishirPatil/gorilla
(Apache 2.0). 884 examples, sourced from real user queries rather than
synthesized — pull this file over the smaller 239-example synthetic
BFCL_v4_irrelevance.json for the same reason this project trusts
POST /outcome traffic over benchmark scores: real queries beat invented
ones. This is the single best-matched source of the four — it's a large pool
of exactly the pattern this project already hand-writes 2 of 3 tasks in:
offer a tool, correct behavior is not calling it.
Schema gap to close (confirmed by fetching a live sample): BFCL's own
function schema is not the OpenAI tool-call schema this project uses —
"type": "dict" where OpenAI/this project's YAML says "type": "object", no
{"type": "function", "function": {...}} wrapper, and question is a list of
turns (list of message dicts) rather than this project's flat prompt
string. Translation per selected example:
tools = [
{"type": "function", "function": {
"name": fn["name"],
"description": fn["description"],
"parameters": {**fn["parameters"], "type": "object"}, # dict -> object
}}
for fn in example["function"]
]
prompt = example["question"][0][0]["content"] # single-turn BFCL cases only
The balance trap. Every BFCL irrelevance example is, by definition, a
"correctly abstain" case. If the new batch is 100% abstain-cases, a model
that never calls a tool scores 1.0 on the enlarged category despite being
just as broken as one that always calls a tool — the category would stop
measuring over-triggering and start measuring only under-triggering.
score_tool already has both directions (expect_tool: null vs. a named
tool), and the current 3-task set is already 1:2 (call : abstain). Keep new
additions roughly in that band or better: pull irrelevance examples for the
new abstain cases, and pull a matching number of positive call cases from
BFCL's live_simple/live_multiple files (same repo, same schema, translate
the same way) so the enlarged category still rewards correct discrimination,
not just caution.
Worked example, adapted from a real fetched entry:
- id: tool_geocode_irrelevant
category: tool_use_agentic
kind: tool
prompt: |
Can you provide the address for latitude 37.4224764 and longitude
-122.0842499 using the Geocoding API?
expect_tool: null # no offered function does geocoding
tools:
- type: function
function:
name: requests.get
description: Sends a GET request to the specified URL to retrieve data.
parameters:
type: object
properties:
url: {type: string, description: The URL to send the GET request to.}
required: [url]
coding_general ← CRUXEval-O
Source: facebookresearch/cruxeval
(MIT). 800 short Python functions with input/output pairs, split into
CRUXEval-I (predict an input that produces a given output) and CRUXEval-O
(predict the output of running the function on a given input). Use O only.
I is a genuine structural mismatch: an input that produces a target output is
often not unique, so scoring it needs to actually execute the candidate input
through the real function and compare — that's neither exact (no single
golden string to normalize against) nor code (the model isn't writing a
function) as this harness defines them. Not worth a third scoring kind for
half of one source; skip I, use O, which is a clean fit for exact once the
ast.literal_eval fix above lands.
At release, best-of-field models scored 63-67% on CRUXEval-O — meaning it
was built to sit in a discriminating band, not to saturate the way this
project's own three coding_general tasks did.
Worked example (illustrative shape, not a literal fetched row):
- id: cruxeval_trace_dedup
category: coding_general
kind: exact
prompt: |
What does this function return when called as shown? Reply with ONLY
the literal Python value — no explanation, no fences.
def f(lst):
seen = {}
out = []
for x in lst:
if x not in seen:
seen[x] = True
out.append(x * 2)
return out
f([3, 1, 3, 2, 1])
answer: "[6, 2, 4]"
coding_refactor / debugging ← Exercism, via Aider's selection method
What Aider's polyglot-benchmark actually gives you. I checked: the
Aider-AI/polyglot-benchmark
repo's own GitHub license field comes back blank via the API — don't pull
files from it directly. What's genuinely reusable is the method their
Dec 2024 writeup describes:
225 of Exercism's 697 exercises, filtered to ones solved by 3 or fewer of 7
frontier models, deliberately excluding the 258 "solved by all 7" as too
easy. They didn't publish the 225 names. The underlying exercises are
Exercism's own tracks, confirmed MIT-licensed
(exercism/python, etc.) — go there directly, using Aider's selection
criterion (not their file) as the bar: prefer exercises in Exercism's own
practice/ directories with a track difficulty/reputation for edge-case
traps. bowling, dominoes, and affine-cipher (all present in the
exercism/python practice/ tree, confirmed via the GitHub API) are
reasonable starting picks — each is well known for one nasty, non-obvious
edge case rather than raw difficulty, which is exactly this project's own
stated design principle ("at least one edge case a plausible-looking solution
gets wrong").
Why this can't be mechanical, unlike the other two sources. Exercism ships
a spec + a test suite, not a "here's buggy/ugly code, fix/refactor it"
prompt — that shape doesn't exist upstream. coding_refactor and debugging
tasks in this project embed the broken or ugly-but-working code directly in
the prompt (confirmed from apply_settings, first_match, describe in the
current set) and the checks pin exact edge-case behavior. So the real value
of going to Exercism isn't a copyable prompt, it's borrowing the hard edge
case: read the exercise's canonical solution and its test suite (that's the
whole reason the exercise is hard), then hand-author a variant:
coding_refactorvariant: take the canonical solution, deliberately restructure it into duplicated/nested-conditional/flag-variable style (same pattern as the existing three) while preserving every edge case exactly — the model refactors it back, checks re-run the exercise's own tricky test cases.debuggingvariant: take the canonical solution, introduce one subtle off-by-one/boundary bug that breaks exactly the input the exercise is famous for tripping people up on (bowling's 10th-frame strike/spare bonus scoring is the textbook example), leave everything else correct — checks must include that exact edge case, pertests/test_task_set.py's existing requirement that every debugging target actually fails pre-fix.
This is real authoring work, not translation — budget it accordingly (see below). What Exercism/Aider buys is which edge cases are worth encoding, pre-validated by the fact that most frontier models already fail them, rather than inventing new traps from scratch the way the original 23-task set's "touching intervals / full semver / late-binding closures" hardening pass did.
License / attribution handling
Nothing here is substantial verbatim reproduction, so full license-file
vendoring isn't needed — but this project already has a citation precedent
(session-classification-cache-ttl.md's LLMRouter attribution, and
PinchConfig's docstring crediting llmrouter directly in code). Match it: a
one-line comment above each newly-added task block in evals/tasks.yaml
naming the source and license, e.g.:
# --- tool_use_agentic (added) — adapted from BFCL live_irrelevance /
# live_simple, Apache 2.0, github.com/ShishirPatil/gorilla ---
Validation obligation (not optional)
tests/test_task_set.py requires, per kind:
code: a reference solution that passes every check.exact: an independent (brute-force or hand-derived) recomputation of the answer — seetest_counting_answer_matches_brute_forcefor the pattern.debugging: the buggy target must fail its own checks pre-fix.coding_refactor: the ugly target must pass its own checks pre-refactor.
This is what already caught one wrong expected-value check in the existing set (CLAUDE.md: "would have docked every model on a task and been indistinguishable from genuine difficulty"). Every task added from this spec needs its matching test — this roughly doubles the authoring cost per task versus just writing the YAML, and should be budgeted as such rather than treated as a follow-up.
Suggested first-pass batch size
This session's full run (23 tasks x 11 models, some judge calls) was 216
billed calls / $0.79. A reasonable first batch — 8 new tool_use_agentic
tasks (5 abstain + 3 positive-call), 6 new coding_general (CRUXEval-O), 3
new coding_refactor, 3 new debugging — is +20 tasks, roughly matching
the size of the existing set. Expect a proportional cost per pass (~$1.50-2),
and two passes (not one) to clear self_eval_min_samples: 10 for the
newly-added tasks, since one pass only gives n=1 per new task per model.
Recommendation
Build in this order: (1) the score_exact/ast.literal_eval fix plus its
test, since CRUXEval-O tasks are silently wrong without it; (2) BFCL-derived
tool_use_agentic tasks, since that category already has the strongest
reproduced signal from this session's run and BFCL requires no hand-authoring
beyond schema translation; (3) CRUXEval-O coding_general tasks; (4) the
Exercism-sourced coding_refactor/debugging variants last, since they're
the only genuinely bespoke authoring work and benefit from the other three
being done first as a format reference. Run eval_proficiency twice after
each category lands rather than once at the end, so a broken check (per the
validation section above) is caught against one category's blast radius, not
all four at once.