Files
6krrt/plans/benchmark-sourced-eval-tasks.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

14 KiB

Spec: sourcing harder self-eval tasks from external benchmarks

Status: done -- evals/tasks.yaml, 43 tasks

Origin. PYTHONPATH=src python -m eval_proficiency was re-run in full this session (11 models x 23 tasks, 216 billed calls, $0.79 / 0.098 kWh, 0 call failures) specifically to see whether more samples of the existing task set would break the ties CLAUDE.md already flagged as suspicious. Result, doubling n from 2-3 to 4-6 per model per category:

  • tool_use_agentic: deepseek-v4-flash held at exactly 2/6, the same ratio as the earlier 1/3 — real signal, not n=3 noise, but still short of self_eval_min_samples: 10.
  • coding_general, coding_refactor, debugging: still perfectly flat at 1.00 for every model. This is no longer "not enough samples" — it's "these 3 tasks per category don't discriminate anything, at any sample count." More passes over the same 23 tasks won't fix that; the task set itself needs to get harder.

This spec is the result of "where do we get harder tasks without inventing plausible-looking difficulty from nothing" — the same objection this project already raised against leaderboards.yaml's empty priors. No code changes accompany this document — this is the spec opencode builds from.


Non-goals, stated up front

  • No new categories. proficiency.categories in config.py is read throughout scoring/tiering/routing; adding one is a real migration, not a task-authoring change, and nothing here needs it. Every source below gets mapped onto one of the 9 existing categories.
  • No new scoring kinds. Everything maps onto code, exact, or tool. (judge categories — docs_writing, general_chat, summarization, translation, general_chat — aren't the problem this session found; skip them for this pass.)
  • No repo-context benchmarks. This is the reason SWE-bench (and anything else shaped like "here's a git checkout, go fix an issue") is not on the list below despite being the best-known real-bug-fixing benchmark. Every task in evals/tasks.yaml is one self-contained prompt scored by one subprocess call or one structural check — the harness has no notion of a repo checkout, and building one is a much bigger lift than the actual problem (three flat categories) justifies.
  • No verbatim copying. Every source below gets adapted — either translated into this project's YAML+checks shape, or (for coding_refactor/debugging) used only to identify which underlying problems are hard enough to be worth hand-authoring a variant of. Nothing here is "download file, paste into tasks.yaml."

A code change this needs, before any task uses it

score_exact (eval_proficiency.py:191) normalizes by regexing out the last number in the reply (normalize_answer). That's correct for this project's current exact tasks — they're all arithmetic word problems with a scalar numeric answer (math_percent_trap, math_rate_trap, math_counting). It is not correct for CRUXEval-style tasks (below), where the expected value is a Python literal that can be a list, dict, tuple, string, or None — e.g. an answer of [1, 2, 3] would get reduced to 3, silently discarding the structure and scoring against the wrong thing.

Fix, scoped and small: in normalize_answer, try ast.literal_eval on the stripped text first (covers numbers, strings, lists, dicts, tuples, bools, None uniformly); fall back to the existing regex-last-number heuristic only on a ValueError/SyntaxError, which is exactly what a numeric word-problem answer with prose around it produces. This is additive — every existing exact task still gets scored the same way, since literal_eval will fail on "3 minutes" and fall through to the current path unchanged. Needs its own test alongside the existing test_counting_answer_matches_brute_force-style ones (tests/test_task_set.py) before any CRUXEval-derived task is added, per the project's own rule that a check the reference solution can't pass "scores the task set rather than the model."

tool_use_agentic ← BFCL live_irrelevance

Source: BFCL_v4_live_irrelevance.json in ShishirPatil/gorilla (Apache 2.0). 884 examples, sourced from real user queries rather than synthesized — pull this file over the smaller 239-example synthetic BFCL_v4_irrelevance.json for the same reason this project trusts POST /outcome traffic over benchmark scores: real queries beat invented ones. This is the single best-matched source of the four — it's a large pool of exactly the pattern this project already hand-writes 2 of 3 tasks in: offer a tool, correct behavior is not calling it.

Schema gap to close (confirmed by fetching a live sample): BFCL's own function schema is not the OpenAI tool-call schema this project uses — "type": "dict" where OpenAI/this project's YAML says "type": "object", no {"type": "function", "function": {...}} wrapper, and question is a list of turns (list of message dicts) rather than this project's flat prompt string. Translation per selected example:

tools = [
    {"type": "function", "function": {
        "name": fn["name"],
        "description": fn["description"],
        "parameters": {**fn["parameters"], "type": "object"},  # dict -> object
    }}
    for fn in example["function"]
]
prompt = example["question"][0][0]["content"]  # single-turn BFCL cases only

The balance trap. Every BFCL irrelevance example is, by definition, a "correctly abstain" case. If the new batch is 100% abstain-cases, a model that never calls a tool scores 1.0 on the enlarged category despite being just as broken as one that always calls a tool — the category would stop measuring over-triggering and start measuring only under-triggering. score_tool already has both directions (expect_tool: null vs. a named tool), and the current 3-task set is already 1:2 (call : abstain). Keep new additions roughly in that band or better: pull irrelevance examples for the new abstain cases, and pull a matching number of positive call cases from BFCL's live_simple/live_multiple files (same repo, same schema, translate the same way) so the enlarged category still rewards correct discrimination, not just caution.

Worked example, adapted from a real fetched entry:

  - id: tool_geocode_irrelevant
    category: tool_use_agentic
    kind: tool
    prompt: |
      Can you provide the address for latitude 37.4224764 and longitude
      -122.0842499 using the Geocoding API?
    expect_tool: null   # no offered function does geocoding
    tools:
      - type: function
        function:
          name: requests.get
          description: Sends a GET request to the specified URL to retrieve data.
          parameters:
            type: object
            properties:
              url: {type: string, description: The URL to send the GET request to.}
            required: [url]

coding_general ← CRUXEval-O

Source: facebookresearch/cruxeval (MIT). 800 short Python functions with input/output pairs, split into CRUXEval-I (predict an input that produces a given output) and CRUXEval-O (predict the output of running the function on a given input). Use O only. I is a genuine structural mismatch: an input that produces a target output is often not unique, so scoring it needs to actually execute the candidate input through the real function and compare — that's neither exact (no single golden string to normalize against) nor code (the model isn't writing a function) as this harness defines them. Not worth a third scoring kind for half of one source; skip I, use O, which is a clean fit for exact once the ast.literal_eval fix above lands.

At release, best-of-field models scored 63-67% on CRUXEval-O — meaning it was built to sit in a discriminating band, not to saturate the way this project's own three coding_general tasks did.

Worked example (illustrative shape, not a literal fetched row):

  - id: cruxeval_trace_dedup
    category: coding_general
    kind: exact
    prompt: |
      What does this function return when called as shown? Reply with ONLY
      the literal Python value — no explanation, no fences.

          def f(lst):
              seen = {}
              out = []
              for x in lst:
                  if x not in seen:
                      seen[x] = True
                      out.append(x * 2)
              return out

          f([3, 1, 3, 2, 1])
    answer: "[6, 2, 4]"

coding_refactor / debugging ← Exercism, via Aider's selection method

What Aider's polyglot-benchmark actually gives you. I checked: the Aider-AI/polyglot-benchmark repo's own GitHub license field comes back blank via the API — don't pull files from it directly. What's genuinely reusable is the method their Dec 2024 writeup describes: 225 of Exercism's 697 exercises, filtered to ones solved by 3 or fewer of 7 frontier models, deliberately excluding the 258 "solved by all 7" as too easy. They didn't publish the 225 names. The underlying exercises are Exercism's own tracks, confirmed MIT-licensed (exercism/python, etc.) — go there directly, using Aider's selection criterion (not their file) as the bar: prefer exercises in Exercism's own practice/ directories with a track difficulty/reputation for edge-case traps. bowling, dominoes, and affine-cipher (all present in the exercism/python practice/ tree, confirmed via the GitHub API) are reasonable starting picks — each is well known for one nasty, non-obvious edge case rather than raw difficulty, which is exactly this project's own stated design principle ("at least one edge case a plausible-looking solution gets wrong").

Why this can't be mechanical, unlike the other two sources. Exercism ships a spec + a test suite, not a "here's buggy/ugly code, fix/refactor it" prompt — that shape doesn't exist upstream. coding_refactor and debugging tasks in this project embed the broken or ugly-but-working code directly in the prompt (confirmed from apply_settings, first_match, describe in the current set) and the checks pin exact edge-case behavior. So the real value of going to Exercism isn't a copyable prompt, it's borrowing the hard edge case: read the exercise's canonical solution and its test suite (that's the whole reason the exercise is hard), then hand-author a variant:

  • coding_refactor variant: take the canonical solution, deliberately restructure it into duplicated/nested-conditional/flag-variable style (same pattern as the existing three) while preserving every edge case exactly — the model refactors it back, checks re-run the exercise's own tricky test cases.
  • debugging variant: take the canonical solution, introduce one subtle off-by-one/boundary bug that breaks exactly the input the exercise is famous for tripping people up on (bowling's 10th-frame strike/spare bonus scoring is the textbook example), leave everything else correct — checks must include that exact edge case, per tests/test_task_set.py's existing requirement that every debugging target actually fails pre-fix.

This is real authoring work, not translation — budget it accordingly (see below). What Exercism/Aider buys is which edge cases are worth encoding, pre-validated by the fact that most frontier models already fail them, rather than inventing new traps from scratch the way the original 23-task set's "touching intervals / full semver / late-binding closures" hardening pass did.

License / attribution handling

Nothing here is substantial verbatim reproduction, so full license-file vendoring isn't needed — but this project already has a citation precedent (session-classification-cache-ttl.md's LLMRouter attribution, and PinchConfig's docstring crediting llmrouter directly in code). Match it: a one-line comment above each newly-added task block in evals/tasks.yaml naming the source and license, e.g.:

  # --- tool_use_agentic (added) — adapted from BFCL live_irrelevance /
  # live_simple, Apache 2.0, github.com/ShishirPatil/gorilla ---

Validation obligation (not optional)

tests/test_task_set.py requires, per kind:

  • code: a reference solution that passes every check.
  • exact: an independent (brute-force or hand-derived) recomputation of the answer — see test_counting_answer_matches_brute_force for the pattern.
  • debugging: the buggy target must fail its own checks pre-fix.
  • coding_refactor: the ugly target must pass its own checks pre-refactor.

This is what already caught one wrong expected-value check in the existing set (CLAUDE.md: "would have docked every model on a task and been indistinguishable from genuine difficulty"). Every task added from this spec needs its matching test — this roughly doubles the authoring cost per task versus just writing the YAML, and should be budgeted as such rather than treated as a follow-up.

Suggested first-pass batch size

This session's full run (23 tasks x 11 models, some judge calls) was 216 billed calls / $0.79. A reasonable first batch — 8 new tool_use_agentic tasks (5 abstain + 3 positive-call), 6 new coding_general (CRUXEval-O), 3 new coding_refactor, 3 new debugging — is +20 tasks, roughly matching the size of the existing set. Expect a proportional cost per pass (~$1.50-2), and two passes (not one) to clear self_eval_min_samples: 10 for the newly-added tasks, since one pass only gives n=1 per new task per model.

Recommendation

Build in this order: (1) the score_exact/ast.literal_eval fix plus its test, since CRUXEval-O tasks are silently wrong without it; (2) BFCL-derived tool_use_agentic tasks, since that category already has the strongest reproduced signal from this session's run and BFCL requires no hand-authoring beyond schema translation; (3) CRUXEval-O coding_general tasks; (4) the Exercism-sourced coding_refactor/debugging variants last, since they're the only genuinely bespoke authoring work and benefit from the other three being done first as a format reference. Run eval_proficiency twice after each category lands rather than once at the end, so a broken check (per the validation section above) is caught against one category's blast radius, not all four at once.