Files
6krrt/plans/benchmark-sourced-eval-tasks.md
adlee-was-taken 3523dcf93e docs(plans): give every plan a Status line so the queue is greppable
plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.

Now `grep -H '^Status:' plans/*.md` is the answer:

    50 done   3 in progress   2 planned   2 reference   1 parked

Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.

The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.

test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.

Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
2026-09-08 18:55:16 -04:00

283 lines
14 KiB
Markdown

# Spec: sourcing harder self-eval tasks from external benchmarks
Status: done -- evals/tasks.yaml, 43 tasks
**Origin.** `PYTHONPATH=src python -m eval_proficiency` was re-run in full this
session (11 models x 23 tasks, 216 billed calls, $0.79 / 0.098 kWh, 0 call
failures) specifically to see whether more samples of the *existing* task set
would break the ties CLAUDE.md already flagged as suspicious. Result, doubling
n from 2-3 to 4-6 per model per category:
- `tool_use_agentic`: `deepseek-v4-flash` held at exactly 2/6, the same ratio
as the earlier 1/3 — real signal, not n=3 noise, but still short of
`self_eval_min_samples: 10`.
- `coding_general`, `coding_refactor`, `debugging`: **still perfectly flat at
1.00** for every model. This is no longer "not enough samples" — it's
"these 3 tasks per category don't discriminate anything, at any sample
count." More passes over the same 23 tasks won't fix that; the task set
itself needs to get harder.
This spec is the result of "where do we get harder tasks without inventing
plausible-looking difficulty from nothing" — the same objection this project
already raised against `leaderboards.yaml`'s empty priors. No code changes
accompany this document — this is the spec opencode builds from.
---
## Non-goals, stated up front
- **No new categories.** `proficiency.categories` in `config.py` is read
throughout scoring/tiering/routing; adding one is a real migration, not a
task-authoring change, and nothing here needs it. Every source below gets
mapped onto one of the 9 existing categories.
- **No new scoring kinds.** Everything maps onto `code`, `exact`, or `tool`.
(`judge` categories — `docs_writing`, `general_chat`, `summarization`,
`translation`, `general_chat` — aren't the problem this session found; skip
them for this pass.)
- **No repo-context benchmarks.** This is the reason SWE-bench (and anything
else shaped like "here's a git checkout, go fix an issue") is not on the
list below despite being the best-known real-bug-fixing benchmark. Every
task in `evals/tasks.yaml` is one self-contained prompt scored by one
subprocess call or one structural check — the harness has no notion of a
repo checkout, and building one is a much bigger lift than the actual
problem (three flat categories) justifies.
- **No verbatim copying.** Every source below gets *adapted* — either
translated into this project's YAML+checks shape, or (for
`coding_refactor`/`debugging`) used only to identify which underlying
problems are hard enough to be worth hand-authoring a variant of. Nothing
here is "download file, paste into tasks.yaml."
## A code change this needs, before any task uses it
`score_exact` (`eval_proficiency.py:191`) normalizes by regexing out the
**last number** in the reply (`normalize_answer`). That's correct for this
project's current `exact` tasks — they're all arithmetic word problems with a
scalar numeric answer (`math_percent_trap`, `math_rate_trap`,
`math_counting`). It is **not** correct for CRUXEval-style tasks (below),
where the expected value is a Python literal that can be a list, dict, tuple,
string, or `None` — e.g. an answer of `[1, 2, 3]` would get reduced to `3`,
silently discarding the structure and scoring against the wrong thing.
Fix, scoped and small: in `normalize_answer`, try `ast.literal_eval` on the
stripped text first (covers numbers, strings, lists, dicts, tuples, bools,
`None` uniformly); fall back to the existing regex-last-number heuristic only
on a `ValueError`/`SyntaxError`, which is exactly what a numeric word-problem
answer with prose around it produces. This is additive — every existing
`exact` task still gets scored the same way, since `literal_eval` will fail on
"3 minutes" and fall through to the current path unchanged. Needs its own
test alongside the existing `test_counting_answer_matches_brute_force`-style
ones (`tests/test_task_set.py`) before any CRUXEval-derived task is added, per
the project's own rule that a check the reference solution can't pass "scores
the task set rather than the model."
## `tool_use_agentic` ← BFCL live_irrelevance
**Source:** `BFCL_v4_live_irrelevance.json` in
[`ShishirPatil/gorilla`](https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard/bfcl_eval/data)
(Apache 2.0). 884 examples, sourced from real user queries rather than
synthesized — pull this file over the smaller 239-example synthetic
`BFCL_v4_irrelevance.json` for the same reason this project trusts
`POST /outcome` traffic over benchmark scores: real queries beat invented
ones. This is the single best-matched source of the four — it's a large pool
of *exactly* the pattern this project already hand-writes 2 of 3 tasks in:
offer a tool, correct behavior is not calling it.
**Schema gap to close** (confirmed by fetching a live sample): BFCL's own
function schema is not the OpenAI tool-call schema this project uses —
`"type": "dict"` where OpenAI/this project's YAML says `"type": "object"`, no
`{"type": "function", "function": {...}}` wrapper, and `question` is a list of
turns (list of message dicts) rather than this project's flat `prompt`
string. Translation per selected example:
```python
tools = [
{"type": "function", "function": {
"name": fn["name"],
"description": fn["description"],
"parameters": {**fn["parameters"], "type": "object"}, # dict -> object
}}
for fn in example["function"]
]
prompt = example["question"][0][0]["content"] # single-turn BFCL cases only
```
**The balance trap.** Every BFCL *irrelevance* example is, by definition, a
"correctly abstain" case. If the new batch is 100% abstain-cases, a model
that never calls a tool scores 1.0 on the enlarged category despite being
just as broken as one that always calls a tool — the category would stop
measuring over-triggering and start measuring only under-triggering.
`score_tool` already has both directions (`expect_tool: null` vs. a named
tool), and the current 3-task set is already 1:2 (call : abstain). Keep new
additions roughly in that band or better: pull irrelevance examples for the
new abstain cases, and pull a matching number of **positive** call cases from
BFCL's `live_simple`/`live_multiple` files (same repo, same schema, translate
the same way) so the enlarged category still rewards correct discrimination,
not just caution.
**Worked example**, adapted from a real fetched entry:
```yaml
- id: tool_geocode_irrelevant
category: tool_use_agentic
kind: tool
prompt: |
Can you provide the address for latitude 37.4224764 and longitude
-122.0842499 using the Geocoding API?
expect_tool: null # no offered function does geocoding
tools:
- type: function
function:
name: requests.get
description: Sends a GET request to the specified URL to retrieve data.
parameters:
type: object
properties:
url: {type: string, description: The URL to send the GET request to.}
required: [url]
```
## `coding_general` ← CRUXEval-O
**Source:** [`facebookresearch/cruxeval`](https://github.com/facebookresearch/cruxeval)
(MIT). 800 short Python functions with input/output pairs, split into
CRUXEval-I (predict an input that produces a given output) and CRUXEval-O
(predict the output of running the function on a given input). **Use O only.**
I is a genuine structural mismatch: an input that produces a target output is
often not unique, so scoring it needs to actually execute the candidate input
through the real function and compare — that's neither `exact` (no single
golden string to normalize against) nor `code` (the model isn't writing a
function) as this harness defines them. Not worth a third scoring kind for
half of one source; skip I, use O, which is a clean fit for `exact` once the
`ast.literal_eval` fix above lands.
At release, best-of-field models scored 63-67% on CRUXEval-O — meaning it
was built to sit in a discriminating band, not to saturate the way this
project's own three `coding_general` tasks did.
**Worked example** (illustrative shape, not a literal fetched row):
```yaml
- id: cruxeval_trace_dedup
category: coding_general
kind: exact
prompt: |
What does this function return when called as shown? Reply with ONLY
the literal Python value — no explanation, no fences.
def f(lst):
seen = {}
out = []
for x in lst:
if x not in seen:
seen[x] = True
out.append(x * 2)
return out
f([3, 1, 3, 2, 1])
answer: "[6, 2, 4]"
```
## `coding_refactor` / `debugging` ← Exercism, via Aider's selection method
**What Aider's polyglot-benchmark actually gives you.** I checked: the
[`Aider-AI/polyglot-benchmark`](https://github.com/Aider-AI/polyglot-benchmark)
repo's own GitHub license field comes back blank via the API — don't pull
files from it directly. What's genuinely reusable is the *method* their
[Dec 2024 writeup](https://aider.chat/2024/12/21/polyglot.html) describes:
225 of Exercism's 697 exercises, filtered to ones **solved by 3 or fewer of 7
frontier models**, deliberately excluding the 258 "solved by all 7" as too
easy. They didn't publish the 225 names. The underlying exercises are
Exercism's own tracks, confirmed MIT-licensed
(`exercism/python`, etc.) — go there directly, using Aider's selection
criterion (not their file) as the bar: prefer exercises in Exercism's own
`practice/` directories with a track difficulty/reputation for edge-case
traps. `bowling`, `dominoes`, and `affine-cipher` (all present in the
`exercism/python` `practice/` tree, confirmed via the GitHub API) are
reasonable starting picks — each is well known for one nasty, non-obvious
edge case rather than raw difficulty, which is exactly this project's own
stated design principle ("at least one edge case a plausible-looking solution
gets wrong").
**Why this can't be mechanical, unlike the other two sources.** Exercism ships
a spec + a test suite, not a "here's buggy/ugly code, fix/refactor it"
prompt — that shape doesn't exist upstream. `coding_refactor` and `debugging`
tasks in this project embed the *broken or ugly-but-working* code directly in
the prompt (confirmed from `apply_settings`, `first_match`, `describe` in the
current set) and the checks pin exact edge-case behavior. So the real value
of going to Exercism isn't a copyable prompt, it's **borrowing the hard edge
case**: read the exercise's canonical solution and its test suite (that's the
whole reason the exercise is hard), then hand-author a variant:
- **`coding_refactor` variant:** take the canonical solution, deliberately
restructure it into duplicated/nested-conditional/flag-variable style (same
pattern as the existing three) while preserving every edge case exactly —
the model refactors it back, checks re-run the exercise's own tricky test
cases.
- **`debugging` variant:** take the canonical solution, introduce one subtle
off-by-one/boundary bug that breaks exactly the input the exercise is
famous for tripping people up on (bowling's 10th-frame strike/spare bonus
scoring is the textbook example), leave everything else correct — checks
must include that exact edge case, per `tests/test_task_set.py`'s existing
requirement that every debugging target actually fails pre-fix.
This is real authoring work, not translation — budget it accordingly (see
below). What Exercism/Aider buys is **which** edge cases are worth encoding,
pre-validated by the fact that most frontier models already fail them, rather
than inventing new traps from scratch the way the original 23-task set's
"touching intervals / full semver / late-binding closures" hardening pass
did.
## License / attribution handling
Nothing here is substantial verbatim reproduction, so full license-file
vendoring isn't needed — but this project already has a citation precedent
(`session-classification-cache-ttl.md`'s LLMRouter attribution, and
`PinchConfig`'s docstring crediting llmrouter directly in code). Match it: a
one-line comment above each newly-added task block in `evals/tasks.yaml`
naming the source and license, e.g.:
```yaml
# --- tool_use_agentic (added) — adapted from BFCL live_irrelevance /
# live_simple, Apache 2.0, github.com/ShishirPatil/gorilla ---
```
## Validation obligation (not optional)
`tests/test_task_set.py` requires, per kind:
- `code`: a reference solution that passes every check.
- `exact`: an independent (brute-force or hand-derived) recomputation of the
answer — see `test_counting_answer_matches_brute_force` for the pattern.
- `debugging`: the buggy target must **fail** its own checks pre-fix.
- `coding_refactor`: the ugly target must **pass** its own checks pre-refactor.
This is what already caught one wrong expected-value check in the existing
set (CLAUDE.md: "would have docked every model on a task and been
indistinguishable from genuine difficulty"). Every task added from this spec
needs its matching test — this roughly doubles the authoring cost per task
versus just writing the YAML, and should be budgeted as such rather than
treated as a follow-up.
## Suggested first-pass batch size
This session's full run (23 tasks x 11 models, some judge calls) was 216
billed calls / $0.79. A reasonable first batch — **8 new `tool_use_agentic`
tasks (5 abstain + 3 positive-call), 6 new `coding_general` (CRUXEval-O), 3
new `coding_refactor`, 3 new `debugging`** — is +20 tasks, roughly matching
the size of the existing set. Expect a proportional cost per pass (~$1.50-2),
and **two passes** (not one) to clear `self_eval_min_samples: 10` for the
newly-added tasks, since one pass only gives n=1 per new task per model.
## Recommendation
Build in this order: (1) the `score_exact`/`ast.literal_eval` fix plus its
test, since CRUXEval-O tasks are silently wrong without it; (2) BFCL-derived
`tool_use_agentic` tasks, since that category already has the strongest
reproduced signal from this session's run and BFCL requires no hand-authoring
beyond schema translation; (3) CRUXEval-O `coding_general` tasks; (4) the
Exercism-sourced `coding_refactor`/`debugging` variants last, since they're
the only genuinely bespoke authoring work and benefit from the other three
being done first as a format reference. Run `eval_proficiency` twice after
each category lands rather than once at the end, so a broken check (per the
validation section above) is caught against one category's blast radius, not
all four at once.