# Spec: sourcing harder self-eval tasks from external benchmarks Status: done -- evals/tasks.yaml, 43 tasks **Origin.** `PYTHONPATH=src python -m eval_proficiency` was re-run in full this session (11 models x 23 tasks, 216 billed calls, $0.79 / 0.098 kWh, 0 call failures) specifically to see whether more samples of the *existing* task set would break the ties CLAUDE.md already flagged as suspicious. Result, doubling n from 2-3 to 4-6 per model per category: - `tool_use_agentic`: `deepseek-v4-flash` held at exactly 2/6, the same ratio as the earlier 1/3 — real signal, not n=3 noise, but still short of `self_eval_min_samples: 10`. - `coding_general`, `coding_refactor`, `debugging`: **still perfectly flat at 1.00** for every model. This is no longer "not enough samples" — it's "these 3 tasks per category don't discriminate anything, at any sample count." More passes over the same 23 tasks won't fix that; the task set itself needs to get harder. This spec is the result of "where do we get harder tasks without inventing plausible-looking difficulty from nothing" — the same objection this project already raised against `leaderboards.yaml`'s empty priors. No code changes accompany this document — this is the spec opencode builds from. --- ## Non-goals, stated up front - **No new categories.** `proficiency.categories` in `config.py` is read throughout scoring/tiering/routing; adding one is a real migration, not a task-authoring change, and nothing here needs it. Every source below gets mapped onto one of the 9 existing categories. - **No new scoring kinds.** Everything maps onto `code`, `exact`, or `tool`. (`judge` categories — `docs_writing`, `general_chat`, `summarization`, `translation`, `general_chat` — aren't the problem this session found; skip them for this pass.) - **No repo-context benchmarks.** This is the reason SWE-bench (and anything else shaped like "here's a git checkout, go fix an issue") is not on the list below despite being the best-known real-bug-fixing benchmark. Every task in `evals/tasks.yaml` is one self-contained prompt scored by one subprocess call or one structural check — the harness has no notion of a repo checkout, and building one is a much bigger lift than the actual problem (three flat categories) justifies. - **No verbatim copying.** Every source below gets *adapted* — either translated into this project's YAML+checks shape, or (for `coding_refactor`/`debugging`) used only to identify which underlying problems are hard enough to be worth hand-authoring a variant of. Nothing here is "download file, paste into tasks.yaml." ## A code change this needs, before any task uses it `score_exact` (`eval_proficiency.py:191`) normalizes by regexing out the **last number** in the reply (`normalize_answer`). That's correct for this project's current `exact` tasks — they're all arithmetic word problems with a scalar numeric answer (`math_percent_trap`, `math_rate_trap`, `math_counting`). It is **not** correct for CRUXEval-style tasks (below), where the expected value is a Python literal that can be a list, dict, tuple, string, or `None` — e.g. an answer of `[1, 2, 3]` would get reduced to `3`, silently discarding the structure and scoring against the wrong thing. Fix, scoped and small: in `normalize_answer`, try `ast.literal_eval` on the stripped text first (covers numbers, strings, lists, dicts, tuples, bools, `None` uniformly); fall back to the existing regex-last-number heuristic only on a `ValueError`/`SyntaxError`, which is exactly what a numeric word-problem answer with prose around it produces. This is additive — every existing `exact` task still gets scored the same way, since `literal_eval` will fail on "3 minutes" and fall through to the current path unchanged. Needs its own test alongside the existing `test_counting_answer_matches_brute_force`-style ones (`tests/test_task_set.py`) before any CRUXEval-derived task is added, per the project's own rule that a check the reference solution can't pass "scores the task set rather than the model." ## `tool_use_agentic` ← BFCL live_irrelevance **Source:** `BFCL_v4_live_irrelevance.json` in [`ShishirPatil/gorilla`](https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard/bfcl_eval/data) (Apache 2.0). 884 examples, sourced from real user queries rather than synthesized — pull this file over the smaller 239-example synthetic `BFCL_v4_irrelevance.json` for the same reason this project trusts `POST /outcome` traffic over benchmark scores: real queries beat invented ones. This is the single best-matched source of the four — it's a large pool of *exactly* the pattern this project already hand-writes 2 of 3 tasks in: offer a tool, correct behavior is not calling it. **Schema gap to close** (confirmed by fetching a live sample): BFCL's own function schema is not the OpenAI tool-call schema this project uses — `"type": "dict"` where OpenAI/this project's YAML says `"type": "object"`, no `{"type": "function", "function": {...}}` wrapper, and `question` is a list of turns (list of message dicts) rather than this project's flat `prompt` string. Translation per selected example: ```python tools = [ {"type": "function", "function": { "name": fn["name"], "description": fn["description"], "parameters": {**fn["parameters"], "type": "object"}, # dict -> object }} for fn in example["function"] ] prompt = example["question"][0][0]["content"] # single-turn BFCL cases only ``` **The balance trap.** Every BFCL *irrelevance* example is, by definition, a "correctly abstain" case. If the new batch is 100% abstain-cases, a model that never calls a tool scores 1.0 on the enlarged category despite being just as broken as one that always calls a tool — the category would stop measuring over-triggering and start measuring only under-triggering. `score_tool` already has both directions (`expect_tool: null` vs. a named tool), and the current 3-task set is already 1:2 (call : abstain). Keep new additions roughly in that band or better: pull irrelevance examples for the new abstain cases, and pull a matching number of **positive** call cases from BFCL's `live_simple`/`live_multiple` files (same repo, same schema, translate the same way) so the enlarged category still rewards correct discrimination, not just caution. **Worked example**, adapted from a real fetched entry: ```yaml - id: tool_geocode_irrelevant category: tool_use_agentic kind: tool prompt: | Can you provide the address for latitude 37.4224764 and longitude -122.0842499 using the Geocoding API? expect_tool: null # no offered function does geocoding tools: - type: function function: name: requests.get description: Sends a GET request to the specified URL to retrieve data. parameters: type: object properties: url: {type: string, description: The URL to send the GET request to.} required: [url] ``` ## `coding_general` ← CRUXEval-O **Source:** [`facebookresearch/cruxeval`](https://github.com/facebookresearch/cruxeval) (MIT). 800 short Python functions with input/output pairs, split into CRUXEval-I (predict an input that produces a given output) and CRUXEval-O (predict the output of running the function on a given input). **Use O only.** I is a genuine structural mismatch: an input that produces a target output is often not unique, so scoring it needs to actually execute the candidate input through the real function and compare — that's neither `exact` (no single golden string to normalize against) nor `code` (the model isn't writing a function) as this harness defines them. Not worth a third scoring kind for half of one source; skip I, use O, which is a clean fit for `exact` once the `ast.literal_eval` fix above lands. At release, best-of-field models scored 63-67% on CRUXEval-O — meaning it was built to sit in a discriminating band, not to saturate the way this project's own three `coding_general` tasks did. **Worked example** (illustrative shape, not a literal fetched row): ```yaml - id: cruxeval_trace_dedup category: coding_general kind: exact prompt: | What does this function return when called as shown? Reply with ONLY the literal Python value — no explanation, no fences. def f(lst): seen = {} out = [] for x in lst: if x not in seen: seen[x] = True out.append(x * 2) return out f([3, 1, 3, 2, 1]) answer: "[6, 2, 4]" ``` ## `coding_refactor` / `debugging` ← Exercism, via Aider's selection method **What Aider's polyglot-benchmark actually gives you.** I checked: the [`Aider-AI/polyglot-benchmark`](https://github.com/Aider-AI/polyglot-benchmark) repo's own GitHub license field comes back blank via the API — don't pull files from it directly. What's genuinely reusable is the *method* their [Dec 2024 writeup](https://aider.chat/2024/12/21/polyglot.html) describes: 225 of Exercism's 697 exercises, filtered to ones **solved by 3 or fewer of 7 frontier models**, deliberately excluding the 258 "solved by all 7" as too easy. They didn't publish the 225 names. The underlying exercises are Exercism's own tracks, confirmed MIT-licensed (`exercism/python`, etc.) — go there directly, using Aider's selection criterion (not their file) as the bar: prefer exercises in Exercism's own `practice/` directories with a track difficulty/reputation for edge-case traps. `bowling`, `dominoes`, and `affine-cipher` (all present in the `exercism/python` `practice/` tree, confirmed via the GitHub API) are reasonable starting picks — each is well known for one nasty, non-obvious edge case rather than raw difficulty, which is exactly this project's own stated design principle ("at least one edge case a plausible-looking solution gets wrong"). **Why this can't be mechanical, unlike the other two sources.** Exercism ships a spec + a test suite, not a "here's buggy/ugly code, fix/refactor it" prompt — that shape doesn't exist upstream. `coding_refactor` and `debugging` tasks in this project embed the *broken or ugly-but-working* code directly in the prompt (confirmed from `apply_settings`, `first_match`, `describe` in the current set) and the checks pin exact edge-case behavior. So the real value of going to Exercism isn't a copyable prompt, it's **borrowing the hard edge case**: read the exercise's canonical solution and its test suite (that's the whole reason the exercise is hard), then hand-author a variant: - **`coding_refactor` variant:** take the canonical solution, deliberately restructure it into duplicated/nested-conditional/flag-variable style (same pattern as the existing three) while preserving every edge case exactly — the model refactors it back, checks re-run the exercise's own tricky test cases. - **`debugging` variant:** take the canonical solution, introduce one subtle off-by-one/boundary bug that breaks exactly the input the exercise is famous for tripping people up on (bowling's 10th-frame strike/spare bonus scoring is the textbook example), leave everything else correct — checks must include that exact edge case, per `tests/test_task_set.py`'s existing requirement that every debugging target actually fails pre-fix. This is real authoring work, not translation — budget it accordingly (see below). What Exercism/Aider buys is **which** edge cases are worth encoding, pre-validated by the fact that most frontier models already fail them, rather than inventing new traps from scratch the way the original 23-task set's "touching intervals / full semver / late-binding closures" hardening pass did. ## License / attribution handling Nothing here is substantial verbatim reproduction, so full license-file vendoring isn't needed — but this project already has a citation precedent (`session-classification-cache-ttl.md`'s LLMRouter attribution, and `PinchConfig`'s docstring crediting llmrouter directly in code). Match it: a one-line comment above each newly-added task block in `evals/tasks.yaml` naming the source and license, e.g.: ```yaml # --- tool_use_agentic (added) — adapted from BFCL live_irrelevance / # live_simple, Apache 2.0, github.com/ShishirPatil/gorilla --- ``` ## Validation obligation (not optional) `tests/test_task_set.py` requires, per kind: - `code`: a reference solution that passes every check. - `exact`: an independent (brute-force or hand-derived) recomputation of the answer — see `test_counting_answer_matches_brute_force` for the pattern. - `debugging`: the buggy target must **fail** its own checks pre-fix. - `coding_refactor`: the ugly target must **pass** its own checks pre-refactor. This is what already caught one wrong expected-value check in the existing set (CLAUDE.md: "would have docked every model on a task and been indistinguishable from genuine difficulty"). Every task added from this spec needs its matching test — this roughly doubles the authoring cost per task versus just writing the YAML, and should be budgeted as such rather than treated as a follow-up. ## Suggested first-pass batch size This session's full run (23 tasks x 11 models, some judge calls) was 216 billed calls / $0.79. A reasonable first batch — **8 new `tool_use_agentic` tasks (5 abstain + 3 positive-call), 6 new `coding_general` (CRUXEval-O), 3 new `coding_refactor`, 3 new `debugging`** — is +20 tasks, roughly matching the size of the existing set. Expect a proportional cost per pass (~$1.50-2), and **two passes** (not one) to clear `self_eval_min_samples: 10` for the newly-added tasks, since one pass only gives n=1 per new task per model. ## Recommendation Build in this order: (1) the `score_exact`/`ast.literal_eval` fix plus its test, since CRUXEval-O tasks are silently wrong without it; (2) BFCL-derived `tool_use_agentic` tasks, since that category already has the strongest reproduced signal from this session's run and BFCL requires no hand-authoring beyond schema translation; (3) CRUXEval-O `coding_general` tasks; (4) the Exercism-sourced `coding_refactor`/`debugging` variants last, since they're the only genuinely bespoke authoring work and benefit from the other three being done first as a format reference. Run `eval_proficiency` twice after each category lands rather than once at the end, so a broken check (per the validation section above) is caught against one category's blast radius, not all four at once.