plans/ held 58 documents and exactly one said whether it was open. The rest
mixed finished work, reviews of shipped work, parked specs and genuinely
pending ones, with nothing distinguishing them, so "how many plans are in
the queue" had no answer short of reading all 58.
Now `grep -H '^Status:' plans/*.md` is the answer:
50 done 3 in progress 2 planned 2 reference 1 parked
Statuses were derived rather than guessed: CLAUDE.md's own built list and
"What's NOT built yet" section, plus checking the subject exists in the
code. A review of work that shipped counts as done -- it records what was
found, it is not a request for anything. `reference` separates the two docs
that are conventions rather than work items (admin-design-standards,
admin-work-framework), which otherwise read as permanently-open plans.
The vocabulary is deliberately five words. A larger one invites "mostly
done" and "blocked-ish", which is how the directory became unreadable.
test_plans_declare_status.py keeps it from rotting: a new plan without a
marker fails, as does an unknown status, one buried below the eighth line,
or an open status with no reason -- "planned" alone is the state that rots,
since nobody can tell later whether it waits on a decision, a dependency,
or just nobody's turn.
Also updates the sweep plan with what landed and what did not, including
that #9 was not a defect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
283 lines
14 KiB
Markdown
283 lines
14 KiB
Markdown
# Spec: sourcing harder self-eval tasks from external benchmarks
|
|
|
|
Status: done -- evals/tasks.yaml, 43 tasks
|
|
|
|
**Origin.** `PYTHONPATH=src python -m eval_proficiency` was re-run in full this
|
|
session (11 models x 23 tasks, 216 billed calls, $0.79 / 0.098 kWh, 0 call
|
|
failures) specifically to see whether more samples of the *existing* task set
|
|
would break the ties CLAUDE.md already flagged as suspicious. Result, doubling
|
|
n from 2-3 to 4-6 per model per category:
|
|
|
|
- `tool_use_agentic`: `deepseek-v4-flash` held at exactly 2/6, the same ratio
|
|
as the earlier 1/3 — real signal, not n=3 noise, but still short of
|
|
`self_eval_min_samples: 10`.
|
|
- `coding_general`, `coding_refactor`, `debugging`: **still perfectly flat at
|
|
1.00** for every model. This is no longer "not enough samples" — it's
|
|
"these 3 tasks per category don't discriminate anything, at any sample
|
|
count." More passes over the same 23 tasks won't fix that; the task set
|
|
itself needs to get harder.
|
|
|
|
This spec is the result of "where do we get harder tasks without inventing
|
|
plausible-looking difficulty from nothing" — the same objection this project
|
|
already raised against `leaderboards.yaml`'s empty priors. No code changes
|
|
accompany this document — this is the spec opencode builds from.
|
|
|
|
---
|
|
|
|
## Non-goals, stated up front
|
|
|
|
- **No new categories.** `proficiency.categories` in `config.py` is read
|
|
throughout scoring/tiering/routing; adding one is a real migration, not a
|
|
task-authoring change, and nothing here needs it. Every source below gets
|
|
mapped onto one of the 9 existing categories.
|
|
- **No new scoring kinds.** Everything maps onto `code`, `exact`, or `tool`.
|
|
(`judge` categories — `docs_writing`, `general_chat`, `summarization`,
|
|
`translation`, `general_chat` — aren't the problem this session found; skip
|
|
them for this pass.)
|
|
- **No repo-context benchmarks.** This is the reason SWE-bench (and anything
|
|
else shaped like "here's a git checkout, go fix an issue") is not on the
|
|
list below despite being the best-known real-bug-fixing benchmark. Every
|
|
task in `evals/tasks.yaml` is one self-contained prompt scored by one
|
|
subprocess call or one structural check — the harness has no notion of a
|
|
repo checkout, and building one is a much bigger lift than the actual
|
|
problem (three flat categories) justifies.
|
|
- **No verbatim copying.** Every source below gets *adapted* — either
|
|
translated into this project's YAML+checks shape, or (for
|
|
`coding_refactor`/`debugging`) used only to identify which underlying
|
|
problems are hard enough to be worth hand-authoring a variant of. Nothing
|
|
here is "download file, paste into tasks.yaml."
|
|
|
|
## A code change this needs, before any task uses it
|
|
|
|
`score_exact` (`eval_proficiency.py:191`) normalizes by regexing out the
|
|
**last number** in the reply (`normalize_answer`). That's correct for this
|
|
project's current `exact` tasks — they're all arithmetic word problems with a
|
|
scalar numeric answer (`math_percent_trap`, `math_rate_trap`,
|
|
`math_counting`). It is **not** correct for CRUXEval-style tasks (below),
|
|
where the expected value is a Python literal that can be a list, dict, tuple,
|
|
string, or `None` — e.g. an answer of `[1, 2, 3]` would get reduced to `3`,
|
|
silently discarding the structure and scoring against the wrong thing.
|
|
|
|
Fix, scoped and small: in `normalize_answer`, try `ast.literal_eval` on the
|
|
stripped text first (covers numbers, strings, lists, dicts, tuples, bools,
|
|
`None` uniformly); fall back to the existing regex-last-number heuristic only
|
|
on a `ValueError`/`SyntaxError`, which is exactly what a numeric word-problem
|
|
answer with prose around it produces. This is additive — every existing
|
|
`exact` task still gets scored the same way, since `literal_eval` will fail on
|
|
"3 minutes" and fall through to the current path unchanged. Needs its own
|
|
test alongside the existing `test_counting_answer_matches_brute_force`-style
|
|
ones (`tests/test_task_set.py`) before any CRUXEval-derived task is added, per
|
|
the project's own rule that a check the reference solution can't pass "scores
|
|
the task set rather than the model."
|
|
|
|
## `tool_use_agentic` ← BFCL live_irrelevance
|
|
|
|
**Source:** `BFCL_v4_live_irrelevance.json` in
|
|
[`ShishirPatil/gorilla`](https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard/bfcl_eval/data)
|
|
(Apache 2.0). 884 examples, sourced from real user queries rather than
|
|
synthesized — pull this file over the smaller 239-example synthetic
|
|
`BFCL_v4_irrelevance.json` for the same reason this project trusts
|
|
`POST /outcome` traffic over benchmark scores: real queries beat invented
|
|
ones. This is the single best-matched source of the four — it's a large pool
|
|
of *exactly* the pattern this project already hand-writes 2 of 3 tasks in:
|
|
offer a tool, correct behavior is not calling it.
|
|
|
|
**Schema gap to close** (confirmed by fetching a live sample): BFCL's own
|
|
function schema is not the OpenAI tool-call schema this project uses —
|
|
`"type": "dict"` where OpenAI/this project's YAML says `"type": "object"`, no
|
|
`{"type": "function", "function": {...}}` wrapper, and `question` is a list of
|
|
turns (list of message dicts) rather than this project's flat `prompt`
|
|
string. Translation per selected example:
|
|
|
|
```python
|
|
tools = [
|
|
{"type": "function", "function": {
|
|
"name": fn["name"],
|
|
"description": fn["description"],
|
|
"parameters": {**fn["parameters"], "type": "object"}, # dict -> object
|
|
}}
|
|
for fn in example["function"]
|
|
]
|
|
prompt = example["question"][0][0]["content"] # single-turn BFCL cases only
|
|
```
|
|
|
|
**The balance trap.** Every BFCL *irrelevance* example is, by definition, a
|
|
"correctly abstain" case. If the new batch is 100% abstain-cases, a model
|
|
that never calls a tool scores 1.0 on the enlarged category despite being
|
|
just as broken as one that always calls a tool — the category would stop
|
|
measuring over-triggering and start measuring only under-triggering.
|
|
`score_tool` already has both directions (`expect_tool: null` vs. a named
|
|
tool), and the current 3-task set is already 1:2 (call : abstain). Keep new
|
|
additions roughly in that band or better: pull irrelevance examples for the
|
|
new abstain cases, and pull a matching number of **positive** call cases from
|
|
BFCL's `live_simple`/`live_multiple` files (same repo, same schema, translate
|
|
the same way) so the enlarged category still rewards correct discrimination,
|
|
not just caution.
|
|
|
|
**Worked example**, adapted from a real fetched entry:
|
|
|
|
```yaml
|
|
- id: tool_geocode_irrelevant
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: |
|
|
Can you provide the address for latitude 37.4224764 and longitude
|
|
-122.0842499 using the Geocoding API?
|
|
expect_tool: null # no offered function does geocoding
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: requests.get
|
|
description: Sends a GET request to the specified URL to retrieve data.
|
|
parameters:
|
|
type: object
|
|
properties:
|
|
url: {type: string, description: The URL to send the GET request to.}
|
|
required: [url]
|
|
```
|
|
|
|
## `coding_general` ← CRUXEval-O
|
|
|
|
**Source:** [`facebookresearch/cruxeval`](https://github.com/facebookresearch/cruxeval)
|
|
(MIT). 800 short Python functions with input/output pairs, split into
|
|
CRUXEval-I (predict an input that produces a given output) and CRUXEval-O
|
|
(predict the output of running the function on a given input). **Use O only.**
|
|
I is a genuine structural mismatch: an input that produces a target output is
|
|
often not unique, so scoring it needs to actually execute the candidate input
|
|
through the real function and compare — that's neither `exact` (no single
|
|
golden string to normalize against) nor `code` (the model isn't writing a
|
|
function) as this harness defines them. Not worth a third scoring kind for
|
|
half of one source; skip I, use O, which is a clean fit for `exact` once the
|
|
`ast.literal_eval` fix above lands.
|
|
|
|
At release, best-of-field models scored 63-67% on CRUXEval-O — meaning it
|
|
was built to sit in a discriminating band, not to saturate the way this
|
|
project's own three `coding_general` tasks did.
|
|
|
|
**Worked example** (illustrative shape, not a literal fetched row):
|
|
|
|
```yaml
|
|
- id: cruxeval_trace_dedup
|
|
category: coding_general
|
|
kind: exact
|
|
prompt: |
|
|
What does this function return when called as shown? Reply with ONLY
|
|
the literal Python value — no explanation, no fences.
|
|
|
|
def f(lst):
|
|
seen = {}
|
|
out = []
|
|
for x in lst:
|
|
if x not in seen:
|
|
seen[x] = True
|
|
out.append(x * 2)
|
|
return out
|
|
|
|
f([3, 1, 3, 2, 1])
|
|
answer: "[6, 2, 4]"
|
|
```
|
|
|
|
## `coding_refactor` / `debugging` ← Exercism, via Aider's selection method
|
|
|
|
**What Aider's polyglot-benchmark actually gives you.** I checked: the
|
|
[`Aider-AI/polyglot-benchmark`](https://github.com/Aider-AI/polyglot-benchmark)
|
|
repo's own GitHub license field comes back blank via the API — don't pull
|
|
files from it directly. What's genuinely reusable is the *method* their
|
|
[Dec 2024 writeup](https://aider.chat/2024/12/21/polyglot.html) describes:
|
|
225 of Exercism's 697 exercises, filtered to ones **solved by 3 or fewer of 7
|
|
frontier models**, deliberately excluding the 258 "solved by all 7" as too
|
|
easy. They didn't publish the 225 names. The underlying exercises are
|
|
Exercism's own tracks, confirmed MIT-licensed
|
|
(`exercism/python`, etc.) — go there directly, using Aider's selection
|
|
criterion (not their file) as the bar: prefer exercises in Exercism's own
|
|
`practice/` directories with a track difficulty/reputation for edge-case
|
|
traps. `bowling`, `dominoes`, and `affine-cipher` (all present in the
|
|
`exercism/python` `practice/` tree, confirmed via the GitHub API) are
|
|
reasonable starting picks — each is well known for one nasty, non-obvious
|
|
edge case rather than raw difficulty, which is exactly this project's own
|
|
stated design principle ("at least one edge case a plausible-looking solution
|
|
gets wrong").
|
|
|
|
**Why this can't be mechanical, unlike the other two sources.** Exercism ships
|
|
a spec + a test suite, not a "here's buggy/ugly code, fix/refactor it"
|
|
prompt — that shape doesn't exist upstream. `coding_refactor` and `debugging`
|
|
tasks in this project embed the *broken or ugly-but-working* code directly in
|
|
the prompt (confirmed from `apply_settings`, `first_match`, `describe` in the
|
|
current set) and the checks pin exact edge-case behavior. So the real value
|
|
of going to Exercism isn't a copyable prompt, it's **borrowing the hard edge
|
|
case**: read the exercise's canonical solution and its test suite (that's the
|
|
whole reason the exercise is hard), then hand-author a variant:
|
|
|
|
- **`coding_refactor` variant:** take the canonical solution, deliberately
|
|
restructure it into duplicated/nested-conditional/flag-variable style (same
|
|
pattern as the existing three) while preserving every edge case exactly —
|
|
the model refactors it back, checks re-run the exercise's own tricky test
|
|
cases.
|
|
- **`debugging` variant:** take the canonical solution, introduce one subtle
|
|
off-by-one/boundary bug that breaks exactly the input the exercise is
|
|
famous for tripping people up on (bowling's 10th-frame strike/spare bonus
|
|
scoring is the textbook example), leave everything else correct — checks
|
|
must include that exact edge case, per `tests/test_task_set.py`'s existing
|
|
requirement that every debugging target actually fails pre-fix.
|
|
|
|
This is real authoring work, not translation — budget it accordingly (see
|
|
below). What Exercism/Aider buys is **which** edge cases are worth encoding,
|
|
pre-validated by the fact that most frontier models already fail them, rather
|
|
than inventing new traps from scratch the way the original 23-task set's
|
|
"touching intervals / full semver / late-binding closures" hardening pass
|
|
did.
|
|
|
|
## License / attribution handling
|
|
|
|
Nothing here is substantial verbatim reproduction, so full license-file
|
|
vendoring isn't needed — but this project already has a citation precedent
|
|
(`session-classification-cache-ttl.md`'s LLMRouter attribution, and
|
|
`PinchConfig`'s docstring crediting llmrouter directly in code). Match it: a
|
|
one-line comment above each newly-added task block in `evals/tasks.yaml`
|
|
naming the source and license, e.g.:
|
|
|
|
```yaml
|
|
# --- tool_use_agentic (added) — adapted from BFCL live_irrelevance /
|
|
# live_simple, Apache 2.0, github.com/ShishirPatil/gorilla ---
|
|
```
|
|
|
|
## Validation obligation (not optional)
|
|
|
|
`tests/test_task_set.py` requires, per kind:
|
|
- `code`: a reference solution that passes every check.
|
|
- `exact`: an independent (brute-force or hand-derived) recomputation of the
|
|
answer — see `test_counting_answer_matches_brute_force` for the pattern.
|
|
- `debugging`: the buggy target must **fail** its own checks pre-fix.
|
|
- `coding_refactor`: the ugly target must **pass** its own checks pre-refactor.
|
|
|
|
This is what already caught one wrong expected-value check in the existing
|
|
set (CLAUDE.md: "would have docked every model on a task and been
|
|
indistinguishable from genuine difficulty"). Every task added from this spec
|
|
needs its matching test — this roughly doubles the authoring cost per task
|
|
versus just writing the YAML, and should be budgeted as such rather than
|
|
treated as a follow-up.
|
|
|
|
## Suggested first-pass batch size
|
|
|
|
This session's full run (23 tasks x 11 models, some judge calls) was 216
|
|
billed calls / $0.79. A reasonable first batch — **8 new `tool_use_agentic`
|
|
tasks (5 abstain + 3 positive-call), 6 new `coding_general` (CRUXEval-O), 3
|
|
new `coding_refactor`, 3 new `debugging`** — is +20 tasks, roughly matching
|
|
the size of the existing set. Expect a proportional cost per pass (~$1.50-2),
|
|
and **two passes** (not one) to clear `self_eval_min_samples: 10` for the
|
|
newly-added tasks, since one pass only gives n=1 per new task per model.
|
|
|
|
## Recommendation
|
|
|
|
Build in this order: (1) the `score_exact`/`ast.literal_eval` fix plus its
|
|
test, since CRUXEval-O tasks are silently wrong without it; (2) BFCL-derived
|
|
`tool_use_agentic` tasks, since that category already has the strongest
|
|
reproduced signal from this session's run and BFCL requires no hand-authoring
|
|
beyond schema translation; (3) CRUXEval-O `coding_general` tasks; (4) the
|
|
Exercism-sourced `coding_refactor`/`debugging` variants last, since they're
|
|
the only genuinely bespoke authoring work and benefit from the other three
|
|
being done first as a format reference. Run `eval_proficiency` twice after
|
|
each category lands rather than once at the end, so a broken check (per the
|
|
validation section above) is caught against one category's blast radius, not
|
|
all four at once.
|