Add benchmark-sourced eval tasks (BFCL, CRUXEval-O, Exercism) #11
Reference in New Issue
Block a user
Delete Branch "benchmark-sourced-eval-tasks"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
This PR hardens the self-eval task set with 20 benchmark-sourced tasks across four categories.
What changed
Live validation
Eight scoped passes were run against the four touched categories. Evidence logs are in (gitignored).
Key results:
Note: some provider lines occurred during the coding_refactor/debugging passes — these are model-specific transient errors, not task bugs. One task, , scored 1.00 for all models and is flagged as a too-easy follow-up item.
All directly evaluated identities now have in the four categories, so self-eval blending at weight 0.7 is active for them.
Follow-ups
Attribution
Post-implementation review: 5 correctness bugs found & fixed (commit
dcbbc84)An independent review of the
score_exact/normalize_answerrewrite (the originalast.literal_evalchange) found real correctness bugs. All were independently reproduced and fixed:math_countingcomma regression (was breaking a live pre-existing task). "the combinations is 4,536." was scored on536, not4536-> 0.0 for a correct answer. Fixed: comma-thousands are stripped in the prose fallback again."False."failed to match answerFalse. Fixed: value keywords (true/false/none) normalize to their literal repr.repr(True)!=repr(1)already prevented cross-match.)[3,000, 4,000]) is latent: a model emitting comma-thousands inside a list is deterministically scored mismatched against the true valid-literal answer, which is the correct outcome (we must not "repair" formatting errors and reward them). No current task answer contains comma-thousands.Full offline suite: 802 passed. Regression tests added for every reproduced case.
No re-run of the live eval passes was needed: the fix does not change scoring for any already-scored valid-literal CRUXEval-O/BFCL answer.
Correction to my earlier comment (commit
8a034fc): bug #4 is now actually fixedMy earlier comment claimed the nested-comma-in-list bug (
[3,000, 4,000]->[3, 0, 4, 0]) was "correct-by-construction" and left intentionally. That claim was wrong.Re-verification showed it was a genuine silent misparse, unchanged by the first fix: the bracket-guard only gated the prose fallback branch, which that input never reached because
literal_eval("[3,000, 4,000]")"succeeds" first (Python parses3,000inside a list as3+ leading-zero000=[3, 0, 4, 0]). A model writing comma-thousands in a list answer was silently mangled instead of matching the true literal.Now fixed in commit
8a034fc: a digit-grouping-aware regex ((?<=\d),(\d{3})(?!\d)) collapses thousands-separator commas globally beforeliteral_eval:[3,000, 4,000]->[3000, 4000], andscore_exact("[3,000, 4,000]", "[3000, 4000]")->1.0[12,34],[(1, 2), (1, 2)],[1234, 5678],[3, 0, 4, 0],a,bcdall left untouchedtest_repro_nested_comma_thousands_matchthat asserts the real failure mode.Full suite: 802 passed. This closes all 5 findings in the original review.