Add benchmark-sourced eval tasks (BFCL, CRUXEval-O, Exercism) #11

Merged
alee merged 12 commits from benchmark-sourced-eval-tasks into neuralwatt-router-service 2026-08-30 21:22:38 +00:00
Owner

This PR hardens the self-eval task set with 20 benchmark-sourced tasks across four categories.

What changed

  • fix: now parses Python literals via first, with the legacy number-extraction fallback preserved. Without this, CRUXEval-O list/tuple/dict answers would be silently mis-scored.
  • feat: 8 new tasks from BFCL v4 (5 abstain + 3 positive-call), translated with its sibling ground-truth file.
  • feat: 6 new exact tasks from CRUXEval-O output prediction.
  • feat: 6 new refactor/debug tasks from Exercism (bowling, dominoes, affine-cipher).
  • docs: committed the source brief and synced task/test counts.

Live validation

Eight scoped passes were run against the four touched categories. Evidence logs are in (gitignored).

Key results:

  • : strong spread; several new tasks discriminate models.
  • : broke the historical 1.00 flatness (e.g. deepseek 0.758, kimi-k3-fast 0.818).
  • / : no longer flat at 1.00; new tasks discriminate.

Note: some provider lines occurred during the coding_refactor/debugging passes — these are model-specific transient errors, not task bugs. One task, , scored 1.00 for all models and is flagged as a too-easy follow-up item.

All directly evaluated identities now have in the four categories, so self-eval blending at weight 0.7 is active for them.

Follow-ups

  • does not discriminate; a stronger bug (e.g. off-by-one in the coprime modulus instead of an inverted guard) should replace it.
  • The legacy base rows and show low inherited samples; this is pre-existing variant propagation behavior.

Attribution

  • BFCL: ShishirPatil/gorilla, Apache 2.0.
  • CRUXEval-O: facebookresearch/cruxeval, MIT.
  • Exercism exercises: exercism/python, MIT.
This PR hardens the self-eval task set with 20 benchmark-sourced tasks across four categories. ## What changed - **fix:** now parses Python literals via first, with the legacy number-extraction fallback preserved. Without this, CRUXEval-O list/tuple/dict answers would be silently mis-scored. - **feat:** 8 new tasks from BFCL v4 (5 abstain + 3 positive-call), translated with its sibling ground-truth file. - **feat:** 6 new exact tasks from CRUXEval-O output prediction. - **feat:** 6 new refactor/debug tasks from Exercism (bowling, dominoes, affine-cipher). - **docs:** committed the source brief and synced task/test counts. ## Live validation Eight scoped passes were run against the four touched categories. Evidence logs are in (gitignored). Key results: - : strong spread; several new tasks discriminate models. - : broke the historical 1.00 flatness (e.g. deepseek 0.758, kimi-k3-fast 0.818). - / : no longer flat at 1.00; new tasks discriminate. Note: some provider lines occurred during the coding_refactor/debugging passes — these are model-specific transient errors, not task bugs. One task, , scored 1.00 for all models and is flagged as a too-easy follow-up item. All directly evaluated identities now have in the four categories, so self-eval blending at weight 0.7 is active for them. ## Follow-ups - does not discriminate; a stronger bug (e.g. off-by-one in the coprime modulus instead of an inverted guard) should replace it. - The legacy base rows and show low inherited samples; this is pre-existing variant propagation behavior. ## Attribution - BFCL: ShishirPatil/gorilla, Apache 2.0. - CRUXEval-O: facebookresearch/cruxeval, MIT. - Exercism exercises: exercism/python, MIT.
alee added 8 commits 2026-08-30 20:05:22 +00:00
alee added 2 commits 2026-08-30 20:36:59 +00:00
Adds the missing feature writeup for circuit_breaker.py (enabled by default
in PR #10 but never documented) plus the mechanism finding that eval calls
correctly bypass it: pinned dispatch has no `alternatives` to fail over to,
so routing eval through the dispatcher would add no benefit and one real
downside — eval-induced failures poisoning circuit-breaker state for
production traffic.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
Author
Owner

Post-implementation review: 5 correctness bugs found & fixed (commit dcbbc84)

An independent review of the score_exact/normalize_answer rewrite (the original ast.literal_eval change) found real correctness bugs. All were independently reproduced and fixed:

  1. math_counting comma regression (was breaking a live pre-existing task). "the combinations is 4,536." was scored on 536, not 4536 -> 0.0 for a correct answer. Fixed: comma-thousands are stripped in the prose fallback again.
  2. Unconditional fence stripping. A model that shows fenced scratch work then states the answer in prose was scored on the scratch. Fixed: fences are honored only when they wrap the whole reply.
  3. Bool case-mismatch. "False." failed to match answer False. Fixed: value keywords (true/false/none) normalize to their literal repr.
  4. Dead bool/number guard removed. (repr(True) != repr(1) already prevented cross-match.)
  5. Nested comma-thousands (e.g. [3,000, 4,000]) is latent: a model emitting comma-thousands inside a list is deterministically scored mismatched against the true valid-literal answer, which is the correct outcome (we must not "repair" formatting errors and reward them). No current task answer contains comma-thousands.

Full offline suite: 802 passed. Regression tests added for every reproduced case.

No re-run of the live eval passes was needed: the fix does not change scoring for any already-scored valid-literal CRUXEval-O/BFCL answer.

## Post-implementation review: 5 correctness bugs found & fixed (commit dcbbc84) An independent review of the `score_exact`/`normalize_answer` rewrite (the original `ast.literal_eval` change) found real correctness bugs. All were independently reproduced and fixed: 1. **`math_counting` comma regression (was breaking a live pre-existing task).** "the combinations is 4,536." was scored on `536`, not `4536` -> 0.0 for a correct answer. Fixed: comma-thousands are stripped in the prose fallback again. 2. **Unconditional fence stripping.** A model that shows fenced scratch work then states the answer in prose was scored on the scratch. Fixed: fences are honored only when they wrap the whole reply. 3. **Bool case-mismatch.** `"False."` failed to match answer `False`. Fixed: value keywords (`true`/`false`/`none`) normalize to their literal repr. 4. **Dead bool/number guard removed.** (`repr(True)` != `repr(1)` already prevented cross-match.) 5. **Nested comma-thousands** (e.g. `[3,000, 4,000]`) is latent: a model emitting comma-thousands inside a list is deterministically scored mismatched against the true valid-literal answer, which is the correct outcome (we must not "repair" formatting errors and reward them). No current task answer contains comma-thousands. Full offline suite: 802 passed. Regression tests added for every reproduced case. No re-run of the live eval passes was needed: the fix does not change scoring for any already-scored valid-literal CRUXEval-O/BFCL answer.
alee added 1 commit 2026-08-30 21:06:33 +00:00
Author
Owner

Correction to my earlier comment (commit 8a034fc): bug #4 is now actually fixed

My earlier comment claimed the nested-comma-in-list bug ([3,000, 4,000] -> [3, 0, 4, 0]) was "correct-by-construction" and left intentionally. That claim was wrong.

Re-verification showed it was a genuine silent misparse, unchanged by the first fix: the bracket-guard only gated the prose fallback branch, which that input never reached because literal_eval("[3,000, 4,000]") "succeeds" first (Python parses 3,000 inside a list as 3 + leading-zero 000 = [3, 0, 4, 0]). A model writing comma-thousands in a list answer was silently mangled instead of matching the true literal.

Now fixed in commit 8a034fc: a digit-grouping-aware regex ((?<=\d),(\d{3})(?!\d)) collapses thousands-separator commas globally before literal_eval:

  • [3,000, 4,000] -> [3000, 4000], and score_exact("[3,000, 4,000]", "[3000, 4000]") -> 1.0
  • No false positives: [12,34], [(1, 2), (1, 2)], [1234, 5678], [3, 0, 4, 0], a,bcd all left untouched
  • The weak regression test was replaced with test_repro_nested_comma_thousands_match that asserts the real failure mode.

Full suite: 802 passed. This closes all 5 findings in the original review.

## Correction to my earlier comment (commit 8a034fc): bug #4 is now actually fixed My earlier comment claimed the nested-comma-in-list bug (`[3,000, 4,000]` -> `[3, 0, 4, 0]`) was "correct-by-construction" and left intentionally. **That claim was wrong.** Re-verification showed it was a genuine silent misparse, unchanged by the first fix: the bracket-guard only gated the prose fallback branch, which that input never reached because `literal_eval("[3,000, 4,000]")` "succeeds" first (Python parses `3,000` inside a list as `3` + leading-zero `000` = `[3, 0, 4, 0]`). A model writing comma-thousands in a list answer was silently mangled instead of matching the true literal. **Now fixed in commit 8a034fc:** a digit-grouping-aware regex (`(?<=\d),(\d{3})(?!\d)`) collapses thousands-separator commas globally before `literal_eval`: - `[3,000, 4,000]` -> `[3000, 4000]`, and `score_exact("[3,000, 4,000]", "[3000, 4000]")` -> `1.0` - No false positives: `[12,34]`, `[(1, 2), (1, 2)]`, `[1234, 5678]`, `[3, 0, 4, 0]`, `a,bcd` all left untouched - The weak regression test was replaced with `test_repro_nested_comma_thousands_match` that asserts the real failure mode. Full suite: 802 passed. This closes all 5 findings in the original review.
alee added 1 commit 2026-08-30 21:17:37 +00:00
alee merged commit 8d1b223051 into neuralwatt-router-service 2026-08-30 21:22:38 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: alee/6krrt#11