neuralwatt-router-service #12
@@ -21,7 +21,7 @@ code.
|
||||
| Config | `config/config.yaml` + Pydantic (`src/config.py`), `extra="forbid"` |
|
||||
| HTTP client | `requests` (pinned in `requirements.txt`) — do NOT add httpx2/aiohttp without a requirements bump |
|
||||
| TUI | `textual==8.2.8` — imported only by `tui*.py` modules, never by the dispatch path |
|
||||
| Testing | `pytest`, 733 tests, all offline (no provider or local-model calls) |
|
||||
| Testing | `pytest`, 796 tests, all offline (no provider or local-model calls) |
|
||||
| Dependencies | Pinned. Bump deliberately, never use `>=` |
|
||||
|
||||
## Module map and file boundaries
|
||||
@@ -214,7 +214,7 @@ See `CLAUDE.md` → "What's NOT built yet — pick up here" for the full list.
|
||||
|
||||
## After a code change
|
||||
|
||||
- Run `python -m pytest` — 733 tests, ~35s.
|
||||
- Run `python -m pytest` — 796 tests, ~35s.
|
||||
- If you changed `dispatcher.py`, restart the systemd service:
|
||||
`systemctl --user restart llm-router.service` (it doesn't auto-reload code).
|
||||
- If you changed the TUI, run `PYTHONPATH=src python -m tui` to verify it starts.
|
||||
|
||||
64
CLAUDE.md
64
CLAUDE.md
@@ -216,6 +216,11 @@ rather than from months of history.
|
||||
(13 models x 5 = 65 calls, and the whole sweep cost **under a cent**).
|
||||
- `tiering.py` / `tier.py` — pure tier resolver + the DB pass that applies it.
|
||||
- `routing.py` — pure hard filters and ranking.
|
||||
- `circuit_breaker.py` — passive availability skip for a model that starts
|
||||
5xxing. Pure in-memory module (mirrors `session_cache.py`'s shape:
|
||||
module-level dict, injected time for tests, never imports
|
||||
`dispatcher`/`config`). On by default (`circuit_breaker.enabled: true`).
|
||||
See below for what it does and, deliberately, does not cover.
|
||||
- `dispatcher.py` — FastAPI service. `GET /health`, `POST /route` (classify
|
||||
and pick, no provider call), `POST /dispatch` (route, call, log), plus an
|
||||
OpenAI-compatible `GET /v1/models` and `POST /v1/chat/completions`, and a
|
||||
@@ -282,10 +287,42 @@ rather than from months of history.
|
||||
reset on restart. Persisted config edits limited to an allowlist. Model
|
||||
availability overrides feed routing hard filters. Bucketed history endpoint over
|
||||
`energy_observations` and `route_decisions`.
|
||||
- `tests/` — 733 tests across 41 files, all passing, all offline. Verified on
|
||||
- `tests/` — 796 tests across 40 files, all passing, all offline. Verified on
|
||||
Python 3.10 and 3.14; nothing declares `requires-python`, so 3.10 is the
|
||||
tested floor rather than a promised one.
|
||||
|
||||
### Circuit breaker: passive availability skip, and why the eval harness stays outside it
|
||||
|
||||
`circuit_breaker.py` records nothing until a model actually fails: a 5xx
|
||||
from the upstream call marks `(model_id, provider)` down for
|
||||
`initial_cooldown_seconds` (30s default), doubling on each further failure
|
||||
(`backoff_multiplier`, capped at `max_cooldown_seconds`, 600s default) and
|
||||
clearing on the next success. Recovery is passive by design — no background
|
||||
poller, no health-check loop. `is_down` just compares against `down_until`,
|
||||
so the next real request that would have picked the down model becomes its
|
||||
own recovery probe once the cooldown has passed. Two call sites:
|
||||
`_open_circuits` (`dispatcher.py:1845`) excludes down models from the
|
||||
candidate set during routing/selection, and the dispatch retry loop
|
||||
(`dispatcher.py:2560-2605`) records the failure/success on every upstream
|
||||
call and fails over to the next-ranked candidate on a 5xx — but only when
|
||||
`alternatives` is non-empty, which is exactly the `auto`-routed case.
|
||||
|
||||
**`eval_proficiency.py` intentionally stays outside all of this** — it calls
|
||||
NeuralWatt directly (`requests.post`, not through the dispatcher), and that
|
||||
is correct, not a gap. Its whole design is pinning one specific `model_id`
|
||||
per call; a pinned request has `wants_routing = False`, which means
|
||||
`alternatives = []` at the top of the retry loop — there is no candidate to
|
||||
fail over to even if the call went through the router, so circuit_breaker's
|
||||
recovery machinery would have nothing to do for it. Routing eval traffic
|
||||
through the dispatcher would be actively worse, not merely useless: a
|
||||
transient eval-harness failure would still call `circuit_breaker.
|
||||
record_failure`, which could trip the breaker against real production
|
||||
traffic based on nothing but the benchmark exercising a model's edges.
|
||||
Observed live 2026-08-30: `glm-5.3` (newly released) threw two `CALL FAILED
|
||||
HTTPError`s during a `coding_refactor` eval pass, almost certainly transient
|
||||
provider-side capacity ramp, with zero effect on production routing — which
|
||||
is the isolation working as intended, not a missing integration.
|
||||
|
||||
### Monitoring: route_decisions persistence
|
||||
|
||||
`route_decisions` is the newest observability table (not a scoring input). It
|
||||
@@ -510,11 +547,22 @@ right outcome.
|
||||
`coding_general` now spans 0.86-1.00: `glm-5.2-fast` fell to 0.862 over 29
|
||||
samples folded in by `feedback.py` from an actual agent session, and crossed
|
||||
`self_eval_min_samples` on the way, so it reads `self_eval` rather than
|
||||
`self_eval_thin`. That is the intended shape of this system — the 23-task
|
||||
`self_eval_thin`. That is the intended shape of this system — the 43-task
|
||||
benchmark establishes a floor, and your own traffic is what refines it.
|
||||
`coding_refactor` and `debugging` are still flat at 1.00, awaiting the same
|
||||
treatment.
|
||||
|
||||
**Benchmark-sourced hardening landed, and it broke both remaining ties.** The
|
||||
task set grew from 23 to 43: 8 BFCL tool tasks, 6 CRUXEval-O exact tasks (after
|
||||
the `score_exact` literal-eval fix), and three Exercism refactor/debug pairs.
|
||||
The CRUXEval-O rows split `coding_general` into a 0-1 mix across models — five
|
||||
of six now fail at least one model — and the Exercism refactor rows moved
|
||||
`coding_refactor` off its flat 1.00: bowling 0.25-1.00, dominoes 0.10-1.00,
|
||||
affine 0.56-1.00. The debug pairs split `debugging` the same way (bowling
|
||||
0.80-1.00, dominoes a sharp 0/1 split). Two follow-ups recorded, not deleted:
|
||||
`debug_affine_coprime` and most of the BFCL tasks sat flat at 1.00, so they do
|
||||
not discriminate.
|
||||
|
||||
**A 1.00 can also be a sampling artifact, and `docs_writing` was one.** At 2
|
||||
samples per model the category read 0.70-1.00 with a model at the ceiling, and
|
||||
the router paid for that ceiling: `kimi-k3-fast` won every docs route. Six more
|
||||
@@ -810,7 +858,7 @@ explains a `truncated` verdict and nothing else; it is now scoped to exactly
|
||||
that.
|
||||
|
||||
`feedback.py` folds observed failures into `proficiency`, so routing learns
|
||||
from your traffic rather than only the 23-task benchmark. Only **failures**
|
||||
from your traffic rather than only the 43-task benchmark. Only **failures**
|
||||
are folded in: a structural 'ok' means the code parsed, not that it was
|
||||
correct, and recording those as 1.0 would flatten every score toward the
|
||||
ceiling. Failures the model did not cause — a client's own tight `max_tokens`
|
||||
@@ -1012,11 +1060,11 @@ that request was two orders of magnitude low.
|
||||
There is no composite any more — ranking is quality first, cost as the
|
||||
tiebreak inside `quality_tolerance` — so nothing normalizes and nothing
|
||||
compresses.
|
||||
- Answered: the eval set exists (`evals/tasks.yaml`, 23 tasks, four scoring
|
||||
kinds) and `tests/test_task_set.py` keeps it honest. The open part is
|
||||
narrower now — `coding_refactor` and `debugging` are still flat at 1.00
|
||||
across every model, so those tasks discriminate nothing and either need
|
||||
hardening again or should be conceded as non-discriminating. **Try samples
|
||||
- Answered: the eval set exists (`evals/tasks.yaml`, 43 tasks, four scoring
|
||||
kinds) and `tests/test_task_set.py` keeps it honest. The benchmark-sourced
|
||||
rows now split the coding categories, but two tasks are still flat at 1.00
|
||||
(`debug_affine_coprime` and most of the BFCL set) and either need hardening
|
||||
again or should be conceded as non-discriminating. **Try samples
|
||||
before hardening.** `docs_writing` looked flat at the top too, and six more
|
||||
passes spread it 0.66-0.97 without touching a task; two samples per model is
|
||||
not enough to tell a saturated task from an unsampled one.
|
||||
|
||||
@@ -457,7 +457,7 @@ Gating and safeguards:
|
||||
|
||||
Observation → learning. `feedback.py` folds verification failures into
|
||||
`proficiency` so routing improves on **your traffic**, not just the fixed
|
||||
23-task benchmark:
|
||||
43-task benchmark:
|
||||
|
||||
```bash
|
||||
PYTHONPATH=src python -m feedback --dry-run # preview what would change
|
||||
@@ -544,7 +544,7 @@ Key behaviors:
|
||||
| **Config** | `config/config.yaml` loaded & validated by Pydantic (`src/config.py`) |
|
||||
| **OpenAI Client** | `openai==3.0.0` (official SDK) |
|
||||
| **HTTP** | `requests` for poller, `httpx` (via openai/uvicorn) |
|
||||
| **Testing** | `pytest` — 733 tests across 41 files, all offline |
|
||||
| **Testing** | `pytest` — 796 tests across 40 files, all offline |
|
||||
| **Config Files** | `config/config.yaml`, `config/leaderboards.yaml`, `evals/tasks.yaml` |
|
||||
| **Deployment** | systemd user units (`.service` + `.timer` files in `deploy/`) |
|
||||
| **Integration** | `opencode.json` in the repo routes through it by default; any OpenAI-compatible client works |
|
||||
@@ -1153,7 +1153,7 @@ carries the same modality block.
|
||||
## Testing
|
||||
|
||||
```bash
|
||||
python -m pytest # 733 tests
|
||||
python -m pytest # 796 tests
|
||||
python -m pytest --cov # with coverage
|
||||
```
|
||||
|
||||
|
||||
752
evals/tasks.yaml
752
evals/tasks.yaml
@@ -88,6 +88,119 @@ tasks:
|
||||
- 'word_wrap("a b c", 3) == ["a b", "c"]'
|
||||
- 'word_wrap("aa bb cc", 5) == ["aa bb", "cc"]'
|
||||
|
||||
# --- coding_general (added) — adapted from CRUXEval-O output-prediction,
|
||||
# facebookresearch/cruxeval, MIT License. Data fetched once to /tmp,
|
||||
# never vendored. Filter: has loop, ≤400 LOC chars, no unsafe imports,
|
||||
# safe stdlib only. 6 of 6 tasks use `eval("f(<input>)")` for
|
||||
# offline verification (input is the source arg list, not a single
|
||||
# literal).
|
||||
- id: crux_o_sample_9
|
||||
category: coding_general
|
||||
kind: exact
|
||||
answer: "False"
|
||||
prompt: |
|
||||
What does this function return when called as shown? Reply with ONLY
|
||||
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
||||
'text', None), no explanation, no fences.
|
||||
|
||||
def f(t):
|
||||
for c in t:
|
||||
if not c.isnumeric():
|
||||
return False
|
||||
return True
|
||||
|
||||
f('#284376598')
|
||||
|
||||
- id: crux_o_sample_0
|
||||
category: coding_general
|
||||
kind: exact
|
||||
answer: "[(4, 1), (4, 1), (4, 1), (4, 1), (2, 3), (2, 3)]"
|
||||
prompt: |
|
||||
What does this function return when called as shown? Reply with ONLY
|
||||
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
||||
'text', None), no explanation, no fences.
|
||||
|
||||
def f(nums):
|
||||
output = []
|
||||
for n in nums:
|
||||
output.append((nums.count(n), n))
|
||||
output.sort(reverse=True)
|
||||
return output
|
||||
|
||||
f([1, 1, 3, 1, 3, 1])
|
||||
|
||||
- id: crux_o_sample_1
|
||||
category: coding_general
|
||||
kind: exact
|
||||
answer: "{1: None, 2: None}"
|
||||
prompt: |
|
||||
What does this function return when called as shown? Reply with ONLY
|
||||
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
||||
'text', None), no explanation, no fences.
|
||||
|
||||
def f(a, b, c):
|
||||
result = {}
|
||||
for d in a, b, c:
|
||||
result.update(dict.fromkeys(d))
|
||||
return result
|
||||
|
||||
f((1, ), (1, ), (1, 2))
|
||||
|
||||
- id: crux_o_sample_2
|
||||
category: coding_general
|
||||
kind: exact
|
||||
answer: "'hbtofdeiequ'"
|
||||
prompt: |
|
||||
What does this function return when called as shown? Reply with ONLY
|
||||
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
||||
'text', None), no explanation, no fences.
|
||||
|
||||
def f(text):
|
||||
new_text = list(text)
|
||||
for i in '+':
|
||||
if i in new_text:
|
||||
new_text.remove(i)
|
||||
return ''.join(new_text)
|
||||
|
||||
f('hbtofdeiequ')
|
||||
|
||||
- id: crux_o_sample_5
|
||||
category: coding_general
|
||||
kind: exact
|
||||
answer: "(0, 'xxxxxxxxxxxxxxxxxx')"
|
||||
prompt: |
|
||||
What does this function return when called as shown? Reply with ONLY
|
||||
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
||||
'text', None), no explanation, no fences.
|
||||
|
||||
def f(text, lower, upper):
|
||||
count = 0
|
||||
new_text = list()
|
||||
for char in text:
|
||||
char = lower if char.isdecimal() else upper
|
||||
if char in ['p', 'C']:
|
||||
count += 1
|
||||
new_text.append(char)
|
||||
return count, ''.join(new_text)
|
||||
|
||||
f('DSUWeqExTQdCMGpqur', 'a', 'x')
|
||||
|
||||
- id: crux_o_sample_6
|
||||
category: coding_general
|
||||
kind: exact
|
||||
answer: "[('74', 31)]"
|
||||
prompt: |
|
||||
What does this function return when called as shown? Reply with ONLY
|
||||
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
||||
'text', None), no explanation, no fences.
|
||||
|
||||
def f(dic):
|
||||
for k,v in sorted(dic.items(), key=lambda x: len(str(x)))[:-1]:
|
||||
dic.pop(k)
|
||||
return list(dic.items())
|
||||
|
||||
f({'11': 52, '65': 34, 'a': 12, '4': 52, '74': 31})
|
||||
|
||||
# --- coding_refactor ----------------------------------------------------
|
||||
- id: refactor_falsy_defaults
|
||||
category: coding_refactor
|
||||
@@ -178,6 +291,430 @@ tasks:
|
||||
- 'describe(0) == "unknown"'
|
||||
- 'describe(None) == "unknown"'
|
||||
|
||||
# --- bowling (Exercism-derived) -----------------------------------------
|
||||
# Exercism python bowling exercise — MIT licensed, canonical source at
|
||||
# https://github.com/exercism/python/tree/1f6aab8667bf653b10cc3799f94352fcdb749db6/exercises/practice/bowling
|
||||
- id: refactor_bowling_frames
|
||||
category: coding_refactor
|
||||
kind: code
|
||||
entrypoint: BowlingGame
|
||||
prompt: |
|
||||
Refactor this BowlingGame to remove the duplication and nested
|
||||
conditions. Behaviour must be preserved EXACTLY, including scoring,
|
||||
bonuses, and error cases. Reply with ONLY the rewritten class — no
|
||||
explanation, no fences.
|
||||
|
||||
class BowlingGame:
|
||||
def __init__(self):
|
||||
self._frames = []
|
||||
self._current = 0
|
||||
self._bonus = []
|
||||
|
||||
def roll(self, pins):
|
||||
if not (0 <= pins <= 10):
|
||||
raise ValueError('invalid pins')
|
||||
if self._current < 10:
|
||||
if len(self._frames) == self._current:
|
||||
self._frames.append([pins])
|
||||
else:
|
||||
self._frames[self._current].append(pins)
|
||||
current = self._frames[self._current]
|
||||
if sum(current) > 10:
|
||||
raise ValueError("a frame's rolls cannot exceed 10")
|
||||
strike = (len(current) == 1 and current[0] == 10)
|
||||
if strike or len(current) == 2:
|
||||
self._current += 1
|
||||
else:
|
||||
last = self._frames[-1]
|
||||
last_total = sum(last)
|
||||
strike10 = len(last) == 1 and last[0] == 10
|
||||
spare10 = len(last) == 2 and last_total == 10
|
||||
if strike10:
|
||||
if len(self._bonus) >= 2:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a strike')
|
||||
self._bonus.append(pins)
|
||||
if len(self._bonus) == 2 and self._bonus[0] != 10 and sum(self._bonus) > 10:
|
||||
raise ValueError('invalid fill balls')
|
||||
if len(self._bonus) > 2:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a strike')
|
||||
elif spare10:
|
||||
if len(self._bonus) >= 1:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a spare')
|
||||
self._bonus.append(pins)
|
||||
else:
|
||||
raise IndexError('cannot throw bonus with an open tenth frame')
|
||||
|
||||
def score(self):
|
||||
if self._current < 10:
|
||||
raise IndexError('frame less than 10')
|
||||
last = self._frames[-1]
|
||||
if len(last) == 2 and sum(last) == 10 and len(self._bonus) != 1:
|
||||
raise IndexError('one bonus must be rolled when the tenth frame is spare')
|
||||
if len(last) == 1 and last[0] == 10 and len(self._bonus) != 2:
|
||||
raise IndexError('two bonuses must be rolled when the tenth frame is strike')
|
||||
total = 0
|
||||
for i in range(10):
|
||||
frame = self._frames[i]
|
||||
frame_sum = sum(frame)
|
||||
strike = (len(frame) == 1 and frame[0] == 10)
|
||||
spare = (len(frame) == 2 and frame_sum == 10)
|
||||
if strike or spare:
|
||||
nxt = []
|
||||
for j in range(i + 1, 10):
|
||||
nxt.extend(self._frames[j])
|
||||
nxt.extend(self._bonus)
|
||||
if strike:
|
||||
frame_sum += sum(nxt[:2])
|
||||
else:
|
||||
frame_sum += sum(nxt[:1])
|
||||
total += frame_sum
|
||||
return total
|
||||
checks:
|
||||
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 7, 1]], g)[2].score() == 18)"
|
||||
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [7, 3, 7]], g)[2].score() == 17)"
|
||||
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 10, 10]], g)[2].score() == 30)"
|
||||
- "((g := BowlingGame(), [g.roll(r) for r in [6, 4, 3] + [0]*17], g)[2].score() == 16)"
|
||||
- "((g := BowlingGame(), [g.roll(r) for r in [3, 6]*10], g)[2].score() == 90)"
|
||||
- "((g := BowlingGame(), [g.roll(r) for r in [10]*12], g)[2].score() == 300)"
|
||||
- "raises(Exception, lambda: BowlingGame().roll(-1))"
|
||||
- "raises(Exception, lambda: ((g := BowlingGame()), [g.roll(r) for r in [0]*20], g[2].roll(0)))"
|
||||
|
||||
- id: debug_bowling_tenth_frame
|
||||
category: debugging
|
||||
kind: code
|
||||
entrypoint: BowlingGame
|
||||
prompt: |
|
||||
This BowlingGame is wrong on one subtle tenth-frame case. Fix ONLY the
|
||||
bug; do not change anything else. Reply with ONLY the corrected class —
|
||||
no explanation, no fences.
|
||||
|
||||
class BowlingGame:
|
||||
def __init__(self):
|
||||
self.current_frame_idx = 0
|
||||
self.bonus_throws = []
|
||||
self.frames = [Frame(idx) for idx in range(10)]
|
||||
|
||||
@property
|
||||
def current_frame(self):
|
||||
return self.frames[self.current_frame_idx]
|
||||
|
||||
def next_throws(self, frame_idx):
|
||||
throws = []
|
||||
for idx in range(frame_idx + 1, 10):
|
||||
throws.extend(self.frames[idx].throws)
|
||||
throws.extend(self.bonus_throws)
|
||||
return throws
|
||||
|
||||
def roll_bonus(self, pins):
|
||||
tenth_frame = self.frames[-1]
|
||||
if tenth_frame.is_open():
|
||||
raise IndexError('cannot throw bonus with an open tenth frame')
|
||||
self.bonus_throws.append(pins)
|
||||
if tenth_frame.is_strike() and len(self.bonus_throws) > 2:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a strike')
|
||||
elif tenth_frame.is_spare() and len(self.bonus_throws) > 1:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a spare')
|
||||
|
||||
def roll(self, pins):
|
||||
if not 0 <= pins <= 10:
|
||||
raise ValueError('invalid pins')
|
||||
elif self.current_frame_idx == 10:
|
||||
self.roll_bonus(pins)
|
||||
else:
|
||||
self.current_frame.throw(pins)
|
||||
if self.current_frame.is_closed():
|
||||
self.current_frame_idx += 1
|
||||
|
||||
def score(self):
|
||||
if self.current_frame_idx < 10:
|
||||
raise IndexError('frame less than 10')
|
||||
if self.frames[-1].is_spare() and len(self.bonus_throws) != 1:
|
||||
raise IndexError(
|
||||
'one bonus must be rolled when the tenth frame is spare')
|
||||
if self.frames[-1].is_strike() and len(self.bonus_throws) != 2:
|
||||
raise IndexError(
|
||||
'two bonuses must be rolled when the tenth frame is strike')
|
||||
return sum(frame.score(self.next_throws(frame.idx))
|
||||
for frame in self.frames)
|
||||
|
||||
|
||||
class Frame:
|
||||
def __init__(self, idx):
|
||||
self.idx = idx
|
||||
self.throws = []
|
||||
|
||||
@property
|
||||
def total_pins(self):
|
||||
return sum(self.throws)
|
||||
|
||||
def is_strike(self):
|
||||
return self.total_pins == 10 and len(self.throws) == 1
|
||||
|
||||
def is_spare(self):
|
||||
return self.total_pins == 10 and len(self.throws) == 2
|
||||
|
||||
def is_open(self):
|
||||
return self.total_pins < 10 and len(self.throws) == 2
|
||||
|
||||
def is_closed(self):
|
||||
return self.total_pins == 10 or len(self.throws) == 2
|
||||
|
||||
def throw(self, pins):
|
||||
if self.total_pins + pins > 10:
|
||||
raise ValueError("a frame's rolls cannot exceed 10")
|
||||
self.throws.append(pins)
|
||||
|
||||
def score(self, next_throws):
|
||||
result = self.total_pins
|
||||
if self.is_strike():
|
||||
result += sum(next_throws[:2])
|
||||
elif self.is_spare():
|
||||
result += sum(next_throws[:1])
|
||||
return result
|
||||
checks:
|
||||
- "raises(Exception, (lambda: ((g := BowlingGame()), [g.roll(r) for r in [0]*18 + [10, 5, 6]], g)[2].score()))"
|
||||
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 10, 6]], g)[2].score() == 26)"
|
||||
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 7, 1]], g)[2].score() == 18)"
|
||||
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [7, 3, 7]], g)[2].score() == 17)"
|
||||
- "((g := BowlingGame(), [g.roll(r) for r in [10]*12], g)[2].score() == 300)"
|
||||
|
||||
# --- dominoes (Exercism-derived) ----------------------------------------
|
||||
# Exercism python dominoes exercise — MIT licensed, canonical source at
|
||||
# https://github.com/exercism/python/tree/1f6aab8667bf653b10cc3799f94352fcdb749db6/exercises/practice/dominoes
|
||||
- id: refactor_dominoes_chain
|
||||
category: coding_refactor
|
||||
kind: code
|
||||
entrypoint: can_chain
|
||||
prompt: |
|
||||
Refactor this can_chain to remove the duplicated chain-building
|
||||
conditions and the flag variable. Behaviour must be preserved EXACTLY:
|
||||
for a set of dominoes that can form a valid chain it returns a valid
|
||||
chain (ANY valid chain — not a fixed one), and None when no chain is
|
||||
possible. Reply with ONLY the rewritten function — no explanation, no
|
||||
fences.
|
||||
|
||||
from itertools import permutations
|
||||
|
||||
|
||||
def can_chain(dominoes):
|
||||
if not any(dominoes):
|
||||
return []
|
||||
for perm in permutations(dominoes):
|
||||
chain = [perm[0]]
|
||||
complete = True
|
||||
for domino in perm[1:]:
|
||||
prev = chain[-1]
|
||||
if len(chain) == 1 and prev[0] == domino[0]:
|
||||
chain = [(prev[1], prev[0]), domino]
|
||||
elif len(chain) == 1 and prev[0] == domino[1]:
|
||||
chain = [(prev[1], prev[0]), (domino[1], domino[0])]
|
||||
elif prev[1] == domino[0]:
|
||||
chain = chain + [domino]
|
||||
elif prev[1] == domino[1]:
|
||||
chain = chain + [(domino[1], domino[0])]
|
||||
else:
|
||||
complete = False
|
||||
break
|
||||
if complete and chain[0][0] == chain[-1][1]:
|
||||
return chain
|
||||
return None
|
||||
checks:
|
||||
- 'can_chain([]) == []'
|
||||
- '((d := [(1, 1)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
|
||||
- '((d := [(1, 2), (3, 1), (2, 3)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
|
||||
- '((d := [(1, 2), (2, 3), (3, 1), (2, 4), (2, 4)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
|
||||
- 'can_chain([(1, 2)]) is None'
|
||||
- 'can_chain([(1, 2), (4, 1), (2, 3)]) is None'
|
||||
- 'can_chain([(1, 1), (2, 2)]) is None'
|
||||
- 'can_chain([(1, 2), (2, 1), (3, 4), (4, 3)]) is None'
|
||||
- 'can_chain([(1, 2), (2, 3), (3, 1), (4, 4)]) is None'
|
||||
- 'can_chain([(1, 2), (2, 3), (3, 1), (4, 5), (5, 6), (6, 4)]) is None'
|
||||
|
||||
- id: debug_dominoes_no_chain
|
||||
category: debugging
|
||||
kind: code
|
||||
entrypoint: can_chain
|
||||
prompt: |
|
||||
This can_chain returns a bogus "chain" for inputs that cannot be
|
||||
chained — it returns a list instead of None when no valid chain exists.
|
||||
Fix ONLY the one subtle bug; do not change anything else, and do not
|
||||
change the can_chain signature. Reply with ONLY the corrected function
|
||||
— no explanation, no fences.
|
||||
|
||||
from itertools import permutations
|
||||
from functools import reduce
|
||||
|
||||
|
||||
def swap(item_1, item_2):
|
||||
return (item_2, item_1)
|
||||
|
||||
|
||||
def build_chain(chain, domino):
|
||||
if chain is not None:
|
||||
last = chain[-1]
|
||||
if len(chain) == 1 and last[0] == domino[0]:
|
||||
return [swap(*last), domino]
|
||||
elif len(chain) == 1 and last[0] == domino[1]:
|
||||
return [swap(*last), swap(*domino)]
|
||||
elif last[1] == domino[0]:
|
||||
return chain + [domino]
|
||||
elif last[1] == domino[1]:
|
||||
return chain + [swap(*domino)]
|
||||
return None
|
||||
|
||||
|
||||
def can_chain(dominoes):
|
||||
if not any(dominoes):
|
||||
return []
|
||||
for perm in permutations(dominoes):
|
||||
chain = reduce(build_chain, perm[1:], [perm[0]])
|
||||
# BUG: the circular-closure check (chain[0][0] == chain[-1][1])
|
||||
# is missing, so a line that merely matches end-to-start is
|
||||
# returned even when it does not close into a loop.
|
||||
if chain is not None:
|
||||
return chain
|
||||
return None
|
||||
checks:
|
||||
- 'can_chain([(1, 2)]) is None'
|
||||
- 'can_chain([(1, 2), (4, 1), (2, 3)]) is None'
|
||||
- 'can_chain([(1, 1), (2, 2)]) is None'
|
||||
- '((d := [(1, 2), (3, 1), (2, 3)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
|
||||
|
||||
# --- affine-cipher (Exercism-derived) -----------------------------------
|
||||
# Exercism python affine-cipher exercise — MIT licensed, canonical source at
|
||||
# https://github.com/exercism/python/tree/1f6aab8667bf653b10cc3799f94352fcdb749db6/exercises/practice/affine-cipher
|
||||
- id: refactor_affine_encode_decode
|
||||
category: coding_refactor
|
||||
kind: code
|
||||
entrypoint: encode
|
||||
prompt: |
|
||||
Refactor this affine cipher to remove the duplicated cipher math. The
|
||||
same letter-to-index transform and the coprime guard appear inline in
|
||||
both encode and decode; behavioural duplicates like these are where bugs
|
||||
hide. Consolidate them. Behaviour must be preserved EXACTLY, including
|
||||
the ValueError raised when `a` is not coprime with the alphabet size and
|
||||
the 5-character block grouping in encode. Keep the module functions
|
||||
`encode(plain, a, b)` and `decode(ciphered, a, b)`. Reply with ONLY the
|
||||
rewritten module — no explanation, no fences.
|
||||
|
||||
BLOCK_SIZE = 5
|
||||
ALPHABET = 26
|
||||
|
||||
|
||||
def mod_inverse(a_key, alphabet):
|
||||
a_key = a_key % alphabet
|
||||
for idx in range(1, alphabet):
|
||||
if (a_key * idx) % alphabet == 1:
|
||||
return idx
|
||||
return 1
|
||||
|
||||
|
||||
def encode(plain, a, b):
|
||||
inverse = mod_inverse(a, ALPHABET)
|
||||
if inverse == 1:
|
||||
raise ValueError('a and m must be coprime.')
|
||||
chars = []
|
||||
for character in plain:
|
||||
if character.isalnum():
|
||||
origin = ord(character.lower()) - 97
|
||||
if origin < 0:
|
||||
chars.append(character)
|
||||
continue
|
||||
new = (a * origin + b) % ALPHABET
|
||||
chars.append(chr(new + 97))
|
||||
cipher = ''.join(chars)
|
||||
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
|
||||
for idx in range(0, len(cipher), BLOCK_SIZE)])
|
||||
|
||||
|
||||
def decode(ciphered, a, b):
|
||||
inverse = mod_inverse(a, ALPHABET)
|
||||
if inverse == 1:
|
||||
raise ValueError('a and m must be coprime.')
|
||||
chars = []
|
||||
for character in ciphered:
|
||||
if character.isalnum():
|
||||
origin = ord(character.lower()) - 97
|
||||
if origin < 0:
|
||||
chars.append(character)
|
||||
continue
|
||||
new = (inverse * (origin - b)) % ALPHABET
|
||||
chars.append(chr(new + 97))
|
||||
return ''.join(chars)
|
||||
checks:
|
||||
- 'encode("yes", 5, 7) == "xbt"'
|
||||
- 'encode("no", 15, 18) == "fu"'
|
||||
- 'encode("OMG", 21, 3) == "lvz"'
|
||||
- 'encode("O M G", 25, 47) == "hjp"'
|
||||
- 'encode("Testing,1 2 3, testing.", 3, 4) == "jqgjc rw123 jqgjc rw"'
|
||||
- 'decode("tytgn fjr", 3, 7) == "exercism"'
|
||||
- 'decode("qdwju nqcro muwhn odqun oppmd aunwd o", 19, 16) == "anobstacleisoftenasteppingstone"'
|
||||
- 'raises(ValueError, encode, "This is a test.", 6, 17)'
|
||||
- 'raises(ValueError, decode, "Test", 13, 5)'
|
||||
|
||||
- id: debug_affine_coprime
|
||||
category: debugging
|
||||
kind: code
|
||||
entrypoint: encode
|
||||
prompt: |
|
||||
This affine cipher fails to reject keys where `a` is not coprime with the
|
||||
alphabet size. It should raise ValueError('a and m must be coprime.')
|
||||
when `a` shares a factor with 26, but it lets those keys through. Fix
|
||||
ONLY the one subtle bug in the coprime guard; do not change anything
|
||||
else, and do not change the signatures of `encode(plain, a, b)` or
|
||||
`decode(ciphered, a, b)`. Reply with ONLY the corrected module — no
|
||||
explanation, no fences.
|
||||
|
||||
BLOCK_SIZE = 5
|
||||
ALPHABET = 26
|
||||
|
||||
|
||||
def mod_inverse(a_key, alphabet):
|
||||
a_key = a_key % alphabet
|
||||
for idx in range(1, alphabet):
|
||||
if (a_key * idx) % alphabet == 1:
|
||||
return idx
|
||||
return 1
|
||||
|
||||
|
||||
def translate(text, a_key, b_key, mode):
|
||||
inverse = mod_inverse(a_key, ALPHABET)
|
||||
if inverse < 1:
|
||||
raise ValueError('a and m must be coprime.')
|
||||
chars = []
|
||||
for character in text:
|
||||
if character.isalnum():
|
||||
origin = ord(character.lower()) - 97
|
||||
if origin < 0:
|
||||
chars.append(character)
|
||||
continue
|
||||
if mode == 0:
|
||||
new = (a_key * origin + b_key) % ALPHABET
|
||||
elif mode == 1:
|
||||
new = (inverse * (origin - b_key)) % ALPHABET
|
||||
chars.append(chr(new + 97))
|
||||
return ''.join(chars)
|
||||
|
||||
|
||||
def encode(plain, a, b):
|
||||
cipher = translate(plain, a, b, 0)
|
||||
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
|
||||
for idx in range(0, len(cipher), BLOCK_SIZE)])
|
||||
|
||||
|
||||
def decode(ciphered, a, b):
|
||||
return translate(ciphered, a, b, 1)
|
||||
checks:
|
||||
- 'encode("yes", 5, 7) == "xbt"'
|
||||
- 'decode("tytgn fjr", 3, 7) == "exercism"'
|
||||
- 'raises(ValueError, encode, "This is a test.", 6, 17)'
|
||||
- 'raises(ValueError, decode, "Test", 13, 5)'
|
||||
|
||||
# --- debugging ----------------------------------------------------------
|
||||
- id: debug_late_binding
|
||||
category: debugging
|
||||
@@ -346,6 +883,221 @@ tasks:
|
||||
location: {type: string}
|
||||
required: [location]
|
||||
|
||||
# BFCL v4 derived (ShishirPatil/gorilla, MIT): 5 abstain from
|
||||
# BFCL_v4_live_irrelevance.json, 3 positive from BFCL_v4_live_simple.json +
|
||||
# possible_answer ground truth. Data fetched once to /tmp, translated
|
||||
# mechanically, never vendored. Irrelevant prompts are deliberately
|
||||
# unanswerable by the offered weather tool (real abstain cases).
|
||||
- id: tool_bfcl_live_irrelevance_12-2-0
|
||||
category: tool_use_agentic
|
||||
kind: tool
|
||||
prompt: I'd appreciate if you could fetch the DNS resolution info for the domain mapped to IP 255.255.255.0 from VirusTotal. My key for this operation is 'sample_key4'.
|
||||
expect_tool: null
|
||||
tools:
|
||||
- type: function
|
||||
function:
|
||||
name: get_current_weather
|
||||
description: Retrieves the current weather conditions for a specified city and state.
|
||||
parameters:
|
||||
type: object
|
||||
required:
|
||||
- location
|
||||
properties:
|
||||
location:
|
||||
type: string
|
||||
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
|
||||
exist.
|
||||
unit:
|
||||
type: string
|
||||
description: The unit of temperature for the weather report.
|
||||
enum:
|
||||
- celsius
|
||||
- fahrenheit
|
||||
default: fahrenheit
|
||||
- id: tool_bfcl_live_irrelevance_13-2-1
|
||||
category: tool_use_agentic
|
||||
kind: tool
|
||||
prompt: What is diffrence between cpu and gpu?
|
||||
expect_tool: null
|
||||
tools:
|
||||
- type: function
|
||||
function:
|
||||
name: get_current_weather
|
||||
description: Retrieves the current weather conditions for a specified city and state.
|
||||
parameters:
|
||||
type: object
|
||||
required:
|
||||
- location
|
||||
properties:
|
||||
location:
|
||||
type: string
|
||||
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
|
||||
exist.
|
||||
unit:
|
||||
type: string
|
||||
description: The unit of temperature for the weather report.
|
||||
enum:
|
||||
- celsius
|
||||
- fahrenheit
|
||||
default: fahrenheit
|
||||
- id: tool_bfcl_live_irrelevance_14-2-2
|
||||
category: tool_use_agentic
|
||||
kind: tool
|
||||
prompt: Please help me get the votes associated with the IP of http://digdeep.io.
|
||||
expect_tool: null
|
||||
tools:
|
||||
- type: function
|
||||
function:
|
||||
name: get_current_weather
|
||||
description: Retrieves the current weather conditions for a specified city and state.
|
||||
parameters:
|
||||
type: object
|
||||
required:
|
||||
- location
|
||||
properties:
|
||||
location:
|
||||
type: string
|
||||
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
|
||||
exist.
|
||||
unit:
|
||||
type: string
|
||||
description: The unit of temperature for the weather report.
|
||||
enum:
|
||||
- celsius
|
||||
- fahrenheit
|
||||
default: fahrenheit
|
||||
- id: tool_bfcl_live_irrelevance_15-2-3
|
||||
category: tool_use_agentic
|
||||
kind: tool
|
||||
prompt: Using 'api_key_2', retrieve the IDs of graphs containing IP 145.34.45.56 on VirusTotal. Don't forget to set the cursor as 'cursor_b' and limit the results to 8.
|
||||
expect_tool: null
|
||||
tools:
|
||||
- type: function
|
||||
function:
|
||||
name: get_current_weather
|
||||
description: Retrieves the current weather conditions for a specified city and state.
|
||||
parameters:
|
||||
type: object
|
||||
required:
|
||||
- location
|
||||
properties:
|
||||
location:
|
||||
type: string
|
||||
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
|
||||
exist.
|
||||
unit:
|
||||
type: string
|
||||
description: The unit of temperature for the weather report.
|
||||
enum:
|
||||
- celsius
|
||||
- fahrenheit
|
||||
default: fahrenheit
|
||||
- id: tool_bfcl_live_irrelevance_16-2-4
|
||||
category: tool_use_agentic
|
||||
kind: tool
|
||||
prompt: 'How do I pull the domain info of twitter.com from VirusTotal? Using this API key: twt_key_abc.'
|
||||
expect_tool: null
|
||||
tools:
|
||||
- type: function
|
||||
function:
|
||||
name: get_current_weather
|
||||
description: Retrieves the current weather conditions for a specified city and state.
|
||||
parameters:
|
||||
type: object
|
||||
required:
|
||||
- location
|
||||
properties:
|
||||
location:
|
||||
type: string
|
||||
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
|
||||
exist.
|
||||
unit:
|
||||
type: string
|
||||
description: The unit of temperature for the weather report.
|
||||
enum:
|
||||
- celsius
|
||||
- fahrenheit
|
||||
default: fahrenheit
|
||||
- id: tool_bfcl_live_simple_0-0-0
|
||||
category: tool_use_agentic
|
||||
kind: tool
|
||||
prompt: Can you retrieve the details for the user with the ID 7890, who has black as their special request?
|
||||
expect_tool: get_user_info
|
||||
tools:
|
||||
- type: function
|
||||
function:
|
||||
name: get_user_info
|
||||
description: Retrieve details for a specific user by their unique identifier.
|
||||
parameters:
|
||||
type: object
|
||||
required:
|
||||
- user_id
|
||||
properties:
|
||||
user_id:
|
||||
type: integer
|
||||
description: The unique identifier of the user. It is used to fetch the specific user details from the database.
|
||||
special:
|
||||
type: string
|
||||
description: Any special information or parameters that need to be considered while fetching user details.
|
||||
default: none
|
||||
expect_args:
|
||||
user_id: 7890
|
||||
special: black
|
||||
- id: tool_bfcl_live_simple_1-1-0
|
||||
category: tool_use_agentic
|
||||
kind: tool
|
||||
prompt: I want to see the star history of ShishirPatil/gorilla and gorilla-llm/gorilla-cli, with the timelines aligned, so that I can more clearly observe the rate of change from their initial releases.
|
||||
expect_tool: github_star
|
||||
tools:
|
||||
- type: function
|
||||
function:
|
||||
name: github_star
|
||||
description: Generates a URL for tracking the star history of specified GitHub repositories, with the option to align them on the same timeline.
|
||||
parameters:
|
||||
type: object
|
||||
required:
|
||||
- repos
|
||||
properties:
|
||||
repos:
|
||||
type: string
|
||||
description: A comma-separated list of GitHub repositories to track, each in the 'owner/repo' format, such as 'octocat/Hello-World,octo-org/octo-repo'.
|
||||
aligned:
|
||||
type: boolean
|
||||
description: Whether to align the repositories on the same timeline for comparison. If true, the star history of all repositories will start from the same point.
|
||||
default: false
|
||||
expect_args:
|
||||
repos: ShishirPatil/gorilla,gorilla-llm/gorilla-cli
|
||||
aligned: true
|
||||
- id: tool_bfcl_live_simple_4-3-0
|
||||
category: tool_use_agentic
|
||||
kind: tool
|
||||
prompt: What are the current weather conditions in Tel Aviv, and could you provide that in Fahrenheit, please?
|
||||
expect_tool: get_current_weather
|
||||
tools:
|
||||
- type: function
|
||||
function:
|
||||
name: get_current_weather
|
||||
description: Retrieves the current weather conditions for a specified city and state. If using state, then use short form like CA.
|
||||
parameters:
|
||||
type: object
|
||||
required:
|
||||
- location
|
||||
properties:
|
||||
location:
|
||||
type: string
|
||||
description: The location for which to get the weather, in the format of 'City, State (abbr)', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city
|
||||
doesn't exist.
|
||||
unit:
|
||||
type: string
|
||||
description: The unit of temperature for the weather report.
|
||||
enum:
|
||||
- celsius
|
||||
- fahrenheit
|
||||
default: fahrenheit
|
||||
expect_args:
|
||||
location: Tel Aviv, Israel
|
||||
unit: fahrenheit
|
||||
|
||||
# --- docs_writing -------------------------------------------------------
|
||||
- id: docs_function
|
||||
category: docs_writing
|
||||
|
||||
280
plans/benchmark-sourced-eval-tasks.md
Normal file
280
plans/benchmark-sourced-eval-tasks.md
Normal file
@@ -0,0 +1,280 @@
|
||||
# Spec: sourcing harder self-eval tasks from external benchmarks
|
||||
|
||||
**Origin.** `PYTHONPATH=src python -m eval_proficiency` was re-run in full this
|
||||
session (11 models x 23 tasks, 216 billed calls, $0.79 / 0.098 kWh, 0 call
|
||||
failures) specifically to see whether more samples of the *existing* task set
|
||||
would break the ties CLAUDE.md already flagged as suspicious. Result, doubling
|
||||
n from 2-3 to 4-6 per model per category:
|
||||
|
||||
- `tool_use_agentic`: `deepseek-v4-flash` held at exactly 2/6, the same ratio
|
||||
as the earlier 1/3 — real signal, not n=3 noise, but still short of
|
||||
`self_eval_min_samples: 10`.
|
||||
- `coding_general`, `coding_refactor`, `debugging`: **still perfectly flat at
|
||||
1.00** for every model. This is no longer "not enough samples" — it's
|
||||
"these 3 tasks per category don't discriminate anything, at any sample
|
||||
count." More passes over the same 23 tasks won't fix that; the task set
|
||||
itself needs to get harder.
|
||||
|
||||
This spec is the result of "where do we get harder tasks without inventing
|
||||
plausible-looking difficulty from nothing" — the same objection this project
|
||||
already raised against `leaderboards.yaml`'s empty priors. No code changes
|
||||
accompany this document — this is the spec opencode builds from.
|
||||
|
||||
---
|
||||
|
||||
## Non-goals, stated up front
|
||||
|
||||
- **No new categories.** `proficiency.categories` in `config.py` is read
|
||||
throughout scoring/tiering/routing; adding one is a real migration, not a
|
||||
task-authoring change, and nothing here needs it. Every source below gets
|
||||
mapped onto one of the 9 existing categories.
|
||||
- **No new scoring kinds.** Everything maps onto `code`, `exact`, or `tool`.
|
||||
(`judge` categories — `docs_writing`, `general_chat`, `summarization`,
|
||||
`translation`, `general_chat` — aren't the problem this session found; skip
|
||||
them for this pass.)
|
||||
- **No repo-context benchmarks.** This is the reason SWE-bench (and anything
|
||||
else shaped like "here's a git checkout, go fix an issue") is not on the
|
||||
list below despite being the best-known real-bug-fixing benchmark. Every
|
||||
task in `evals/tasks.yaml` is one self-contained prompt scored by one
|
||||
subprocess call or one structural check — the harness has no notion of a
|
||||
repo checkout, and building one is a much bigger lift than the actual
|
||||
problem (three flat categories) justifies.
|
||||
- **No verbatim copying.** Every source below gets *adapted* — either
|
||||
translated into this project's YAML+checks shape, or (for
|
||||
`coding_refactor`/`debugging`) used only to identify which underlying
|
||||
problems are hard enough to be worth hand-authoring a variant of. Nothing
|
||||
here is "download file, paste into tasks.yaml."
|
||||
|
||||
## A code change this needs, before any task uses it
|
||||
|
||||
`score_exact` (`eval_proficiency.py:191`) normalizes by regexing out the
|
||||
**last number** in the reply (`normalize_answer`). That's correct for this
|
||||
project's current `exact` tasks — they're all arithmetic word problems with a
|
||||
scalar numeric answer (`math_percent_trap`, `math_rate_trap`,
|
||||
`math_counting`). It is **not** correct for CRUXEval-style tasks (below),
|
||||
where the expected value is a Python literal that can be a list, dict, tuple,
|
||||
string, or `None` — e.g. an answer of `[1, 2, 3]` would get reduced to `3`,
|
||||
silently discarding the structure and scoring against the wrong thing.
|
||||
|
||||
Fix, scoped and small: in `normalize_answer`, try `ast.literal_eval` on the
|
||||
stripped text first (covers numbers, strings, lists, dicts, tuples, bools,
|
||||
`None` uniformly); fall back to the existing regex-last-number heuristic only
|
||||
on a `ValueError`/`SyntaxError`, which is exactly what a numeric word-problem
|
||||
answer with prose around it produces. This is additive — every existing
|
||||
`exact` task still gets scored the same way, since `literal_eval` will fail on
|
||||
"3 minutes" and fall through to the current path unchanged. Needs its own
|
||||
test alongside the existing `test_counting_answer_matches_brute_force`-style
|
||||
ones (`tests/test_task_set.py`) before any CRUXEval-derived task is added, per
|
||||
the project's own rule that a check the reference solution can't pass "scores
|
||||
the task set rather than the model."
|
||||
|
||||
## `tool_use_agentic` ← BFCL live_irrelevance
|
||||
|
||||
**Source:** `BFCL_v4_live_irrelevance.json` in
|
||||
[`ShishirPatil/gorilla`](https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard/bfcl_eval/data)
|
||||
(Apache 2.0). 884 examples, sourced from real user queries rather than
|
||||
synthesized — pull this file over the smaller 239-example synthetic
|
||||
`BFCL_v4_irrelevance.json` for the same reason this project trusts
|
||||
`POST /outcome` traffic over benchmark scores: real queries beat invented
|
||||
ones. This is the single best-matched source of the four — it's a large pool
|
||||
of *exactly* the pattern this project already hand-writes 2 of 3 tasks in:
|
||||
offer a tool, correct behavior is not calling it.
|
||||
|
||||
**Schema gap to close** (confirmed by fetching a live sample): BFCL's own
|
||||
function schema is not the OpenAI tool-call schema this project uses —
|
||||
`"type": "dict"` where OpenAI/this project's YAML says `"type": "object"`, no
|
||||
`{"type": "function", "function": {...}}` wrapper, and `question` is a list of
|
||||
turns (list of message dicts) rather than this project's flat `prompt`
|
||||
string. Translation per selected example:
|
||||
|
||||
```python
|
||||
tools = [
|
||||
{"type": "function", "function": {
|
||||
"name": fn["name"],
|
||||
"description": fn["description"],
|
||||
"parameters": {**fn["parameters"], "type": "object"}, # dict -> object
|
||||
}}
|
||||
for fn in example["function"]
|
||||
]
|
||||
prompt = example["question"][0][0]["content"] # single-turn BFCL cases only
|
||||
```
|
||||
|
||||
**The balance trap.** Every BFCL *irrelevance* example is, by definition, a
|
||||
"correctly abstain" case. If the new batch is 100% abstain-cases, a model
|
||||
that never calls a tool scores 1.0 on the enlarged category despite being
|
||||
just as broken as one that always calls a tool — the category would stop
|
||||
measuring over-triggering and start measuring only under-triggering.
|
||||
`score_tool` already has both directions (`expect_tool: null` vs. a named
|
||||
tool), and the current 3-task set is already 1:2 (call : abstain). Keep new
|
||||
additions roughly in that band or better: pull irrelevance examples for the
|
||||
new abstain cases, and pull a matching number of **positive** call cases from
|
||||
BFCL's `live_simple`/`live_multiple` files (same repo, same schema, translate
|
||||
the same way) so the enlarged category still rewards correct discrimination,
|
||||
not just caution.
|
||||
|
||||
**Worked example**, adapted from a real fetched entry:
|
||||
|
||||
```yaml
|
||||
- id: tool_geocode_irrelevant
|
||||
category: tool_use_agentic
|
||||
kind: tool
|
||||
prompt: |
|
||||
Can you provide the address for latitude 37.4224764 and longitude
|
||||
-122.0842499 using the Geocoding API?
|
||||
expect_tool: null # no offered function does geocoding
|
||||
tools:
|
||||
- type: function
|
||||
function:
|
||||
name: requests.get
|
||||
description: Sends a GET request to the specified URL to retrieve data.
|
||||
parameters:
|
||||
type: object
|
||||
properties:
|
||||
url: {type: string, description: The URL to send the GET request to.}
|
||||
required: [url]
|
||||
```
|
||||
|
||||
## `coding_general` ← CRUXEval-O
|
||||
|
||||
**Source:** [`facebookresearch/cruxeval`](https://github.com/facebookresearch/cruxeval)
|
||||
(MIT). 800 short Python functions with input/output pairs, split into
|
||||
CRUXEval-I (predict an input that produces a given output) and CRUXEval-O
|
||||
(predict the output of running the function on a given input). **Use O only.**
|
||||
I is a genuine structural mismatch: an input that produces a target output is
|
||||
often not unique, so scoring it needs to actually execute the candidate input
|
||||
through the real function and compare — that's neither `exact` (no single
|
||||
golden string to normalize against) nor `code` (the model isn't writing a
|
||||
function) as this harness defines them. Not worth a third scoring kind for
|
||||
half of one source; skip I, use O, which is a clean fit for `exact` once the
|
||||
`ast.literal_eval` fix above lands.
|
||||
|
||||
At release, best-of-field models scored 63-67% on CRUXEval-O — meaning it
|
||||
was built to sit in a discriminating band, not to saturate the way this
|
||||
project's own three `coding_general` tasks did.
|
||||
|
||||
**Worked example** (illustrative shape, not a literal fetched row):
|
||||
|
||||
```yaml
|
||||
- id: cruxeval_trace_dedup
|
||||
category: coding_general
|
||||
kind: exact
|
||||
prompt: |
|
||||
What does this function return when called as shown? Reply with ONLY
|
||||
the literal Python value — no explanation, no fences.
|
||||
|
||||
def f(lst):
|
||||
seen = {}
|
||||
out = []
|
||||
for x in lst:
|
||||
if x not in seen:
|
||||
seen[x] = True
|
||||
out.append(x * 2)
|
||||
return out
|
||||
|
||||
f([3, 1, 3, 2, 1])
|
||||
answer: "[6, 2, 4]"
|
||||
```
|
||||
|
||||
## `coding_refactor` / `debugging` ← Exercism, via Aider's selection method
|
||||
|
||||
**What Aider's polyglot-benchmark actually gives you.** I checked: the
|
||||
[`Aider-AI/polyglot-benchmark`](https://github.com/Aider-AI/polyglot-benchmark)
|
||||
repo's own GitHub license field comes back blank via the API — don't pull
|
||||
files from it directly. What's genuinely reusable is the *method* their
|
||||
[Dec 2024 writeup](https://aider.chat/2024/12/21/polyglot.html) describes:
|
||||
225 of Exercism's 697 exercises, filtered to ones **solved by 3 or fewer of 7
|
||||
frontier models**, deliberately excluding the 258 "solved by all 7" as too
|
||||
easy. They didn't publish the 225 names. The underlying exercises are
|
||||
Exercism's own tracks, confirmed MIT-licensed
|
||||
(`exercism/python`, etc.) — go there directly, using Aider's selection
|
||||
criterion (not their file) as the bar: prefer exercises in Exercism's own
|
||||
`practice/` directories with a track difficulty/reputation for edge-case
|
||||
traps. `bowling`, `dominoes`, and `affine-cipher` (all present in the
|
||||
`exercism/python` `practice/` tree, confirmed via the GitHub API) are
|
||||
reasonable starting picks — each is well known for one nasty, non-obvious
|
||||
edge case rather than raw difficulty, which is exactly this project's own
|
||||
stated design principle ("at least one edge case a plausible-looking solution
|
||||
gets wrong").
|
||||
|
||||
**Why this can't be mechanical, unlike the other two sources.** Exercism ships
|
||||
a spec + a test suite, not a "here's buggy/ugly code, fix/refactor it"
|
||||
prompt — that shape doesn't exist upstream. `coding_refactor` and `debugging`
|
||||
tasks in this project embed the *broken or ugly-but-working* code directly in
|
||||
the prompt (confirmed from `apply_settings`, `first_match`, `describe` in the
|
||||
current set) and the checks pin exact edge-case behavior. So the real value
|
||||
of going to Exercism isn't a copyable prompt, it's **borrowing the hard edge
|
||||
case**: read the exercise's canonical solution and its test suite (that's the
|
||||
whole reason the exercise is hard), then hand-author a variant:
|
||||
|
||||
- **`coding_refactor` variant:** take the canonical solution, deliberately
|
||||
restructure it into duplicated/nested-conditional/flag-variable style (same
|
||||
pattern as the existing three) while preserving every edge case exactly —
|
||||
the model refactors it back, checks re-run the exercise's own tricky test
|
||||
cases.
|
||||
- **`debugging` variant:** take the canonical solution, introduce one subtle
|
||||
off-by-one/boundary bug that breaks exactly the input the exercise is
|
||||
famous for tripping people up on (bowling's 10th-frame strike/spare bonus
|
||||
scoring is the textbook example), leave everything else correct — checks
|
||||
must include that exact edge case, per `tests/test_task_set.py`'s existing
|
||||
requirement that every debugging target actually fails pre-fix.
|
||||
|
||||
This is real authoring work, not translation — budget it accordingly (see
|
||||
below). What Exercism/Aider buys is **which** edge cases are worth encoding,
|
||||
pre-validated by the fact that most frontier models already fail them, rather
|
||||
than inventing new traps from scratch the way the original 23-task set's
|
||||
"touching intervals / full semver / late-binding closures" hardening pass
|
||||
did.
|
||||
|
||||
## License / attribution handling
|
||||
|
||||
Nothing here is substantial verbatim reproduction, so full license-file
|
||||
vendoring isn't needed — but this project already has a citation precedent
|
||||
(`session-classification-cache-ttl.md`'s LLMRouter attribution, and
|
||||
`PinchConfig`'s docstring crediting llmrouter directly in code). Match it: a
|
||||
one-line comment above each newly-added task block in `evals/tasks.yaml`
|
||||
naming the source and license, e.g.:
|
||||
|
||||
```yaml
|
||||
# --- tool_use_agentic (added) — adapted from BFCL live_irrelevance /
|
||||
# live_simple, Apache 2.0, github.com/ShishirPatil/gorilla ---
|
||||
```
|
||||
|
||||
## Validation obligation (not optional)
|
||||
|
||||
`tests/test_task_set.py` requires, per kind:
|
||||
- `code`: a reference solution that passes every check.
|
||||
- `exact`: an independent (brute-force or hand-derived) recomputation of the
|
||||
answer — see `test_counting_answer_matches_brute_force` for the pattern.
|
||||
- `debugging`: the buggy target must **fail** its own checks pre-fix.
|
||||
- `coding_refactor`: the ugly target must **pass** its own checks pre-refactor.
|
||||
|
||||
This is what already caught one wrong expected-value check in the existing
|
||||
set (CLAUDE.md: "would have docked every model on a task and been
|
||||
indistinguishable from genuine difficulty"). Every task added from this spec
|
||||
needs its matching test — this roughly doubles the authoring cost per task
|
||||
versus just writing the YAML, and should be budgeted as such rather than
|
||||
treated as a follow-up.
|
||||
|
||||
## Suggested first-pass batch size
|
||||
|
||||
This session's full run (23 tasks x 11 models, some judge calls) was 216
|
||||
billed calls / $0.79. A reasonable first batch — **8 new `tool_use_agentic`
|
||||
tasks (5 abstain + 3 positive-call), 6 new `coding_general` (CRUXEval-O), 3
|
||||
new `coding_refactor`, 3 new `debugging`** — is +20 tasks, roughly matching
|
||||
the size of the existing set. Expect a proportional cost per pass (~$1.50-2),
|
||||
and **two passes** (not one) to clear `self_eval_min_samples: 10` for the
|
||||
newly-added tasks, since one pass only gives n=1 per new task per model.
|
||||
|
||||
## Recommendation
|
||||
|
||||
Build in this order: (1) the `score_exact`/`ast.literal_eval` fix plus its
|
||||
test, since CRUXEval-O tasks are silently wrong without it; (2) BFCL-derived
|
||||
`tool_use_agentic` tasks, since that category already has the strongest
|
||||
reproduced signal from this session's run and BFCL requires no hand-authoring
|
||||
beyond schema translation; (3) CRUXEval-O `coding_general` tasks; (4) the
|
||||
Exercism-sourced `coding_refactor`/`debugging` variants last, since they're
|
||||
the only genuinely bespoke authoring work and benefit from the other three
|
||||
being done first as a format reference. Run `eval_proficiency` twice after
|
||||
each category lands rather than once at the end, so a broken check (per the
|
||||
validation section above) is caught against one category's blast radius, not
|
||||
all four at once.
|
||||
@@ -28,6 +28,7 @@ functions, so nothing invites filesystem or network use, but treat this as
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import ast
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
@@ -78,6 +79,11 @@ def raises(exc, fn, *a, **kw):
|
||||
|
||||
FENCE_RE = re.compile(r"^\s*```[a-zA-Z]*\n(.*?)```", re.DOTALL | re.MULTILINE)
|
||||
|
||||
# A standalone number with comma thousands separators (e.g. "3,000"). Matched
|
||||
# on the WHOLE stripped string so "3,000" is treated as the number 3000 rather
|
||||
# than literal_eval'ing to the tuple (3, 0).
|
||||
BARE_COMMA_NUMBER_RE = re.compile(r"-?\d[\d,]*(?:\.\d+)?")
|
||||
|
||||
|
||||
# --- extraction and scoring ----------------------------------------------
|
||||
|
||||
@@ -174,26 +180,85 @@ def score_code(model_output: str, checks: list[str]) -> tuple[float, str]:
|
||||
return passed / len(checks), f"{passed}/{len(checks)} checks"
|
||||
|
||||
|
||||
_LITERAL_WORD = {"true": "True", "false": "False", "none": "None"}
|
||||
|
||||
|
||||
def normalize_answer(text: str) -> str:
|
||||
"""Reduce a free-text reply to a comparable token.
|
||||
|
||||
Models wrap a number in prose, punctuation, or thousands separators even
|
||||
when told not to; none of that is the thing being measured.
|
||||
Structured answers (lists, tuples, dicts, bools, None, quoted strings)
|
||||
parse as Python literals first, so a model told to reply with a literal
|
||||
value is scored against the value, not its string form. Prose-wrapped
|
||||
numbers fall back to the last number in the reply, with thousands
|
||||
separators handled, so "the combinations is 4,536." still matches "4536".
|
||||
|
||||
Markdown fences are honored only when they wrap the ENTIRE reply; a fence
|
||||
followed by prose is treated as prose, because the concluding prose value
|
||||
is what is being measured.
|
||||
"""
|
||||
cleaned = (text or "").strip().replace(",", "")
|
||||
stripped = (text or "").strip()
|
||||
|
||||
# Honor a whole-reply single fenced block; never discard trailing prose.
|
||||
if FENCE_RE.fullmatch(stripped):
|
||||
match = FENCE_RE.search(stripped)
|
||||
if match:
|
||||
stripped = match.group(1).strip()
|
||||
|
||||
# Collapse thousands-separator commas BEFORE literal_eval. Without this,
|
||||
# ast.literal_eval("[3,000, 4,000]") silently misparses to [3, 0, 4, 0]
|
||||
# because Python treats "3,000" inside a list as "3" followed by leading-
|
||||
# zero "000". The look-behind / look-ahead ensure only groups of exactly
|
||||
# three digits after a comma are collapsed, so "[12,34]" and tuples are
|
||||
# left untouched.
|
||||
# KNOWN RESIDUAL AMBIGUITY (accepted, not chased): "[100,200,300]" (three
|
||||
# adjacent 3-digit list elements, no spaces) is character-identical to a
|
||||
# thousands-grouped number, so this collapse merges it to "[100200300]".
|
||||
# It cannot be disambiguated from the raw text alone. No current task's
|
||||
# answer contains such text, so this is latent; do not add a fourth comma
|
||||
# heuristic to guess structural intent — regex composition over raw text
|
||||
# just relocates this ambiguity instead of resolving it.
|
||||
collapsed = re.sub(r"(?<=\d),(\d{3})(?!\d)", r"\1", stripped)
|
||||
|
||||
# A bare comma-thousands number like "3,000" would otherwise literal_eval
|
||||
# to the tuple (3, 0). (With the collapse above, this branch is now
|
||||
# mostly unreachable for comma-numbers, but the regex stays for safety.)
|
||||
|
||||
value: object
|
||||
try:
|
||||
value = ast.literal_eval(collapsed)
|
||||
except (ValueError, SyntaxError):
|
||||
pass
|
||||
else:
|
||||
if isinstance(value, float):
|
||||
if value.is_integer():
|
||||
return str(int(value))
|
||||
return str(value)
|
||||
return repr(value)
|
||||
|
||||
# Prose fallback: the concluding value is the last number, with thousands
|
||||
# separators stripped - but never inside bracket-wrapped structured text,
|
||||
# where commas are element separators, not separators to remove.
|
||||
if collapsed and not any(c in collapsed for c in "[](){}"):
|
||||
cleaned = collapsed.replace(",", "")
|
||||
numbers = re.findall(r"-?\d+(?:\.\d+)?", cleaned)
|
||||
if numbers:
|
||||
value = numbers[-1] # the conclusion, if it reasoned out loud
|
||||
return value.rstrip("0").rstrip(".") if "." in value else value
|
||||
return cleaned.lower().strip(" .!\"'")
|
||||
candidate = numbers[-1]
|
||||
return (
|
||||
candidate.rstrip("0").rstrip(".")
|
||||
if "." in candidate
|
||||
else candidate
|
||||
)
|
||||
|
||||
# Value keywords ("False.", "true.") match their literal repr
|
||||
# case-insensitively, so "False." still equals the answer "False".
|
||||
cleaned = collapsed.replace(",", "").lower().strip(" .!\"'")
|
||||
return _LITERAL_WORD.get(cleaned, cleaned)
|
||||
|
||||
def score_exact(model_output: str, expected: str) -> tuple[float, str]:
|
||||
got = normalize_answer(model_output)
|
||||
want = normalize_answer(expected)
|
||||
return (1.0, f"{got!r}") if got == want else (0.0, f"got {got!r} want {want!r}")
|
||||
|
||||
|
||||
def score_tool(tool_calls: list, task: dict) -> tuple[float, str]:
|
||||
"""Score a tool-use task structurally — no judge required.
|
||||
|
||||
|
||||
@@ -11,12 +11,14 @@ These tests run offline and take milliseconds — they are the cheap guard
|
||||
against spending an hour of API calls measuring a typo.
|
||||
"""
|
||||
|
||||
import re
|
||||
import textwrap
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
import yaml
|
||||
|
||||
from eval_proficiency import score_code, score_exact
|
||||
from eval_proficiency import score_code, score_exact, score_tool
|
||||
|
||||
TASKS = yaml.safe_load((Path(__file__).resolve().parent.parent / "evals" / "tasks.yaml").read_text())["tasks"]
|
||||
BY_ID = {t["id"]: t for t in TASKS}
|
||||
@@ -92,6 +94,327 @@ _TABLE = {200: "ok", 201: "created", 404: "not found", 500: "server error"}
|
||||
|
||||
def describe(code):
|
||||
return _TABLE.get(code, "unknown")
|
||||
''',
|
||||
"refactor_bowling_frames": '''
|
||||
class Frame:
|
||||
def __init__(self, idx):
|
||||
self.idx = idx
|
||||
self.throws = []
|
||||
|
||||
@property
|
||||
def total_pins(self):
|
||||
return sum(self.throws)
|
||||
|
||||
def is_strike(self):
|
||||
return self.total_pins == 10 and len(self.throws) == 1
|
||||
|
||||
def is_spare(self):
|
||||
return self.total_pins == 10 and len(self.throws) == 2
|
||||
|
||||
def is_open(self):
|
||||
return self.total_pins < 10 and len(self.throws) == 2
|
||||
|
||||
def is_closed(self):
|
||||
return self.total_pins == 10 or len(self.throws) == 2
|
||||
|
||||
def throw(self, pins):
|
||||
if self.total_pins + pins > 10:
|
||||
raise ValueError("a frame's rolls cannot exceed 10")
|
||||
self.throws.append(pins)
|
||||
|
||||
def score(self, next_throws):
|
||||
result = self.total_pins
|
||||
if self.is_strike():
|
||||
result += sum(next_throws[:2])
|
||||
elif self.is_spare():
|
||||
result += sum(next_throws[:1])
|
||||
return result
|
||||
|
||||
|
||||
class BowlingGame:
|
||||
def __init__(self):
|
||||
self.current_frame_idx = 0
|
||||
self.bonus_throws = []
|
||||
self.frames = [Frame(idx) for idx in range(10)]
|
||||
|
||||
@property
|
||||
def current_frame(self):
|
||||
return self.frames[self.current_frame_idx]
|
||||
|
||||
def next_throws(self, frame_idx):
|
||||
throws = []
|
||||
for idx in range(frame_idx + 1, 10):
|
||||
throws.extend(self.frames[idx].throws)
|
||||
throws.extend(self.bonus_throws)
|
||||
return throws
|
||||
|
||||
def roll_bonus(self, pins):
|
||||
tenth_frame = self.frames[-1]
|
||||
if tenth_frame.is_open():
|
||||
raise IndexError('cannot throw bonus with an open tenth frame')
|
||||
self.bonus_throws.append(pins)
|
||||
if (len(self.bonus_throws) == 2 and self.bonus_throws[0] != 10 and
|
||||
sum(self.bonus_throws) > 10):
|
||||
raise ValueError('invalid fill balls')
|
||||
if tenth_frame.is_strike() and len(self.bonus_throws) > 2:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a strike')
|
||||
elif tenth_frame.is_spare() and len(self.bonus_throws) > 1:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a spare')
|
||||
|
||||
def roll(self, pins):
|
||||
if not 0 <= pins <= 10:
|
||||
raise ValueError('invalid pins')
|
||||
elif self.current_frame_idx == 10:
|
||||
self.roll_bonus(pins)
|
||||
else:
|
||||
self.current_frame.throw(pins)
|
||||
if self.current_frame.is_closed():
|
||||
self.current_frame_idx += 1
|
||||
|
||||
def score(self):
|
||||
if self.current_frame_idx < 10:
|
||||
raise IndexError('frame less than 10')
|
||||
if self.frames[-1].is_spare() and len(self.bonus_throws) != 1:
|
||||
raise IndexError(
|
||||
'one bonus must be rolled when the tenth frame is spare')
|
||||
if self.frames[-1].is_strike() and len(self.bonus_throws) != 2:
|
||||
raise IndexError(
|
||||
'two bonuses must be rolled when the tenth frame is strike')
|
||||
return sum(frame.score(self.next_throws(frame.idx))
|
||||
for frame in self.frames)
|
||||
''',
|
||||
"debug_bowling_tenth_frame": '''
|
||||
class Frame:
|
||||
def __init__(self, idx):
|
||||
self.idx = idx
|
||||
self.throws = []
|
||||
|
||||
@property
|
||||
def total_pins(self):
|
||||
return sum(self.throws)
|
||||
|
||||
def is_strike(self):
|
||||
return self.total_pins == 10 and len(self.throws) == 1
|
||||
|
||||
def is_spare(self):
|
||||
return self.total_pins == 10 and len(self.throws) == 2
|
||||
|
||||
def is_open(self):
|
||||
return self.total_pins < 10 and len(self.throws) == 2
|
||||
|
||||
def is_closed(self):
|
||||
return self.total_pins == 10 or len(self.throws) == 2
|
||||
|
||||
def throw(self, pins):
|
||||
if self.total_pins + pins > 10:
|
||||
raise ValueError("a frame's rolls cannot exceed 10")
|
||||
self.throws.append(pins)
|
||||
|
||||
def score(self, next_throws):
|
||||
result = self.total_pins
|
||||
if self.is_strike():
|
||||
result += sum(next_throws[:2])
|
||||
elif self.is_spare():
|
||||
result += sum(next_throws[:1])
|
||||
return result
|
||||
|
||||
|
||||
class BowlingGame:
|
||||
def __init__(self):
|
||||
self.current_frame_idx = 0
|
||||
self.bonus_throws = []
|
||||
self.frames = [Frame(idx) for idx in range(10)]
|
||||
|
||||
@property
|
||||
def current_frame(self):
|
||||
return self.frames[self.current_frame_idx]
|
||||
|
||||
def next_throws(self, frame_idx):
|
||||
throws = []
|
||||
for idx in range(frame_idx + 1, 10):
|
||||
throws.extend(self.frames[idx].throws)
|
||||
throws.extend(self.bonus_throws)
|
||||
return throws
|
||||
|
||||
def roll_bonus(self, pins):
|
||||
tenth_frame = self.frames[-1]
|
||||
if tenth_frame.is_open():
|
||||
raise IndexError('cannot throw bonus with an open tenth frame')
|
||||
self.bonus_throws.append(pins)
|
||||
if (len(self.bonus_throws) == 2 and self.bonus_throws[0] != 10 and
|
||||
sum(self.bonus_throws) > 10):
|
||||
raise ValueError('invalid fill balls')
|
||||
if tenth_frame.is_strike() and len(self.bonus_throws) > 2:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a strike')
|
||||
elif tenth_frame.is_spare() and len(self.bonus_throws) > 1:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a spare')
|
||||
|
||||
def roll(self, pins):
|
||||
if not 0 <= pins <= 10:
|
||||
raise ValueError('invalid pins')
|
||||
elif self.current_frame_idx == 10:
|
||||
self.roll_bonus(pins)
|
||||
else:
|
||||
self.current_frame.throw(pins)
|
||||
if self.current_frame.is_closed():
|
||||
self.current_frame_idx += 1
|
||||
|
||||
def score(self):
|
||||
if self.current_frame_idx < 10:
|
||||
raise IndexError('frame less than 10')
|
||||
if self.frames[-1].is_spare() and len(self.bonus_throws) != 1:
|
||||
raise IndexError(
|
||||
'one bonus must be rolled when the tenth frame is spare')
|
||||
if self.frames[-1].is_strike() and len(self.bonus_throws) != 2:
|
||||
raise IndexError(
|
||||
'two bonuses must be rolled when the tenth frame is strike')
|
||||
return sum(frame.score(self.next_throws(frame.idx))
|
||||
for frame in self.frames)
|
||||
''',
|
||||
"debug_dominoes_no_chain": '''
|
||||
from itertools import permutations
|
||||
from functools import reduce
|
||||
|
||||
|
||||
def swap(item_1, item_2):
|
||||
return (item_2, item_1)
|
||||
|
||||
|
||||
def build_chain(chain, domino):
|
||||
if chain is not None:
|
||||
last = chain[-1]
|
||||
if len(chain) == 1 and last[0] == domino[0]:
|
||||
return [swap(*last), domino]
|
||||
elif len(chain) == 1 and last[0] == domino[1]:
|
||||
return [swap(*last), swap(*domino)]
|
||||
elif last[1] == domino[0]:
|
||||
return chain + [domino]
|
||||
elif last[1] == domino[1]:
|
||||
return chain + [swap(*domino)]
|
||||
return None
|
||||
|
||||
|
||||
def can_chain(dominoes):
|
||||
if not any(dominoes):
|
||||
return []
|
||||
for perm in permutations(dominoes):
|
||||
chain = reduce(build_chain, perm[1:], [perm[0]])
|
||||
if chain is not None and chain[0][0] == chain[-1][1]:
|
||||
return chain
|
||||
return None
|
||||
''',
|
||||
"refactor_dominoes_chain": '''
|
||||
from itertools import permutations
|
||||
|
||||
|
||||
def can_chain(dominoes):
|
||||
if not any(dominoes):
|
||||
return []
|
||||
for perm in permutations(dominoes):
|
||||
chain = [perm[0]]
|
||||
complete = True
|
||||
for domino in perm[1:]:
|
||||
prev = chain[-1]
|
||||
if len(chain) == 1 and prev[0] == domino[0]:
|
||||
chain = [(prev[1], prev[0]), domino]
|
||||
elif len(chain) == 1 and prev[0] == domino[1]:
|
||||
chain = [(prev[1], prev[0]), (domino[1], domino[0])]
|
||||
elif prev[1] == domino[0]:
|
||||
chain = chain + [domino]
|
||||
elif prev[1] == domino[1]:
|
||||
chain = chain + [(domino[1], domino[0])]
|
||||
else:
|
||||
complete = False
|
||||
break
|
||||
if complete and chain[0][0] == chain[-1][1]:
|
||||
return chain
|
||||
return None
|
||||
''',
|
||||
"refactor_affine_encode_decode": '''
|
||||
BLOCK_SIZE = 5
|
||||
ALPHABET = 26
|
||||
|
||||
|
||||
def mod_inverse(a_key, alphabet):
|
||||
a_key = a_key % alphabet
|
||||
for idx in range(1, alphabet):
|
||||
if (a_key * idx) % alphabet == 1:
|
||||
return idx
|
||||
return 1
|
||||
|
||||
|
||||
def translate(text, a_key, b_key, mode):
|
||||
inverse = mod_inverse(a_key, ALPHABET)
|
||||
if inverse == 1:
|
||||
raise ValueError('a and m must be coprime.')
|
||||
chars = []
|
||||
for character in text:
|
||||
if character.isalnum():
|
||||
origin = ord(character.lower()) - 97
|
||||
if origin < 0:
|
||||
chars.append(character)
|
||||
continue
|
||||
if mode == 0:
|
||||
new = (a_key * origin + b_key) % ALPHABET
|
||||
elif mode == 1:
|
||||
new = (inverse * (origin - b_key)) % ALPHABET
|
||||
chars.append(chr(new + 97))
|
||||
return ''.join(chars)
|
||||
|
||||
|
||||
def encode(plain, a, b):
|
||||
cipher = translate(plain, a, b, 0)
|
||||
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
|
||||
for idx in range(0, len(cipher), BLOCK_SIZE)])
|
||||
|
||||
|
||||
def decode(ciphered, a, b):
|
||||
return translate(ciphered, a, b, 1)
|
||||
''',
|
||||
"debug_affine_coprime": '''
|
||||
BLOCK_SIZE = 5
|
||||
ALPHABET = 26
|
||||
|
||||
|
||||
def mod_inverse(a_key, alphabet):
|
||||
a_key = a_key % alphabet
|
||||
for idx in range(1, alphabet):
|
||||
if (a_key * idx) % alphabet == 1:
|
||||
return idx
|
||||
return 1
|
||||
|
||||
|
||||
def translate(text, a_key, b_key, mode):
|
||||
inverse = mod_inverse(a_key, ALPHABET)
|
||||
if inverse == 1:
|
||||
raise ValueError('a and m must be coprime.')
|
||||
chars = []
|
||||
for character in text:
|
||||
if character.isalnum():
|
||||
origin = ord(character.lower()) - 97
|
||||
if origin < 0:
|
||||
chars.append(character)
|
||||
continue
|
||||
if mode == 0:
|
||||
new = (a_key * origin + b_key) % ALPHABET
|
||||
elif mode == 1:
|
||||
new = (inverse * (origin - b_key)) % ALPHABET
|
||||
chars.append(chr(new + 97))
|
||||
return ''.join(chars)
|
||||
|
||||
|
||||
def encode(plain, a, b):
|
||||
cipher = translate(plain, a, b, 0)
|
||||
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
|
||||
for idx in range(0, len(cipher), BLOCK_SIZE)])
|
||||
|
||||
|
||||
def decode(ciphered, a, b):
|
||||
return translate(ciphered, a, b, 1)
|
||||
''',
|
||||
"debug_late_binding": '''
|
||||
def make_multipliers(factors):
|
||||
@@ -151,6 +474,318 @@ def first_match(items, predicates):
|
||||
done = True
|
||||
break
|
||||
return found
|
||||
''',
|
||||
"refactor_bowling_frames": '''
|
||||
class BowlingGame:
|
||||
def __init__(self):
|
||||
self._frames = []
|
||||
self._current = 0
|
||||
self._bonus = []
|
||||
|
||||
def roll(self, pins):
|
||||
if not (0 <= pins <= 10):
|
||||
raise ValueError('invalid pins')
|
||||
if self._current < 10:
|
||||
if len(self._frames) == self._current:
|
||||
self._frames.append([pins])
|
||||
else:
|
||||
self._frames[self._current].append(pins)
|
||||
current = self._frames[self._current]
|
||||
if sum(current) > 10:
|
||||
raise ValueError("a frame's rolls cannot exceed 10")
|
||||
strike = (len(current) == 1 and current[0] == 10)
|
||||
if strike or len(current) == 2:
|
||||
self._current += 1
|
||||
else:
|
||||
last = self._frames[-1]
|
||||
strike10 = len(last) == 1 and last[0] == 10
|
||||
spare10 = len(last) == 2 and sum(last) == 10
|
||||
if not (strike10 or spare10):
|
||||
raise IndexError('cannot throw bonus with an open tenth frame')
|
||||
if strike10:
|
||||
if len(self._bonus) >= 2:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a strike')
|
||||
self._bonus.append(pins)
|
||||
if len(self._bonus) == 2 and self._bonus[0] != 10 and sum(self._bonus) > 10:
|
||||
raise ValueError('invalid fill balls')
|
||||
if len(self._bonus) > 2:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a strike')
|
||||
elif spare10:
|
||||
if len(self._bonus) >= 1:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a spare')
|
||||
self._bonus.append(pins)
|
||||
|
||||
def score(self):
|
||||
if self._current < 10:
|
||||
raise IndexError('frame less than 10')
|
||||
last = self._frames[-1]
|
||||
if len(last) == 2 and sum(last) == 10 and len(self._bonus) != 1:
|
||||
raise IndexError('one bonus must be rolled when the tenth frame is spare')
|
||||
if len(last) == 1 and last[0] == 10 and len(self._bonus) != 2:
|
||||
raise IndexError('two bonuses must be rolled when the tenth frame is strike')
|
||||
total = 0
|
||||
for i in range(10):
|
||||
frame = self._frames[i]
|
||||
frame_sum = sum(frame)
|
||||
strike = (len(frame) == 1 and frame[0] == 10)
|
||||
spare = (len(frame) == 2 and frame_sum == 10)
|
||||
if strike or spare:
|
||||
nxt = []
|
||||
for j in range(i + 1, 10):
|
||||
nxt.extend(self._frames[j])
|
||||
nxt.extend(self._bonus)
|
||||
if strike:
|
||||
frame_sum += sum(nxt[:2])
|
||||
else:
|
||||
frame_sum += sum(nxt[:1])
|
||||
total += frame_sum
|
||||
return total
|
||||
''',
|
||||
"debug_bowling_tenth_frame": '''
|
||||
class Frame:
|
||||
def __init__(self, idx):
|
||||
self.idx = idx
|
||||
self.throws = []
|
||||
|
||||
@property
|
||||
def total_pins(self):
|
||||
return sum(self.throws)
|
||||
|
||||
def is_strike(self):
|
||||
return self.total_pins == 10 and len(self.throws) == 1
|
||||
|
||||
def is_spare(self):
|
||||
return self.total_pins == 10 and len(self.throws) == 2
|
||||
|
||||
def is_open(self):
|
||||
return self.total_pins < 10 and len(self.throws) == 2
|
||||
|
||||
def is_closed(self):
|
||||
return self.total_pins == 10 or len(self.throws) == 2
|
||||
|
||||
def throw(self, pins):
|
||||
if self.total_pins + pins > 10:
|
||||
raise ValueError("a frame's rolls cannot exceed 10")
|
||||
self.throws.append(pins)
|
||||
|
||||
def score(self, next_throws):
|
||||
result = self.total_pins
|
||||
if self.is_strike():
|
||||
result += sum(next_throws[:2])
|
||||
elif self.is_spare():
|
||||
result += sum(next_throws[:1])
|
||||
return result
|
||||
|
||||
|
||||
class BowlingGame:
|
||||
def __init__(self):
|
||||
self.current_frame_idx = 0
|
||||
self.bonus_throws = []
|
||||
self.frames = [Frame(idx) for idx in range(10)]
|
||||
|
||||
@property
|
||||
def current_frame(self):
|
||||
return self.frames[self.current_frame_idx]
|
||||
|
||||
def next_throws(self, frame_idx):
|
||||
throws = []
|
||||
for idx in range(frame_idx + 1, 10):
|
||||
throws.extend(self.frames[idx].throws)
|
||||
throws.extend(self.bonus_throws)
|
||||
return throws
|
||||
|
||||
def roll_bonus(self, pins):
|
||||
tenth_frame = self.frames[-1]
|
||||
if tenth_frame.is_open():
|
||||
raise IndexError('cannot throw bonus with an open tenth frame')
|
||||
self.bonus_throws.append(pins)
|
||||
# BUG: the invalid fill-balls guard below has been removed.
|
||||
# if (len(self.bonus_throws) == 2 and self.bonus_throws[0] != 10 and
|
||||
# sum(self.bonus_throws) > 10):
|
||||
# raise ValueError('invalid fill balls')
|
||||
if tenth_frame.is_strike() and len(self.bonus_throws) > 2:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a strike')
|
||||
elif tenth_frame.is_spare() and len(self.bonus_throws) > 1:
|
||||
raise IndexError(
|
||||
'wrong number of fill balls when the tenth frame is a spare')
|
||||
|
||||
def roll(self, pins):
|
||||
if not 0 <= pins <= 10:
|
||||
raise ValueError('invalid pins')
|
||||
elif self.current_frame_idx == 10:
|
||||
self.roll_bonus(pins)
|
||||
else:
|
||||
self.current_frame.throw(pins)
|
||||
if self.current_frame.is_closed():
|
||||
self.current_frame_idx += 1
|
||||
|
||||
def score(self):
|
||||
if self.current_frame_idx < 10:
|
||||
raise IndexError('frame less than 10')
|
||||
if self.frames[-1].is_spare() and len(self.bonus_throws) != 1:
|
||||
raise IndexError(
|
||||
'one bonus must be rolled when the tenth frame is spare')
|
||||
if self.frames[-1].is_strike() and len(self.bonus_throws) != 2:
|
||||
raise IndexError(
|
||||
'two bonuses must be rolled when the tenth frame is strike')
|
||||
return sum(frame.score(self.next_throws(frame.idx))
|
||||
for frame in self.frames)
|
||||
''',
|
||||
"debug_dominoes_no_chain": '''
|
||||
from itertools import permutations
|
||||
from functools import reduce
|
||||
|
||||
|
||||
def swap(item_1, item_2):
|
||||
return (item_2, item_1)
|
||||
|
||||
|
||||
def build_chain(chain, domino):
|
||||
if chain is not None:
|
||||
last = chain[-1]
|
||||
if len(chain) == 1 and last[0] == domino[0]:
|
||||
return [swap(*last), domino]
|
||||
elif len(chain) == 1 and last[0] == domino[1]:
|
||||
return [swap(*last), swap(*domino)]
|
||||
elif last[1] == domino[0]:
|
||||
return chain + [domino]
|
||||
elif last[1] == domino[1]:
|
||||
return chain + [swap(*domino)]
|
||||
return None
|
||||
|
||||
|
||||
def can_chain(dominoes):
|
||||
if not any(dominoes):
|
||||
return []
|
||||
for perm in permutations(dominoes):
|
||||
chain = reduce(build_chain, perm[1:], [perm[0]])
|
||||
# BUG: the circular-closure check (chain[0][0] == chain[-1][1]) is
|
||||
# missing, so a line that merely matches end-to-start is returned
|
||||
# even when it does not close into a loop.
|
||||
if chain is not None:
|
||||
return chain
|
||||
return None
|
||||
''',
|
||||
"refactor_dominoes_chain": '''
|
||||
from itertools import permutations
|
||||
|
||||
|
||||
def can_chain(dominoes):
|
||||
if not any(dominoes):
|
||||
return []
|
||||
for perm in permutations(dominoes):
|
||||
chain = [perm[0]]
|
||||
complete = True
|
||||
for domino in perm[1:]:
|
||||
prev = chain[-1]
|
||||
if len(chain) == 1 and prev[0] == domino[0]:
|
||||
chain = [(prev[1], prev[0]), domino]
|
||||
elif len(chain) == 1 and prev[0] == domino[1]:
|
||||
chain = [(prev[1], prev[0]), (domino[1], domino[0])]
|
||||
elif prev[1] == domino[0]:
|
||||
chain = chain + [domino]
|
||||
elif prev[1] == domino[1]:
|
||||
chain = chain + [(domino[1], domino[0])]
|
||||
else:
|
||||
complete = False
|
||||
break
|
||||
if complete and chain[0][0] == chain[-1][1]:
|
||||
return chain
|
||||
return None
|
||||
''',
|
||||
"refactor_affine_encode_decode": '''
|
||||
BLOCK_SIZE = 5
|
||||
ALPHABET = 26
|
||||
|
||||
|
||||
def mod_inverse(a_key, alphabet):
|
||||
a_key = a_key % alphabet
|
||||
for idx in range(1, alphabet):
|
||||
if (a_key * idx) % alphabet == 1:
|
||||
return idx
|
||||
return 1
|
||||
|
||||
|
||||
def encode(plain, a, b):
|
||||
inverse = mod_inverse(a, ALPHABET)
|
||||
if inverse == 1:
|
||||
raise ValueError('a and m must be coprime.')
|
||||
chars = []
|
||||
for character in plain:
|
||||
if character.isalnum():
|
||||
origin = ord(character.lower()) - 97
|
||||
if origin < 0:
|
||||
chars.append(character)
|
||||
continue
|
||||
new = (a * origin + b) % ALPHABET
|
||||
chars.append(chr(new + 97))
|
||||
cipher = ''.join(chars)
|
||||
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
|
||||
for idx in range(0, len(cipher), BLOCK_SIZE)])
|
||||
|
||||
|
||||
def decode(ciphered, a, b):
|
||||
inverse = mod_inverse(a, ALPHABET)
|
||||
if inverse == 1:
|
||||
raise ValueError('a and m must be coprime.')
|
||||
chars = []
|
||||
for character in ciphered:
|
||||
if character.isalnum():
|
||||
origin = ord(character.lower()) - 97
|
||||
if origin < 0:
|
||||
chars.append(character)
|
||||
continue
|
||||
new = (inverse * (origin - b)) % ALPHABET
|
||||
chars.append(chr(new + 97))
|
||||
return ''.join(chars)
|
||||
''',
|
||||
"debug_affine_coprime": '''
|
||||
BLOCK_SIZE = 5
|
||||
ALPHABET = 26
|
||||
|
||||
|
||||
def mod_inverse(a_key, alphabet):
|
||||
a_key = a_key % alphabet
|
||||
for idx in range(1, alphabet):
|
||||
if (a_key * idx) % alphabet == 1:
|
||||
return idx
|
||||
return 1
|
||||
|
||||
|
||||
def translate(text, a_key, b_key, mode):
|
||||
inverse = mod_inverse(a_key, ALPHABET)
|
||||
# BUG: compares `inverse < 1` instead of `inverse == 1`, so the coprime
|
||||
# guard never fires and a key whose `a` is not coprime with 26 slips
|
||||
# through without raising ValueError.
|
||||
if inverse < 1:
|
||||
raise ValueError('a and m must be coprime.')
|
||||
chars = []
|
||||
for character in text:
|
||||
if character.isalnum():
|
||||
origin = ord(character.lower()) - 97
|
||||
if origin < 0:
|
||||
chars.append(character)
|
||||
continue
|
||||
if mode == 0:
|
||||
new = (a_key * origin + b_key) % ALPHABET
|
||||
elif mode == 1:
|
||||
new = (inverse * (origin - b_key)) % ALPHABET
|
||||
chars.append(chr(new + 97))
|
||||
return ''.join(chars)
|
||||
|
||||
|
||||
def encode(plain, a, b):
|
||||
cipher = translate(plain, a, b, 0)
|
||||
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
|
||||
for idx in range(0, len(cipher), BLOCK_SIZE)])
|
||||
|
||||
|
||||
def decode(ciphered, a, b):
|
||||
return translate(ciphered, a, b, 1)
|
||||
''',
|
||||
"debug_late_binding": '''
|
||||
def make_multipliers(factors):
|
||||
@@ -168,14 +803,15 @@ def extract_tags(text):
|
||||
}
|
||||
|
||||
|
||||
def test_refactor_target_already_passes_its_own_checks():
|
||||
@pytest.mark.parametrize("task_id", [tid for tid in sorted(BUGGY) if tid.startswith("refactor_")])
|
||||
def test_refactor_target_already_passes_its_own_checks(task_id):
|
||||
# A refactor task's ORIGINAL code must pass, or the task is secretly a
|
||||
# debugging task and "behaviour must not change" is a lie
|
||||
score, detail = score_code(BUGGY["refactor_first_match"], BY_ID["refactor_first_match"]["checks"])
|
||||
assert score == 1.0, f"refactor target fails its own checks: {detail}"
|
||||
score, detail = score_code(BUGGY[task_id], BY_ID[task_id]["checks"])
|
||||
assert score == 1.0, f"{task_id}: refactor target fails its own checks: {detail}"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("task_id", ["debug_late_binding", "debug_greedy_regex"])
|
||||
@pytest.mark.parametrize("task_id", ["debug_late_binding", "debug_greedy_regex", "debug_bowling_tenth_frame", "debug_dominoes_no_chain", "debug_affine_coprime"])
|
||||
def test_debugging_tasks_start_broken(task_id):
|
||||
# The whole point is that the given code fails. If it passes, the task
|
||||
# measures nothing — a model could return the input unchanged.
|
||||
@@ -210,6 +846,160 @@ def test_counting_answer_matches_brute_force():
|
||||
assert str(count) == BY_ID["math_counting"]["answer"]
|
||||
|
||||
|
||||
# --- exact answers as Python literals -------------------------------------
|
||||
|
||||
def test_exact_scalar_does_not_match_list():
|
||||
# A scalar number answer must never cross-match a list literal (nor the
|
||||
# reverse): the model said "3", the task wanted "[1, 2, 3]",
|
||||
# and that is a wrong answer, not a formatting variant.
|
||||
assert score_exact("3", "[1, 2, 3]")[0] == 0.0
|
||||
assert score_exact("[1, 2, 3]", "3")[0] == 0.0
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"literal",
|
||||
["[1, 2, 3]", "{'a': 1}", "(4, 5)", "None", "True", "'text'", "3.5", "-12"],
|
||||
)
|
||||
def test_exact_structured_answers_match_their_literal(literal):
|
||||
# Structured/literal answers match themselves: lists, dicts, tuples,
|
||||
# None, bools, quoted strings, floats.
|
||||
assert score_exact(literal, literal)[0] == 1.0
|
||||
|
||||
|
||||
def test_exact_fenced_literal():
|
||||
# A reply wrapped in ``` fences is just formatting, not a different answer.
|
||||
assert score_exact("```\n[1, 2, 3]\n```", "[1, 2, 3]")[0] == 1.0
|
||||
|
||||
|
||||
def test_exact_comma_number_is_not_a_tuple():
|
||||
# "3,000" is a comma-separated thousands number, not the tuple (3, 0).
|
||||
assert score_exact("3,000", "3000")[0] == 1.0
|
||||
assert score_exact("1,000,000", "1000000")[0] == 1.0
|
||||
|
||||
|
||||
def test_exact_prose_number_still_falls_back():
|
||||
# Prose-wrapped numbers keep matching the bare number.
|
||||
assert score_exact("The answer is 3 minutes.", "3")[0] == 1.0
|
||||
assert score_exact("there are 100 widgets", "100")[0] == 1.0
|
||||
|
||||
|
||||
def test_exact_float_and_int_agree():
|
||||
# Numerically-equal floats/ints and currency strings stay equivalent.
|
||||
assert score_exact("3.0", "3")[0] == 1.0
|
||||
assert score_exact("3.50", "3.5")[0] == 1.0
|
||||
assert score_exact("$100", "100")[0] == 1.0
|
||||
|
||||
|
||||
def test_exact_bool_does_not_cross_match_number():
|
||||
# True is not the number 1, in either order.
|
||||
assert score_exact("True", "1")[0] == 0.0
|
||||
assert score_exact("1", "True")[0] == 0.0
|
||||
|
||||
|
||||
# --- CRUXEval-O exact answers (recomputed) --------------------------------
|
||||
|
||||
# Mapping: task id → (code, input) where the model is asked what
|
||||
# `f(<input>)` returns. The answer is verified by `eval("f(" + input + ")")`
|
||||
# NOT `literal_eval(input)` because CRUXEval-O input is the source arg list,
|
||||
# not a single literal.
|
||||
|
||||
_CRUX_ROWS = {
|
||||
"crux_o_sample_9": (
|
||||
textwrap.dedent("""\
|
||||
def f(t):
|
||||
for c in t:
|
||||
if not c.isnumeric():
|
||||
return False
|
||||
return True
|
||||
""").strip(),
|
||||
"'#284376598'",
|
||||
),
|
||||
"crux_o_sample_0": (
|
||||
textwrap.dedent("""\
|
||||
def f(nums):
|
||||
output = []
|
||||
for n in nums:
|
||||
output.append((nums.count(n), n))
|
||||
output.sort(reverse=True)
|
||||
return output
|
||||
""").strip(),
|
||||
"[1, 1, 3, 1, 3, 1]",
|
||||
),
|
||||
"crux_o_sample_1": (
|
||||
textwrap.dedent("""\
|
||||
def f(a, b, c):
|
||||
result = {}
|
||||
for d in a, b, c:
|
||||
result.update(dict.fromkeys(d))
|
||||
return result
|
||||
""").strip(),
|
||||
"(1, ), (1, ), (1, 2)",
|
||||
),
|
||||
"crux_o_sample_2": (
|
||||
textwrap.dedent("""\
|
||||
def f(text):
|
||||
new_text = list(text)
|
||||
for i in '+':
|
||||
if i in new_text:
|
||||
new_text.remove(i)
|
||||
return ''.join(new_text)
|
||||
""").strip(),
|
||||
"'hbtofdeiequ'",
|
||||
),
|
||||
"crux_o_sample_5": (
|
||||
textwrap.dedent("""\
|
||||
def f(text, lower, upper):
|
||||
count = 0
|
||||
new_text = list()
|
||||
for char in text:
|
||||
char = lower if char.isdecimal() else upper
|
||||
if char in ['p', 'C']:
|
||||
count += 1
|
||||
new_text.append(char)
|
||||
return count, ''.join(new_text)
|
||||
""").strip(),
|
||||
"'DSUWeqExTQdCMGpqur', 'a', 'x'",
|
||||
),
|
||||
"crux_o_sample_6": (
|
||||
textwrap.dedent("""\
|
||||
def f(dic):
|
||||
for k,v in sorted(dic.items(), key=lambda x: len(str(x)))[:-1]:
|
||||
dic.pop(k)
|
||||
return list(dic.items())
|
||||
""").strip(),
|
||||
"{'11': 52, '65': 34, 'a': 12, '4': 52, '74': 31}",
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def _crux_execute(code, input_str):
|
||||
"""Execute `f(<input>)` and return `repr(result)` matching the YAML answer."""
|
||||
ns = {}
|
||||
exec(code, ns) # noqa: S102
|
||||
result = eval("f(" + input_str + ")", ns)
|
||||
return repr(result)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("task_id", sorted(_CRUX_ROWS))
|
||||
def test_crux_answer_matches_execution(task_id):
|
||||
code, input_str = _CRUX_ROWS[task_id]
|
||||
expected = _crux_execute(code, input_str)
|
||||
actual = BY_ID[task_id]["answer"]
|
||||
assert actual == expected, (
|
||||
f"{task_id}: yaml answer={actual!r} != execution repr={expected!r} "
|
||||
f"(eval(\"f({input_str})\"))"
|
||||
)
|
||||
|
||||
|
||||
def test_crux_prompts_request_literal_only():
|
||||
for task_id in _CRUX_ROWS:
|
||||
task = BY_ID[task_id]
|
||||
prompt = task["prompt"].lower()
|
||||
assert "literal" in prompt or "only" in prompt, (
|
||||
f"{task_id}: prompt does not restrict to literal-only reply"
|
||||
)
|
||||
|
||||
|
||||
# --- task set hygiene -----------------------------------------------------
|
||||
|
||||
def test_every_task_has_the_fields_its_kind_needs():
|
||||
@@ -230,3 +1020,130 @@ def test_every_task_has_the_fields_its_kind_needs():
|
||||
def test_task_ids_are_unique():
|
||||
ids = [t["id"] for t in TASKS]
|
||||
assert len(ids) == len(set(ids))
|
||||
|
||||
|
||||
# --- BFCL-derived tool tasks ----------------------------------------------
|
||||
|
||||
_TOOL_NAME_RE = re.compile(r"^[A-Za-z0-9_-]{1,64}$")
|
||||
|
||||
|
||||
def _bfcl_positive():
|
||||
return [t for t in TASKS if t["id"].startswith("tool_bfcl_") and t.get("expect_tool")]
|
||||
|
||||
|
||||
def _bfcl_abstain():
|
||||
return [
|
||||
t
|
||||
for t in TASKS
|
||||
if t["id"].startswith("tool_bfcl_") and t.get("expect_tool") is None
|
||||
]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("task_id", [t["id"] for t in TASKS if t["id"].startswith("tool_bfcl_")])
|
||||
def test_bfcl_tools_are_openai_shaped(task_id):
|
||||
# The provider receives `tools` verbatim at request time; every BFCL tool
|
||||
# must already be in OpenAI shape (type:function wrapper, object params).
|
||||
task = BY_ID[task_id]
|
||||
for tool in task["tools"]:
|
||||
assert tool["type"] == "function"
|
||||
fn = tool["function"]
|
||||
assert fn["parameters"]["type"] == "object"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("task_id", [t["id"] for t in TASKS if t["id"].startswith("tool_bfcl_")])
|
||||
def test_bfcl_expect_args_are_scalar(task_id):
|
||||
# Ground-truth args that survive translation must be scalars (str/int/
|
||||
# float/bool). An array/object arg would have been skipped at translation,
|
||||
# so a container here means the selection filter regressed.
|
||||
task = BY_ID[task_id]
|
||||
for arg, value in (task.get("expect_args") or {}).items():
|
||||
assert isinstance(value, (str, int, float, bool)), (
|
||||
f"{task_id}: expect_args[{arg!r}] is not scalar"
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("task_id", [t["id"] for t in TASKS if t["kind"] == "tool"])
|
||||
def test_all_tool_names_are_provider_safe(task_id):
|
||||
# OpenAI tool names only allow A-Za-z0-9_- (1-64 chars). A dotted BFCL name
|
||||
# like uber.ride/aws.* would be rejected by the provider — the selection
|
||||
# filter must have elided those rows.
|
||||
task = BY_ID[task_id]
|
||||
for tool in task["tools"]:
|
||||
name = tool["function"]["name"]
|
||||
assert _TOOL_NAME_RE.fullmatch(name), f"{task_id}: unsafe tool name {name!r}"
|
||||
|
||||
|
||||
def test_bfcl_positive_scores_correct_call():
|
||||
import json
|
||||
|
||||
for task in _bfcl_positive():
|
||||
assert task["expect_args"], f"{task['id']} has no expect_args"
|
||||
args = json.dumps(task["expect_args"])
|
||||
call = [{"function": {"name": task["expect_tool"], "arguments": args}}]
|
||||
score, detail = score_tool(call, task)
|
||||
assert score == 1.0, f"{task['id']}: correct call scored {score} ({detail})"
|
||||
# A wrong tool name must not score: the task discriminates the call.
|
||||
wrong = [{"function": {"name": "some_other_tool", "arguments": args}}]
|
||||
assert score_tool(wrong, task)[0] == 0.0, f"{task['id']}: wrong tool scored"
|
||||
|
||||
|
||||
def test_bfcl_irrelevant_scores_abstention():
|
||||
# The 5 BFCL abstain tasks pair a weather tool with a VirusTotal/DNS/CPU
|
||||
# question. A model that reaches for the offered tool must score 0.0;
|
||||
# abstaining must score 1.0.
|
||||
for task in _bfcl_abstain():
|
||||
assert score_tool([], task)[0] == 1.0, f"{task['id']} did not reward abstention"
|
||||
for tool in task["tools"]:
|
||||
name = tool["function"]["name"]
|
||||
call = [{"function": {"name": name, "arguments": "{}"}}]
|
||||
assert score_tool(call, task)[0] == 0.0, f"{task['id']}: calling {name} scored"
|
||||
|
||||
|
||||
# --- EV-01: Regression tests for normalize_answer / score_exact correctness ---
|
||||
|
||||
def test_repro_math_counting_comma_prose():
|
||||
# Bug 1: "the combinations is 4,536." must match answer "4536"
|
||||
assert score_exact("the combinations is 4,536.", "4536")[0] == 1.0
|
||||
# Standalone comma-number must also match
|
||||
assert score_exact("4,536", "4536")[0] == 1.0
|
||||
|
||||
|
||||
def test_repro_fence_with_trailing_prose():
|
||||
# Bug 2: Trailing prose after a code fence must NOT be discarded
|
||||
assert score_exact("```python\nx = 5 * 4\n```\nThe answer is 20.", "20")[0] == 1.0
|
||||
|
||||
|
||||
def test_repro_fence_only_still_literal():
|
||||
# Whole-reply fenced block should still extract the literal
|
||||
assert score_exact("```\n[1, 2, 3]\n```", "[1, 2, 3]")[0] == 1.0
|
||||
|
||||
|
||||
def test_repro_bool_punctuated_matches():
|
||||
# Bug 3: "False." must normalize to "False" (literal repr)
|
||||
assert score_exact("False.", "False")[0] == 1.0
|
||||
assert score_exact("True.", "True")[0] == 1.0
|
||||
|
||||
|
||||
def test_repro_none_matches_none_dot():
|
||||
# "None." prose should normalize to "None" literal repr
|
||||
assert score_exact("None.", "None")[0] == 1.0
|
||||
|
||||
|
||||
def test_repro_nested_comma_thousands_match():
|
||||
# Bug 4: comma-thousands inside a list must collapse so a model's
|
||||
# [3,000, 4,000] matches the true answer [3000, 4000], not misparse
|
||||
# to [3, 0, 4, 0].
|
||||
from eval_proficiency import normalize_answer
|
||||
|
||||
# Core bug: thousands-separator commas inside a list must collapse
|
||||
assert normalize_answer("[3,000, 4,000]") == "[3000, 4000]"
|
||||
assert score_exact("[3,000, 4,000]", "[3000, 4000]")[0] == 1.0
|
||||
|
||||
# Standalone comma-number still works via collapse -> literal_eval
|
||||
assert normalize_answer("3,000") == "3000"
|
||||
assert score_exact("total is 4,536.", "4536")[0] == 1.0
|
||||
|
||||
# Non-thousands commas must NOT be collapsed
|
||||
assert normalize_answer("[12,34]") == "[12, 34]" # repr-style spacing
|
||||
assert normalize_answer("[abc, def]") == "[abc def]"
|
||||
assert normalize_answer("[(1, 2), (1, 2)]") == "[(1, 2), (1, 2)]"
|
||||
|
||||
Reference in New Issue
Block a user