neuralwatt-router-service #12

Merged
alee merged 13 commits from neuralwatt-router-service into main 2026-08-30 21:23:28 +00:00
7 changed files with 2090 additions and 28 deletions

View File

@@ -21,7 +21,7 @@ code.
| Config | `config/config.yaml` + Pydantic (`src/config.py`), `extra="forbid"` | | Config | `config/config.yaml` + Pydantic (`src/config.py`), `extra="forbid"` |
| HTTP client | `requests` (pinned in `requirements.txt`) — do NOT add httpx2/aiohttp without a requirements bump | | HTTP client | `requests` (pinned in `requirements.txt`) — do NOT add httpx2/aiohttp without a requirements bump |
| TUI | `textual==8.2.8` — imported only by `tui*.py` modules, never by the dispatch path | | TUI | `textual==8.2.8` — imported only by `tui*.py` modules, never by the dispatch path |
| Testing | `pytest`, 733 tests, all offline (no provider or local-model calls) | | Testing | `pytest`, 796 tests, all offline (no provider or local-model calls) |
| Dependencies | Pinned. Bump deliberately, never use `>=` | | Dependencies | Pinned. Bump deliberately, never use `>=` |
## Module map and file boundaries ## Module map and file boundaries
@@ -214,7 +214,7 @@ See `CLAUDE.md` → "What's NOT built yet — pick up here" for the full list.
## After a code change ## After a code change
- Run `python -m pytest` — 733 tests, ~35s. - Run `python -m pytest` — 796 tests, ~35s.
- If you changed `dispatcher.py`, restart the systemd service: - If you changed `dispatcher.py`, restart the systemd service:
`systemctl --user restart llm-router.service` (it doesn't auto-reload code). `systemctl --user restart llm-router.service` (it doesn't auto-reload code).
- If you changed the TUI, run `PYTHONPATH=src python -m tui` to verify it starts. - If you changed the TUI, run `PYTHONPATH=src python -m tui` to verify it starts.

View File

@@ -216,6 +216,11 @@ rather than from months of history.
(13 models x 5 = 65 calls, and the whole sweep cost **under a cent**). (13 models x 5 = 65 calls, and the whole sweep cost **under a cent**).
- `tiering.py` / `tier.py` — pure tier resolver + the DB pass that applies it. - `tiering.py` / `tier.py` — pure tier resolver + the DB pass that applies it.
- `routing.py` — pure hard filters and ranking. - `routing.py` — pure hard filters and ranking.
- `circuit_breaker.py` — passive availability skip for a model that starts
5xxing. Pure in-memory module (mirrors `session_cache.py`'s shape:
module-level dict, injected time for tests, never imports
`dispatcher`/`config`). On by default (`circuit_breaker.enabled: true`).
See below for what it does and, deliberately, does not cover.
- `dispatcher.py` — FastAPI service. `GET /health`, `POST /route` (classify - `dispatcher.py` — FastAPI service. `GET /health`, `POST /route` (classify
and pick, no provider call), `POST /dispatch` (route, call, log), plus an and pick, no provider call), `POST /dispatch` (route, call, log), plus an
OpenAI-compatible `GET /v1/models` and `POST /v1/chat/completions`, and a OpenAI-compatible `GET /v1/models` and `POST /v1/chat/completions`, and a
@@ -282,10 +287,42 @@ rather than from months of history.
reset on restart. Persisted config edits limited to an allowlist. Model reset on restart. Persisted config edits limited to an allowlist. Model
availability overrides feed routing hard filters. Bucketed history endpoint over availability overrides feed routing hard filters. Bucketed history endpoint over
`energy_observations` and `route_decisions`. `energy_observations` and `route_decisions`.
- `tests/` — 733 tests across 41 files, all passing, all offline. Verified on - `tests/` — 796 tests across 40 files, all passing, all offline. Verified on
Python 3.10 and 3.14; nothing declares `requires-python`, so 3.10 is the Python 3.10 and 3.14; nothing declares `requires-python`, so 3.10 is the
tested floor rather than a promised one. tested floor rather than a promised one.
### Circuit breaker: passive availability skip, and why the eval harness stays outside it
`circuit_breaker.py` records nothing until a model actually fails: a 5xx
from the upstream call marks `(model_id, provider)` down for
`initial_cooldown_seconds` (30s default), doubling on each further failure
(`backoff_multiplier`, capped at `max_cooldown_seconds`, 600s default) and
clearing on the next success. Recovery is passive by design — no background
poller, no health-check loop. `is_down` just compares against `down_until`,
so the next real request that would have picked the down model becomes its
own recovery probe once the cooldown has passed. Two call sites:
`_open_circuits` (`dispatcher.py:1845`) excludes down models from the
candidate set during routing/selection, and the dispatch retry loop
(`dispatcher.py:2560-2605`) records the failure/success on every upstream
call and fails over to the next-ranked candidate on a 5xx — but only when
`alternatives` is non-empty, which is exactly the `auto`-routed case.
**`eval_proficiency.py` intentionally stays outside all of this** — it calls
NeuralWatt directly (`requests.post`, not through the dispatcher), and that
is correct, not a gap. Its whole design is pinning one specific `model_id`
per call; a pinned request has `wants_routing = False`, which means
`alternatives = []` at the top of the retry loop — there is no candidate to
fail over to even if the call went through the router, so circuit_breaker's
recovery machinery would have nothing to do for it. Routing eval traffic
through the dispatcher would be actively worse, not merely useless: a
transient eval-harness failure would still call `circuit_breaker.
record_failure`, which could trip the breaker against real production
traffic based on nothing but the benchmark exercising a model's edges.
Observed live 2026-08-30: `glm-5.3` (newly released) threw two `CALL FAILED
HTTPError`s during a `coding_refactor` eval pass, almost certainly transient
provider-side capacity ramp, with zero effect on production routing — which
is the isolation working as intended, not a missing integration.
### Monitoring: route_decisions persistence ### Monitoring: route_decisions persistence
`route_decisions` is the newest observability table (not a scoring input). It `route_decisions` is the newest observability table (not a scoring input). It
@@ -510,11 +547,22 @@ right outcome.
`coding_general` now spans 0.86-1.00: `glm-5.2-fast` fell to 0.862 over 29 `coding_general` now spans 0.86-1.00: `glm-5.2-fast` fell to 0.862 over 29
samples folded in by `feedback.py` from an actual agent session, and crossed samples folded in by `feedback.py` from an actual agent session, and crossed
`self_eval_min_samples` on the way, so it reads `self_eval` rather than `self_eval_min_samples` on the way, so it reads `self_eval` rather than
`self_eval_thin`. That is the intended shape of this system — the 23-task `self_eval_thin`. That is the intended shape of this system — the 43-task
benchmark establishes a floor, and your own traffic is what refines it. benchmark establishes a floor, and your own traffic is what refines it.
`coding_refactor` and `debugging` are still flat at 1.00, awaiting the same `coding_refactor` and `debugging` are still flat at 1.00, awaiting the same
treatment. treatment.
**Benchmark-sourced hardening landed, and it broke both remaining ties.** The
task set grew from 23 to 43: 8 BFCL tool tasks, 6 CRUXEval-O exact tasks (after
the `score_exact` literal-eval fix), and three Exercism refactor/debug pairs.
The CRUXEval-O rows split `coding_general` into a 0-1 mix across models — five
of six now fail at least one model — and the Exercism refactor rows moved
`coding_refactor` off its flat 1.00: bowling 0.25-1.00, dominoes 0.10-1.00,
affine 0.56-1.00. The debug pairs split `debugging` the same way (bowling
0.80-1.00, dominoes a sharp 0/1 split). Two follow-ups recorded, not deleted:
`debug_affine_coprime` and most of the BFCL tasks sat flat at 1.00, so they do
not discriminate.
**A 1.00 can also be a sampling artifact, and `docs_writing` was one.** At 2 **A 1.00 can also be a sampling artifact, and `docs_writing` was one.** At 2
samples per model the category read 0.70-1.00 with a model at the ceiling, and samples per model the category read 0.70-1.00 with a model at the ceiling, and
the router paid for that ceiling: `kimi-k3-fast` won every docs route. Six more the router paid for that ceiling: `kimi-k3-fast` won every docs route. Six more
@@ -810,7 +858,7 @@ explains a `truncated` verdict and nothing else; it is now scoped to exactly
that. that.
`feedback.py` folds observed failures into `proficiency`, so routing learns `feedback.py` folds observed failures into `proficiency`, so routing learns
from your traffic rather than only the 23-task benchmark. Only **failures** from your traffic rather than only the 43-task benchmark. Only **failures**
are folded in: a structural 'ok' means the code parsed, not that it was are folded in: a structural 'ok' means the code parsed, not that it was
correct, and recording those as 1.0 would flatten every score toward the correct, and recording those as 1.0 would flatten every score toward the
ceiling. Failures the model did not cause — a client's own tight `max_tokens` ceiling. Failures the model did not cause — a client's own tight `max_tokens`
@@ -1012,11 +1060,11 @@ that request was two orders of magnitude low.
There is no composite any more — ranking is quality first, cost as the There is no composite any more — ranking is quality first, cost as the
tiebreak inside `quality_tolerance` — so nothing normalizes and nothing tiebreak inside `quality_tolerance` — so nothing normalizes and nothing
compresses. compresses.
- Answered: the eval set exists (`evals/tasks.yaml`, 23 tasks, four scoring - Answered: the eval set exists (`evals/tasks.yaml`, 43 tasks, four scoring
kinds) and `tests/test_task_set.py` keeps it honest. The open part is kinds) and `tests/test_task_set.py` keeps it honest. The benchmark-sourced
narrower now — `coding_refactor` and `debugging` are still flat at 1.00 rows now split the coding categories, but two tasks are still flat at 1.00
across every model, so those tasks discriminate nothing and either need (`debug_affine_coprime` and most of the BFCL set) and either need hardening
hardening again or should be conceded as non-discriminating. **Try samples again or should be conceded as non-discriminating. **Try samples
before hardening.** `docs_writing` looked flat at the top too, and six more before hardening.** `docs_writing` looked flat at the top too, and six more
passes spread it 0.66-0.97 without touching a task; two samples per model is passes spread it 0.66-0.97 without touching a task; two samples per model is
not enough to tell a saturated task from an unsampled one. not enough to tell a saturated task from an unsampled one.

View File

@@ -457,7 +457,7 @@ Gating and safeguards:
Observation → learning. `feedback.py` folds verification failures into Observation → learning. `feedback.py` folds verification failures into
`proficiency` so routing improves on **your traffic**, not just the fixed `proficiency` so routing improves on **your traffic**, not just the fixed
23-task benchmark: 43-task benchmark:
```bash ```bash
PYTHONPATH=src python -m feedback --dry-run # preview what would change PYTHONPATH=src python -m feedback --dry-run # preview what would change
@@ -544,7 +544,7 @@ Key behaviors:
| **Config** | `config/config.yaml` loaded & validated by Pydantic (`src/config.py`) | | **Config** | `config/config.yaml` loaded & validated by Pydantic (`src/config.py`) |
| **OpenAI Client** | `openai==3.0.0` (official SDK) | | **OpenAI Client** | `openai==3.0.0` (official SDK) |
| **HTTP** | `requests` for poller, `httpx` (via openai/uvicorn) | | **HTTP** | `requests` for poller, `httpx` (via openai/uvicorn) |
| **Testing** | `pytest` — 733 tests across 41 files, all offline | | **Testing** | `pytest` — 796 tests across 40 files, all offline |
| **Config Files** | `config/config.yaml`, `config/leaderboards.yaml`, `evals/tasks.yaml` | | **Config Files** | `config/config.yaml`, `config/leaderboards.yaml`, `evals/tasks.yaml` |
| **Deployment** | systemd user units (`.service` + `.timer` files in `deploy/`) | | **Deployment** | systemd user units (`.service` + `.timer` files in `deploy/`) |
| **Integration** | `opencode.json` in the repo routes through it by default; any OpenAI-compatible client works | | **Integration** | `opencode.json` in the repo routes through it by default; any OpenAI-compatible client works |
@@ -1153,7 +1153,7 @@ carries the same modality block.
## Testing ## Testing
```bash ```bash
python -m pytest # 733 tests python -m pytest # 796 tests
python -m pytest --cov # with coverage python -m pytest --cov # with coverage
``` ```

View File

@@ -88,6 +88,119 @@ tasks:
- 'word_wrap("a b c", 3) == ["a b", "c"]' - 'word_wrap("a b c", 3) == ["a b", "c"]'
- 'word_wrap("aa bb cc", 5) == ["aa bb", "cc"]' - 'word_wrap("aa bb cc", 5) == ["aa bb", "cc"]'
# --- coding_general (added) — adapted from CRUXEval-O output-prediction,
# facebookresearch/cruxeval, MIT License. Data fetched once to /tmp,
# never vendored. Filter: has loop, ≤400 LOC chars, no unsafe imports,
# safe stdlib only. 6 of 6 tasks use `eval("f(<input>)")` for
# offline verification (input is the source arg list, not a single
# literal).
- id: crux_o_sample_9
category: coding_general
kind: exact
answer: "False"
prompt: |
What does this function return when called as shown? Reply with ONLY
the literal Python value — strings in quotes (for example: 42, [1, 2],
'text', None), no explanation, no fences.
def f(t):
for c in t:
if not c.isnumeric():
return False
return True
f('#284376598')
- id: crux_o_sample_0
category: coding_general
kind: exact
answer: "[(4, 1), (4, 1), (4, 1), (4, 1), (2, 3), (2, 3)]"
prompt: |
What does this function return when called as shown? Reply with ONLY
the literal Python value — strings in quotes (for example: 42, [1, 2],
'text', None), no explanation, no fences.
def f(nums):
output = []
for n in nums:
output.append((nums.count(n), n))
output.sort(reverse=True)
return output
f([1, 1, 3, 1, 3, 1])
- id: crux_o_sample_1
category: coding_general
kind: exact
answer: "{1: None, 2: None}"
prompt: |
What does this function return when called as shown? Reply with ONLY
the literal Python value — strings in quotes (for example: 42, [1, 2],
'text', None), no explanation, no fences.
def f(a, b, c):
result = {}
for d in a, b, c:
result.update(dict.fromkeys(d))
return result
f((1, ), (1, ), (1, 2))
- id: crux_o_sample_2
category: coding_general
kind: exact
answer: "'hbtofdeiequ'"
prompt: |
What does this function return when called as shown? Reply with ONLY
the literal Python value — strings in quotes (for example: 42, [1, 2],
'text', None), no explanation, no fences.
def f(text):
new_text = list(text)
for i in '+':
if i in new_text:
new_text.remove(i)
return ''.join(new_text)
f('hbtofdeiequ')
- id: crux_o_sample_5
category: coding_general
kind: exact
answer: "(0, 'xxxxxxxxxxxxxxxxxx')"
prompt: |
What does this function return when called as shown? Reply with ONLY
the literal Python value — strings in quotes (for example: 42, [1, 2],
'text', None), no explanation, no fences.
def f(text, lower, upper):
count = 0
new_text = list()
for char in text:
char = lower if char.isdecimal() else upper
if char in ['p', 'C']:
count += 1
new_text.append(char)
return count, ''.join(new_text)
f('DSUWeqExTQdCMGpqur', 'a', 'x')
- id: crux_o_sample_6
category: coding_general
kind: exact
answer: "[('74', 31)]"
prompt: |
What does this function return when called as shown? Reply with ONLY
the literal Python value — strings in quotes (for example: 42, [1, 2],
'text', None), no explanation, no fences.
def f(dic):
for k,v in sorted(dic.items(), key=lambda x: len(str(x)))[:-1]:
dic.pop(k)
return list(dic.items())
f({'11': 52, '65': 34, 'a': 12, '4': 52, '74': 31})
# --- coding_refactor ---------------------------------------------------- # --- coding_refactor ----------------------------------------------------
- id: refactor_falsy_defaults - id: refactor_falsy_defaults
category: coding_refactor category: coding_refactor
@@ -178,6 +291,430 @@ tasks:
- 'describe(0) == "unknown"' - 'describe(0) == "unknown"'
- 'describe(None) == "unknown"' - 'describe(None) == "unknown"'
# --- bowling (Exercism-derived) -----------------------------------------
# Exercism python bowling exercise — MIT licensed, canonical source at
# https://github.com/exercism/python/tree/1f6aab8667bf653b10cc3799f94352fcdb749db6/exercises/practice/bowling
- id: refactor_bowling_frames
category: coding_refactor
kind: code
entrypoint: BowlingGame
prompt: |
Refactor this BowlingGame to remove the duplication and nested
conditions. Behaviour must be preserved EXACTLY, including scoring,
bonuses, and error cases. Reply with ONLY the rewritten class — no
explanation, no fences.
class BowlingGame:
def __init__(self):
self._frames = []
self._current = 0
self._bonus = []
def roll(self, pins):
if not (0 <= pins <= 10):
raise ValueError('invalid pins')
if self._current < 10:
if len(self._frames) == self._current:
self._frames.append([pins])
else:
self._frames[self._current].append(pins)
current = self._frames[self._current]
if sum(current) > 10:
raise ValueError("a frame's rolls cannot exceed 10")
strike = (len(current) == 1 and current[0] == 10)
if strike or len(current) == 2:
self._current += 1
else:
last = self._frames[-1]
last_total = sum(last)
strike10 = len(last) == 1 and last[0] == 10
spare10 = len(last) == 2 and last_total == 10
if strike10:
if len(self._bonus) >= 2:
raise IndexError(
'wrong number of fill balls when the tenth frame is a strike')
self._bonus.append(pins)
if len(self._bonus) == 2 and self._bonus[0] != 10 and sum(self._bonus) > 10:
raise ValueError('invalid fill balls')
if len(self._bonus) > 2:
raise IndexError(
'wrong number of fill balls when the tenth frame is a strike')
elif spare10:
if len(self._bonus) >= 1:
raise IndexError(
'wrong number of fill balls when the tenth frame is a spare')
self._bonus.append(pins)
else:
raise IndexError('cannot throw bonus with an open tenth frame')
def score(self):
if self._current < 10:
raise IndexError('frame less than 10')
last = self._frames[-1]
if len(last) == 2 and sum(last) == 10 and len(self._bonus) != 1:
raise IndexError('one bonus must be rolled when the tenth frame is spare')
if len(last) == 1 and last[0] == 10 and len(self._bonus) != 2:
raise IndexError('two bonuses must be rolled when the tenth frame is strike')
total = 0
for i in range(10):
frame = self._frames[i]
frame_sum = sum(frame)
strike = (len(frame) == 1 and frame[0] == 10)
spare = (len(frame) == 2 and frame_sum == 10)
if strike or spare:
nxt = []
for j in range(i + 1, 10):
nxt.extend(self._frames[j])
nxt.extend(self._bonus)
if strike:
frame_sum += sum(nxt[:2])
else:
frame_sum += sum(nxt[:1])
total += frame_sum
return total
checks:
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 7, 1]], g)[2].score() == 18)"
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [7, 3, 7]], g)[2].score() == 17)"
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 10, 10]], g)[2].score() == 30)"
- "((g := BowlingGame(), [g.roll(r) for r in [6, 4, 3] + [0]*17], g)[2].score() == 16)"
- "((g := BowlingGame(), [g.roll(r) for r in [3, 6]*10], g)[2].score() == 90)"
- "((g := BowlingGame(), [g.roll(r) for r in [10]*12], g)[2].score() == 300)"
- "raises(Exception, lambda: BowlingGame().roll(-1))"
- "raises(Exception, lambda: ((g := BowlingGame()), [g.roll(r) for r in [0]*20], g[2].roll(0)))"
- id: debug_bowling_tenth_frame
category: debugging
kind: code
entrypoint: BowlingGame
prompt: |
This BowlingGame is wrong on one subtle tenth-frame case. Fix ONLY the
bug; do not change anything else. Reply with ONLY the corrected class —
no explanation, no fences.
class BowlingGame:
def __init__(self):
self.current_frame_idx = 0
self.bonus_throws = []
self.frames = [Frame(idx) for idx in range(10)]
@property
def current_frame(self):
return self.frames[self.current_frame_idx]
def next_throws(self, frame_idx):
throws = []
for idx in range(frame_idx + 1, 10):
throws.extend(self.frames[idx].throws)
throws.extend(self.bonus_throws)
return throws
def roll_bonus(self, pins):
tenth_frame = self.frames[-1]
if tenth_frame.is_open():
raise IndexError('cannot throw bonus with an open tenth frame')
self.bonus_throws.append(pins)
if tenth_frame.is_strike() and len(self.bonus_throws) > 2:
raise IndexError(
'wrong number of fill balls when the tenth frame is a strike')
elif tenth_frame.is_spare() and len(self.bonus_throws) > 1:
raise IndexError(
'wrong number of fill balls when the tenth frame is a spare')
def roll(self, pins):
if not 0 <= pins <= 10:
raise ValueError('invalid pins')
elif self.current_frame_idx == 10:
self.roll_bonus(pins)
else:
self.current_frame.throw(pins)
if self.current_frame.is_closed():
self.current_frame_idx += 1
def score(self):
if self.current_frame_idx < 10:
raise IndexError('frame less than 10')
if self.frames[-1].is_spare() and len(self.bonus_throws) != 1:
raise IndexError(
'one bonus must be rolled when the tenth frame is spare')
if self.frames[-1].is_strike() and len(self.bonus_throws) != 2:
raise IndexError(
'two bonuses must be rolled when the tenth frame is strike')
return sum(frame.score(self.next_throws(frame.idx))
for frame in self.frames)
class Frame:
def __init__(self, idx):
self.idx = idx
self.throws = []
@property
def total_pins(self):
return sum(self.throws)
def is_strike(self):
return self.total_pins == 10 and len(self.throws) == 1
def is_spare(self):
return self.total_pins == 10 and len(self.throws) == 2
def is_open(self):
return self.total_pins < 10 and len(self.throws) == 2
def is_closed(self):
return self.total_pins == 10 or len(self.throws) == 2
def throw(self, pins):
if self.total_pins + pins > 10:
raise ValueError("a frame's rolls cannot exceed 10")
self.throws.append(pins)
def score(self, next_throws):
result = self.total_pins
if self.is_strike():
result += sum(next_throws[:2])
elif self.is_spare():
result += sum(next_throws[:1])
return result
checks:
- "raises(Exception, (lambda: ((g := BowlingGame()), [g.roll(r) for r in [0]*18 + [10, 5, 6]], g)[2].score()))"
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 10, 6]], g)[2].score() == 26)"
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 7, 1]], g)[2].score() == 18)"
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [7, 3, 7]], g)[2].score() == 17)"
- "((g := BowlingGame(), [g.roll(r) for r in [10]*12], g)[2].score() == 300)"
# --- dominoes (Exercism-derived) ----------------------------------------
# Exercism python dominoes exercise — MIT licensed, canonical source at
# https://github.com/exercism/python/tree/1f6aab8667bf653b10cc3799f94352fcdb749db6/exercises/practice/dominoes
- id: refactor_dominoes_chain
category: coding_refactor
kind: code
entrypoint: can_chain
prompt: |
Refactor this can_chain to remove the duplicated chain-building
conditions and the flag variable. Behaviour must be preserved EXACTLY:
for a set of dominoes that can form a valid chain it returns a valid
chain (ANY valid chain — not a fixed one), and None when no chain is
possible. Reply with ONLY the rewritten function — no explanation, no
fences.
from itertools import permutations
def can_chain(dominoes):
if not any(dominoes):
return []
for perm in permutations(dominoes):
chain = [perm[0]]
complete = True
for domino in perm[1:]:
prev = chain[-1]
if len(chain) == 1 and prev[0] == domino[0]:
chain = [(prev[1], prev[0]), domino]
elif len(chain) == 1 and prev[0] == domino[1]:
chain = [(prev[1], prev[0]), (domino[1], domino[0])]
elif prev[1] == domino[0]:
chain = chain + [domino]
elif prev[1] == domino[1]:
chain = chain + [(domino[1], domino[0])]
else:
complete = False
break
if complete and chain[0][0] == chain[-1][1]:
return chain
return None
checks:
- 'can_chain([]) == []'
- '((d := [(1, 1)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
- '((d := [(1, 2), (3, 1), (2, 3)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
- '((d := [(1, 2), (2, 3), (3, 1), (2, 4), (2, 4)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
- 'can_chain([(1, 2)]) is None'
- 'can_chain([(1, 2), (4, 1), (2, 3)]) is None'
- 'can_chain([(1, 1), (2, 2)]) is None'
- 'can_chain([(1, 2), (2, 1), (3, 4), (4, 3)]) is None'
- 'can_chain([(1, 2), (2, 3), (3, 1), (4, 4)]) is None'
- 'can_chain([(1, 2), (2, 3), (3, 1), (4, 5), (5, 6), (6, 4)]) is None'
- id: debug_dominoes_no_chain
category: debugging
kind: code
entrypoint: can_chain
prompt: |
This can_chain returns a bogus "chain" for inputs that cannot be
chained — it returns a list instead of None when no valid chain exists.
Fix ONLY the one subtle bug; do not change anything else, and do not
change the can_chain signature. Reply with ONLY the corrected function
— no explanation, no fences.
from itertools import permutations
from functools import reduce
def swap(item_1, item_2):
return (item_2, item_1)
def build_chain(chain, domino):
if chain is not None:
last = chain[-1]
if len(chain) == 1 and last[0] == domino[0]:
return [swap(*last), domino]
elif len(chain) == 1 and last[0] == domino[1]:
return [swap(*last), swap(*domino)]
elif last[1] == domino[0]:
return chain + [domino]
elif last[1] == domino[1]:
return chain + [swap(*domino)]
return None
def can_chain(dominoes):
if not any(dominoes):
return []
for perm in permutations(dominoes):
chain = reduce(build_chain, perm[1:], [perm[0]])
# BUG: the circular-closure check (chain[0][0] == chain[-1][1])
# is missing, so a line that merely matches end-to-start is
# returned even when it does not close into a loop.
if chain is not None:
return chain
return None
checks:
- 'can_chain([(1, 2)]) is None'
- 'can_chain([(1, 2), (4, 1), (2, 3)]) is None'
- 'can_chain([(1, 1), (2, 2)]) is None'
- '((d := [(1, 2), (3, 1), (2, 3)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
# --- affine-cipher (Exercism-derived) -----------------------------------
# Exercism python affine-cipher exercise — MIT licensed, canonical source at
# https://github.com/exercism/python/tree/1f6aab8667bf653b10cc3799f94352fcdb749db6/exercises/practice/affine-cipher
- id: refactor_affine_encode_decode
category: coding_refactor
kind: code
entrypoint: encode
prompt: |
Refactor this affine cipher to remove the duplicated cipher math. The
same letter-to-index transform and the coprime guard appear inline in
both encode and decode; behavioural duplicates like these are where bugs
hide. Consolidate them. Behaviour must be preserved EXACTLY, including
the ValueError raised when `a` is not coprime with the alphabet size and
the 5-character block grouping in encode. Keep the module functions
`encode(plain, a, b)` and `decode(ciphered, a, b)`. Reply with ONLY the
rewritten module — no explanation, no fences.
BLOCK_SIZE = 5
ALPHABET = 26
def mod_inverse(a_key, alphabet):
a_key = a_key % alphabet
for idx in range(1, alphabet):
if (a_key * idx) % alphabet == 1:
return idx
return 1
def encode(plain, a, b):
inverse = mod_inverse(a, ALPHABET)
if inverse == 1:
raise ValueError('a and m must be coprime.')
chars = []
for character in plain:
if character.isalnum():
origin = ord(character.lower()) - 97
if origin < 0:
chars.append(character)
continue
new = (a * origin + b) % ALPHABET
chars.append(chr(new + 97))
cipher = ''.join(chars)
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
for idx in range(0, len(cipher), BLOCK_SIZE)])
def decode(ciphered, a, b):
inverse = mod_inverse(a, ALPHABET)
if inverse == 1:
raise ValueError('a and m must be coprime.')
chars = []
for character in ciphered:
if character.isalnum():
origin = ord(character.lower()) - 97
if origin < 0:
chars.append(character)
continue
new = (inverse * (origin - b)) % ALPHABET
chars.append(chr(new + 97))
return ''.join(chars)
checks:
- 'encode("yes", 5, 7) == "xbt"'
- 'encode("no", 15, 18) == "fu"'
- 'encode("OMG", 21, 3) == "lvz"'
- 'encode("O M G", 25, 47) == "hjp"'
- 'encode("Testing,1 2 3, testing.", 3, 4) == "jqgjc rw123 jqgjc rw"'
- 'decode("tytgn fjr", 3, 7) == "exercism"'
- 'decode("qdwju nqcro muwhn odqun oppmd aunwd o", 19, 16) == "anobstacleisoftenasteppingstone"'
- 'raises(ValueError, encode, "This is a test.", 6, 17)'
- 'raises(ValueError, decode, "Test", 13, 5)'
- id: debug_affine_coprime
category: debugging
kind: code
entrypoint: encode
prompt: |
This affine cipher fails to reject keys where `a` is not coprime with the
alphabet size. It should raise ValueError('a and m must be coprime.')
when `a` shares a factor with 26, but it lets those keys through. Fix
ONLY the one subtle bug in the coprime guard; do not change anything
else, and do not change the signatures of `encode(plain, a, b)` or
`decode(ciphered, a, b)`. Reply with ONLY the corrected module — no
explanation, no fences.
BLOCK_SIZE = 5
ALPHABET = 26
def mod_inverse(a_key, alphabet):
a_key = a_key % alphabet
for idx in range(1, alphabet):
if (a_key * idx) % alphabet == 1:
return idx
return 1
def translate(text, a_key, b_key, mode):
inverse = mod_inverse(a_key, ALPHABET)
if inverse < 1:
raise ValueError('a and m must be coprime.')
chars = []
for character in text:
if character.isalnum():
origin = ord(character.lower()) - 97
if origin < 0:
chars.append(character)
continue
if mode == 0:
new = (a_key * origin + b_key) % ALPHABET
elif mode == 1:
new = (inverse * (origin - b_key)) % ALPHABET
chars.append(chr(new + 97))
return ''.join(chars)
def encode(plain, a, b):
cipher = translate(plain, a, b, 0)
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
for idx in range(0, len(cipher), BLOCK_SIZE)])
def decode(ciphered, a, b):
return translate(ciphered, a, b, 1)
checks:
- 'encode("yes", 5, 7) == "xbt"'
- 'decode("tytgn fjr", 3, 7) == "exercism"'
- 'raises(ValueError, encode, "This is a test.", 6, 17)'
- 'raises(ValueError, decode, "Test", 13, 5)'
# --- debugging ---------------------------------------------------------- # --- debugging ----------------------------------------------------------
- id: debug_late_binding - id: debug_late_binding
category: debugging category: debugging
@@ -346,6 +883,221 @@ tasks:
location: {type: string} location: {type: string}
required: [location] required: [location]
# BFCL v4 derived (ShishirPatil/gorilla, MIT): 5 abstain from
# BFCL_v4_live_irrelevance.json, 3 positive from BFCL_v4_live_simple.json +
# possible_answer ground truth. Data fetched once to /tmp, translated
# mechanically, never vendored. Irrelevant prompts are deliberately
# unanswerable by the offered weather tool (real abstain cases).
- id: tool_bfcl_live_irrelevance_12-2-0
category: tool_use_agentic
kind: tool
prompt: I'd appreciate if you could fetch the DNS resolution info for the domain mapped to IP 255.255.255.0 from VirusTotal. My key for this operation is 'sample_key4'.
expect_tool: null
tools:
- type: function
function:
name: get_current_weather
description: Retrieves the current weather conditions for a specified city and state.
parameters:
type: object
required:
- location
properties:
location:
type: string
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
exist.
unit:
type: string
description: The unit of temperature for the weather report.
enum:
- celsius
- fahrenheit
default: fahrenheit
- id: tool_bfcl_live_irrelevance_13-2-1
category: tool_use_agentic
kind: tool
prompt: What is diffrence between cpu and gpu?
expect_tool: null
tools:
- type: function
function:
name: get_current_weather
description: Retrieves the current weather conditions for a specified city and state.
parameters:
type: object
required:
- location
properties:
location:
type: string
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
exist.
unit:
type: string
description: The unit of temperature for the weather report.
enum:
- celsius
- fahrenheit
default: fahrenheit
- id: tool_bfcl_live_irrelevance_14-2-2
category: tool_use_agentic
kind: tool
prompt: Please help me get the votes associated with the IP of http://digdeep.io.
expect_tool: null
tools:
- type: function
function:
name: get_current_weather
description: Retrieves the current weather conditions for a specified city and state.
parameters:
type: object
required:
- location
properties:
location:
type: string
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
exist.
unit:
type: string
description: The unit of temperature for the weather report.
enum:
- celsius
- fahrenheit
default: fahrenheit
- id: tool_bfcl_live_irrelevance_15-2-3
category: tool_use_agentic
kind: tool
prompt: Using 'api_key_2', retrieve the IDs of graphs containing IP 145.34.45.56 on VirusTotal. Don't forget to set the cursor as 'cursor_b' and limit the results to 8.
expect_tool: null
tools:
- type: function
function:
name: get_current_weather
description: Retrieves the current weather conditions for a specified city and state.
parameters:
type: object
required:
- location
properties:
location:
type: string
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
exist.
unit:
type: string
description: The unit of temperature for the weather report.
enum:
- celsius
- fahrenheit
default: fahrenheit
- id: tool_bfcl_live_irrelevance_16-2-4
category: tool_use_agentic
kind: tool
prompt: 'How do I pull the domain info of twitter.com from VirusTotal? Using this API key: twt_key_abc.'
expect_tool: null
tools:
- type: function
function:
name: get_current_weather
description: Retrieves the current weather conditions for a specified city and state.
parameters:
type: object
required:
- location
properties:
location:
type: string
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
exist.
unit:
type: string
description: The unit of temperature for the weather report.
enum:
- celsius
- fahrenheit
default: fahrenheit
- id: tool_bfcl_live_simple_0-0-0
category: tool_use_agentic
kind: tool
prompt: Can you retrieve the details for the user with the ID 7890, who has black as their special request?
expect_tool: get_user_info
tools:
- type: function
function:
name: get_user_info
description: Retrieve details for a specific user by their unique identifier.
parameters:
type: object
required:
- user_id
properties:
user_id:
type: integer
description: The unique identifier of the user. It is used to fetch the specific user details from the database.
special:
type: string
description: Any special information or parameters that need to be considered while fetching user details.
default: none
expect_args:
user_id: 7890
special: black
- id: tool_bfcl_live_simple_1-1-0
category: tool_use_agentic
kind: tool
prompt: I want to see the star history of ShishirPatil/gorilla and gorilla-llm/gorilla-cli, with the timelines aligned, so that I can more clearly observe the rate of change from their initial releases.
expect_tool: github_star
tools:
- type: function
function:
name: github_star
description: Generates a URL for tracking the star history of specified GitHub repositories, with the option to align them on the same timeline.
parameters:
type: object
required:
- repos
properties:
repos:
type: string
description: A comma-separated list of GitHub repositories to track, each in the 'owner/repo' format, such as 'octocat/Hello-World,octo-org/octo-repo'.
aligned:
type: boolean
description: Whether to align the repositories on the same timeline for comparison. If true, the star history of all repositories will start from the same point.
default: false
expect_args:
repos: ShishirPatil/gorilla,gorilla-llm/gorilla-cli
aligned: true
- id: tool_bfcl_live_simple_4-3-0
category: tool_use_agentic
kind: tool
prompt: What are the current weather conditions in Tel Aviv, and could you provide that in Fahrenheit, please?
expect_tool: get_current_weather
tools:
- type: function
function:
name: get_current_weather
description: Retrieves the current weather conditions for a specified city and state. If using state, then use short form like CA.
parameters:
type: object
required:
- location
properties:
location:
type: string
description: The location for which to get the weather, in the format of 'City, State (abbr)', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city
doesn't exist.
unit:
type: string
description: The unit of temperature for the weather report.
enum:
- celsius
- fahrenheit
default: fahrenheit
expect_args:
location: Tel Aviv, Israel
unit: fahrenheit
# --- docs_writing ------------------------------------------------------- # --- docs_writing -------------------------------------------------------
- id: docs_function - id: docs_function
category: docs_writing category: docs_writing

View File

@@ -0,0 +1,280 @@
# Spec: sourcing harder self-eval tasks from external benchmarks
**Origin.** `PYTHONPATH=src python -m eval_proficiency` was re-run in full this
session (11 models x 23 tasks, 216 billed calls, $0.79 / 0.098 kWh, 0 call
failures) specifically to see whether more samples of the *existing* task set
would break the ties CLAUDE.md already flagged as suspicious. Result, doubling
n from 2-3 to 4-6 per model per category:
- `tool_use_agentic`: `deepseek-v4-flash` held at exactly 2/6, the same ratio
as the earlier 1/3 — real signal, not n=3 noise, but still short of
`self_eval_min_samples: 10`.
- `coding_general`, `coding_refactor`, `debugging`: **still perfectly flat at
1.00** for every model. This is no longer "not enough samples" — it's
"these 3 tasks per category don't discriminate anything, at any sample
count." More passes over the same 23 tasks won't fix that; the task set
itself needs to get harder.
This spec is the result of "where do we get harder tasks without inventing
plausible-looking difficulty from nothing" — the same objection this project
already raised against `leaderboards.yaml`'s empty priors. No code changes
accompany this document — this is the spec opencode builds from.
---
## Non-goals, stated up front
- **No new categories.** `proficiency.categories` in `config.py` is read
throughout scoring/tiering/routing; adding one is a real migration, not a
task-authoring change, and nothing here needs it. Every source below gets
mapped onto one of the 9 existing categories.
- **No new scoring kinds.** Everything maps onto `code`, `exact`, or `tool`.
(`judge` categories — `docs_writing`, `general_chat`, `summarization`,
`translation`, `general_chat` — aren't the problem this session found; skip
them for this pass.)
- **No repo-context benchmarks.** This is the reason SWE-bench (and anything
else shaped like "here's a git checkout, go fix an issue") is not on the
list below despite being the best-known real-bug-fixing benchmark. Every
task in `evals/tasks.yaml` is one self-contained prompt scored by one
subprocess call or one structural check — the harness has no notion of a
repo checkout, and building one is a much bigger lift than the actual
problem (three flat categories) justifies.
- **No verbatim copying.** Every source below gets *adapted* — either
translated into this project's YAML+checks shape, or (for
`coding_refactor`/`debugging`) used only to identify which underlying
problems are hard enough to be worth hand-authoring a variant of. Nothing
here is "download file, paste into tasks.yaml."
## A code change this needs, before any task uses it
`score_exact` (`eval_proficiency.py:191`) normalizes by regexing out the
**last number** in the reply (`normalize_answer`). That's correct for this
project's current `exact` tasks — they're all arithmetic word problems with a
scalar numeric answer (`math_percent_trap`, `math_rate_trap`,
`math_counting`). It is **not** correct for CRUXEval-style tasks (below),
where the expected value is a Python literal that can be a list, dict, tuple,
string, or `None` — e.g. an answer of `[1, 2, 3]` would get reduced to `3`,
silently discarding the structure and scoring against the wrong thing.
Fix, scoped and small: in `normalize_answer`, try `ast.literal_eval` on the
stripped text first (covers numbers, strings, lists, dicts, tuples, bools,
`None` uniformly); fall back to the existing regex-last-number heuristic only
on a `ValueError`/`SyntaxError`, which is exactly what a numeric word-problem
answer with prose around it produces. This is additive — every existing
`exact` task still gets scored the same way, since `literal_eval` will fail on
"3 minutes" and fall through to the current path unchanged. Needs its own
test alongside the existing `test_counting_answer_matches_brute_force`-style
ones (`tests/test_task_set.py`) before any CRUXEval-derived task is added, per
the project's own rule that a check the reference solution can't pass "scores
the task set rather than the model."
## `tool_use_agentic` ← BFCL live_irrelevance
**Source:** `BFCL_v4_live_irrelevance.json` in
[`ShishirPatil/gorilla`](https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard/bfcl_eval/data)
(Apache 2.0). 884 examples, sourced from real user queries rather than
synthesized — pull this file over the smaller 239-example synthetic
`BFCL_v4_irrelevance.json` for the same reason this project trusts
`POST /outcome` traffic over benchmark scores: real queries beat invented
ones. This is the single best-matched source of the four — it's a large pool
of *exactly* the pattern this project already hand-writes 2 of 3 tasks in:
offer a tool, correct behavior is not calling it.
**Schema gap to close** (confirmed by fetching a live sample): BFCL's own
function schema is not the OpenAI tool-call schema this project uses —
`"type": "dict"` where OpenAI/this project's YAML says `"type": "object"`, no
`{"type": "function", "function": {...}}` wrapper, and `question` is a list of
turns (list of message dicts) rather than this project's flat `prompt`
string. Translation per selected example:
```python
tools = [
{"type": "function", "function": {
"name": fn["name"],
"description": fn["description"],
"parameters": {**fn["parameters"], "type": "object"}, # dict -> object
}}
for fn in example["function"]
]
prompt = example["question"][0][0]["content"] # single-turn BFCL cases only
```
**The balance trap.** Every BFCL *irrelevance* example is, by definition, a
"correctly abstain" case. If the new batch is 100% abstain-cases, a model
that never calls a tool scores 1.0 on the enlarged category despite being
just as broken as one that always calls a tool — the category would stop
measuring over-triggering and start measuring only under-triggering.
`score_tool` already has both directions (`expect_tool: null` vs. a named
tool), and the current 3-task set is already 1:2 (call : abstain). Keep new
additions roughly in that band or better: pull irrelevance examples for the
new abstain cases, and pull a matching number of **positive** call cases from
BFCL's `live_simple`/`live_multiple` files (same repo, same schema, translate
the same way) so the enlarged category still rewards correct discrimination,
not just caution.
**Worked example**, adapted from a real fetched entry:
```yaml
- id: tool_geocode_irrelevant
category: tool_use_agentic
kind: tool
prompt: |
Can you provide the address for latitude 37.4224764 and longitude
-122.0842499 using the Geocoding API?
expect_tool: null # no offered function does geocoding
tools:
- type: function
function:
name: requests.get
description: Sends a GET request to the specified URL to retrieve data.
parameters:
type: object
properties:
url: {type: string, description: The URL to send the GET request to.}
required: [url]
```
## `coding_general` ← CRUXEval-O
**Source:** [`facebookresearch/cruxeval`](https://github.com/facebookresearch/cruxeval)
(MIT). 800 short Python functions with input/output pairs, split into
CRUXEval-I (predict an input that produces a given output) and CRUXEval-O
(predict the output of running the function on a given input). **Use O only.**
I is a genuine structural mismatch: an input that produces a target output is
often not unique, so scoring it needs to actually execute the candidate input
through the real function and compare — that's neither `exact` (no single
golden string to normalize against) nor `code` (the model isn't writing a
function) as this harness defines them. Not worth a third scoring kind for
half of one source; skip I, use O, which is a clean fit for `exact` once the
`ast.literal_eval` fix above lands.
At release, best-of-field models scored 63-67% on CRUXEval-O — meaning it
was built to sit in a discriminating band, not to saturate the way this
project's own three `coding_general` tasks did.
**Worked example** (illustrative shape, not a literal fetched row):
```yaml
- id: cruxeval_trace_dedup
category: coding_general
kind: exact
prompt: |
What does this function return when called as shown? Reply with ONLY
the literal Python value — no explanation, no fences.
def f(lst):
seen = {}
out = []
for x in lst:
if x not in seen:
seen[x] = True
out.append(x * 2)
return out
f([3, 1, 3, 2, 1])
answer: "[6, 2, 4]"
```
## `coding_refactor` / `debugging` ← Exercism, via Aider's selection method
**What Aider's polyglot-benchmark actually gives you.** I checked: the
[`Aider-AI/polyglot-benchmark`](https://github.com/Aider-AI/polyglot-benchmark)
repo's own GitHub license field comes back blank via the API — don't pull
files from it directly. What's genuinely reusable is the *method* their
[Dec 2024 writeup](https://aider.chat/2024/12/21/polyglot.html) describes:
225 of Exercism's 697 exercises, filtered to ones **solved by 3 or fewer of 7
frontier models**, deliberately excluding the 258 "solved by all 7" as too
easy. They didn't publish the 225 names. The underlying exercises are
Exercism's own tracks, confirmed MIT-licensed
(`exercism/python`, etc.) — go there directly, using Aider's selection
criterion (not their file) as the bar: prefer exercises in Exercism's own
`practice/` directories with a track difficulty/reputation for edge-case
traps. `bowling`, `dominoes`, and `affine-cipher` (all present in the
`exercism/python` `practice/` tree, confirmed via the GitHub API) are
reasonable starting picks — each is well known for one nasty, non-obvious
edge case rather than raw difficulty, which is exactly this project's own
stated design principle ("at least one edge case a plausible-looking solution
gets wrong").
**Why this can't be mechanical, unlike the other two sources.** Exercism ships
a spec + a test suite, not a "here's buggy/ugly code, fix/refactor it"
prompt — that shape doesn't exist upstream. `coding_refactor` and `debugging`
tasks in this project embed the *broken or ugly-but-working* code directly in
the prompt (confirmed from `apply_settings`, `first_match`, `describe` in the
current set) and the checks pin exact edge-case behavior. So the real value
of going to Exercism isn't a copyable prompt, it's **borrowing the hard edge
case**: read the exercise's canonical solution and its test suite (that's the
whole reason the exercise is hard), then hand-author a variant:
- **`coding_refactor` variant:** take the canonical solution, deliberately
restructure it into duplicated/nested-conditional/flag-variable style (same
pattern as the existing three) while preserving every edge case exactly —
the model refactors it back, checks re-run the exercise's own tricky test
cases.
- **`debugging` variant:** take the canonical solution, introduce one subtle
off-by-one/boundary bug that breaks exactly the input the exercise is
famous for tripping people up on (bowling's 10th-frame strike/spare bonus
scoring is the textbook example), leave everything else correct — checks
must include that exact edge case, per `tests/test_task_set.py`'s existing
requirement that every debugging target actually fails pre-fix.
This is real authoring work, not translation — budget it accordingly (see
below). What Exercism/Aider buys is **which** edge cases are worth encoding,
pre-validated by the fact that most frontier models already fail them, rather
than inventing new traps from scratch the way the original 23-task set's
"touching intervals / full semver / late-binding closures" hardening pass
did.
## License / attribution handling
Nothing here is substantial verbatim reproduction, so full license-file
vendoring isn't needed — but this project already has a citation precedent
(`session-classification-cache-ttl.md`'s LLMRouter attribution, and
`PinchConfig`'s docstring crediting llmrouter directly in code). Match it: a
one-line comment above each newly-added task block in `evals/tasks.yaml`
naming the source and license, e.g.:
```yaml
# --- tool_use_agentic (added) — adapted from BFCL live_irrelevance /
# live_simple, Apache 2.0, github.com/ShishirPatil/gorilla ---
```
## Validation obligation (not optional)
`tests/test_task_set.py` requires, per kind:
- `code`: a reference solution that passes every check.
- `exact`: an independent (brute-force or hand-derived) recomputation of the
answer — see `test_counting_answer_matches_brute_force` for the pattern.
- `debugging`: the buggy target must **fail** its own checks pre-fix.
- `coding_refactor`: the ugly target must **pass** its own checks pre-refactor.
This is what already caught one wrong expected-value check in the existing
set (CLAUDE.md: "would have docked every model on a task and been
indistinguishable from genuine difficulty"). Every task added from this spec
needs its matching test — this roughly doubles the authoring cost per task
versus just writing the YAML, and should be budgeted as such rather than
treated as a follow-up.
## Suggested first-pass batch size
This session's full run (23 tasks x 11 models, some judge calls) was 216
billed calls / $0.79. A reasonable first batch — **8 new `tool_use_agentic`
tasks (5 abstain + 3 positive-call), 6 new `coding_general` (CRUXEval-O), 3
new `coding_refactor`, 3 new `debugging`** — is +20 tasks, roughly matching
the size of the existing set. Expect a proportional cost per pass (~$1.50-2),
and **two passes** (not one) to clear `self_eval_min_samples: 10` for the
newly-added tasks, since one pass only gives n=1 per new task per model.
## Recommendation
Build in this order: (1) the `score_exact`/`ast.literal_eval` fix plus its
test, since CRUXEval-O tasks are silently wrong without it; (2) BFCL-derived
`tool_use_agentic` tasks, since that category already has the strongest
reproduced signal from this session's run and BFCL requires no hand-authoring
beyond schema translation; (3) CRUXEval-O `coding_general` tasks; (4) the
Exercism-sourced `coding_refactor`/`debugging` variants last, since they're
the only genuinely bespoke authoring work and benefit from the other three
being done first as a format reference. Run `eval_proficiency` twice after
each category lands rather than once at the end, so a broken check (per the
validation section above) is caught against one category's blast radius, not
all four at once.

View File

@@ -28,6 +28,7 @@ functions, so nothing invites filesystem or network use, but treat this as
from __future__ import annotations from __future__ import annotations
import argparse import argparse
import ast
import json import json
import os import os
import re import re
@@ -78,6 +79,11 @@ def raises(exc, fn, *a, **kw):
FENCE_RE = re.compile(r"^\s*```[a-zA-Z]*\n(.*?)```", re.DOTALL | re.MULTILINE) FENCE_RE = re.compile(r"^\s*```[a-zA-Z]*\n(.*?)```", re.DOTALL | re.MULTILINE)
# A standalone number with comma thousands separators (e.g. "3,000"). Matched
# on the WHOLE stripped string so "3,000" is treated as the number 3000 rather
# than literal_eval'ing to the tuple (3, 0).
BARE_COMMA_NUMBER_RE = re.compile(r"-?\d[\d,]*(?:\.\d+)?")
# --- extraction and scoring ---------------------------------------------- # --- extraction and scoring ----------------------------------------------
@@ -174,26 +180,85 @@ def score_code(model_output: str, checks: list[str]) -> tuple[float, str]:
return passed / len(checks), f"{passed}/{len(checks)} checks" return passed / len(checks), f"{passed}/{len(checks)} checks"
_LITERAL_WORD = {"true": "True", "false": "False", "none": "None"}
def normalize_answer(text: str) -> str: def normalize_answer(text: str) -> str:
"""Reduce a free-text reply to a comparable token. """Reduce a free-text reply to a comparable token.
Models wrap a number in prose, punctuation, or thousands separators even Structured answers (lists, tuples, dicts, bools, None, quoted strings)
when told not to; none of that is the thing being measured. parse as Python literals first, so a model told to reply with a literal
""" value is scored against the value, not its string form. Prose-wrapped
cleaned = (text or "").strip().replace(",", "") numbers fall back to the last number in the reply, with thousands
numbers = re.findall(r"-?\d+(?:\.\d+)?", cleaned) separators handled, so "the combinations is 4,536." still matches "4536".
if numbers:
value = numbers[-1] # the conclusion, if it reasoned out loud
return value.rstrip("0").rstrip(".") if "." in value else value
return cleaned.lower().strip(" .!\"'")
Markdown fences are honored only when they wrap the ENTIRE reply; a fence
followed by prose is treated as prose, because the concluding prose value
is what is being measured.
"""
stripped = (text or "").strip()
# Honor a whole-reply single fenced block; never discard trailing prose.
if FENCE_RE.fullmatch(stripped):
match = FENCE_RE.search(stripped)
if match:
stripped = match.group(1).strip()
# Collapse thousands-separator commas BEFORE literal_eval. Without this,
# ast.literal_eval("[3,000, 4,000]") silently misparses to [3, 0, 4, 0]
# because Python treats "3,000" inside a list as "3" followed by leading-
# zero "000". The look-behind / look-ahead ensure only groups of exactly
# three digits after a comma are collapsed, so "[12,34]" and tuples are
# left untouched.
# KNOWN RESIDUAL AMBIGUITY (accepted, not chased): "[100,200,300]" (three
# adjacent 3-digit list elements, no spaces) is character-identical to a
# thousands-grouped number, so this collapse merges it to "[100200300]".
# It cannot be disambiguated from the raw text alone. No current task's
# answer contains such text, so this is latent; do not add a fourth comma
# heuristic to guess structural intent — regex composition over raw text
# just relocates this ambiguity instead of resolving it.
collapsed = re.sub(r"(?<=\d),(\d{3})(?!\d)", r"\1", stripped)
# A bare comma-thousands number like "3,000" would otherwise literal_eval
# to the tuple (3, 0). (With the collapse above, this branch is now
# mostly unreachable for comma-numbers, but the regex stays for safety.)
value: object
try:
value = ast.literal_eval(collapsed)
except (ValueError, SyntaxError):
pass
else:
if isinstance(value, float):
if value.is_integer():
return str(int(value))
return str(value)
return repr(value)
# Prose fallback: the concluding value is the last number, with thousands
# separators stripped - but never inside bracket-wrapped structured text,
# where commas are element separators, not separators to remove.
if collapsed and not any(c in collapsed for c in "[](){}"):
cleaned = collapsed.replace(",", "")
numbers = re.findall(r"-?\d+(?:\.\d+)?", cleaned)
if numbers:
candidate = numbers[-1]
return (
candidate.rstrip("0").rstrip(".")
if "." in candidate
else candidate
)
# Value keywords ("False.", "true.") match their literal repr
# case-insensitively, so "False." still equals the answer "False".
cleaned = collapsed.replace(",", "").lower().strip(" .!\"'")
return _LITERAL_WORD.get(cleaned, cleaned)
def score_exact(model_output: str, expected: str) -> tuple[float, str]: def score_exact(model_output: str, expected: str) -> tuple[float, str]:
got = normalize_answer(model_output) got = normalize_answer(model_output)
want = normalize_answer(expected) want = normalize_answer(expected)
return (1.0, f"{got!r}") if got == want else (0.0, f"got {got!r} want {want!r}") return (1.0, f"{got!r}") if got == want else (0.0, f"got {got!r} want {want!r}")
def score_tool(tool_calls: list, task: dict) -> tuple[float, str]: def score_tool(tool_calls: list, task: dict) -> tuple[float, str]:
"""Score a tool-use task structurally — no judge required. """Score a tool-use task structurally — no judge required.

View File

@@ -11,12 +11,14 @@ These tests run offline and take milliseconds — they are the cheap guard
against spending an hour of API calls measuring a typo. against spending an hour of API calls measuring a typo.
""" """
import re
import textwrap
from pathlib import Path from pathlib import Path
import pytest import pytest
import yaml import yaml
from eval_proficiency import score_code, score_exact from eval_proficiency import score_code, score_exact, score_tool
TASKS = yaml.safe_load((Path(__file__).resolve().parent.parent / "evals" / "tasks.yaml").read_text())["tasks"] TASKS = yaml.safe_load((Path(__file__).resolve().parent.parent / "evals" / "tasks.yaml").read_text())["tasks"]
BY_ID = {t["id"]: t for t in TASKS} BY_ID = {t["id"]: t for t in TASKS}
@@ -92,6 +94,327 @@ _TABLE = {200: "ok", 201: "created", 404: "not found", 500: "server error"}
def describe(code): def describe(code):
return _TABLE.get(code, "unknown") return _TABLE.get(code, "unknown")
''',
"refactor_bowling_frames": '''
class Frame:
def __init__(self, idx):
self.idx = idx
self.throws = []
@property
def total_pins(self):
return sum(self.throws)
def is_strike(self):
return self.total_pins == 10 and len(self.throws) == 1
def is_spare(self):
return self.total_pins == 10 and len(self.throws) == 2
def is_open(self):
return self.total_pins < 10 and len(self.throws) == 2
def is_closed(self):
return self.total_pins == 10 or len(self.throws) == 2
def throw(self, pins):
if self.total_pins + pins > 10:
raise ValueError("a frame's rolls cannot exceed 10")
self.throws.append(pins)
def score(self, next_throws):
result = self.total_pins
if self.is_strike():
result += sum(next_throws[:2])
elif self.is_spare():
result += sum(next_throws[:1])
return result
class BowlingGame:
def __init__(self):
self.current_frame_idx = 0
self.bonus_throws = []
self.frames = [Frame(idx) for idx in range(10)]
@property
def current_frame(self):
return self.frames[self.current_frame_idx]
def next_throws(self, frame_idx):
throws = []
for idx in range(frame_idx + 1, 10):
throws.extend(self.frames[idx].throws)
throws.extend(self.bonus_throws)
return throws
def roll_bonus(self, pins):
tenth_frame = self.frames[-1]
if tenth_frame.is_open():
raise IndexError('cannot throw bonus with an open tenth frame')
self.bonus_throws.append(pins)
if (len(self.bonus_throws) == 2 and self.bonus_throws[0] != 10 and
sum(self.bonus_throws) > 10):
raise ValueError('invalid fill balls')
if tenth_frame.is_strike() and len(self.bonus_throws) > 2:
raise IndexError(
'wrong number of fill balls when the tenth frame is a strike')
elif tenth_frame.is_spare() and len(self.bonus_throws) > 1:
raise IndexError(
'wrong number of fill balls when the tenth frame is a spare')
def roll(self, pins):
if not 0 <= pins <= 10:
raise ValueError('invalid pins')
elif self.current_frame_idx == 10:
self.roll_bonus(pins)
else:
self.current_frame.throw(pins)
if self.current_frame.is_closed():
self.current_frame_idx += 1
def score(self):
if self.current_frame_idx < 10:
raise IndexError('frame less than 10')
if self.frames[-1].is_spare() and len(self.bonus_throws) != 1:
raise IndexError(
'one bonus must be rolled when the tenth frame is spare')
if self.frames[-1].is_strike() and len(self.bonus_throws) != 2:
raise IndexError(
'two bonuses must be rolled when the tenth frame is strike')
return sum(frame.score(self.next_throws(frame.idx))
for frame in self.frames)
''',
"debug_bowling_tenth_frame": '''
class Frame:
def __init__(self, idx):
self.idx = idx
self.throws = []
@property
def total_pins(self):
return sum(self.throws)
def is_strike(self):
return self.total_pins == 10 and len(self.throws) == 1
def is_spare(self):
return self.total_pins == 10 and len(self.throws) == 2
def is_open(self):
return self.total_pins < 10 and len(self.throws) == 2
def is_closed(self):
return self.total_pins == 10 or len(self.throws) == 2
def throw(self, pins):
if self.total_pins + pins > 10:
raise ValueError("a frame's rolls cannot exceed 10")
self.throws.append(pins)
def score(self, next_throws):
result = self.total_pins
if self.is_strike():
result += sum(next_throws[:2])
elif self.is_spare():
result += sum(next_throws[:1])
return result
class BowlingGame:
def __init__(self):
self.current_frame_idx = 0
self.bonus_throws = []
self.frames = [Frame(idx) for idx in range(10)]
@property
def current_frame(self):
return self.frames[self.current_frame_idx]
def next_throws(self, frame_idx):
throws = []
for idx in range(frame_idx + 1, 10):
throws.extend(self.frames[idx].throws)
throws.extend(self.bonus_throws)
return throws
def roll_bonus(self, pins):
tenth_frame = self.frames[-1]
if tenth_frame.is_open():
raise IndexError('cannot throw bonus with an open tenth frame')
self.bonus_throws.append(pins)
if (len(self.bonus_throws) == 2 and self.bonus_throws[0] != 10 and
sum(self.bonus_throws) > 10):
raise ValueError('invalid fill balls')
if tenth_frame.is_strike() and len(self.bonus_throws) > 2:
raise IndexError(
'wrong number of fill balls when the tenth frame is a strike')
elif tenth_frame.is_spare() and len(self.bonus_throws) > 1:
raise IndexError(
'wrong number of fill balls when the tenth frame is a spare')
def roll(self, pins):
if not 0 <= pins <= 10:
raise ValueError('invalid pins')
elif self.current_frame_idx == 10:
self.roll_bonus(pins)
else:
self.current_frame.throw(pins)
if self.current_frame.is_closed():
self.current_frame_idx += 1
def score(self):
if self.current_frame_idx < 10:
raise IndexError('frame less than 10')
if self.frames[-1].is_spare() and len(self.bonus_throws) != 1:
raise IndexError(
'one bonus must be rolled when the tenth frame is spare')
if self.frames[-1].is_strike() and len(self.bonus_throws) != 2:
raise IndexError(
'two bonuses must be rolled when the tenth frame is strike')
return sum(frame.score(self.next_throws(frame.idx))
for frame in self.frames)
''',
"debug_dominoes_no_chain": '''
from itertools import permutations
from functools import reduce
def swap(item_1, item_2):
return (item_2, item_1)
def build_chain(chain, domino):
if chain is not None:
last = chain[-1]
if len(chain) == 1 and last[0] == domino[0]:
return [swap(*last), domino]
elif len(chain) == 1 and last[0] == domino[1]:
return [swap(*last), swap(*domino)]
elif last[1] == domino[0]:
return chain + [domino]
elif last[1] == domino[1]:
return chain + [swap(*domino)]
return None
def can_chain(dominoes):
if not any(dominoes):
return []
for perm in permutations(dominoes):
chain = reduce(build_chain, perm[1:], [perm[0]])
if chain is not None and chain[0][0] == chain[-1][1]:
return chain
return None
''',
"refactor_dominoes_chain": '''
from itertools import permutations
def can_chain(dominoes):
if not any(dominoes):
return []
for perm in permutations(dominoes):
chain = [perm[0]]
complete = True
for domino in perm[1:]:
prev = chain[-1]
if len(chain) == 1 and prev[0] == domino[0]:
chain = [(prev[1], prev[0]), domino]
elif len(chain) == 1 and prev[0] == domino[1]:
chain = [(prev[1], prev[0]), (domino[1], domino[0])]
elif prev[1] == domino[0]:
chain = chain + [domino]
elif prev[1] == domino[1]:
chain = chain + [(domino[1], domino[0])]
else:
complete = False
break
if complete and chain[0][0] == chain[-1][1]:
return chain
return None
''',
"refactor_affine_encode_decode": '''
BLOCK_SIZE = 5
ALPHABET = 26
def mod_inverse(a_key, alphabet):
a_key = a_key % alphabet
for idx in range(1, alphabet):
if (a_key * idx) % alphabet == 1:
return idx
return 1
def translate(text, a_key, b_key, mode):
inverse = mod_inverse(a_key, ALPHABET)
if inverse == 1:
raise ValueError('a and m must be coprime.')
chars = []
for character in text:
if character.isalnum():
origin = ord(character.lower()) - 97
if origin < 0:
chars.append(character)
continue
if mode == 0:
new = (a_key * origin + b_key) % ALPHABET
elif mode == 1:
new = (inverse * (origin - b_key)) % ALPHABET
chars.append(chr(new + 97))
return ''.join(chars)
def encode(plain, a, b):
cipher = translate(plain, a, b, 0)
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
for idx in range(0, len(cipher), BLOCK_SIZE)])
def decode(ciphered, a, b):
return translate(ciphered, a, b, 1)
''',
"debug_affine_coprime": '''
BLOCK_SIZE = 5
ALPHABET = 26
def mod_inverse(a_key, alphabet):
a_key = a_key % alphabet
for idx in range(1, alphabet):
if (a_key * idx) % alphabet == 1:
return idx
return 1
def translate(text, a_key, b_key, mode):
inverse = mod_inverse(a_key, ALPHABET)
if inverse == 1:
raise ValueError('a and m must be coprime.')
chars = []
for character in text:
if character.isalnum():
origin = ord(character.lower()) - 97
if origin < 0:
chars.append(character)
continue
if mode == 0:
new = (a_key * origin + b_key) % ALPHABET
elif mode == 1:
new = (inverse * (origin - b_key)) % ALPHABET
chars.append(chr(new + 97))
return ''.join(chars)
def encode(plain, a, b):
cipher = translate(plain, a, b, 0)
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
for idx in range(0, len(cipher), BLOCK_SIZE)])
def decode(ciphered, a, b):
return translate(ciphered, a, b, 1)
''', ''',
"debug_late_binding": ''' "debug_late_binding": '''
def make_multipliers(factors): def make_multipliers(factors):
@@ -151,6 +474,318 @@ def first_match(items, predicates):
done = True done = True
break break
return found return found
''',
"refactor_bowling_frames": '''
class BowlingGame:
def __init__(self):
self._frames = []
self._current = 0
self._bonus = []
def roll(self, pins):
if not (0 <= pins <= 10):
raise ValueError('invalid pins')
if self._current < 10:
if len(self._frames) == self._current:
self._frames.append([pins])
else:
self._frames[self._current].append(pins)
current = self._frames[self._current]
if sum(current) > 10:
raise ValueError("a frame's rolls cannot exceed 10")
strike = (len(current) == 1 and current[0] == 10)
if strike or len(current) == 2:
self._current += 1
else:
last = self._frames[-1]
strike10 = len(last) == 1 and last[0] == 10
spare10 = len(last) == 2 and sum(last) == 10
if not (strike10 or spare10):
raise IndexError('cannot throw bonus with an open tenth frame')
if strike10:
if len(self._bonus) >= 2:
raise IndexError(
'wrong number of fill balls when the tenth frame is a strike')
self._bonus.append(pins)
if len(self._bonus) == 2 and self._bonus[0] != 10 and sum(self._bonus) > 10:
raise ValueError('invalid fill balls')
if len(self._bonus) > 2:
raise IndexError(
'wrong number of fill balls when the tenth frame is a strike')
elif spare10:
if len(self._bonus) >= 1:
raise IndexError(
'wrong number of fill balls when the tenth frame is a spare')
self._bonus.append(pins)
def score(self):
if self._current < 10:
raise IndexError('frame less than 10')
last = self._frames[-1]
if len(last) == 2 and sum(last) == 10 and len(self._bonus) != 1:
raise IndexError('one bonus must be rolled when the tenth frame is spare')
if len(last) == 1 and last[0] == 10 and len(self._bonus) != 2:
raise IndexError('two bonuses must be rolled when the tenth frame is strike')
total = 0
for i in range(10):
frame = self._frames[i]
frame_sum = sum(frame)
strike = (len(frame) == 1 and frame[0] == 10)
spare = (len(frame) == 2 and frame_sum == 10)
if strike or spare:
nxt = []
for j in range(i + 1, 10):
nxt.extend(self._frames[j])
nxt.extend(self._bonus)
if strike:
frame_sum += sum(nxt[:2])
else:
frame_sum += sum(nxt[:1])
total += frame_sum
return total
''',
"debug_bowling_tenth_frame": '''
class Frame:
def __init__(self, idx):
self.idx = idx
self.throws = []
@property
def total_pins(self):
return sum(self.throws)
def is_strike(self):
return self.total_pins == 10 and len(self.throws) == 1
def is_spare(self):
return self.total_pins == 10 and len(self.throws) == 2
def is_open(self):
return self.total_pins < 10 and len(self.throws) == 2
def is_closed(self):
return self.total_pins == 10 or len(self.throws) == 2
def throw(self, pins):
if self.total_pins + pins > 10:
raise ValueError("a frame's rolls cannot exceed 10")
self.throws.append(pins)
def score(self, next_throws):
result = self.total_pins
if self.is_strike():
result += sum(next_throws[:2])
elif self.is_spare():
result += sum(next_throws[:1])
return result
class BowlingGame:
def __init__(self):
self.current_frame_idx = 0
self.bonus_throws = []
self.frames = [Frame(idx) for idx in range(10)]
@property
def current_frame(self):
return self.frames[self.current_frame_idx]
def next_throws(self, frame_idx):
throws = []
for idx in range(frame_idx + 1, 10):
throws.extend(self.frames[idx].throws)
throws.extend(self.bonus_throws)
return throws
def roll_bonus(self, pins):
tenth_frame = self.frames[-1]
if tenth_frame.is_open():
raise IndexError('cannot throw bonus with an open tenth frame')
self.bonus_throws.append(pins)
# BUG: the invalid fill-balls guard below has been removed.
# if (len(self.bonus_throws) == 2 and self.bonus_throws[0] != 10 and
# sum(self.bonus_throws) > 10):
# raise ValueError('invalid fill balls')
if tenth_frame.is_strike() and len(self.bonus_throws) > 2:
raise IndexError(
'wrong number of fill balls when the tenth frame is a strike')
elif tenth_frame.is_spare() and len(self.bonus_throws) > 1:
raise IndexError(
'wrong number of fill balls when the tenth frame is a spare')
def roll(self, pins):
if not 0 <= pins <= 10:
raise ValueError('invalid pins')
elif self.current_frame_idx == 10:
self.roll_bonus(pins)
else:
self.current_frame.throw(pins)
if self.current_frame.is_closed():
self.current_frame_idx += 1
def score(self):
if self.current_frame_idx < 10:
raise IndexError('frame less than 10')
if self.frames[-1].is_spare() and len(self.bonus_throws) != 1:
raise IndexError(
'one bonus must be rolled when the tenth frame is spare')
if self.frames[-1].is_strike() and len(self.bonus_throws) != 2:
raise IndexError(
'two bonuses must be rolled when the tenth frame is strike')
return sum(frame.score(self.next_throws(frame.idx))
for frame in self.frames)
''',
"debug_dominoes_no_chain": '''
from itertools import permutations
from functools import reduce
def swap(item_1, item_2):
return (item_2, item_1)
def build_chain(chain, domino):
if chain is not None:
last = chain[-1]
if len(chain) == 1 and last[0] == domino[0]:
return [swap(*last), domino]
elif len(chain) == 1 and last[0] == domino[1]:
return [swap(*last), swap(*domino)]
elif last[1] == domino[0]:
return chain + [domino]
elif last[1] == domino[1]:
return chain + [swap(*domino)]
return None
def can_chain(dominoes):
if not any(dominoes):
return []
for perm in permutations(dominoes):
chain = reduce(build_chain, perm[1:], [perm[0]])
# BUG: the circular-closure check (chain[0][0] == chain[-1][1]) is
# missing, so a line that merely matches end-to-start is returned
# even when it does not close into a loop.
if chain is not None:
return chain
return None
''',
"refactor_dominoes_chain": '''
from itertools import permutations
def can_chain(dominoes):
if not any(dominoes):
return []
for perm in permutations(dominoes):
chain = [perm[0]]
complete = True
for domino in perm[1:]:
prev = chain[-1]
if len(chain) == 1 and prev[0] == domino[0]:
chain = [(prev[1], prev[0]), domino]
elif len(chain) == 1 and prev[0] == domino[1]:
chain = [(prev[1], prev[0]), (domino[1], domino[0])]
elif prev[1] == domino[0]:
chain = chain + [domino]
elif prev[1] == domino[1]:
chain = chain + [(domino[1], domino[0])]
else:
complete = False
break
if complete and chain[0][0] == chain[-1][1]:
return chain
return None
''',
"refactor_affine_encode_decode": '''
BLOCK_SIZE = 5
ALPHABET = 26
def mod_inverse(a_key, alphabet):
a_key = a_key % alphabet
for idx in range(1, alphabet):
if (a_key * idx) % alphabet == 1:
return idx
return 1
def encode(plain, a, b):
inverse = mod_inverse(a, ALPHABET)
if inverse == 1:
raise ValueError('a and m must be coprime.')
chars = []
for character in plain:
if character.isalnum():
origin = ord(character.lower()) - 97
if origin < 0:
chars.append(character)
continue
new = (a * origin + b) % ALPHABET
chars.append(chr(new + 97))
cipher = ''.join(chars)
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
for idx in range(0, len(cipher), BLOCK_SIZE)])
def decode(ciphered, a, b):
inverse = mod_inverse(a, ALPHABET)
if inverse == 1:
raise ValueError('a and m must be coprime.')
chars = []
for character in ciphered:
if character.isalnum():
origin = ord(character.lower()) - 97
if origin < 0:
chars.append(character)
continue
new = (inverse * (origin - b)) % ALPHABET
chars.append(chr(new + 97))
return ''.join(chars)
''',
"debug_affine_coprime": '''
BLOCK_SIZE = 5
ALPHABET = 26
def mod_inverse(a_key, alphabet):
a_key = a_key % alphabet
for idx in range(1, alphabet):
if (a_key * idx) % alphabet == 1:
return idx
return 1
def translate(text, a_key, b_key, mode):
inverse = mod_inverse(a_key, ALPHABET)
# BUG: compares `inverse < 1` instead of `inverse == 1`, so the coprime
# guard never fires and a key whose `a` is not coprime with 26 slips
# through without raising ValueError.
if inverse < 1:
raise ValueError('a and m must be coprime.')
chars = []
for character in text:
if character.isalnum():
origin = ord(character.lower()) - 97
if origin < 0:
chars.append(character)
continue
if mode == 0:
new = (a_key * origin + b_key) % ALPHABET
elif mode == 1:
new = (inverse * (origin - b_key)) % ALPHABET
chars.append(chr(new + 97))
return ''.join(chars)
def encode(plain, a, b):
cipher = translate(plain, a, b, 0)
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
for idx in range(0, len(cipher), BLOCK_SIZE)])
def decode(ciphered, a, b):
return translate(ciphered, a, b, 1)
''', ''',
"debug_late_binding": ''' "debug_late_binding": '''
def make_multipliers(factors): def make_multipliers(factors):
@@ -168,14 +803,15 @@ def extract_tags(text):
} }
def test_refactor_target_already_passes_its_own_checks(): @pytest.mark.parametrize("task_id", [tid for tid in sorted(BUGGY) if tid.startswith("refactor_")])
def test_refactor_target_already_passes_its_own_checks(task_id):
# A refactor task's ORIGINAL code must pass, or the task is secretly a # A refactor task's ORIGINAL code must pass, or the task is secretly a
# debugging task and "behaviour must not change" is a lie # debugging task and "behaviour must not change" is a lie
score, detail = score_code(BUGGY["refactor_first_match"], BY_ID["refactor_first_match"]["checks"]) score, detail = score_code(BUGGY[task_id], BY_ID[task_id]["checks"])
assert score == 1.0, f"refactor target fails its own checks: {detail}" assert score == 1.0, f"{task_id}: refactor target fails its own checks: {detail}"
@pytest.mark.parametrize("task_id", ["debug_late_binding", "debug_greedy_regex"]) @pytest.mark.parametrize("task_id", ["debug_late_binding", "debug_greedy_regex", "debug_bowling_tenth_frame", "debug_dominoes_no_chain", "debug_affine_coprime"])
def test_debugging_tasks_start_broken(task_id): def test_debugging_tasks_start_broken(task_id):
# The whole point is that the given code fails. If it passes, the task # The whole point is that the given code fails. If it passes, the task
# measures nothing — a model could return the input unchanged. # measures nothing — a model could return the input unchanged.
@@ -210,6 +846,160 @@ def test_counting_answer_matches_brute_force():
assert str(count) == BY_ID["math_counting"]["answer"] assert str(count) == BY_ID["math_counting"]["answer"]
# --- exact answers as Python literals -------------------------------------
def test_exact_scalar_does_not_match_list():
# A scalar number answer must never cross-match a list literal (nor the
# reverse): the model said "3", the task wanted "[1, 2, 3]",
# and that is a wrong answer, not a formatting variant.
assert score_exact("3", "[1, 2, 3]")[0] == 0.0
assert score_exact("[1, 2, 3]", "3")[0] == 0.0
@pytest.mark.parametrize(
"literal",
["[1, 2, 3]", "{'a': 1}", "(4, 5)", "None", "True", "'text'", "3.5", "-12"],
)
def test_exact_structured_answers_match_their_literal(literal):
# Structured/literal answers match themselves: lists, dicts, tuples,
# None, bools, quoted strings, floats.
assert score_exact(literal, literal)[0] == 1.0
def test_exact_fenced_literal():
# A reply wrapped in ``` fences is just formatting, not a different answer.
assert score_exact("```\n[1, 2, 3]\n```", "[1, 2, 3]")[0] == 1.0
def test_exact_comma_number_is_not_a_tuple():
# "3,000" is a comma-separated thousands number, not the tuple (3, 0).
assert score_exact("3,000", "3000")[0] == 1.0
assert score_exact("1,000,000", "1000000")[0] == 1.0
def test_exact_prose_number_still_falls_back():
# Prose-wrapped numbers keep matching the bare number.
assert score_exact("The answer is 3 minutes.", "3")[0] == 1.0
assert score_exact("there are 100 widgets", "100")[0] == 1.0
def test_exact_float_and_int_agree():
# Numerically-equal floats/ints and currency strings stay equivalent.
assert score_exact("3.0", "3")[0] == 1.0
assert score_exact("3.50", "3.5")[0] == 1.0
assert score_exact("$100", "100")[0] == 1.0
def test_exact_bool_does_not_cross_match_number():
# True is not the number 1, in either order.
assert score_exact("True", "1")[0] == 0.0
assert score_exact("1", "True")[0] == 0.0
# --- CRUXEval-O exact answers (recomputed) --------------------------------
# Mapping: task id → (code, input) where the model is asked what
# `f(<input>)` returns. The answer is verified by `eval("f(" + input + ")")`
# NOT `literal_eval(input)` because CRUXEval-O input is the source arg list,
# not a single literal.
_CRUX_ROWS = {
"crux_o_sample_9": (
textwrap.dedent("""\
def f(t):
for c in t:
if not c.isnumeric():
return False
return True
""").strip(),
"'#284376598'",
),
"crux_o_sample_0": (
textwrap.dedent("""\
def f(nums):
output = []
for n in nums:
output.append((nums.count(n), n))
output.sort(reverse=True)
return output
""").strip(),
"[1, 1, 3, 1, 3, 1]",
),
"crux_o_sample_1": (
textwrap.dedent("""\
def f(a, b, c):
result = {}
for d in a, b, c:
result.update(dict.fromkeys(d))
return result
""").strip(),
"(1, ), (1, ), (1, 2)",
),
"crux_o_sample_2": (
textwrap.dedent("""\
def f(text):
new_text = list(text)
for i in '+':
if i in new_text:
new_text.remove(i)
return ''.join(new_text)
""").strip(),
"'hbtofdeiequ'",
),
"crux_o_sample_5": (
textwrap.dedent("""\
def f(text, lower, upper):
count = 0
new_text = list()
for char in text:
char = lower if char.isdecimal() else upper
if char in ['p', 'C']:
count += 1
new_text.append(char)
return count, ''.join(new_text)
""").strip(),
"'DSUWeqExTQdCMGpqur', 'a', 'x'",
),
"crux_o_sample_6": (
textwrap.dedent("""\
def f(dic):
for k,v in sorted(dic.items(), key=lambda x: len(str(x)))[:-1]:
dic.pop(k)
return list(dic.items())
""").strip(),
"{'11': 52, '65': 34, 'a': 12, '4': 52, '74': 31}",
),
}
def _crux_execute(code, input_str):
"""Execute `f(<input>)` and return `repr(result)` matching the YAML answer."""
ns = {}
exec(code, ns) # noqa: S102
result = eval("f(" + input_str + ")", ns)
return repr(result)
@pytest.mark.parametrize("task_id", sorted(_CRUX_ROWS))
def test_crux_answer_matches_execution(task_id):
code, input_str = _CRUX_ROWS[task_id]
expected = _crux_execute(code, input_str)
actual = BY_ID[task_id]["answer"]
assert actual == expected, (
f"{task_id}: yaml answer={actual!r} != execution repr={expected!r} "
f"(eval(\"f({input_str})\"))"
)
def test_crux_prompts_request_literal_only():
for task_id in _CRUX_ROWS:
task = BY_ID[task_id]
prompt = task["prompt"].lower()
assert "literal" in prompt or "only" in prompt, (
f"{task_id}: prompt does not restrict to literal-only reply"
)
# --- task set hygiene ----------------------------------------------------- # --- task set hygiene -----------------------------------------------------
def test_every_task_has_the_fields_its_kind_needs(): def test_every_task_has_the_fields_its_kind_needs():
@@ -230,3 +1020,130 @@ def test_every_task_has_the_fields_its_kind_needs():
def test_task_ids_are_unique(): def test_task_ids_are_unique():
ids = [t["id"] for t in TASKS] ids = [t["id"] for t in TASKS]
assert len(ids) == len(set(ids)) assert len(ids) == len(set(ids))
# --- BFCL-derived tool tasks ----------------------------------------------
_TOOL_NAME_RE = re.compile(r"^[A-Za-z0-9_-]{1,64}$")
def _bfcl_positive():
return [t for t in TASKS if t["id"].startswith("tool_bfcl_") and t.get("expect_tool")]
def _bfcl_abstain():
return [
t
for t in TASKS
if t["id"].startswith("tool_bfcl_") and t.get("expect_tool") is None
]
@pytest.mark.parametrize("task_id", [t["id"] for t in TASKS if t["id"].startswith("tool_bfcl_")])
def test_bfcl_tools_are_openai_shaped(task_id):
# The provider receives `tools` verbatim at request time; every BFCL tool
# must already be in OpenAI shape (type:function wrapper, object params).
task = BY_ID[task_id]
for tool in task["tools"]:
assert tool["type"] == "function"
fn = tool["function"]
assert fn["parameters"]["type"] == "object"
@pytest.mark.parametrize("task_id", [t["id"] for t in TASKS if t["id"].startswith("tool_bfcl_")])
def test_bfcl_expect_args_are_scalar(task_id):
# Ground-truth args that survive translation must be scalars (str/int/
# float/bool). An array/object arg would have been skipped at translation,
# so a container here means the selection filter regressed.
task = BY_ID[task_id]
for arg, value in (task.get("expect_args") or {}).items():
assert isinstance(value, (str, int, float, bool)), (
f"{task_id}: expect_args[{arg!r}] is not scalar"
)
@pytest.mark.parametrize("task_id", [t["id"] for t in TASKS if t["kind"] == "tool"])
def test_all_tool_names_are_provider_safe(task_id):
# OpenAI tool names only allow A-Za-z0-9_- (1-64 chars). A dotted BFCL name
# like uber.ride/aws.* would be rejected by the provider — the selection
# filter must have elided those rows.
task = BY_ID[task_id]
for tool in task["tools"]:
name = tool["function"]["name"]
assert _TOOL_NAME_RE.fullmatch(name), f"{task_id}: unsafe tool name {name!r}"
def test_bfcl_positive_scores_correct_call():
import json
for task in _bfcl_positive():
assert task["expect_args"], f"{task['id']} has no expect_args"
args = json.dumps(task["expect_args"])
call = [{"function": {"name": task["expect_tool"], "arguments": args}}]
score, detail = score_tool(call, task)
assert score == 1.0, f"{task['id']}: correct call scored {score} ({detail})"
# A wrong tool name must not score: the task discriminates the call.
wrong = [{"function": {"name": "some_other_tool", "arguments": args}}]
assert score_tool(wrong, task)[0] == 0.0, f"{task['id']}: wrong tool scored"
def test_bfcl_irrelevant_scores_abstention():
# The 5 BFCL abstain tasks pair a weather tool with a VirusTotal/DNS/CPU
# question. A model that reaches for the offered tool must score 0.0;
# abstaining must score 1.0.
for task in _bfcl_abstain():
assert score_tool([], task)[0] == 1.0, f"{task['id']} did not reward abstention"
for tool in task["tools"]:
name = tool["function"]["name"]
call = [{"function": {"name": name, "arguments": "{}"}}]
assert score_tool(call, task)[0] == 0.0, f"{task['id']}: calling {name} scored"
# --- EV-01: Regression tests for normalize_answer / score_exact correctness ---
def test_repro_math_counting_comma_prose():
# Bug 1: "the combinations is 4,536." must match answer "4536"
assert score_exact("the combinations is 4,536.", "4536")[0] == 1.0
# Standalone comma-number must also match
assert score_exact("4,536", "4536")[0] == 1.0
def test_repro_fence_with_trailing_prose():
# Bug 2: Trailing prose after a code fence must NOT be discarded
assert score_exact("```python\nx = 5 * 4\n```\nThe answer is 20.", "20")[0] == 1.0
def test_repro_fence_only_still_literal():
# Whole-reply fenced block should still extract the literal
assert score_exact("```\n[1, 2, 3]\n```", "[1, 2, 3]")[0] == 1.0
def test_repro_bool_punctuated_matches():
# Bug 3: "False." must normalize to "False" (literal repr)
assert score_exact("False.", "False")[0] == 1.0
assert score_exact("True.", "True")[0] == 1.0
def test_repro_none_matches_none_dot():
# "None." prose should normalize to "None" literal repr
assert score_exact("None.", "None")[0] == 1.0
def test_repro_nested_comma_thousands_match():
# Bug 4: comma-thousands inside a list must collapse so a model's
# [3,000, 4,000] matches the true answer [3000, 4000], not misparse
# to [3, 0, 4, 0].
from eval_proficiency import normalize_answer
# Core bug: thousands-separator commas inside a list must collapse
assert normalize_answer("[3,000, 4,000]") == "[3000, 4000]"
assert score_exact("[3,000, 4,000]", "[3000, 4000]")[0] == 1.0
# Standalone comma-number still works via collapse -> literal_eval
assert normalize_answer("3,000") == "3000"
assert score_exact("total is 4,536.", "4536")[0] == 1.0
# Non-thousands commas must NOT be collapsed
assert normalize_answer("[12,34]") == "[12, 34]" # repr-style spacing
assert normalize_answer("[abc, def]") == "[abc def]"
assert normalize_answer("[(1, 2), (1, 2)]") == "[(1, 2), (1, 2)]"