Three defects, all found by actually running the local-dispatch runbook for
the first time. None could have been caught offline.
1. diff_checking was logically unanswerable (evals/tasks.yaml)
All four safe/buggy pairs are exact MIRROR IMAGES: BEFORE_safe ==
AFTER_buggy and AFTER_safe == BEFORE_buggy. The question asked "Does the
AFTER version change behaviour for any valid input?" -- which is SYMMETRIC:
if A->B changes behaviour, so does B->A. But the pairs carry OPPOSITE
labels, so four of the eight tasks were wrong no matter what any model
answered.
Measured before the fix: the maximum achievable score was 0.50, and a model
that blindly answered "no" also scored 0.50. deepseek-v4-flash, which scores
1.00 on all three coding categories, got 0.25 -- punished for engaging with
the question. After the fix it scores 0.625 and nemotron-mini:4b's true
profile is visible (0/4 bugs detected).
The question is now antisymmetric ("does the AFTER version introduce a bug
that the BEFORE version does not have?"), which is what opposite labels
require. tests/test_task_set.py grows a regression test that pins the
invariant; it was verified to FAIL against the old phrasing, not merely to
pass against the new one.
This is the fifth harness bug in this project that scored the rig rather
than the model.
2. seed_local_dispatch_energy closed its DB connection mid-run
main() closed conn right after reading the catalog, then used it four more
times. Every run died on the first sample with "Cannot operate on a closed
database". The step sits behind the user-tariff gate, so it had never been
executed and the defect shipped unseen.
3. seed_local_dispatch_energy read its measurement one line too early
ctx.avg_power_watts was read INSIDE the `with measure(...)` block, but
measure finalizes on __exit__ (that is where the sampler thread is joined
and the average computed). It was therefore always None, and the script
wrote cost_per_1m_prompt = cost_per_1m_completion = $0.0000 -- pricing local
compute as FREE, the exact failure this feature exists to prevent.
With both fixed, the measured rates on a Quadro RTX 6000 at $0.159/kWh are
$0.0054/1M prompt and $0.2286/1M completion (r^2 = 0.9985).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VRQXz5SYZYVWscxS1QqF6U
1580 lines
60 KiB
YAML
1580 lines
60 KiB
YAML
# Self-eval task set. One score per task per model; scores accumulate into
|
|
# proficiency.self_eval_score and self_eval_samples.
|
|
#
|
|
# Four kinds, chosen so each category is scored the most objective way it
|
|
# admits:
|
|
#
|
|
# code model writes Python; each check is eval'd against it in a
|
|
# subprocess. Score = fraction of checks passing. Fully objective.
|
|
# exact model replies with one value; compared after normalization.
|
|
# tool model is given a tool schema; scored on whether it calls the right
|
|
# tool with the right arguments. Structural, no judge needed.
|
|
# judge a strong model scores the output against a rubric. Only for the
|
|
# prose categories, where nothing checkable exists.
|
|
#
|
|
# DIFFICULTY: the happy path is not worth testing. Every current model passes
|
|
# "reverse a list", and a category where everyone scores 1.00 discriminates no
|
|
# better than the constant 0.5 it replaced. Each task here carries at least one
|
|
# edge case a plausible-looking solution gets wrong: touching vs overlapping
|
|
# intervals, present-but-falsy values, late-binding closures, greedy regexes,
|
|
# empty input, or a trap in the arithmetic.
|
|
#
|
|
# Every `code` task is validated against a reference solution by
|
|
# tests/test_task_set.py. A check my own reference cannot pass is a broken
|
|
# check, and would score the task set rather than the model — which has
|
|
# already happened here once.
|
|
#
|
|
# Keep prompts tight and self-contained. An ambiguous task scores the prompt.
|
|
|
|
tasks:
|
|
# --- coding_general -----------------------------------------------------
|
|
- id: merge_intervals
|
|
category: coding_general
|
|
kind: code
|
|
entrypoint: merge_intervals
|
|
prompt: |
|
|
Write a Python function `merge_intervals(intervals)` where intervals is
|
|
a list of [start, end] lists. Merge all overlapping intervals and return
|
|
a new list of [start, end] lists sorted by start. Intervals that merely
|
|
touch (one ends exactly where the next begins) must be merged. Input may
|
|
be unsorted and may contain intervals fully nested inside others.
|
|
Reply with ONLY the function definition — no explanation, no fences.
|
|
checks:
|
|
- 'merge_intervals([[1,3],[2,6],[8,10]]) == [[1,6],[8,10]]'
|
|
- 'merge_intervals([]) == []'
|
|
- 'merge_intervals([[1,4],[4,5]]) == [[1,5]]'
|
|
- 'merge_intervals([[1,10],[2,3]]) == [[1,10]]'
|
|
- 'merge_intervals([[5,6],[1,2]]) == [[1,2],[5,6]]'
|
|
- 'merge_intervals([[1,2]]) == [[1,2]]'
|
|
|
|
- id: parse_semver
|
|
category: coding_general
|
|
kind: code
|
|
entrypoint: parse_semver
|
|
prompt: |
|
|
Write a Python function `parse_semver(version)` that parses a semantic
|
|
version string into a dict with keys: major, minor, patch (ints), and
|
|
prerelease, build (strings, or None when absent). Valid examples:
|
|
"1.2.3", "1.2.3-alpha.1", "1.2.3+build.5", "1.2.3-rc.1+exp.sha.5114f85".
|
|
Raise ValueError if the string is not a valid semantic version, for
|
|
example "1.2" or "1.2.x".
|
|
Reply with ONLY the function definition — no explanation, no fences.
|
|
checks:
|
|
- 'parse_semver("1.2.3") == {"major":1,"minor":2,"patch":3,"prerelease":None,"build":None}'
|
|
- 'parse_semver("1.2.3-alpha.1")["prerelease"] == "alpha.1"'
|
|
- 'parse_semver("1.2.3+build.5")["build"] == "build.5"'
|
|
- 'parse_semver("1.2.3-rc.1+exp.sha.5114f85")["prerelease"] == "rc.1"'
|
|
- 'parse_semver("1.2.3-rc.1+exp.sha.5114f85")["build"] == "exp.sha.5114f85"'
|
|
- 'parse_semver("0.0.0")["major"] == 0'
|
|
- 'raises(ValueError, parse_semver, "1.2")'
|
|
- 'raises(ValueError, parse_semver, "1.2.x")'
|
|
|
|
- id: word_wrap
|
|
category: coding_general
|
|
kind: code
|
|
entrypoint: word_wrap
|
|
prompt: |
|
|
Write a Python function `word_wrap(text, width)` returning a list of
|
|
lines. Split on whitespace and pack as many words per line as fit within
|
|
`width` characters, joining words with a single space. Never split a
|
|
word: a word longer than `width` gets its own line. Runs of whitespace
|
|
collapse. Empty or whitespace-only text returns an empty list.
|
|
Reply with ONLY the function definition — no explanation, no fences.
|
|
checks:
|
|
- 'word_wrap("the quick brown fox", 10) == ["the quick", "brown fox"]'
|
|
- 'word_wrap("", 5) == []'
|
|
- 'word_wrap(" ", 5) == []'
|
|
- 'word_wrap("supercalifragilistic", 5) == ["supercalifragilistic"]'
|
|
- 'word_wrap("a b c", 3) == ["a b", "c"]'
|
|
- 'word_wrap("aa bb cc", 5) == ["aa bb", "cc"]'
|
|
|
|
# --- coding_general (added) — adapted from CRUXEval-O output-prediction,
|
|
# facebookresearch/cruxeval, MIT License. Data fetched once to /tmp,
|
|
# never vendored. Filter: has loop, ≤400 LOC chars, no unsafe imports,
|
|
# safe stdlib only. 6 of 6 tasks use `eval("f(<input>)")` for
|
|
# offline verification (input is the source arg list, not a single
|
|
# literal).
|
|
- id: crux_o_sample_9
|
|
category: coding_general
|
|
kind: exact
|
|
answer: "False"
|
|
prompt: |
|
|
What does this function return when called as shown? Reply with ONLY
|
|
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
|
'text', None), no explanation, no fences.
|
|
|
|
def f(t):
|
|
for c in t:
|
|
if not c.isnumeric():
|
|
return False
|
|
return True
|
|
|
|
f('#284376598')
|
|
|
|
- id: crux_o_sample_0
|
|
category: coding_general
|
|
kind: exact
|
|
answer: "[(4, 1), (4, 1), (4, 1), (4, 1), (2, 3), (2, 3)]"
|
|
prompt: |
|
|
What does this function return when called as shown? Reply with ONLY
|
|
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
|
'text', None), no explanation, no fences.
|
|
|
|
def f(nums):
|
|
output = []
|
|
for n in nums:
|
|
output.append((nums.count(n), n))
|
|
output.sort(reverse=True)
|
|
return output
|
|
|
|
f([1, 1, 3, 1, 3, 1])
|
|
|
|
- id: crux_o_sample_1
|
|
category: coding_general
|
|
kind: exact
|
|
answer: "{1: None, 2: None}"
|
|
prompt: |
|
|
What does this function return when called as shown? Reply with ONLY
|
|
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
|
'text', None), no explanation, no fences.
|
|
|
|
def f(a, b, c):
|
|
result = {}
|
|
for d in a, b, c:
|
|
result.update(dict.fromkeys(d))
|
|
return result
|
|
|
|
f((1, ), (1, ), (1, 2))
|
|
|
|
- id: crux_o_sample_2
|
|
category: coding_general
|
|
kind: exact
|
|
answer: "'hbtofdeiequ'"
|
|
prompt: |
|
|
What does this function return when called as shown? Reply with ONLY
|
|
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
|
'text', None), no explanation, no fences.
|
|
|
|
def f(text):
|
|
new_text = list(text)
|
|
for i in '+':
|
|
if i in new_text:
|
|
new_text.remove(i)
|
|
return ''.join(new_text)
|
|
|
|
f('hbtofdeiequ')
|
|
|
|
- id: crux_o_sample_5
|
|
category: coding_general
|
|
kind: exact
|
|
answer: "(0, 'xxxxxxxxxxxxxxxxxx')"
|
|
prompt: |
|
|
What does this function return when called as shown? Reply with ONLY
|
|
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
|
'text', None), no explanation, no fences.
|
|
|
|
def f(text, lower, upper):
|
|
count = 0
|
|
new_text = list()
|
|
for char in text:
|
|
char = lower if char.isdecimal() else upper
|
|
if char in ['p', 'C']:
|
|
count += 1
|
|
new_text.append(char)
|
|
return count, ''.join(new_text)
|
|
|
|
f('DSUWeqExTQdCMGpqur', 'a', 'x')
|
|
|
|
- id: crux_o_sample_6
|
|
category: coding_general
|
|
kind: exact
|
|
answer: "[('74', 31)]"
|
|
prompt: |
|
|
What does this function return when called as shown? Reply with ONLY
|
|
the literal Python value — strings in quotes (for example: 42, [1, 2],
|
|
'text', None), no explanation, no fences.
|
|
|
|
def f(dic):
|
|
for k,v in sorted(dic.items(), key=lambda x: len(str(x)))[:-1]:
|
|
dic.pop(k)
|
|
return list(dic.items())
|
|
|
|
f({'11': 52, '65': 34, 'a': 12, '4': 52, '74': 31})
|
|
|
|
# --- coding_refactor ----------------------------------------------------
|
|
- id: refactor_falsy_defaults
|
|
category: coding_refactor
|
|
kind: code
|
|
entrypoint: apply_settings
|
|
prompt: |
|
|
Refactor this function to remove the repetition. Behaviour must be
|
|
preserved EXACTLY, including for values that are present but falsy.
|
|
Reply with ONLY the rewritten function — no explanation, no fences.
|
|
|
|
def apply_settings(overrides):
|
|
result = {}
|
|
if "retries" in overrides:
|
|
result["retries"] = overrides["retries"]
|
|
else:
|
|
result["retries"] = 3
|
|
if "timeout" in overrides:
|
|
result["timeout"] = overrides["timeout"]
|
|
else:
|
|
result["timeout"] = 30
|
|
if "verbose" in overrides:
|
|
result["verbose"] = overrides["verbose"]
|
|
else:
|
|
result["verbose"] = False
|
|
return result
|
|
checks:
|
|
- 'apply_settings({}) == {"retries":3,"timeout":30,"verbose":False}'
|
|
- 'apply_settings({"retries":0})["retries"] == 0'
|
|
- 'apply_settings({"timeout":0})["timeout"] == 0'
|
|
- 'apply_settings({"verbose":True})["verbose"] is True'
|
|
- 'apply_settings({"retries":5}) == {"retries":5,"timeout":30,"verbose":False}'
|
|
|
|
- id: refactor_first_match
|
|
category: coding_refactor
|
|
kind: code
|
|
entrypoint: first_match
|
|
prompt: |
|
|
Refactor this to remove the nested loops and the flag variable.
|
|
Behaviour must be preserved exactly, including which item wins when
|
|
several match. Reply with ONLY the rewritten function — no explanation,
|
|
no fences.
|
|
|
|
def first_match(items, predicates):
|
|
found = None
|
|
done = False
|
|
for item in items:
|
|
if done:
|
|
break
|
|
for p in predicates:
|
|
if p(item):
|
|
found = item
|
|
done = True
|
|
break
|
|
return found
|
|
checks:
|
|
- 'first_match([1,2,3,4], [lambda x: x > 2]) == 3'
|
|
- 'first_match([1,2,3], [lambda x: x > 10]) is None'
|
|
- 'first_match([], [lambda x: True]) is None'
|
|
- 'first_match([1,2,3], []) is None'
|
|
- 'first_match([5,2,9], [lambda x: x > 8, lambda x: x < 3]) == 2'
|
|
- 'first_match([0,1], [lambda x: x == 0]) == 0'
|
|
|
|
- id: refactor_dispatch
|
|
category: coding_refactor
|
|
kind: code
|
|
entrypoint: describe
|
|
prompt: |
|
|
Refactor this if/elif chain into a table-driven lookup. Behaviour must be
|
|
preserved exactly for every input, including inputs that match no case.
|
|
Reply with ONLY the rewritten code — no explanation, no fences.
|
|
|
|
def describe(code):
|
|
if code == 200:
|
|
return "ok"
|
|
elif code == 201:
|
|
return "created"
|
|
elif code == 404:
|
|
return "not found"
|
|
elif code == 500:
|
|
return "server error"
|
|
else:
|
|
return "unknown"
|
|
checks:
|
|
- 'describe(200) == "ok"'
|
|
- 'describe(404) == "not found"'
|
|
- 'describe(500) == "server error"'
|
|
- 'describe(418) == "unknown"'
|
|
- 'describe(0) == "unknown"'
|
|
- 'describe(None) == "unknown"'
|
|
|
|
# --- bowling (Exercism-derived) -----------------------------------------
|
|
# Exercism python bowling exercise — MIT licensed, canonical source at
|
|
# https://github.com/exercism/python/tree/1f6aab8667bf653b10cc3799f94352fcdb749db6/exercises/practice/bowling
|
|
- id: refactor_bowling_frames
|
|
category: coding_refactor
|
|
kind: code
|
|
entrypoint: BowlingGame
|
|
prompt: |
|
|
Refactor this BowlingGame to remove the duplication and nested
|
|
conditions. Behaviour must be preserved EXACTLY, including scoring,
|
|
bonuses, and error cases. Reply with ONLY the rewritten class — no
|
|
explanation, no fences.
|
|
|
|
class BowlingGame:
|
|
def __init__(self):
|
|
self._frames = []
|
|
self._current = 0
|
|
self._bonus = []
|
|
|
|
def roll(self, pins):
|
|
if not (0 <= pins <= 10):
|
|
raise ValueError('invalid pins')
|
|
if self._current < 10:
|
|
if len(self._frames) == self._current:
|
|
self._frames.append([pins])
|
|
else:
|
|
self._frames[self._current].append(pins)
|
|
current = self._frames[self._current]
|
|
if sum(current) > 10:
|
|
raise ValueError("a frame's rolls cannot exceed 10")
|
|
strike = (len(current) == 1 and current[0] == 10)
|
|
if strike or len(current) == 2:
|
|
self._current += 1
|
|
else:
|
|
last = self._frames[-1]
|
|
last_total = sum(last)
|
|
strike10 = len(last) == 1 and last[0] == 10
|
|
spare10 = len(last) == 2 and last_total == 10
|
|
if strike10:
|
|
if len(self._bonus) >= 2:
|
|
raise IndexError(
|
|
'wrong number of fill balls when the tenth frame is a strike')
|
|
self._bonus.append(pins)
|
|
if len(self._bonus) == 2 and self._bonus[0] != 10 and sum(self._bonus) > 10:
|
|
raise ValueError('invalid fill balls')
|
|
if len(self._bonus) > 2:
|
|
raise IndexError(
|
|
'wrong number of fill balls when the tenth frame is a strike')
|
|
elif spare10:
|
|
if len(self._bonus) >= 1:
|
|
raise IndexError(
|
|
'wrong number of fill balls when the tenth frame is a spare')
|
|
self._bonus.append(pins)
|
|
else:
|
|
raise IndexError('cannot throw bonus with an open tenth frame')
|
|
|
|
def score(self):
|
|
if self._current < 10:
|
|
raise IndexError('frame less than 10')
|
|
last = self._frames[-1]
|
|
if len(last) == 2 and sum(last) == 10 and len(self._bonus) != 1:
|
|
raise IndexError('one bonus must be rolled when the tenth frame is spare')
|
|
if len(last) == 1 and last[0] == 10 and len(self._bonus) != 2:
|
|
raise IndexError('two bonuses must be rolled when the tenth frame is strike')
|
|
total = 0
|
|
for i in range(10):
|
|
frame = self._frames[i]
|
|
frame_sum = sum(frame)
|
|
strike = (len(frame) == 1 and frame[0] == 10)
|
|
spare = (len(frame) == 2 and frame_sum == 10)
|
|
if strike or spare:
|
|
nxt = []
|
|
for j in range(i + 1, 10):
|
|
nxt.extend(self._frames[j])
|
|
nxt.extend(self._bonus)
|
|
if strike:
|
|
frame_sum += sum(nxt[:2])
|
|
else:
|
|
frame_sum += sum(nxt[:1])
|
|
total += frame_sum
|
|
return total
|
|
checks:
|
|
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 7, 1]], g)[2].score() == 18)"
|
|
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [7, 3, 7]], g)[2].score() == 17)"
|
|
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 10, 10]], g)[2].score() == 30)"
|
|
- "((g := BowlingGame(), [g.roll(r) for r in [6, 4, 3] + [0]*17], g)[2].score() == 16)"
|
|
- "((g := BowlingGame(), [g.roll(r) for r in [3, 6]*10], g)[2].score() == 90)"
|
|
- "((g := BowlingGame(), [g.roll(r) for r in [10]*12], g)[2].score() == 300)"
|
|
- "raises(Exception, lambda: BowlingGame().roll(-1))"
|
|
- "raises(Exception, lambda: ((g := BowlingGame()), [g.roll(r) for r in [0]*20], g[2].roll(0)))"
|
|
|
|
- id: debug_bowling_tenth_frame
|
|
category: debugging
|
|
kind: code
|
|
entrypoint: BowlingGame
|
|
prompt: |
|
|
This BowlingGame is wrong on one subtle tenth-frame case. Fix ONLY the
|
|
bug; do not change anything else. Reply with ONLY the corrected class —
|
|
no explanation, no fences.
|
|
|
|
class BowlingGame:
|
|
def __init__(self):
|
|
self.current_frame_idx = 0
|
|
self.bonus_throws = []
|
|
self.frames = [Frame(idx) for idx in range(10)]
|
|
|
|
@property
|
|
def current_frame(self):
|
|
return self.frames[self.current_frame_idx]
|
|
|
|
def next_throws(self, frame_idx):
|
|
throws = []
|
|
for idx in range(frame_idx + 1, 10):
|
|
throws.extend(self.frames[idx].throws)
|
|
throws.extend(self.bonus_throws)
|
|
return throws
|
|
|
|
def roll_bonus(self, pins):
|
|
tenth_frame = self.frames[-1]
|
|
if tenth_frame.is_open():
|
|
raise IndexError('cannot throw bonus with an open tenth frame')
|
|
self.bonus_throws.append(pins)
|
|
if tenth_frame.is_strike() and len(self.bonus_throws) > 2:
|
|
raise IndexError(
|
|
'wrong number of fill balls when the tenth frame is a strike')
|
|
elif tenth_frame.is_spare() and len(self.bonus_throws) > 1:
|
|
raise IndexError(
|
|
'wrong number of fill balls when the tenth frame is a spare')
|
|
|
|
def roll(self, pins):
|
|
if not 0 <= pins <= 10:
|
|
raise ValueError('invalid pins')
|
|
elif self.current_frame_idx == 10:
|
|
self.roll_bonus(pins)
|
|
else:
|
|
self.current_frame.throw(pins)
|
|
if self.current_frame.is_closed():
|
|
self.current_frame_idx += 1
|
|
|
|
def score(self):
|
|
if self.current_frame_idx < 10:
|
|
raise IndexError('frame less than 10')
|
|
if self.frames[-1].is_spare() and len(self.bonus_throws) != 1:
|
|
raise IndexError(
|
|
'one bonus must be rolled when the tenth frame is spare')
|
|
if self.frames[-1].is_strike() and len(self.bonus_throws) != 2:
|
|
raise IndexError(
|
|
'two bonuses must be rolled when the tenth frame is strike')
|
|
return sum(frame.score(self.next_throws(frame.idx))
|
|
for frame in self.frames)
|
|
|
|
|
|
class Frame:
|
|
def __init__(self, idx):
|
|
self.idx = idx
|
|
self.throws = []
|
|
|
|
@property
|
|
def total_pins(self):
|
|
return sum(self.throws)
|
|
|
|
def is_strike(self):
|
|
return self.total_pins == 10 and len(self.throws) == 1
|
|
|
|
def is_spare(self):
|
|
return self.total_pins == 10 and len(self.throws) == 2
|
|
|
|
def is_open(self):
|
|
return self.total_pins < 10 and len(self.throws) == 2
|
|
|
|
def is_closed(self):
|
|
return self.total_pins == 10 or len(self.throws) == 2
|
|
|
|
def throw(self, pins):
|
|
if self.total_pins + pins > 10:
|
|
raise ValueError("a frame's rolls cannot exceed 10")
|
|
self.throws.append(pins)
|
|
|
|
def score(self, next_throws):
|
|
result = self.total_pins
|
|
if self.is_strike():
|
|
result += sum(next_throws[:2])
|
|
elif self.is_spare():
|
|
result += sum(next_throws[:1])
|
|
return result
|
|
checks:
|
|
- "raises(Exception, (lambda: ((g := BowlingGame()), [g.roll(r) for r in [0]*18 + [10, 5, 6]], g)[2].score()))"
|
|
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 10, 6]], g)[2].score() == 26)"
|
|
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [10, 7, 1]], g)[2].score() == 18)"
|
|
- "((g := BowlingGame(), [g.roll(r) for r in [0]*18 + [7, 3, 7]], g)[2].score() == 17)"
|
|
- "((g := BowlingGame(), [g.roll(r) for r in [10]*12], g)[2].score() == 300)"
|
|
|
|
# --- dominoes (Exercism-derived) ----------------------------------------
|
|
# Exercism python dominoes exercise — MIT licensed, canonical source at
|
|
# https://github.com/exercism/python/tree/1f6aab8667bf653b10cc3799f94352fcdb749db6/exercises/practice/dominoes
|
|
- id: refactor_dominoes_chain
|
|
category: coding_refactor
|
|
kind: code
|
|
entrypoint: can_chain
|
|
prompt: |
|
|
Refactor this can_chain to remove the duplicated chain-building
|
|
conditions and the flag variable. Behaviour must be preserved EXACTLY:
|
|
for a set of dominoes that can form a valid chain it returns a valid
|
|
chain (ANY valid chain — not a fixed one), and None when no chain is
|
|
possible. Reply with ONLY the rewritten function — no explanation, no
|
|
fences.
|
|
|
|
from itertools import permutations
|
|
|
|
|
|
def can_chain(dominoes):
|
|
if not any(dominoes):
|
|
return []
|
|
for perm in permutations(dominoes):
|
|
chain = [perm[0]]
|
|
complete = True
|
|
for domino in perm[1:]:
|
|
prev = chain[-1]
|
|
if len(chain) == 1 and prev[0] == domino[0]:
|
|
chain = [(prev[1], prev[0]), domino]
|
|
elif len(chain) == 1 and prev[0] == domino[1]:
|
|
chain = [(prev[1], prev[0]), (domino[1], domino[0])]
|
|
elif prev[1] == domino[0]:
|
|
chain = chain + [domino]
|
|
elif prev[1] == domino[1]:
|
|
chain = chain + [(domino[1], domino[0])]
|
|
else:
|
|
complete = False
|
|
break
|
|
if complete and chain[0][0] == chain[-1][1]:
|
|
return chain
|
|
return None
|
|
checks:
|
|
- 'can_chain([]) == []'
|
|
- '((d := [(1, 1)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
|
|
- '((d := [(1, 2), (3, 1), (2, 3)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
|
|
- '((d := [(1, 2), (2, 3), (3, 1), (2, 4), (2, 4)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
|
|
- 'can_chain([(1, 2)]) is None'
|
|
- 'can_chain([(1, 2), (4, 1), (2, 3)]) is None'
|
|
- 'can_chain([(1, 1), (2, 2)]) is None'
|
|
- 'can_chain([(1, 2), (2, 1), (3, 4), (4, 3)]) is None'
|
|
- 'can_chain([(1, 2), (2, 3), (3, 1), (4, 4)]) is None'
|
|
- 'can_chain([(1, 2), (2, 3), (3, 1), (4, 5), (5, 6), (6, 4)]) is None'
|
|
|
|
- id: debug_dominoes_no_chain
|
|
category: debugging
|
|
kind: code
|
|
entrypoint: can_chain
|
|
prompt: |
|
|
This can_chain returns a bogus "chain" for inputs that cannot be
|
|
chained — it returns a list instead of None when no valid chain exists.
|
|
Fix ONLY the one subtle bug; do not change anything else, and do not
|
|
change the can_chain signature. Reply with ONLY the corrected function
|
|
— no explanation, no fences.
|
|
|
|
from itertools import permutations
|
|
from functools import reduce
|
|
|
|
|
|
def swap(item_1, item_2):
|
|
return (item_2, item_1)
|
|
|
|
|
|
def build_chain(chain, domino):
|
|
if chain is not None:
|
|
last = chain[-1]
|
|
if len(chain) == 1 and last[0] == domino[0]:
|
|
return [swap(*last), domino]
|
|
elif len(chain) == 1 and last[0] == domino[1]:
|
|
return [swap(*last), swap(*domino)]
|
|
elif last[1] == domino[0]:
|
|
return chain + [domino]
|
|
elif last[1] == domino[1]:
|
|
return chain + [swap(*domino)]
|
|
return None
|
|
|
|
|
|
def can_chain(dominoes):
|
|
if not any(dominoes):
|
|
return []
|
|
for perm in permutations(dominoes):
|
|
chain = reduce(build_chain, perm[1:], [perm[0]])
|
|
# BUG: the circular-closure check (chain[0][0] == chain[-1][1])
|
|
# is missing, so a line that merely matches end-to-start is
|
|
# returned even when it does not close into a loop.
|
|
if chain is not None:
|
|
return chain
|
|
return None
|
|
checks:
|
|
- 'can_chain([(1, 2)]) is None'
|
|
- 'can_chain([(1, 2), (4, 1), (2, 3)]) is None'
|
|
- 'can_chain([(1, 1), (2, 2)]) is None'
|
|
- '((d := [(1, 2), (3, 1), (2, 3)], c := can_chain(d), c is not None and len(c) == len(d) and all(a[1] == b[0] for a, b in zip(c, c[1:])) and c[0][0] == c[-1][1])[-1])'
|
|
|
|
# --- affine-cipher (Exercism-derived) -----------------------------------
|
|
# Exercism python affine-cipher exercise — MIT licensed, canonical source at
|
|
# https://github.com/exercism/python/tree/1f6aab8667bf653b10cc3799f94352fcdb749db6/exercises/practice/affine-cipher
|
|
- id: refactor_affine_encode_decode
|
|
category: coding_refactor
|
|
kind: code
|
|
entrypoint: encode
|
|
prompt: |
|
|
Refactor this affine cipher to remove the duplicated cipher math. The
|
|
same letter-to-index transform and the coprime guard appear inline in
|
|
both encode and decode; behavioural duplicates like these are where bugs
|
|
hide. Consolidate them. Behaviour must be preserved EXACTLY, including
|
|
the ValueError raised when `a` is not coprime with the alphabet size and
|
|
the 5-character block grouping in encode. Keep the module functions
|
|
`encode(plain, a, b)` and `decode(ciphered, a, b)`. Reply with ONLY the
|
|
rewritten module — no explanation, no fences.
|
|
|
|
BLOCK_SIZE = 5
|
|
ALPHABET = 26
|
|
|
|
|
|
def mod_inverse(a_key, alphabet):
|
|
a_key = a_key % alphabet
|
|
for idx in range(1, alphabet):
|
|
if (a_key * idx) % alphabet == 1:
|
|
return idx
|
|
return 1
|
|
|
|
|
|
def encode(plain, a, b):
|
|
inverse = mod_inverse(a, ALPHABET)
|
|
if inverse == 1:
|
|
raise ValueError('a and m must be coprime.')
|
|
chars = []
|
|
for character in plain:
|
|
if character.isalnum():
|
|
origin = ord(character.lower()) - 97
|
|
if origin < 0:
|
|
chars.append(character)
|
|
continue
|
|
new = (a * origin + b) % ALPHABET
|
|
chars.append(chr(new + 97))
|
|
cipher = ''.join(chars)
|
|
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
|
|
for idx in range(0, len(cipher), BLOCK_SIZE)])
|
|
|
|
|
|
def decode(ciphered, a, b):
|
|
inverse = mod_inverse(a, ALPHABET)
|
|
if inverse == 1:
|
|
raise ValueError('a and m must be coprime.')
|
|
chars = []
|
|
for character in ciphered:
|
|
if character.isalnum():
|
|
origin = ord(character.lower()) - 97
|
|
if origin < 0:
|
|
chars.append(character)
|
|
continue
|
|
new = (inverse * (origin - b)) % ALPHABET
|
|
chars.append(chr(new + 97))
|
|
return ''.join(chars)
|
|
checks:
|
|
- 'encode("yes", 5, 7) == "xbt"'
|
|
- 'encode("no", 15, 18) == "fu"'
|
|
- 'encode("OMG", 21, 3) == "lvz"'
|
|
- 'encode("O M G", 25, 47) == "hjp"'
|
|
- 'encode("Testing,1 2 3, testing.", 3, 4) == "jqgjc rw123 jqgjc rw"'
|
|
- 'decode("tytgn fjr", 3, 7) == "exercism"'
|
|
- 'decode("qdwju nqcro muwhn odqun oppmd aunwd o", 19, 16) == "anobstacleisoftenasteppingstone"'
|
|
- 'raises(ValueError, encode, "This is a test.", 6, 17)'
|
|
- 'raises(ValueError, decode, "Test", 13, 5)'
|
|
|
|
- id: debug_affine_coprime
|
|
category: debugging
|
|
kind: code
|
|
entrypoint: encode
|
|
prompt: |
|
|
This affine cipher fails to reject keys where `a` is not coprime with the
|
|
alphabet size. It should raise ValueError('a and m must be coprime.')
|
|
when `a` shares a factor with 26, but it lets those keys through. Fix
|
|
ONLY the one subtle bug in the coprime guard; do not change anything
|
|
else, and do not change the signatures of `encode(plain, a, b)` or
|
|
`decode(ciphered, a, b)`. Reply with ONLY the corrected module — no
|
|
explanation, no fences.
|
|
|
|
BLOCK_SIZE = 5
|
|
ALPHABET = 26
|
|
|
|
|
|
def mod_inverse(a_key, alphabet):
|
|
a_key = a_key % alphabet
|
|
for idx in range(1, alphabet):
|
|
if (a_key * idx) % alphabet == 1:
|
|
return idx
|
|
return 1
|
|
|
|
|
|
def translate(text, a_key, b_key, mode):
|
|
inverse = mod_inverse(a_key, ALPHABET)
|
|
if inverse < 1:
|
|
raise ValueError('a and m must be coprime.')
|
|
chars = []
|
|
for character in text:
|
|
if character.isalnum():
|
|
origin = ord(character.lower()) - 97
|
|
if origin < 0:
|
|
chars.append(character)
|
|
continue
|
|
if mode == 0:
|
|
new = (a_key * origin + b_key) % ALPHABET
|
|
elif mode == 1:
|
|
new = (inverse * (origin - b_key)) % ALPHABET
|
|
chars.append(chr(new + 97))
|
|
return ''.join(chars)
|
|
|
|
|
|
def encode(plain, a, b):
|
|
cipher = translate(plain, a, b, 0)
|
|
return ' '.join([cipher[idx:idx + BLOCK_SIZE]
|
|
for idx in range(0, len(cipher), BLOCK_SIZE)])
|
|
|
|
|
|
def decode(ciphered, a, b):
|
|
return translate(ciphered, a, b, 1)
|
|
checks:
|
|
- 'encode("yes", 5, 7) == "xbt"'
|
|
- 'decode("tytgn fjr", 3, 7) == "exercism"'
|
|
- 'raises(ValueError, encode, "This is a test.", 6, 17)'
|
|
- 'raises(ValueError, decode, "Test", 13, 5)'
|
|
|
|
# --- debugging ----------------------------------------------------------
|
|
- id: debug_late_binding
|
|
category: debugging
|
|
kind: code
|
|
entrypoint: make_multipliers
|
|
prompt: |
|
|
This should return one multiplier function per factor, but every
|
|
returned function behaves the same. Fix it. Reply with ONLY the
|
|
corrected function — no explanation, no fences.
|
|
|
|
def make_multipliers(factors):
|
|
out = []
|
|
for f in factors:
|
|
out.append(lambda x: x * f)
|
|
return out
|
|
checks:
|
|
- '[m(2) for m in make_multipliers([1,2,3])] == [2,4,6]'
|
|
- '[m(10) for m in make_multipliers([0,1])] == [0,10]'
|
|
- 'make_multipliers([]) == []'
|
|
- 'make_multipliers([7])[0](3) == 21'
|
|
|
|
- id: debug_binary_search
|
|
category: debugging
|
|
kind: code
|
|
entrypoint: bsearch
|
|
prompt: |
|
|
This binary search should return the index of target in a sorted list,
|
|
or -1 if absent. It is wrong for some inputs — one case loops forever.
|
|
Fix it. Reply with ONLY the corrected function — no explanation, no
|
|
fences.
|
|
|
|
def bsearch(items, target):
|
|
lo, hi = 0, len(items)
|
|
while lo < hi:
|
|
mid = (lo + hi) // 2
|
|
if items[mid] == target:
|
|
return mid
|
|
elif items[mid] < target:
|
|
lo = mid
|
|
else:
|
|
hi = mid
|
|
return -1
|
|
checks:
|
|
- 'bsearch([1,3,5,7], 7) == 3'
|
|
- 'bsearch([1,3,5,7], 1) == 0'
|
|
- 'bsearch([], 1) == -1'
|
|
- 'bsearch([1], 1) == 0'
|
|
- 'bsearch([1,3], 2) == -1'
|
|
- 'bsearch([1,2,3,4,5,6], 6) == 5'
|
|
|
|
- id: debug_greedy_regex
|
|
category: debugging
|
|
kind: code
|
|
entrypoint: extract_tags
|
|
prompt: |
|
|
This should return the name inside each angle-bracket tag, in order, but
|
|
it returns the wrong thing when there is more than one tag. Fix it.
|
|
Reply with ONLY the corrected function — no explanation, no fences.
|
|
|
|
import re
|
|
|
|
def extract_tags(text):
|
|
return re.findall(r"<(.+)>", text)
|
|
checks:
|
|
- 'extract_tags("<a><b>") == ["a","b"]'
|
|
- 'extract_tags("<one>") == ["one"]'
|
|
- 'extract_tags("") == []'
|
|
- 'extract_tags("no tags here") == []'
|
|
- 'extract_tags("x <a> y <bc> z") == ["a","bc"]'
|
|
|
|
# --- reasoning_math -----------------------------------------------------
|
|
- id: math_percent_trap
|
|
category: reasoning_math
|
|
kind: exact
|
|
answer: "100"
|
|
prompt: |
|
|
A price rises by 20%, then falls by 20% of its new value. The final
|
|
price is 96. What was the original price? Reply with ONLY the number.
|
|
|
|
- id: math_rate_trap
|
|
category: reasoning_math
|
|
kind: exact
|
|
answer: "3"
|
|
prompt: |
|
|
Three machines take 3 minutes to make 3 widgets, each machine working
|
|
independently at the same constant rate. How many minutes do 100
|
|
machines take to make 100 widgets? Reply with ONLY the number.
|
|
|
|
- id: math_counting
|
|
category: reasoning_math
|
|
kind: exact
|
|
answer: "4536"
|
|
prompt: |
|
|
How many 4-digit whole numbers have four distinct digits and do not
|
|
begin with 0? Reply with ONLY the number.
|
|
|
|
# --- tool_use_agentic ---------------------------------------------------
|
|
- id: tool_multi_arg
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: Convert 250 US dollars into Japanese yen.
|
|
expect_tool: convert_currency
|
|
expect_args:
|
|
amount: 250
|
|
from_currency: USD
|
|
to_currency: JPY
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: get_weather
|
|
description: Get the current weather for a location.
|
|
parameters:
|
|
type: object
|
|
properties:
|
|
location: {type: string}
|
|
required: [location]
|
|
- type: function
|
|
function:
|
|
name: convert_currency
|
|
description: Convert an amount between two currencies.
|
|
parameters:
|
|
type: object
|
|
properties:
|
|
amount: {type: number}
|
|
from_currency: {type: string, description: ISO 4217 code}
|
|
to_currency: {type: string, description: ISO 4217 code}
|
|
required: [amount, from_currency, to_currency]
|
|
|
|
- id: tool_no_tool_needed
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: |
|
|
It is 1:20pm and my meeting starts at 3pm. How many minutes away is it?
|
|
expect_tool: null # plain arithmetic; both times are already given
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: get_calendar_event
|
|
description: Look up a calendar event by title.
|
|
parameters:
|
|
type: object
|
|
properties:
|
|
title: {type: string}
|
|
required: [title]
|
|
- type: function
|
|
function:
|
|
name: get_current_time
|
|
description: Get the current wall-clock time.
|
|
parameters:
|
|
type: object
|
|
properties: {}
|
|
|
|
- id: tool_abstain_creative
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: Write me a haiku about winter.
|
|
expect_tool: null
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: get_weather
|
|
description: Get the current weather for a location.
|
|
parameters:
|
|
type: object
|
|
properties:
|
|
location: {type: string}
|
|
required: [location]
|
|
|
|
# BFCL v4 derived (ShishirPatil/gorilla, MIT): 5 abstain from
|
|
# BFCL_v4_live_irrelevance.json, 3 positive from BFCL_v4_live_simple.json +
|
|
# possible_answer ground truth. Data fetched once to /tmp, translated
|
|
# mechanically, never vendored. Irrelevant prompts are deliberately
|
|
# unanswerable by the offered weather tool (real abstain cases).
|
|
- id: tool_bfcl_live_irrelevance_12-2-0
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: I'd appreciate if you could fetch the DNS resolution info for the domain mapped to IP 255.255.255.0 from VirusTotal. My key for this operation is 'sample_key4'.
|
|
expect_tool: null
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: get_current_weather
|
|
description: Retrieves the current weather conditions for a specified city and state.
|
|
parameters:
|
|
type: object
|
|
required:
|
|
- location
|
|
properties:
|
|
location:
|
|
type: string
|
|
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
|
|
exist.
|
|
unit:
|
|
type: string
|
|
description: The unit of temperature for the weather report.
|
|
enum:
|
|
- celsius
|
|
- fahrenheit
|
|
default: fahrenheit
|
|
- id: tool_bfcl_live_irrelevance_13-2-1
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: What is diffrence between cpu and gpu?
|
|
expect_tool: null
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: get_current_weather
|
|
description: Retrieves the current weather conditions for a specified city and state.
|
|
parameters:
|
|
type: object
|
|
required:
|
|
- location
|
|
properties:
|
|
location:
|
|
type: string
|
|
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
|
|
exist.
|
|
unit:
|
|
type: string
|
|
description: The unit of temperature for the weather report.
|
|
enum:
|
|
- celsius
|
|
- fahrenheit
|
|
default: fahrenheit
|
|
- id: tool_bfcl_live_irrelevance_14-2-2
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: Please help me get the votes associated with the IP of http://digdeep.io.
|
|
expect_tool: null
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: get_current_weather
|
|
description: Retrieves the current weather conditions for a specified city and state.
|
|
parameters:
|
|
type: object
|
|
required:
|
|
- location
|
|
properties:
|
|
location:
|
|
type: string
|
|
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
|
|
exist.
|
|
unit:
|
|
type: string
|
|
description: The unit of temperature for the weather report.
|
|
enum:
|
|
- celsius
|
|
- fahrenheit
|
|
default: fahrenheit
|
|
- id: tool_bfcl_live_irrelevance_15-2-3
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: Using 'api_key_2', retrieve the IDs of graphs containing IP 145.34.45.56 on VirusTotal. Don't forget to set the cursor as 'cursor_b' and limit the results to 8.
|
|
expect_tool: null
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: get_current_weather
|
|
description: Retrieves the current weather conditions for a specified city and state.
|
|
parameters:
|
|
type: object
|
|
required:
|
|
- location
|
|
properties:
|
|
location:
|
|
type: string
|
|
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
|
|
exist.
|
|
unit:
|
|
type: string
|
|
description: The unit of temperature for the weather report.
|
|
enum:
|
|
- celsius
|
|
- fahrenheit
|
|
default: fahrenheit
|
|
- id: tool_bfcl_live_irrelevance_16-2-4
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: 'How do I pull the domain info of twitter.com from VirusTotal? Using this API key: twt_key_abc.'
|
|
expect_tool: null
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: get_current_weather
|
|
description: Retrieves the current weather conditions for a specified city and state.
|
|
parameters:
|
|
type: object
|
|
required:
|
|
- location
|
|
properties:
|
|
location:
|
|
type: string
|
|
description: The location for which to get the weather, in the format of 'City, State', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city doesn't
|
|
exist.
|
|
unit:
|
|
type: string
|
|
description: The unit of temperature for the weather report.
|
|
enum:
|
|
- celsius
|
|
- fahrenheit
|
|
default: fahrenheit
|
|
- id: tool_bfcl_live_simple_0-0-0
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: Can you retrieve the details for the user with the ID 7890, who has black as their special request?
|
|
expect_tool: get_user_info
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: get_user_info
|
|
description: Retrieve details for a specific user by their unique identifier.
|
|
parameters:
|
|
type: object
|
|
required:
|
|
- user_id
|
|
properties:
|
|
user_id:
|
|
type: integer
|
|
description: The unique identifier of the user. It is used to fetch the specific user details from the database.
|
|
special:
|
|
type: string
|
|
description: Any special information or parameters that need to be considered while fetching user details.
|
|
default: none
|
|
expect_args:
|
|
user_id: 7890
|
|
special: black
|
|
- id: tool_bfcl_live_simple_1-1-0
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: I want to see the star history of ShishirPatil/gorilla and gorilla-llm/gorilla-cli, with the timelines aligned, so that I can more clearly observe the rate of change from their initial releases.
|
|
expect_tool: github_star
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: github_star
|
|
description: Generates a URL for tracking the star history of specified GitHub repositories, with the option to align them on the same timeline.
|
|
parameters:
|
|
type: object
|
|
required:
|
|
- repos
|
|
properties:
|
|
repos:
|
|
type: string
|
|
description: A comma-separated list of GitHub repositories to track, each in the 'owner/repo' format, such as 'octocat/Hello-World,octo-org/octo-repo'.
|
|
aligned:
|
|
type: boolean
|
|
description: Whether to align the repositories on the same timeline for comparison. If true, the star history of all repositories will start from the same point.
|
|
default: false
|
|
expect_args:
|
|
repos: ShishirPatil/gorilla,gorilla-llm/gorilla-cli
|
|
aligned: true
|
|
- id: tool_bfcl_live_simple_4-3-0
|
|
category: tool_use_agentic
|
|
kind: tool
|
|
prompt: What are the current weather conditions in Tel Aviv, and could you provide that in Fahrenheit, please?
|
|
expect_tool: get_current_weather
|
|
tools:
|
|
- type: function
|
|
function:
|
|
name: get_current_weather
|
|
description: Retrieves the current weather conditions for a specified city and state. If using state, then use short form like CA.
|
|
parameters:
|
|
type: object
|
|
required:
|
|
- location
|
|
properties:
|
|
location:
|
|
type: string
|
|
description: The location for which to get the weather, in the format of 'City, State (abbr)', such as 'San Francisco, CA' if State for the city exists. 'City, Country' if State for the city
|
|
doesn't exist.
|
|
unit:
|
|
type: string
|
|
description: The unit of temperature for the weather report.
|
|
enum:
|
|
- celsius
|
|
- fahrenheit
|
|
default: fahrenheit
|
|
expect_args:
|
|
location: Tel Aviv, Israel
|
|
unit: fahrenheit
|
|
|
|
# --- docs_writing -------------------------------------------------------
|
|
- id: docs_function
|
|
category: docs_writing
|
|
kind: judge
|
|
prompt: |
|
|
Write a docstring for this function. Reply with ONLY the docstring text.
|
|
|
|
def retry(fn, attempts=3, backoff=2.0):
|
|
delay = 1.0
|
|
for i in range(attempts):
|
|
try:
|
|
return fn()
|
|
except Exception:
|
|
if i == attempts - 1:
|
|
raise
|
|
time.sleep(delay)
|
|
delay *= backoff
|
|
rubric: |
|
|
Score 0-1. Award 1.0 ONLY if it states what the function does, documents
|
|
every parameter including defaults, states the return value, AND states
|
|
that the last exception is re-raised when all attempts fail. Deduct 0.3
|
|
for any invented behaviour the code does not have. Deduct 0.2 if it omits
|
|
that the delay grows by `backoff` between attempts.
|
|
|
|
- id: docs_gotcha
|
|
category: docs_writing
|
|
kind: judge
|
|
prompt: |
|
|
Write a docstring for this function. Reply with ONLY the docstring text.
|
|
|
|
def dedupe(items, key=None):
|
|
seen = set()
|
|
out = []
|
|
for item in items:
|
|
k = key(item) if key else item
|
|
if k in seen:
|
|
continue
|
|
seen.add(k)
|
|
out.append(item)
|
|
return out
|
|
rubric: |
|
|
Score 0-1. Award 1.0 ONLY if it states that ORDER IS PRESERVED and that
|
|
the FIRST occurrence is kept, documents `key`, and notes that elements
|
|
(or their keys) must be hashable. Deduct 0.4 if it omits the
|
|
order-preservation guarantee — that is the whole reason to use this over
|
|
set(). Deduct 0.3 if it omits the hashability requirement.
|
|
|
|
# --- summarization ------------------------------------------------------
|
|
- id: summarize_incident
|
|
category: summarization
|
|
kind: judge
|
|
prompt: |
|
|
Summarize in at most two sentences:
|
|
|
|
At 02:14 UTC the checkout service began returning 502s. The on-call
|
|
engineer found the connection pool exhausted. A deploy at 01:58 had
|
|
lowered the pool size from 50 to 5 through a bad template variable. The
|
|
deploy was rolled back at 02:31 and errors stopped by 02:34. Roughly
|
|
12,000 requests failed. No data was lost.
|
|
rubric: |
|
|
Score 0-1. Award 1.0 ONLY if it names the ROOT CAUSE specifically (a bad
|
|
template variable in the 01:58 deploy cut the pool from 50 to 5), the
|
|
resolution (rollback), and the impact (~12k failed requests, no data
|
|
lost), in two sentences or fewer. Deduct 0.4 for saying only "the
|
|
connection pool was exhausted" — that is the symptom, not the cause.
|
|
Deduct 0.5 for any invented detail.
|
|
|
|
- id: summarize_buried_lede
|
|
category: summarization
|
|
kind: judge
|
|
prompt: |
|
|
Summarize the single most important point in one sentence:
|
|
|
|
The migration ran for six hours. Throughput averaged 4,200 rows per
|
|
second, peaking at 6,100. The team used a rolling window of 5,000 rows
|
|
per batch. Disk usage on the replica grew steadily. Partway through, a
|
|
unique constraint on the accounts table silently rejected 812 rows,
|
|
which were logged to a dead-letter file that nobody has yet processed.
|
|
CPU stayed under 40% throughout.
|
|
rubric: |
|
|
Score 0-1. The important point is that 812 rows were silently dropped and
|
|
remain unprocessed — everything else is routine telemetry. Award 1.0 only
|
|
if the summary leads with that. Score 0.3 or below if it summarizes the
|
|
throughput statistics instead. Must be one sentence.
|
|
|
|
# --- translation --------------------------------------------------------
|
|
- id: translate_technical
|
|
category: translation
|
|
kind: judge
|
|
prompt: |
|
|
Translate into French. Reply with ONLY the translation.
|
|
|
|
"The connection pool was exhausted because a recent deploy reduced its
|
|
size. Roll back the deploy and the errors should stop within a few
|
|
minutes."
|
|
rubric: |
|
|
Score 0-1 on accuracy and fluency. Award 1.0 only for correct technical
|
|
register on "connection pool", "deploy" and "roll back", AND natural
|
|
French rather than a word-for-word calque. Deduct 0.3 per omission or
|
|
untranslated fragment.
|
|
|
|
- id: translate_register
|
|
category: translation
|
|
kind: judge
|
|
prompt: |
|
|
Translate into Spanish, preserving the hedging and the informal tone.
|
|
Reply with ONLY the translation.
|
|
|
|
"I'm not totally sure this is the right call, but I'd lean towards
|
|
shipping it and seeing what breaks — we can always roll it back."
|
|
rubric: |
|
|
Score 0-1. Award 1.0 only if the HEDGING is preserved ("not totally
|
|
sure", "I'd lean towards") rather than flattened into a confident
|
|
statement, the register stays informal, and "roll it back" is rendered
|
|
idiomatically. Deduct 0.4 if the hedging is lost.
|
|
|
|
# --- general_chat -------------------------------------------------------
|
|
- id: chat_explain
|
|
category: general_chat
|
|
kind: judge
|
|
prompt: |
|
|
Explain to a non-programmer, in under 100 words, why a program can be
|
|
correct and still be too slow to use.
|
|
rubric: |
|
|
Score 0-1. Award 1.0 only if it distinguishes correctness from
|
|
performance, gives at least one concrete relatable example, stays under
|
|
100 words, and leaves no jargon unexplained. Deduct 0.3 if over 100
|
|
words.
|
|
|
|
- id: chat_pushback
|
|
category: general_chat
|
|
kind: judge
|
|
prompt: |
|
|
A colleague says "we should rewrite the whole service in Rust, it'll be
|
|
faster." Reply in under 80 words, taking the suggestion seriously but
|
|
identifying what you would want to know first.
|
|
rubric: |
|
|
Score 0-1. Award 1.0 only if it avoids both pure agreement and pure
|
|
dismissal, names at least two specific things worth establishing first
|
|
(for example where time is actually spent, migration cost, team
|
|
familiarity), and stays under 80 words. Score 0.3 or below for a reply
|
|
that simply agrees or simply refuses.
|
|
|
|
# --- diff_checking -------------------------------------------------------
|
|
# Diff-pair exact tasks: model reviews a BEFORE -> AFTER refactor and answers
|
|
# YES if the AFTER changes behaviour for any valid input, NO if it is safe.
|
|
# Each bug is ported from this repo's own known-bug vocabulary.
|
|
|
|
- id: diff_check_late_binding_safe
|
|
category: diff_checking
|
|
kind: exact
|
|
answer: "NO"
|
|
prompt: |
|
|
You are reviewing a refactor. BEFORE: def make_multipliers(factors):
|
|
out = []
|
|
for f in factors:
|
|
out.append(lambda x: x * f)
|
|
return out
|
|
PROPOSED AFTER: def make_multipliers(factors):
|
|
out = []
|
|
for f in factors:
|
|
out.append(lambda x, f=f: x * f)
|
|
return out
|
|
Does the AFTER version introduce a bug that the BEFORE version does not
|
|
have? Reply with only YES or NO.
|
|
|
|
- id: diff_check_late_binding_buggy
|
|
category: diff_checking
|
|
kind: exact
|
|
answer: "YES"
|
|
prompt: |
|
|
You are reviewing a refactor. BEFORE: def make_multipliers(factors):
|
|
out = []
|
|
for f in factors:
|
|
out.append(lambda x, f=f: x * f)
|
|
return out
|
|
PROPOSED AFTER: def make_multipliers(factors):
|
|
out = []
|
|
for f in factors:
|
|
out.append(lambda x: x * f)
|
|
return out
|
|
Does the AFTER version introduce a bug that the BEFORE version does not
|
|
have? Reply with only YES or NO.
|
|
|
|
- id: diff_check_bsearch_boundary_safe
|
|
category: diff_checking
|
|
kind: exact
|
|
answer: "NO"
|
|
prompt: |
|
|
You are reviewing a refactor. BEFORE: def bsearch(items, target):
|
|
lo, hi = 0, len(items)
|
|
while lo < hi:
|
|
mid = (lo + hi) // 2
|
|
if items[mid] == target:
|
|
return mid
|
|
elif items[mid] < target:
|
|
lo = mid
|
|
else:
|
|
hi = mid
|
|
return -1
|
|
PROPOSED AFTER: def bsearch(items, target):
|
|
lo, hi = 0, len(items)
|
|
while lo < hi:
|
|
mid = (lo + hi) // 2
|
|
if items[mid] == target:
|
|
return mid
|
|
elif items[mid] < target:
|
|
lo = mid + 1
|
|
else:
|
|
hi = mid
|
|
return -1
|
|
Does the AFTER version introduce a bug that the BEFORE version does not
|
|
have? Reply with only YES or NO.
|
|
|
|
- id: diff_check_bsearch_boundary_buggy
|
|
category: diff_checking
|
|
kind: exact
|
|
answer: "YES"
|
|
prompt: |
|
|
You are reviewing a refactor. BEFORE: def bsearch(items, target):
|
|
lo, hi = 0, len(items)
|
|
while lo < hi:
|
|
mid = (lo + hi) // 2
|
|
if items[mid] == target:
|
|
return mid
|
|
elif items[mid] < target:
|
|
lo = mid + 1
|
|
else:
|
|
hi = mid
|
|
return -1
|
|
PROPOSED AFTER: def bsearch(items, target):
|
|
lo, hi = 0, len(items)
|
|
while lo < hi:
|
|
mid = (lo + hi) // 2
|
|
if items[mid] == target:
|
|
return mid
|
|
elif items[mid] < target:
|
|
lo = mid
|
|
else:
|
|
hi = mid
|
|
return -1
|
|
Does the AFTER version introduce a bug that the BEFORE version does not
|
|
have? Reply with only YES or NO.
|
|
|
|
- id: diff_check_falsy_default_safe
|
|
category: diff_checking
|
|
kind: exact
|
|
answer: "NO"
|
|
prompt: |
|
|
You are reviewing a refactor. BEFORE: def apply_settings(overrides):
|
|
result = {}
|
|
result["retries"] = overrides.get("retries", 3)
|
|
result["timeout"] = overrides.get("timeout", 30)
|
|
result["verbose"] = overrides.get("verbose", False)
|
|
return result
|
|
PROPOSED AFTER: def apply_settings(overrides):
|
|
result = {}
|
|
if "retries" in overrides:
|
|
result["retries"] = overrides["retries"]
|
|
else:
|
|
result["retries"] = 3
|
|
if "timeout" in overrides:
|
|
result["timeout"] = overrides["timeout"]
|
|
else:
|
|
result["timeout"] = 30
|
|
if "verbose" in overrides:
|
|
result["verbose"] = overrides["verbose"]
|
|
else:
|
|
result["verbose"] = False
|
|
return result
|
|
Does the AFTER version introduce a bug that the BEFORE version does not
|
|
have? Reply with only YES or NO.
|
|
|
|
- id: diff_check_falsy_default_buggy
|
|
category: diff_checking
|
|
kind: exact
|
|
answer: "YES"
|
|
prompt: |
|
|
You are reviewing a refactor. BEFORE: def apply_settings(overrides):
|
|
result = {}
|
|
if "retries" in overrides:
|
|
result["retries"] = overrides["retries"]
|
|
else:
|
|
result["retries"] = 3
|
|
if "timeout" in overrides:
|
|
result["timeout"] = overrides["timeout"]
|
|
else:
|
|
result["timeout"] = 30
|
|
if "verbose" in overrides:
|
|
result["verbose"] = overrides["verbose"]
|
|
else:
|
|
result["verbose"] = False
|
|
return result
|
|
PROPOSED AFTER: def apply_settings(overrides):
|
|
result = {}
|
|
result["retries"] = overrides.get("retries", 3)
|
|
result["timeout"] = overrides.get("timeout", 30)
|
|
result["verbose"] = overrides.get("verbose", False)
|
|
return result
|
|
Does the AFTER version introduce a bug that the BEFORE version does not
|
|
have? Reply with only YES or NO.
|
|
|
|
- id: diff_check_greedy_regex_safe
|
|
category: diff_checking
|
|
kind: exact
|
|
answer: "NO"
|
|
prompt: |
|
|
You are reviewing a refactor. BEFORE: import re
|
|
def extract_tags(text):
|
|
return re.findall(r"<(.+)>", text)
|
|
PROPOSED AFTER: import re
|
|
def extract_tags(text):
|
|
return re.findall(r"<([^<>]+)>", text)
|
|
Does the AFTER version introduce a bug that the BEFORE version does not
|
|
have? Reply with only YES or NO.
|
|
|
|
- id: diff_check_greedy_regex_buggy
|
|
category: diff_checking
|
|
kind: exact
|
|
answer: "YES"
|
|
prompt: |
|
|
You are reviewing a refactor. BEFORE: import re
|
|
def extract_tags(text):
|
|
return re.findall(r"<([^<>]+)>", text)
|
|
PROPOSED AFTER: import re
|
|
def extract_tags(text):
|
|
return re.findall(r"<(.+)>", text)
|
|
Does the AFTER version introduce a bug that the BEFORE version does not
|
|
have? Reply with only YES or NO.
|
|
|
|
# --- file_summarization ---------------------------------------------------
|
|
# Judge tasks: each summarises a small real file whose gotcha is non-obvious.
|
|
# The rubric must hinge on identifying the gotcha, not the happy path.
|
|
|
|
- id: summarize_yaml_falsy_zero
|
|
category: file_summarization
|
|
kind: judge
|
|
prompt: |
|
|
Summarise the most important non-obvious gotcha in this code:
|
|
|
|
iteration:
|
|
retries: 0
|
|
backoff: 2.0
|
|
cache:
|
|
enabled: true
|
|
max_size: 128
|
|
|
|
(config/config.yaml)
|
|
|
|
What does the retries: 0 value signal, and what happens to cache when
|
|
its subkeys are missing? How does a function that uses key-in-d checks
|
|
differ from one that uses d-get-default in each case?
|
|
rubric: |
|
|
Score 0-1. Award 1.0 ONLY if it states that retries: 0 is an
|
|
intentional override meaning "do not retry" — a present-but-falsy value
|
|
is NOT treated as unset — and that a missing cache subkey under a
|
|
present-but-empty cache falls back to the function's defaults. Deduct
|
|
0.4 for saying "retries: 0 means unset."
|
|
|
|
- id: summarize_dedupe_module
|
|
category: file_summarization
|
|
kind: judge
|
|
prompt: |
|
|
Summarise the most important non-obvious gotcha in this code:
|
|
|
|
def dedupe(items, key=None):
|
|
seen = set()
|
|
out = []
|
|
for item in items:
|
|
k = key(item) if key else item
|
|
if k in seen:
|
|
continue
|
|
seen.add(k)
|
|
out.append(item)
|
|
return out
|
|
|
|
(from tests/test_task_set.py)
|
|
|
|
What two guarantees does this function provide beyond what set() offers?
|
|
rubric: |
|
|
Score 0-1. Award 1.0 ONLY if it states that ORDER IS PRESERVED and that
|
|
the FIRST occurrence is kept, and notes that elements (or their keys via
|
|
key=) must be hashable. Deduct 0.4 for omitting order preservation —
|
|
that is the whole reason to use this over set(). Deduct 0.3 for omitting
|
|
the hashability requirement.
|
|
|
|
- id: summarize_retry_raiser
|
|
category: file_summarization
|
|
kind: judge
|
|
prompt: |
|
|
Summarise the most important non-obvious gotcha in this code:
|
|
|
|
def retry(fn, attempts=3, backoff=2.0):
|
|
delay = 1.0
|
|
for i in range(attempts):
|
|
try:
|
|
return fn()
|
|
except Exception:
|
|
if i == attempts - 1:
|
|
raise
|
|
time.sleep(delay)
|
|
delay *= backoff
|
|
|
|
(from src/eval_proficiency.py)
|
|
|
|
What happens to the exception when all attempts are exhausted? How does
|
|
the delay between retries change?
|
|
rubric: |
|
|
Score 0-1. Award 1.0 ONLY if it states the last exception IS re-raised
|
|
(not swallowed) and that the delay multiplies by backoff between
|
|
attempts. Deduct 0.4 for saying the exception is caught and discarded
|
|
or for omitting the backoff multiplication.
|
|
|
|
- id: summarize_single_use_iterator
|
|
category: file_summarization
|
|
kind: judge
|
|
prompt: |
|
|
Summarise the most important non-obvious gotcha in this code:
|
|
|
|
gen = (x * 2 for x in range(5))
|
|
for _ in range(3):
|
|
for v in gen:
|
|
print(v)
|
|
print(list(gen))
|
|
|
|
(from dispatcher.py)
|
|
|
|
What is printed by the inner loop, and what does list(gen) produce
|
|
afterwards?
|
|
rubric: |
|
|
Score 0-1. Award 1.0 ONLY if it states the second outer-iteration
|
|
(or the list(gen) call) yields NOTHING because the generator is
|
|
exhausted after the first inner loop. This is the core gotcha: you
|
|
cannot re-iterate an exhausted generator.
|
|
|
|
- id: summarize_naive_datetime
|
|
category: file_summarization
|
|
kind: judge
|
|
prompt: |
|
|
Summarise the most important non-obvious gotcha in this code:
|
|
|
|
import datetime
|
|
start = datetime.datetime(2026, 3, 8, 2, 0)
|
|
end = start + datetime.timedelta(hours=48)
|
|
# start is timezone-naive
|
|
|
|
(from a schedule module in the router codebase)
|
|
|
|
What is wrong with performing arithmetic on a naive datetime across a
|
|
clock-change boundary?
|
|
rubric: |
|
|
Score 0-1. Award 1.0 ONLY if it states that naive datetimes have NO
|
|
timezone/DST handling — timedelta arithmetic is unreliable across a
|
|
clock change (spring-forward or fall-back) because the wall-clock
|
|
hours may not correspond to real hours. Deduct 0.4 for not mentioning
|
|
DST/timezone specifically.
|
|
|
|
- id: summarize_sqlite_fk_pragma
|
|
category: file_summarization
|
|
kind: judge
|
|
prompt: |
|
|
Summarise the most important non-obvious gotcha in this code:
|
|
|
|
import sqlite3
|
|
conn = sqlite3.connect("mydb.db")
|
|
cursor = conn.cursor()
|
|
cursor.execute(
|
|
"INSERT INTO children (parent_id, name) VALUES (1, 'Alice')")
|
|
|
|
(from a script in the router codebase)
|
|
|
|
Assuming a children table has a foreign-key constraint to a parents
|
|
table, will this insert fail if parent_id=1 doesn't exist? Why or why
|
|
not?
|
|
rubric: |
|
|
Score 0-1. Award 1.0 ONLY if it states that foreign keys are OFF by
|
|
default in SQLite unless a PRAGMA foreign_keys = ON has been executed
|
|
on THAT connection (per-connection). Deduct 0.4 for assuming SQLite
|
|
enforces FKs automatically.
|