llm-as-a-coach/scripts/eval/eval_gpt4o_fuzzy.py:389-450, aggregated at 716-746
_ARENA_VERDICT_PATTERNS = [r"\[\[([AB<>=]+)\]\]", r"\[([AB<>=]+)\]"]
def _parse_arena_verdict(judgment_text):
if not judgment_text:
return None
upper = judgment_text.upper()
for pattern in _ARENA_VERDICT_PATTERNS:
matches = re.findall(pattern, upper)
matches = [m for m in matches if m]
if matches:
return matches[-1].strip("\n")
return None
def _arena_score_round(verdict, flipped):
if verdict is None:
return 0.5
verdict = verdict.replace(" ", "")
if not flipped:
if verdict in ("B>A", "B>>A"): return 1.0
if verdict == "A=B": return 0.5
if verdict in ("A>B", "A>>B"): return 0.0
else:
if verdict in ("A>B", "A>>B"): return 1.0
if verdict == "A=B": return 0.5
if verdict in ("B>A", "B>>A"): return 0.0
return 0.5 # <- unrecognized-but-non-None verdict falls through here
def judge_arena_hard_v2(oai_client, sample, idx):
...
win_rate = (s1 + s2) / 2.0
failed = v1 is None or v2 is None # <- only catches fully-empty output
return idx, win_rate, {"round1": output1, "round2": output2}, failed
_parse_arena_verdict's character class [AB<>=]+ accepts any run of those five characters inside brackets, not just the five literals the judge prompt actually asks for (A>>B, A>B, A=B, B>A, B>>A). A judge reply like [[AB]] or [[A>A]] — no operator, or a self-comparison — passes the regex, so v1/v2 come back non-None, so failed stays False. But the string matches none of the five branches in _arena_score_round, so it falls through to the final return 0.5, the same value the function returns for a genuine, well-formed tie (A=B).
What happens (reproduced on the real module)
_parse_arena_verdict's permissive character class [AB<>=]+ lets any run of A/B/</>/= characters count as a 'parsed' verdict (e.g. [[AB]], [[A>A]]). That makes v1/v2 non-None, so the round is excluded from failed = v1 is None or v2 is None (eval_gpt4o_fuzzy.py:449), even though _arena_score_round (lines 404-424) has no branch for that string and falls through to the same return 0.5 used for a genuine A=B tie.
As a result, a broken-but-bracketed judge call is indistinguishable, in both avg_win_rate and failed_indices, from a real tie. The sibling judge_wildbench (line 334) already closes this gap by checking the parsed value against a known-good vocabulary (choice.strip() not in WILDBENCH_REWARD_MAP) instead of only checking for None.
Suggested fix: define the five valid literals and set failed = v1 not in VALID_VERDICTS or v2 not in VALID_VERDICTS, mirroring the existing in-file pattern.
Happy to open the PR.
llm-as-a-coach/scripts/eval/eval_gpt4o_fuzzy.py:389-450, aggregated at716-746_parse_arena_verdict's character class[AB<>=]+accepts any run of those five characters inside brackets, not just the five literals the judge prompt actually asks for (A>>B,A>B,A=B,B>A,B>>A). A judge reply like[[AB]]or[[A>A]]— no operator, or a self-comparison — passes the regex, sov1/v2come back non-None, sofailedstaysFalse. But the string matches none of the five branches in_arena_score_round, so it falls through to the finalreturn 0.5, the same value the function returns for a genuine, well-formed tie (A=B).What happens (reproduced on the real module)
_parse_arena_verdict's permissive character class[AB<>=]+lets any run of A/B/</>/= characters count as a 'parsed' verdict (e.g.[[AB]],[[A>A]]). That makesv1/v2non-None, so the round is excluded fromfailed = v1 is None or v2 is None(eval_gpt4o_fuzzy.py:449), even though_arena_score_round(lines 404-424) has no branch for that string and falls through to the samereturn 0.5used for a genuineA=Btie.As a result, a broken-but-bracketed judge call is indistinguishable, in both
avg_win_rateandfailed_indices, from a real tie. The siblingjudge_wildbench(line 334) already closes this gap by checking the parsed value against a known-good vocabulary (choice.strip() not in WILDBENCH_REWARD_MAP) instead of only checking forNone.Suggested fix: define the five valid literals and set
failed = v1 not in VALID_VERDICTS or v2 not in VALID_VERDICTS, mirroring the existing in-file pattern.Happy to open the PR.