Skip to content

Arena-Hard-V2 judge: an unrecognized/garbled verdict silently scores as a tie and is never flagged failed #445

Description

llm-as-a-coach/scripts/eval/eval_gpt4o_fuzzy.py:389-450, aggregated at 716-746

_ARENA_VERDICT_PATTERNS = [r"\[\[([AB<>=]+)\]\]", r"\[([AB<>=]+)\]"]

def _parse_arena_verdict(judgment_text):
    if not judgment_text:
        return None
    upper = judgment_text.upper()
    for pattern in _ARENA_VERDICT_PATTERNS:
        matches = re.findall(pattern, upper)
        matches = [m for m in matches if m]
        if matches:
            return matches[-1].strip("\n")
    return None

def _arena_score_round(verdict, flipped):
    if verdict is None:
        return 0.5
    verdict = verdict.replace(" ", "")
    if not flipped:
        if verdict in ("B>A", "B>>A"): return 1.0
        if verdict == "A=B": return 0.5
        if verdict in ("A>B", "A>>B"): return 0.0
    else:
        if verdict in ("A>B", "A>>B"): return 1.0
        if verdict == "A=B": return 0.5
        if verdict in ("B>A", "B>>A"): return 0.0
    return 0.5   # <- unrecognized-but-non-None verdict falls through here

def judge_arena_hard_v2(oai_client, sample, idx):
    ...
    win_rate = (s1 + s2) / 2.0
    failed = v1 is None or v2 is None   # <- only catches fully-empty output
    return idx, win_rate, {"round1": output1, "round2": output2}, failed

_parse_arena_verdict's character class [AB<>=]+ accepts any run of those five characters inside brackets, not just the five literals the judge prompt actually asks for (A>>B, A>B, A=B, B>A, B>>A). A judge reply like [[AB]] or [[A>A]] — no operator, or a self-comparison — passes the regex, so v1/v2 come back non-None, so failed stays False. But the string matches none of the five branches in _arena_score_round, so it falls through to the final return 0.5, the same value the function returns for a genuine, well-formed tie (A=B).

What happens (reproduced on the real module)

_parse_arena_verdict's permissive character class [AB<>=]+ lets any run of A/B/</>/= characters count as a 'parsed' verdict (e.g. [[AB]], [[A>A]]). That makes v1/v2 non-None, so the round is excluded from failed = v1 is None or v2 is None (eval_gpt4o_fuzzy.py:449), even though _arena_score_round (lines 404-424) has no branch for that string and falls through to the same return 0.5 used for a genuine A=B tie.

As a result, a broken-but-bracketed judge call is indistinguishable, in both avg_win_rate and failed_indices, from a real tie. The sibling judge_wildbench (line 334) already closes this gap by checking the parsed value against a known-good vocabulary (choice.strip() not in WILDBENCH_REWARD_MAP) instead of only checking for None.

Suggested fix: define the five valid literals and set failed = v1 not in VALID_VERDICTS or v2 not in VALID_VERDICTS, mirroring the existing in-file pattern.

Happy to open the PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions