Add-models: re-judge or retry a model a run already has - #98
Merged
Merged
Conversation
A judging outage can leave a model's scores in a published run mostly failed, which scores it near zero. With rejudge, the add-models job keeps that model's saved responses, drops its saved scores and judges only it again, publishing a new run. Other models keep their scores untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj
Some providers reject a share of requests, which leaves a model with missing answers in an otherwise complete run. With retryFailed, the add-models job asks a model the run already has again only on the prompts where its answer failed, judges those new answers, and keeps every other answer and score. The asking code is shared with the path that adds new models. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.
Tip: disable this comment in your organization's Code Review settings.
Contributor
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two follow-ups to the Apertus add-models job. Until now, the job skipped any blueprint whose latest run already had the model, so a run with bad data for that model couldn't be repaired.
rejudge). A judging outage near the end of the job (OpenRouter 402s) left Apertus's scores on 3 benchmarks mostly failed: 16 of 40, 107 of 166, and 24 of 25 points. That put it last on each. Withrejudge, the job keeps the model's saved responses, drops its saved scores and judges only that model again.retryFailed). The Current AI router rejected some Apertus requests, up to 30–40% on a few benchmarks. WithretryFailed, the job asks the model again only on prompts where its answer failed, and judges just those new answers. Every other answer and score is kept, and a blueprint with no failed answers is skipped.Each run that changes publishes a new run, the same way adding a model does. Other models' responses and scores are copied unchanged.
Changes
addModelsToLatestRun(configId, models, logger, { rejudge?, retryFailed? }): both options apply only to listed models the latest run already contains. Asking a model is now one shared helper, used for both added models and retries.POST /api/internal/add-models-to-runsacceptsrejudgeandretryFailedbooleans.rejudgeandretry_failedcheckboxes, both off by default.Test plan
tsc --noEmitis clean.rejudgeonvisual__web_design_101,yka-setandyka__disability-rights-accommodation;retry_failedandrebuild_summarieson the full Apertus list.Risks
Both flags are off by default, so existing workflow runs behave exactly as before. Each repaired blueprint gets one new published run, and older runs are untouched.
🤖 Generated with Claude Code
https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj
Generated by Claude Code