Skip to content

Add-models: re-judge or retry a model a run already has - #98

Merged
nojibe merged 3 commits into
mainfrom
claude/rejudge-models
Oct 1, 2026
Merged

nojibe merged 3 commits into
mainfrom
claude/rejudge-models

Conversation

@nojibe

@nojibe nojibe commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Summary

Two follow-ups to the Apertus add-models job. Until now, the job skipped any blueprint whose latest run already had the model, so a run with bad data for that model couldn't be repaired.

  • Re-judge (rejudge). A judging outage near the end of the job (OpenRouter 402s) left Apertus's scores on 3 benchmarks mostly failed: 16 of 40, 107 of 166, and 24 of 25 points. That put it last on each. With rejudge, the job keeps the model's saved responses, drops its saved scores and judges only that model again.
  • Retry failed answers (retryFailed). The Current AI router rejected some Apertus requests, up to 30–40% on a few benchmarks. With retryFailed, the job asks the model again only on prompts where its answer failed, and judges just those new answers. Every other answer and score is kept, and a blueprint with no failed answers is skipped.

Each run that changes publishes a new run, the same way adding a model does. Other models' responses and scores are copied unchanged.

Changes

  • addModelsToLatestRun(configId, models, logger, { rejudge?, retryFailed? }): both options apply only to listed models the latest run already contains. Asking a model is now one shared helper, used for both added models and retries.
  • POST /api/internal/add-models-to-runs accepts rejudge and retryFailed booleans.
  • The Add Models To Runs workflow gets rejudge and retry_failed checkboxes, both off by default.

Test plan

  • Service tests (8): re-judge keeps responses and drops only that model's scores; retry asks only the failed pair in the run's own variant, keeps the good answer's score, and leaves the retried pair to be judged; retry skips a blueprint with no failures.
  • Route tests (9): both flags are passed through.
  • Full suite: 1175 passed, 2 skipped. tsc --noEmit is clean.
  • After merge:
    • run with rejudge on visual__web_design_101, yka-set and yka__disability-rights-accommodation;
    • then run with retry_failed and rebuild_summaries on the full Apertus list.

Risks

Both flags are off by default, so existing workflow runs behave exactly as before. Each repaired blueprint gets one new published run, and older runs are untouched.

🤖 Generated with Claude Code

https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj


Generated by Claude Code

claude added 3 commits October 1, 2026 05:55
A judging outage can leave a model's scores in a published run mostly
failed, which scores it near zero. With rejudge, the add-models job keeps
that model's saved responses, drops its saved scores and judges only it
again, publishing a new run. Other models keep their scores untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj
Some providers reject a share of requests, which leaves a model with missing
answers in an otherwise complete run. With retryFailed, the add-models job
asks a model the run already has again only on the prompts where its answer
failed, judges those new answers, and keeps every other answer and score.
The asking code is shared with the path that adds new models.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015roAwvpcszBvjX1Ce5NFuj

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.

Tip: disable this comment in your organization's Code Review settings.

@railway-app
railway-app Bot temporarily deployed to weval / app-pr-98 October 1, 2026 15:07 Destroyed
@railway-app

railway-app Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

🚅 Deployed to the app-pr-98 environment in weval

Service Status Web Updated
weval-app 🕒 Building (View Logs) Web Oct 1, 2026 at 3:07 pm UTC

@nojibe
nojibe merged commit 39a2a9d into main Oct 1, 2026
1 of 2 checks passed

This branch was successfully deployed

No deployments
weval / app-pr-98 — a99732dc Deployed Oct 1, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants