Skip to content

N=4 validation milestone: top-4 roster vs static baseline #46

Description

@jkbennitt

Goal

The first statistically defensible RLE result: N=4 paired runs of the v0.3.0 spread's top models vs the static no-agent baseline, with bootstrap CIs.

Roster (from the 2026-06-11 v0.3.0 N=1 spread)

model why
x-ai/grok-4.3 repeat champion (0.836 mean, $0.36)
mistralai/mistral-medium-3-5 #2, cheapest frontier-quality ($0.73)
z-ai/glm-5.1 the ONLY model above baseline (+0.028, 7/10 ticks) — the scientifically interesting one
nvidia/nemotron-3-super-120b-a12b open-weight/value anchor, best action success (85%, $0.12)

Decisions to make BEFORE launching

  • Tick count. 10 ticks ends day 2–5; baseline mean time-to-end is 8 days — most of the composite is still the opening decay curve, so N=4 at 10 ticks would just tighten CIs around "everyone tracks baseline." Recommend 20–30 ticks. Cost scales linearly (~$15–25 metered at 10 ticks for this roster).
  • Restraint probe first (cheap, ~$0.50): one N=1 grok-4.3 run with a "no-action must be justified / actions must beat doing nothing" prompt variant, to test whether glm-5.1's baseline-beating restraint is portable. Result shapes the N=4 prompt. (See discussion on Agents must beat unmanaged baseline #6.)

First live exercise of (watch console for both)

Pipeline

Existing spread tooling runs as-is (run_spread_n1.sh pattern, OBS rig, run-analysis skill). This run should also be the first --push-hf artifact (see HF integration issue).

🤖 Generated with Claude Code

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions