Quickstart: full knowledge loop over 3 hard ARC-AGI-1 tasks - #14
Merged
Merged
Conversation
The quickstart previously ran a single generation with both forums off, so only the execution phase fired — the discussion/distill/seed phases the project is actually about were never exercised in the first-run demo. Run 3 generations with the per-task and cross-task forums on. Because the bundled demo tasks all solve on generation 1, add --no-drop-solved so the solved tasks are retained rather than dropped (which would empty the pool and stop the run before seeding fires). All five phases now run end to end: execute -> forum -> distill -> seed -> next generation. Sync the docs that described the old single-generation behavior: - getting-started: intro, run description, completion output, DB-check example, and the note (flipped from "why steps 3-5 don't fire" to "why the demo keeps solved tasks") - faq: cost answer - README: quickstart blurb - experiments: drop the stale "run one generation" claim Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… --drop-solved default
Rather than force the loop with --no-drop-solved, use tasks that don't all
solve on the first attempt so the loop carries forward naturally under the
default --drop-solved (the realistic production setting).
New demo set (examples/custom_tasks/tasks.jsonl), four distinct failure modes:
- calc-eval: arithmetic parser; hidden grader tests '/' truncation toward
zero (naive floor division fails). Visible tests omit negatives.
- range-queries: point-set + range-sum; hidden grader runs 2e5 ops under a
5s internal budget, so an O(n)-per-query scan times out -> 0.
Needs a Fenwick/BIT.
- precise-sum: compensated float summation; hidden grader feeds adversarial
magnitudes where naive accumulation loses precision.
- tsp-heuristic: continuous score min(1.0, ref/len) over hidden instances --
effectively never 1.0, so it never drops and its score climbs
generation over generation.
Each grader was validated end-to-end (reference passes, naive fails / times out
/ scores low) through the exact eval-command strings.
- quickstart.sh: drop --no-drop-solved, bump --max-concurrent-tasks to 4, update
the task-name banner and comments.
- test_custom_tasks_example.py: update to 4 tasks with reference solutions that
prove each grader is satisfiable.
- docs (getting-started, faq, README): reframe the success signal from
solved=3/3 to "attempts run and get scored; scores improve across
generations", update task names/count, and replace the --no-drop-solved note
with why these tasks are hard on purpose.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…eliably Replace the custom-task demo with 3 hard ARC-AGI-1 tasks. Custom tasks calibrated for one model don't transfer: a real gpt-5.4-mini run one-shot all four custom tasks (including the "never 1.0" TSP task, whose reference was beatable), so --drop-solved emptied the pool and the run stopped after generation 1. ARC tasks are hard for every current model, which fixes the calibration problem. - examples/quickstart/arc1_hard/: 3 ARC-1 tasks (97239e3d, d22278a0, 776ffc46) with low observed pass rates, copied from the vendored ARC-1 corpus (they span the eval/training splits, so a single --tasks-path dir holds all three). README documents the selection + pass rates. No dataset download. - quickstart.sh: default to --task-source arc over these 3 tasks; keep 3 generations / forums on / --drop-solved default. TASKS_PATH still switches to custom-task mode; ARC_DATA_DIR / ARC_TASK_MAP knobs added. - Revert the custom-task detour (examples/custom_tasks/tasks.jsonl and its test) back to main. - docs (getting-started, faq, README): reframe the demo around ARC, with real output/DB figures from a verified run. Verified end to end (gpt-5.4-mini): gen 1 solved 2/3, d22278a0 carried through gens 2-3; distillation ran each generation (per_task_distill=3, cross_task=1) and seeding fired between generations (seed_snapshots=2) — the full execute->forum->distill->seed loop runs across all three generations. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Make the quickstart demo exercise the whole knowledge-refinement loop (execute → per-task forum → cross-task forum → distill → seed) across generations, instead of just one execution phase.
Previously the demo ran trivial tasks that every agent solved on attempt 1 — so with
--drop-solved(default) the pool emptied after generation 1 and the run stopped before the discussion/distill/seed phases ran.Why ARC-AGI-1 (and how we got here)
The demo now runs 3 hard ARC-AGI-1 tasks. The journey there matters: custom "hard" tasks calibrated for one model don't transfer. A real
gpt-5.4-minirun one-shot all four custom tasks — including the TSP task that was supposed to be "mathematically never solvable" (its reference tour was beatable) — so the pool emptied and the loop collapsed to a single generation again.ARC-AGI-1 tasks are hard for every current model, which removes the calibration problem entirely. No dataset download — the ARC-1 corpus is already vendored under
benchmarks/arc1/.Changes
examples/quickstart/arc1_hard/— 3 curated hard ARC-1 tasks (97239e3d,d22278a0,776ffc46) with low observed pass rates, copied from the vendored corpus (they span the eval/training splits, so one--tasks-pathdir holds all three). A README documents the selection and pass rates.scripts/quickstart.sh— default to--task-source arcover these 3 tasks; 3 generations, forums on,--drop-solveddefault.TASKS_PATHstill switches to custom-task mode;ARC_DATA_DIR/ARC_TASK_MAPknobs added.examples/custom_tasks/tasks.jsonl+ its test) back tomain.Verified end-to-end (real
gpt-5.4-minirun)The loop fires across all three generations:
Final
completed traces=5 tasks=3 solved=2/3. Knowledge DB proof:per_task_distill=3,cross_task_distill=1, across_task_forumpost, andseed_snapshots=2(seeding fired between every generation).d22278a0stayed unsolved (genuinely hard), so no solved-via-learning flip in this run, but the full execute→forum→distill→seed→next-gen machinery demonstrably runs.Notes for reviewers
--arc-max-trials 2), leaving a single task to drive the loop — so gens 2–3 are single-agent and the forum is thin. On weaker models (e.g. Haiku, the documented default) more would carry forward. The demo still proves the pipeline.🤖 Generated with Claude Code