Skip to content

Quickstart: full knowledge loop over 3 hard ARC-AGI-1 tasks - #14

Merged
xuefei-wang merged 3 commits into
mainfrom
docs/quickstart-full-loop
Jul 22, 2026
Merged

xuefei-wang merged 3 commits into
mainfrom
docs/quickstart-full-loop

Conversation

@xuefei-wang

@xuefei-wang xuefei-wang commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

What

Make the quickstart demo exercise the whole knowledge-refinement loop (execute → per-task forum → cross-task forum → distill → seed) across generations, instead of just one execution phase.

Previously the demo ran trivial tasks that every agent solved on attempt 1 — so with --drop-solved (default) the pool emptied after generation 1 and the run stopped before the discussion/distill/seed phases ran.

Why ARC-AGI-1 (and how we got here)

The demo now runs 3 hard ARC-AGI-1 tasks. The journey there matters: custom "hard" tasks calibrated for one model don't transfer. A real gpt-5.4-mini run one-shot all four custom tasks — including the TSP task that was supposed to be "mathematically never solvable" (its reference tour was beatable) — so the pool emptied and the loop collapsed to a single generation again.

ARC-AGI-1 tasks are hard for every current model, which removes the calibration problem entirely. No dataset download — the ARC-1 corpus is already vendored under benchmarks/arc1/.

Changes

  • examples/quickstart/arc1_hard/ — 3 curated hard ARC-1 tasks (97239e3d, d22278a0, 776ffc46) with low observed pass rates, copied from the vendored corpus (they span the eval/training splits, so one --tasks-path dir holds all three). A README documents the selection and pass rates.
  • scripts/quickstart.sh — default to --task-source arc over these 3 tasks; 3 generations, forums on, --drop-solved default. TASKS_PATH still switches to custom-task mode; ARC_DATA_DIR / ARC_TASK_MAP knobs added.
  • Reverted the custom-task detour (examples/custom_tasks/tasks.jsonl + its test) back to main.
  • docs (getting-started, faq, README) reframed around the ARC demo, with real output and DB figures.

Verified end-to-end (real gpt-5.4-mini run)

The loop fires across all three generations:

gen task(s) score then
1 776ffc46, 97239e3d 1.0 (solved → dropped) distill → seed
1 d22278a0 0.0 (carries)
2 d22278a0 0.0 (carries) distill → seed
3 d22278a0 0.0 distill (+ cross-task)

Final completed traces=5 tasks=3 solved=2/3. Knowledge DB proof: per_task_distill=3, cross_task_distill=1, a cross_task_forum post, and seed_snapshots=2 (seeding fired between every generation). d22278a0 stayed unsolved (genuinely hard), so no solved-via-learning flip in this run, but the full execute→forum→distill→seed→next-gen machinery demonstrably runs.

Notes for reviewers

  • gpt-5.4-mini solved 2 of the 3 "hard" tasks on gen 1 (stronger than the pass-rate data suggested, likely helped by --arc-max-trials 2), leaving a single task to drive the loop — so gens 2–3 are single-agent and the forum is thin. On weaker models (e.g. Haiku, the documented default) more would carry forward. The demo still proves the pipeline.
  • The getting-started sample output and DB counts are now from this real run, not illustrative.

🤖 Generated with Claude Code

xuefei-wang and others added 2 commits July 21, 2026 21:42
The quickstart previously ran a single generation with both forums off, so
only the execution phase fired — the discussion/distill/seed phases the
project is actually about were never exercised in the first-run demo.

Run 3 generations with the per-task and cross-task forums on. Because the
bundled demo tasks all solve on generation 1, add --no-drop-solved so the
solved tasks are retained rather than dropped (which would empty the pool
and stop the run before seeding fires). All five phases now run end to end:
execute -> forum -> distill -> seed -> next generation.

Sync the docs that described the old single-generation behavior:
- getting-started: intro, run description, completion output, DB-check
  example, and the note (flipped from "why steps 3-5 don't fire" to "why
  the demo keeps solved tasks")
- faq: cost answer
- README: quickstart blurb
- experiments: drop the stale "run one generation" claim

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… --drop-solved default

Rather than force the loop with --no-drop-solved, use tasks that don't all
solve on the first attempt so the loop carries forward naturally under the
default --drop-solved (the realistic production setting).

New demo set (examples/custom_tasks/tasks.jsonl), four distinct failure modes:
- calc-eval:      arithmetic parser; hidden grader tests '/' truncation toward
                  zero (naive floor division fails). Visible tests omit negatives.
- range-queries:  point-set + range-sum; hidden grader runs 2e5 ops under a
                  5s internal budget, so an O(n)-per-query scan times out -> 0.
                  Needs a Fenwick/BIT.
- precise-sum:    compensated float summation; hidden grader feeds adversarial
                  magnitudes where naive accumulation loses precision.
- tsp-heuristic:  continuous score min(1.0, ref/len) over hidden instances --
                  effectively never 1.0, so it never drops and its score climbs
                  generation over generation.

Each grader was validated end-to-end (reference passes, naive fails / times out
/ scores low) through the exact eval-command strings.

- quickstart.sh: drop --no-drop-solved, bump --max-concurrent-tasks to 4, update
  the task-name banner and comments.
- test_custom_tasks_example.py: update to 4 tasks with reference solutions that
  prove each grader is satisfiable.
- docs (getting-started, faq, README): reframe the success signal from
  solved=3/3 to "attempts run and get scored; scores improve across
  generations", update task names/count, and replace the --no-drop-solved note
  with why these tasks are hard on purpose.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@xuefei-wang xuefei-wang changed the title Quickstart: run the full knowledge loop by default Quickstart: full knowledge loop over four harder tasks Jul 22, 2026
…eliably

Replace the custom-task demo with 3 hard ARC-AGI-1 tasks. Custom tasks calibrated
for one model don't transfer: a real gpt-5.4-mini run one-shot all four custom
tasks (including the "never 1.0" TSP task, whose reference was beatable), so
--drop-solved emptied the pool and the run stopped after generation 1. ARC tasks
are hard for every current model, which fixes the calibration problem.

- examples/quickstart/arc1_hard/: 3 ARC-1 tasks (97239e3d, d22278a0, 776ffc46)
  with low observed pass rates, copied from the vendored ARC-1 corpus (they span
  the eval/training splits, so a single --tasks-path dir holds all three). README
  documents the selection + pass rates. No dataset download.
- quickstart.sh: default to --task-source arc over these 3 tasks; keep 3
  generations / forums on / --drop-solved default. TASKS_PATH still switches to
  custom-task mode; ARC_DATA_DIR / ARC_TASK_MAP knobs added.
- Revert the custom-task detour (examples/custom_tasks/tasks.jsonl and its test)
  back to main.
- docs (getting-started, faq, README): reframe the demo around ARC, with real
  output/DB figures from a verified run.

Verified end to end (gpt-5.4-mini): gen 1 solved 2/3, d22278a0 carried through
gens 2-3; distillation ran each generation (per_task_distill=3, cross_task=1)
and seeding fired between generations (seed_snapshots=2) — the full
execute->forum->distill->seed loop runs across all three generations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@xuefei-wang xuefei-wang changed the title Quickstart: full knowledge loop over four harder tasks Quickstart: full knowledge loop over 3 hard ARC-AGI-1 tasks Jul 22, 2026
@xuefei-wang
xuefei-wang merged commit 3dec501 into main Jul 22, 2026
3 checks passed
@xuefei-wang
xuefei-wang deleted the docs/quickstart-full-loop branch July 22, 2026 15:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant