From b4dba12e94be7143b866bb97a0368c99bbfbeeb0 Mon Sep 17 00:00:00 2001 From: xuefei-wang Date: Tue, 21 Jul 2026 21:42:26 -0700 Subject: [PATCH 1/3] feat(quickstart): run the full knowledge loop by default MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The quickstart previously ran a single generation with both forums off, so only the execution phase fired — the discussion/distill/seed phases the project is actually about were never exercised in the first-run demo. Run 3 generations with the per-task and cross-task forums on. Because the bundled demo tasks all solve on generation 1, add --no-drop-solved so the solved tasks are retained rather than dropped (which would empty the pool and stop the run before seeding fires). All five phases now run end to end: execute -> forum -> distill -> seed -> next generation. Sync the docs that described the old single-generation behavior: - getting-started: intro, run description, completion output, DB-check example, and the note (flipped from "why steps 3-5 don't fire" to "why the demo keeps solved tasks") - faq: cost answer - README: quickstart blurb - experiments: drop the stale "run one generation" claim Co-Authored-By: Claude Opus 4.8 (1M context) --- README.md | 5 ++-- docs/experiments.md | 6 ++-- docs/faq.md | 3 +- docs/getting-started.md | 61 +++++++++++++++++++++++------------------ scripts/quickstart.sh | 19 +++++++++---- 5 files changed, 56 insertions(+), 38 deletions(-) diff --git a/README.md b/README.md index 5a94517..6c44705 100644 --- a/README.md +++ b/README.md @@ -41,8 +41,9 @@ bash scripts/quickstart.sh The script self-bootstraps everything it needs: it synthesizes a provider profile from your key, builds the `ksi-agent:bench` image on first run, -installs the host Node dependencies, then runs one generation over the -bundled [`examples/custom_tasks/`](./examples/custom_tasks/) demo. If +installs the host Node dependencies, then runs three generations over the +bundled [`examples/custom_tasks/`](./examples/custom_tasks/) demo — with the +forums on, so the full execute → forum → distill → seed loop fires. If anything is missing, `uv run ksi-doctor` prints a ✓/✗ readiness checklist with the exact command to fix it. diff --git a/docs/experiments.md b/docs/experiments.md index 8f05ad0..6c1057d 100644 --- a/docs/experiments.md +++ b/docs/experiments.md @@ -1,9 +1,9 @@ # Running larger runs The [quickstart](getting-started.md) and [your own tasks](your_own_tasks.md) -walkthroughs run one generation over a handful of tiny tasks. This page -covers the flags that matter once you scale up — more tasks, more -generations, or a maintained reference benchmark instead of your own tasks. +walkthroughs run over a handful of tiny tasks. This page covers the flags +that matter once you scale up — more tasks, more generations, or a +maintained reference benchmark instead of your own tasks. There is **one canonical launch surface**: the `ksi.cli` argument parser. Everything else (bash presets, `ksi.run(...)`) is a layer over the same diff --git a/docs/faq.md b/docs/faq.md index 51cda64..de2bd03 100644 --- a/docs/faq.md +++ b/docs/faq.md @@ -84,7 +84,8 @@ Every run makes real LLM API calls billed to the key in your provider profile; there is no built-in spending cap. Cost scales with the number of tasks, generations, and the model you choose. To get a feel before committing, start with the bundled synthetic demo: `bash scripts/quickstart.sh` runs -three tasks, one generation, with Haiku — the fastest and cheapest option. +three tasks across three generations with Haiku — the cheapest way to watch +the full loop. Use `DRY_RUN=true` on any experiment wrapper script to print the full CLI command and DB paths without launching anything. diff --git a/docs/getting-started.md b/docs/getting-started.md index 7e167f2..715d557 100644 --- a/docs/getting-started.md +++ b/docs/getting-started.md @@ -4,9 +4,10 @@ Go from a fresh clone to a solved demo task in one command, then learn what just ## What you'll do -Run one fast generation of agents against three bundled, self-contained tasks and see them score — -no dataset download, no manual setup. The whole demo finishes in a few minutes and leaves you with -a working environment ready to run your own tasks or a reference benchmark. +Run three generations of agents against three bundled, self-contained tasks and watch the full +knowledge loop — execute, discuss, distill, seed — fire end to end. No dataset download, no manual +setup. The demo takes several minutes and leaves you with a working environment ready to run your +own tasks or a reference benchmark. ## Prerequisites @@ -24,10 +25,11 @@ bash scripts/quickstart.sh The script self-bootstraps everything it needs: it synthesizes a provider profile from your key, builds the `ksi-agent:bench` image on first run (this takes a few minutes), installs the host -Node dependencies, then runs one generation over the three bundled tasks under +Node dependencies, then runs three generations over the three bundled tasks under [`examples/custom_tasks/`](https://github.com/recursive-knowledge/KSI/tree/main/examples/custom_tasks) (`fizzbuzz`, `reverse-words`, `anagram-groups`) — each graded by running `python3 tests.py` -against the agent's attempt. +against the agent's attempt, with the per-task and cross-task forums on so every phase of the +loop fires. For the complete benchmark environment (including benchmark preparation and smoke tests), run `bash scripts/setup_all.sh`. Use `--no-test` when you need @@ -61,9 +63,11 @@ The run logs each attempt and its score as it progresses. When it finishes, resu | Execution traces | `analysis/traces//` | For the quickstart, `` defaults to `quickstart_demo`. The run prints -each task's score as it goes and ends with a `completed … solved=3/3 (100.0%)` -line — that's the signal your environment is set up correctly. Elapsed times and -token counts vary by model and run; the task names and `solved=3/3` don't. +each task's score as it goes, and each generation ends with a +`completed … solved=3/3 (100.0%)` line — three in all, one per generation. +Seeing `solved=3/3` is the signal your environment is set up correctly. Elapsed +times and token counts vary by model and run; the task names and `solved=3/3` +don't. ??? note "A closer look — sample output, optional artifacts, and the knowledge DB" @@ -72,30 +76,33 @@ token counts vary by model and run; the task names and `solved=3/3` don't. run preset, which sets it for you) for a score summary on disk. Traces default to `analysis/traces//` — set `KSI_TRACE_DIR` to change the root. - A real excerpt from a run against `claude-haiku-4-5-20251001` (the default - `configs/ksi/.env.haiku` profile), timestamps trimmed: + An illustrative excerpt from a generation's execution phase against + `claude-haiku-4-5-20251001` (the default `configs/ksi/.env.haiku` profile), + timestamps trimmed — each generation logs a block like this, followed by the + forum and distillation phases: ```text INFO ksi.orchestrator.execution_phase: [gen 1] task=reverse-words agent=agent-1 done elapsed=27.4s score=1.0000 INFO ksi.orchestrator.execution_phase: [gen 1] task=fizzbuzz agent=agent-0 done elapsed=28.1s score=1.0000 INFO ksi.orchestrator.execution_phase: [gen 1] task=anagram-groups agent=agent-2 done elapsed=33.0s score=1.0000 INFO ksi.orchestrator.engine: completed traces=3 tasks=3 solved=3/3 (100.0%) - INFO ksi.orchestrator.persistence: [tokens] total=418,329 cached_input=346,149 uncached_input=9,129 output=5,090 cache_create=57,961 ``` - **Knowledge DB check** — every solved attempt writes an `entry_type='attempt'` - row plus an `insight` row. The quickstart turns both forums off for speed - (`--per-task-forum-rounds 0 --cross-task-forum-rounds 0`), so there are no - discussion posts, and with nothing unsolved in this single-generation run - distillation has nothing to write either: + **Knowledge DB check** — because the demo now runs the full loop, the + knowledge DB carries rows from every phase, not just execution. Group by + `entry_type` and `source_phase` to see them: ```console $ sqlite3 runtime_state/knowledge/quickstart_demo/quickstart_demo_knowledge.sqlite \ "select entry_type, source_phase, count(*) from knowledge group by entry_type, source_phase order by entry_type, source_phase;" - attempt|execution|3 - insight|execution|3 ``` + You'll see `attempt` and `insight` rows from `execution`, `post` rows from + `per_task_forum` and `cross_task_forum` (the two discussion phases), and + `distillation` rows from `per_task_distill` and `cross_task_distill`. Exact + counts vary with the model and how much each agent posts, and they grow with + each of the three generations. + ## What just happened? KSI runs a knowledge-refinement loop across generations: @@ -106,14 +113,16 @@ KSI runs a knowledge-refinement loop across generations: 4. The system [*distills*](glossary.md#distillation) those discussions into reusable guidance. 5. The next generation is [*seeded*](glossary.md#seeding) with that guidance. -!!! note "Why the demo doesn't show steps 3–5" - The quickstart runs a single generation with both forums off, so steps 3–5 - don't fire here. And because every task solves on the first attempt, there - would be nothing to learn anyway: a solved task is dropped from later - generations (`--drop-solved`, on by default), so a multi-generation run - **stops early** once everything is solved. To watch the full loop, turn the - forums on, request several generations, and use tasks hard enough that some - fail — see [experiments.md](experiments.md). +!!! note "Why the demo keeps solved tasks (`--no-drop-solved`)" + The quickstart runs three generations with both forums on, so all five steps + fire. There's one wrinkle: every demo task solves on the first attempt, and a + solved task is normally dropped from later generations (`--drop-solved`, on by + default), which would empty the task pool and **stop the run** before seeding + ever fires. So the quickstart passes `--no-drop-solved` to retain the solved + tasks and carry the full loop across all three generations. On your own tasks, + leaving `--drop-solved` on (the default) is usually what you want — the loop + then concentrates each generation on what's still unsolved. See + [experiments.md](experiments.md). ## Next steps diff --git a/scripts/quickstart.sh b/scripts/quickstart.sh index ae78d7e..7a61b3d 100755 --- a/scripts/quickstart.sh +++ b/scripts/quickstart.sh @@ -10,7 +10,8 @@ # from the environment); # - the ksi-agent:bench Docker image (built on first run if missing); # - the host runtime_runner Node dependencies. -# Then it runs a single minimal generation so the demo finishes in a few minutes. +# Then it runs 3 generations with the forums on and solved tasks retained, so +# the full knowledge loop (execute -> forum -> distill -> seed) fires end to end. # # Usage: # ANTHROPIC_API_KEY=sk-ant-... bash scripts/quickstart.sh @@ -130,17 +131,23 @@ if [[ ! -f "$PROFILE" ]]; then exit 1 fi -# Minimal canonical run: only the required flags plus one fast generation with -# the discussion/distillation phases off, so the demo finishes quickly. +# Full-loop demo: 3 generations with the per-task and cross-task forums on and +# solved tasks retained (--no-drop-solved), so every phase fires across +# generations — execute -> forum -> distill -> seed -> next generation. This +# takes longer than a single generation but exercises the whole +# knowledge-refinement loop rather than just the execution phase. (Without +# --no-drop-solved the demo tasks all solve on generation 1 and get dropped, +# so the run would stop before seeding ever fires.) CMD=( "${PYRUN[@]}" -m ksi.cli --task-source custom --tasks-path "$TASKS_PATH" --evaluator command --provider-profile "$PROFILE" - --generations 1 - --per-task-forum-rounds 0 - --cross-task-forum-rounds 0 + --generations 3 + --per-task-forum-rounds 1 + --cross-task-forum-rounds 1 + --no-drop-solved --max-concurrent-tasks 3 --experiment-name "$EXPERIMENT_NAME" ) From 274a04bbe677afb3f34479b33b96cf662facfdfc Mon Sep 17 00:00:00 2001 From: xuefei-wang Date: Tue, 21 Jul 2026 22:31:24 -0700 Subject: [PATCH 2/3] feat(quickstart): replace easy demo tasks with four harder ones; keep --drop-solved default Rather than force the loop with --no-drop-solved, use tasks that don't all solve on the first attempt so the loop carries forward naturally under the default --drop-solved (the realistic production setting). New demo set (examples/custom_tasks/tasks.jsonl), four distinct failure modes: - calc-eval: arithmetic parser; hidden grader tests '/' truncation toward zero (naive floor division fails). Visible tests omit negatives. - range-queries: point-set + range-sum; hidden grader runs 2e5 ops under a 5s internal budget, so an O(n)-per-query scan times out -> 0. Needs a Fenwick/BIT. - precise-sum: compensated float summation; hidden grader feeds adversarial magnitudes where naive accumulation loses precision. - tsp-heuristic: continuous score min(1.0, ref/len) over hidden instances -- effectively never 1.0, so it never drops and its score climbs generation over generation. Each grader was validated end-to-end (reference passes, naive fails / times out / scores low) through the exact eval-command strings. - quickstart.sh: drop --no-drop-solved, bump --max-concurrent-tasks to 4, update the task-name banner and comments. - test_custom_tasks_example.py: update to 4 tasks with reference solutions that prove each grader is satisfiable. - docs (getting-started, faq, README): reframe the success signal from solved=3/3 to "attempts run and get scored; scores improve across generations", update task names/count, and replace the --no-drop-solved note with why these tasks are hard on purpose. Co-Authored-By: Claude Opus 4.8 (1M context) --- README.md | 7 +- docs/faq.md | 4 +- docs/getting-started.md | 62 +++++++++------- examples/custom_tasks/tasks.jsonl | 7 +- scripts/quickstart.sh | 26 +++---- tests/test_custom_tasks_example.py | 110 ++++++++++++++++++++++++++--- 6 files changed, 160 insertions(+), 56 deletions(-) diff --git a/README.md b/README.md index 6c44705..631313d 100644 --- a/README.md +++ b/README.md @@ -41,9 +41,10 @@ bash scripts/quickstart.sh The script self-bootstraps everything it needs: it synthesizes a provider profile from your key, builds the `ksi-agent:bench` image on first run, -installs the host Node dependencies, then runs three generations over the -bundled [`examples/custom_tasks/`](./examples/custom_tasks/) demo — with the -forums on, so the full execute → forum → distill → seed loop fires. If +installs the host Node dependencies, then runs three generations over four +deliberately harder tasks in the bundled +[`examples/custom_tasks/`](./examples/custom_tasks/) demo — with the forums on, +so the full execute → forum → distill → seed loop fires. If anything is missing, `uv run ksi-doctor` prints a ✓/✗ readiness checklist with the exact command to fix it. diff --git a/docs/faq.md b/docs/faq.md index de2bd03..76bc22d 100644 --- a/docs/faq.md +++ b/docs/faq.md @@ -84,8 +84,8 @@ Every run makes real LLM API calls billed to the key in your provider profile; there is no built-in spending cap. Cost scales with the number of tasks, generations, and the model you choose. To get a feel before committing, start with the bundled synthetic demo: `bash scripts/quickstart.sh` runs -three tasks across three generations with Haiku — the cheapest way to watch -the full loop. +four harder tasks across three generations with Haiku — the cheapest way to +watch the full loop. Use `DRY_RUN=true` on any experiment wrapper script to print the full CLI command and DB paths without launching anything. diff --git a/docs/getting-started.md b/docs/getting-started.md index 715d557..0cd408a 100644 --- a/docs/getting-started.md +++ b/docs/getting-started.md @@ -4,10 +4,11 @@ Go from a fresh clone to a solved demo task in one command, then learn what just ## What you'll do -Run three generations of agents against three bundled, self-contained tasks and watch the full -knowledge loop — execute, discuss, distill, seed — fire end to end. No dataset download, no manual -setup. The demo takes several minutes and leaves you with a working environment ready to run your -own tasks or a reference benchmark. +Run three generations of agents against four bundled, self-contained tasks — pitched hard enough +that they don't all solve on the first try — and watch the full knowledge loop — execute, discuss, +distill, seed — fire end to end. No dataset download, no manual setup. The demo takes several +minutes and leaves you with a working environment ready to run your own tasks or a reference +benchmark. ## Prerequisites @@ -25,11 +26,11 @@ bash scripts/quickstart.sh The script self-bootstraps everything it needs: it synthesizes a provider profile from your key, builds the `ksi-agent:bench` image on first run (this takes a few minutes), installs the host -Node dependencies, then runs three generations over the three bundled tasks under +Node dependencies, then runs three generations over the four bundled tasks under [`examples/custom_tasks/`](https://github.com/recursive-knowledge/KSI/tree/main/examples/custom_tasks) -(`fizzbuzz`, `reverse-words`, `anagram-groups`) — each graded by running `python3 tests.py` -against the agent's attempt, with the per-task and cross-task forums on so every phase of the -loop fires. +(`calc-eval`, `range-queries`, `precise-sum`, `tsp-heuristic`) — each graded on the host by the +task's own eval command, with the per-task and cross-task forums on so every phase of the loop +fires. For the complete benchmark environment (including benchmark preparation and smoke tests), run `bash scripts/setup_all.sh`. Use `--no-test` when you need @@ -62,12 +63,15 @@ The run logs each attempt and its score as it progresses. When it finishes, resu | Score summary (optional — only when `--output-json` is set) | `results/.json` | | Execution traces | `analysis/traces//` | -For the quickstart, `` defaults to `quickstart_demo`. The run prints +For the quickstart, `` defaults to `quickstart_demo`. The run logs each task's score as it goes, and each generation ends with a -`completed … solved=3/3 (100.0%)` line — three in all, one per generation. -Seeing `solved=3/3` is the signal your environment is set up correctly. Elapsed -times and token counts vary by model and run; the task names and `solved=3/3` -don't. +`completed … solved=N/M` line — three in all, one per generation. These tasks +are meant to be hard, so expect only some to solve on generation 1, with the +solve count and the `tsp-heuristic` score generally improving across the three +generations as distilled knowledge accumulates (the hardest tasks may stay +unsolved — that's fine). **The signal that your environment is set up correctly +is that attempts run and get scored at all** — not that everything solves. +Elapsed times, token counts, and exact solve counts vary by model and run. ??? note "A closer look — sample output, optional artifacts, and the knowledge DB" @@ -79,13 +83,15 @@ don't. An illustrative excerpt from a generation's execution phase against `claude-haiku-4-5-20251001` (the default `configs/ksi/.env.haiku` profile), timestamps trimmed — each generation logs a block like this, followed by the - forum and distillation phases: + forum and distillation phases. Scores are mixed on generation 1 by design + (`tsp-heuristic` is a continuous score, never a clean `1.0`): ```text - INFO ksi.orchestrator.execution_phase: [gen 1] task=reverse-words agent=agent-1 done elapsed=27.4s score=1.0000 - INFO ksi.orchestrator.execution_phase: [gen 1] task=fizzbuzz agent=agent-0 done elapsed=28.1s score=1.0000 - INFO ksi.orchestrator.execution_phase: [gen 1] task=anagram-groups agent=agent-2 done elapsed=33.0s score=1.0000 - INFO ksi.orchestrator.engine: completed traces=3 tasks=3 solved=3/3 (100.0%) + INFO ksi.orchestrator.execution_phase: [gen 1] task=range-queries agent=agent-1 done elapsed=31.2s score=1.0000 + INFO ksi.orchestrator.execution_phase: [gen 1] task=tsp-heuristic agent=agent-3 done elapsed=44.7s score=0.7800 + INFO ksi.orchestrator.execution_phase: [gen 1] task=calc-eval agent=agent-0 done elapsed=28.1s score=0.0000 + INFO ksi.orchestrator.execution_phase: [gen 1] task=precise-sum agent=agent-2 done elapsed=26.5s score=0.0000 + INFO ksi.orchestrator.engine: completed traces=4 tasks=4 solved=1/4 (25.0%) ``` **Knowledge DB check** — because the demo now runs the full loop, the @@ -113,15 +119,17 @@ KSI runs a knowledge-refinement loop across generations: 4. The system [*distills*](glossary.md#distillation) those discussions into reusable guidance. 5. The next generation is [*seeded*](glossary.md#seeding) with that guidance. -!!! note "Why the demo keeps solved tasks (`--no-drop-solved`)" - The quickstart runs three generations with both forums on, so all five steps - fire. There's one wrinkle: every demo task solves on the first attempt, and a - solved task is normally dropped from later generations (`--drop-solved`, on by - default), which would empty the task pool and **stop the run** before seeding - ever fires. So the quickstart passes `--no-drop-solved` to retain the solved - tasks and carry the full loop across all three generations. On your own tasks, - leaving `--drop-solved` on (the default) is usually what you want — the loop - then concentrates each generation on what's still unsolved. See +!!! note "Why these tasks are hard on purpose" + Earlier versions of this demo used trivial tasks that every agent solved on + the first attempt — so with `--drop-solved` (on by default) the task pool + emptied after generation 1 and the run **stopped** before the forum, distill, + and seed phases could show their value. These four tasks are pitched beyond a + reliable one-shot solve — a truncation-toward-zero parser trap, a range-query + task that needs a Fenwick tree, numerically-stable summation, and a + continuous-score TSP heuristic that is effectively never "perfect" — so + unsolved tasks carry forward under the default `--drop-solved` and the full + loop runs across all three generations. On your own *easy* tasks, expect the + run to stop early once everything is solved; that's the intended behavior. See [experiments.md](experiments.md). ## Next steps diff --git a/examples/custom_tasks/tasks.jsonl b/examples/custom_tasks/tasks.jsonl index bb38fa5..91bffc4 100644 --- a/examples/custom_tasks/tasks.jsonl +++ b/examples/custom_tasks/tasks.jsonl @@ -1,3 +1,4 @@ -{"task_id": "fizzbuzz", "prompt": "Create solution.py defining fizzbuzz(n) that returns a list of strings for 1..n using the classic FizzBuzz rules ('Fizz' for multiples of 3, 'Buzz' for 5, 'FizzBuzz' for both, else the number as a string). Make `python3 tests.py` pass.", "files": {"tests.py": "from solution import fizzbuzz\nout = fizzbuzz(15)\nassert out[0] == '1' and out[2] == 'Fizz' and out[4] == 'Buzz' and out[14] == 'FizzBuzz'\nassert len(out) == 15\nprint('OK')\n"}, "eval": {"command": "python3 tests.py", "timeout_sec": 60}} -{"task_id": "reverse-words", "prompt": "Create solution.py defining reverse_words(s) that reverses the order of words in a whitespace-separated string, collapsing runs of whitespace to single spaces. Make `python3 tests.py` pass.", "files": {"tests.py": "from solution import reverse_words\nassert reverse_words('the sky is blue') == 'blue is sky the'\nassert reverse_words(' hello world ') == 'world hello'\nprint('OK')\n"}, "eval": {"command": "python3 tests.py", "timeout_sec": 60}} -{"task_id": "anagram-groups", "prompt": "Create solution.py defining group_anagrams(words) that groups a list of lowercase words into anagram groups, returned as a list of sorted lists, sorted by each group's first word. Make `python3 tests.py` pass.", "files": {"tests.py": "from solution import group_anagrams\nassert group_anagrams(['eat','tea','tan','ate','nat','bat']) == [['ate','eat','tea'], ['bat'], ['nat','tan']]\nprint('OK')\n"}, "eval": {"command": "python3 tests.py", "timeout_sec": 60}} +{"task_id": "calc-eval", "prompt": "Create solution.py defining calc(s) that evaluates an arithmetic expression string and returns an int. The expression contains non-negative integer literals, the binary operators + - * /, unary minus (e.g. '-3', '2 - -3', '3*-2'), parentheses, and arbitrary spaces. Standard precedence: * and / bind tighter than + and -, evaluated left-to-right within the same precedence. '/' is integer division whose result is the mathematical quotient rounded TOWARD ZERO (so 7/-2 == -3 and -7/2 == -3, NOT Python floor division, which would give -4). A basic tests.py is provided in repo/ for your own checking, but a hidden grader runs additional tests \u2014 self-verify thoroughly against the full specification (especially division with negative operands), not just the sample tests. Make `python3 tests.py` pass.", "files": {"tests.py": "from solution import calc\nassert calc('3+2*2') == 7\nassert calc('2*3+4') == 10\nassert calc('(2+3)*4') == 20\nassert calc('10/3') == 3\nassert calc('100/(2*5)') == 10\nassert calc('2+3*4') == 14\nassert calc('3/2') == 1\nprint('OK')\n"}, "eval": {"command": "python3 -c \"import solution; c=solution.calc; assert c('3+2*2')==7; assert c(' 3/2 ')==1; assert c(' 3+5 / 2 ')==5; assert c('2*3+4')==10; assert c('2+3*4-6/2')==11; assert c('(2+3)*4')==20; assert c('7/-2')==-3; assert c('-7/2')==-3; assert c('-(3+4)')==-7; assert c('2- -3')==5; assert c('10/3')==3; assert c('-10/3')==-3; assert c('100/(2*5)')==10; assert c('3*-2+10')==4; assert c('((1+2)*(3+4))/5')==4; assert c('0-2*2')==-4; assert c('14-3/2')==13; assert c('2*-3*-4')==24; print('OK')\"", "timeout_sec": 60}} +{"task_id": "range-queries", "prompt": "Create solution.py defining process(n, ops). You manage an integer array `a` of length n, initially all zeros. `ops` is a list of operations applied in order, each a tuple:\n - ('set', i, v): assign a[i] = v (0 <= i < n, v may be negative)\n - ('sum', l, r): query a[l] + a[l+1] + ... + a[r], inclusive (0 <= l <= r < n)\nReturn a list with the answer to each ('sum', ...) query, in the order the queries appear. Both n and the number of operations can be as large as 200,000, and the hidden grader runs under a strict time budget \u2014 an O(n) scan per query (e.g. sum(a[l:r+1])) is far too slow and scores 0. Use a data structure that answers updates and range sums in about O(log n) each (a Fenwick / binary indexed tree or a segment tree). A basic tests.py with a tiny case is provided; make `python3 tests.py` pass, then make sure it scales.", "files": {"tests.py": "from solution import process\nout = process(5, [('set', 0, 3), ('set', 4, 10), ('sum', 0, 4),\n ('set', 2, 5), ('sum', 1, 3), ('sum', 4, 4)])\nassert list(out) == [13, 5, 10], out\nprint('OK')\n"}, "eval": {"command": "python3 -c \"$(cat <<'PYEOF'\nimport random, signal, solution\nrng = random.Random(12345)\nn = 200000\nops = []\nfor _ in range(200000):\n if rng.random() < 0.5:\n ops.append(('set', rng.randrange(n), rng.randrange(-1000, 1000)))\n else:\n l = rng.randrange(n); r = rng.randrange(l, n)\n ops.append(('sum', l, r))\n# expected answers via Fenwick\ntree = [0] * (n + 1); cur = [0] * n\ndef upd(i, d):\n i += 1\n while i <= n:\n tree[i] += d; i += i & (-i)\ndef pre(i):\n s = 0\n while i > 0:\n s += tree[i]; i -= i & (-i)\n return s\nexpected = []\nfor op in ops:\n if op[0] == 'set':\n _, i, v = op; upd(i, v - cur[i]); cur[i] = v\n else:\n _, l, r = op; expected.append(pre(r + 1) - pre(l))\nclass _TO(Exception): pass\ndef _h(s, f): raise _TO()\nsignal.signal(signal.SIGALRM, _h)\nsignal.setitimer(signal.ITIMER_REAL, 5.0)\ntry:\n got = solution.process(n, [tuple(o) for o in ops])\nfinally:\n signal.setitimer(signal.ITIMER_REAL, 0)\nassert list(got) == expected, 'wrong answers'\nprint('OK')\nPYEOF\n)\"", "timeout_sec": 60}} +{"task_id": "precise-sum", "prompt": "Create solution.py defining precise_sum(nums) that returns the sum of a list of Python floats as accurately as possible, minimizing floating-point rounding error. The input may mix very large and very small magnitudes (for example 1e16 alongside many 1.0 and -1e16 values), where a naive left-to-right accumulation silently loses the low-order terms. Implement the summation yourself with a compensated / error-tracking technique (e.g. Kahan-Babuska / Neumaier summation); do NOT call math.fsum or use numpy. A basic tests.py with benign sums is provided; a hidden grader checks accuracy on adversarial inputs against an exact reference, so self-verify against the full specification, not just the sample tests. Make `python3 tests.py` pass.", "files": {"tests.py": "import math\nfrom solution import precise_sum\nassert math.isclose(precise_sum([0.1, 0.2, 0.3]), 0.6, rel_tol=1e-12, abs_tol=1e-9)\nassert math.isclose(precise_sum([1.0, 2.0, 3.0, 4.0]), 10.0, rel_tol=1e-12, abs_tol=1e-9)\nassert math.isclose(precise_sum([1e6, 2e6, 3e6]), 6e6, rel_tol=1e-12, abs_tol=1e-9)\nprint('OK')\n"}, "eval": {"command": "python3 -c \"$(cat <<'PYEOF'\nimport math, random, solution\nrng = random.Random(777)\ncases = []\ncases.append([1e16, 1.0, -1e16] + [1.0] * 100)\ncases.append([1e18, 1.0, -1e18, 1.0])\nbig = []\nfor _ in range(2000):\n big += [1e15, 1.0, -1e15, 0.5, -0.25]\ncases.append(big)\ncases.append([10.0 ** e for e in range(-15, 16)] + [-(10.0 ** e) for e in range(-15, 16)] + [3.14])\nfor c in cases:\n want = math.fsum(c)\n got = solution.precise_sum(list(c))\n assert math.isclose(got, want, rel_tol=1e-12, abs_tol=1e-9), (got, want)\nprint('OK')\nPYEOF\n)\"", "timeout_sec": 60}} +{"task_id": "tsp-heuristic", "prompt": "Create solution.py defining solve(points) for the Euclidean Traveling Salesman Problem. `points` is a list of (x, y) float coordinate tuples. Return a tour: a list containing every index 0..len(points)-1 exactly once, giving the visiting order of a closed loop (it returns to the start after the last point). Your goal is to MINIMIZE the total Euclidean length of the closed tour. You are scored continuously in [0,1] as reference_length / your_length against a strong reference tour, averaged over several hidden point sets of 60-180 points, so a naive in-order tour scores well below 1.0 and better heuristics score higher \u2014 invest in local search (e.g. nearest-neighbor construction followed by 2-opt / or-opt improvement, ideally with multiple restarts) within the time limit. A basic tests.py is provided to check your output is a valid permutation.", "files": {"tests.py": "from solution import solve\npts=[(0,0),(0,1),(1,1),(1,0),(0.5,0.5)]\nt=solve(pts)\nassert sorted(t)==list(range(5)), 'must be a permutation of all indices'\nprint('OK')\n"}, "eval": {"command": "python3 -c \"$(cat <<'PYEOF'\nimport math, random, json, solution\nINST=[(0, 60), (1, 120), (2, 180)]\nREF={0:6195.97445814033,1:8743.12899328959,2:10615.721199788375}\ndef gp(seed,n):\n r=random.Random(1000+seed);return [(r.random()*1000,r.random()*1000) for _ in range(n)]\ndef d(a,b):return math.hypot(a[0]-b[0],a[1]-b[1])\ndef tl(p,t):return sum(d(p[t[i]],p[t[(i+1)%len(t)]]) for i in range(len(t)))\nratios=[]\nfor seed,n in INST:\n p=gp(seed,n)\n tour=list(solution.solve([tuple(x) for x in p]))\n assert sorted(tour)==list(range(n)),'not a valid permutation for seed '+str(seed)\n L=tl(p,tour)\n ratios.append(min(1.0, REF[seed]/L) if L>0 else 0.0)\nscore=sum(ratios)/len(ratios)\njson.dump({'score':score},open('score.json','w'))\nprint('score=%.4f ratios=%s'%(score,[round(r,3) for r in ratios]))\nPYEOF\n)\"", "timeout_sec": 120}} diff --git a/scripts/quickstart.sh b/scripts/quickstart.sh index 7a61b3d..df828b1 100755 --- a/scripts/quickstart.sh +++ b/scripts/quickstart.sh @@ -10,8 +10,9 @@ # from the environment); # - the ksi-agent:bench Docker image (built on first run if missing); # - the host runtime_runner Node dependencies. -# Then it runs 3 generations with the forums on and solved tasks retained, so -# the full knowledge loop (execute -> forum -> distill -> seed) fires end to end. +# Then it runs 3 generations with the forums on over four deliberately harder +# tasks (not all solvable in one shot), so the full knowledge loop +# (execute -> forum -> distill -> seed) fires end to end across generations. # # Usage: # ANTHROPIC_API_KEY=sk-ant-... bash scripts/quickstart.sh @@ -131,13 +132,15 @@ if [[ ! -f "$PROFILE" ]]; then exit 1 fi -# Full-loop demo: 3 generations with the per-task and cross-task forums on and -# solved tasks retained (--no-drop-solved), so every phase fires across -# generations — execute -> forum -> distill -> seed -> next generation. This -# takes longer than a single generation but exercises the whole -# knowledge-refinement loop rather than just the execution phase. (Without -# --no-drop-solved the demo tasks all solve on generation 1 and get dropped, -# so the run would stop before seeding ever fires.) +# Full-loop demo: 3 generations with the per-task and cross-task forums on, over +# four tasks deliberately pitched beyond a single-shot solve (an arithmetic +# parser with a truncation trap, a range-query task that needs a Fenwick tree, a +# numerically-stable summation, and a continuous-score TSP heuristic). Because +# they don't all solve on generation 1, unsolved tasks carry forward under the +# default --drop-solved and every phase fires across generations — +# execute -> forum -> distill -> seed -> next generation. This takes longer than +# a single generation but exercises the whole knowledge-refinement loop rather +# than just the execution phase. CMD=( "${PYRUN[@]}" -m ksi.cli --task-source custom @@ -147,12 +150,11 @@ CMD=( --generations 3 --per-task-forum-rounds 1 --cross-task-forum-rounds 1 - --no-drop-solved - --max-concurrent-tasks 3 + --max-concurrent-tasks 4 --experiment-name "$EXPERIMENT_NAME" ) -echo "==> Running quickstart demo (3 custom Python tasks: fizzbuzz, reverse-words, anagram-groups)" +echo "==> Running quickstart demo (4 harder Python tasks: calc-eval, range-queries, precise-sum, tsp-heuristic)" echo " ${CMD[*]}" if [[ "${DRY_RUN:-false}" == "true" ]]; then diff --git a/tests/test_custom_tasks_example.py b/tests/test_custom_tasks_example.py index daeea5b..e505456 100644 --- a/tests/test_custom_tasks_example.py +++ b/tests/test_custom_tasks_example.py @@ -7,27 +7,119 @@ def test_example_tasks_load_and_are_well_formed(): tasks = load_custom_tasks(EXAMPLE) - assert len(tasks) == 3 + assert len(tasks) == 4 for t in tasks: assert t.metadata["eval_command"].startswith("python3 ") +# Reference solutions for each demo task. These deliberately-hard tasks are not +# expected to be one-shot for a small model; the references here are the +# known-good solutions used to prove the graders are satisfiable end-to-end. +_REFERENCE_SOLUTIONS = { + # '/' truncates toward zero (int(a/b)), not Python floor division. + "calc-eval": r""" +import re +def calc(s): + toks = re.findall(r"\d+|[-+*/()]", s.replace(" ", "")) + pos = 0 + def peek(): + return toks[pos] if pos < len(toks) else None + def eat(): + nonlocal pos + t = toks[pos]; pos += 1; return t + def atom(): + t = peek() + if t == "(": + eat(); v = expr(); eat(); return v + if t == "-": + eat(); return -atom() + if t == "+": + eat(); return atom() + return int(eat()) + def term(): + v = atom() + while peek() in ("*", "/"): + op = eat(); r = atom(); v = v * r if op == "*" else int(v / r) + return v + def expr(): + v = term() + while peek() in ("+", "-"): + op = eat(); r = term(); v = v + r if op == "+" else v - r + return v + return expr() +""", + # Fenwick / binary indexed tree: O(log n) point-set and range-sum. + "range-queries": r""" +def process(n, ops): + tree = [0] * (n + 1) + cur = [0] * n + def upd(i, d): + i += 1 + while i <= n: + tree[i] += d + i += i & (-i) + def pre(i): + s = 0 + while i > 0: + s += tree[i] + i -= i & (-i) + return s + out = [] + for op in ops: + if op[0] == "set": + _, i, v = op + upd(i, v - cur[i]); cur[i] = v + else: + _, l, r = op + out.append(pre(r + 1) - pre(l)) + return out +""", + # Neumaier / Kahan-Babuska compensated summation. + "precise-sum": r""" +def precise_sum(nums): + s = 0.0 + c = 0.0 + for x in nums: + t = s + x + if abs(s) >= abs(x): + c += (s - t) + x + else: + c += (x - t) + s + s = t + return s + c +""", + # Nearest-neighbor tour: a valid permutation (score is continuous, so the + # grader only requires a valid tour to exit 0). + "tsp-heuristic": r""" +import math +def solve(points): + n = len(points) + if n <= 1: + return list(range(n)) + def d(a, b): + return math.hypot(a[0] - b[0], a[1] - b[1]) + unvisited = set(range(1, n)) + tour = [0] + cur = 0 + while unvisited: + nxt = min(unvisited, key=lambda j: d(points[cur], points[j])) + tour.append(nxt); unvisited.discard(nxt); cur = nxt + return tour +""", +} + + def test_example_evals_fail_on_starter_and_pass_on_reference(tmp_path): # The demo must be gradeable end-to-end without an agent: seed each # task's files, confirm the eval FAILS pre-solution, then write a # reference solution and confirm it PASSES. import subprocess - solutions = { - "fizzbuzz": "def fizzbuzz(n):\n return ['FizzBuzz' if i%15==0 else 'Fizz' if i%3==0 else 'Buzz' if i%5==0 else str(i) for i in range(1, n+1)]\n", - "reverse-words": "def reverse_words(s):\n return ' '.join(reversed(s.split()))\n", - "anagram-groups": "def group_anagrams(words):\n groups = {}\n for w in words:\n groups.setdefault(''.join(sorted(w)), []).append(w)\n return sorted([sorted(g) for g in groups.values()], key=lambda g: g[0])\n", - } for task in load_custom_tasks(EXAMPLE): seed = Path(task.metadata["repo_path"]) cmd = task.metadata["eval_command"] - pre = subprocess.run(cmd, shell=True, cwd=seed, capture_output=True, timeout=60) + pre = subprocess.run(cmd, shell=True, cwd=seed, capture_output=True, timeout=120) assert pre.returncode != 0, f"{task.id}: eval passed with no solution" - (seed / "solution.py").write_text(solutions[task.id], encoding="utf-8") - post = subprocess.run(cmd, shell=True, cwd=seed, capture_output=True, timeout=60) + (seed / "solution.py").write_text(_REFERENCE_SOLUTIONS[task.id], encoding="utf-8") + post = subprocess.run(cmd, shell=True, cwd=seed, capture_output=True, timeout=120) assert post.returncode == 0, f"{task.id}: reference solution failed: {post.stdout} {post.stderr}" From fef817225beca1790d67ec65d39dc3998970be72 Mon Sep 17 00:00:00 2001 From: xuefei-wang Date: Tue, 21 Jul 2026 23:32:57 -0700 Subject: [PATCH 3/3] feat(quickstart): run 3 hard ARC-AGI-1 tasks so the full loop fires reliably MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Replace the custom-task demo with 3 hard ARC-AGI-1 tasks. Custom tasks calibrated for one model don't transfer: a real gpt-5.4-mini run one-shot all four custom tasks (including the "never 1.0" TSP task, whose reference was beatable), so --drop-solved emptied the pool and the run stopped after generation 1. ARC tasks are hard for every current model, which fixes the calibration problem. - examples/quickstart/arc1_hard/: 3 ARC-1 tasks (97239e3d, d22278a0, 776ffc46) with low observed pass rates, copied from the vendored ARC-1 corpus (they span the eval/training splits, so a single --tasks-path dir holds all three). README documents the selection + pass rates. No dataset download. - quickstart.sh: default to --task-source arc over these 3 tasks; keep 3 generations / forums on / --drop-solved default. TASKS_PATH still switches to custom-task mode; ARC_DATA_DIR / ARC_TASK_MAP knobs added. - Revert the custom-task detour (examples/custom_tasks/tasks.jsonl and its test) back to main. - docs (getting-started, faq, README): reframe the demo around ARC, with real output/DB figures from a verified run. Verified end to end (gpt-5.4-mini): gen 1 solved 2/3, d22278a0 carried through gens 2-3; distillation ran each generation (per_task_distill=3, cross_task=1) and seeding fired between generations (seed_snapshots=2) — the full execute->forum->distill->seed loop runs across all three generations. Co-Authored-By: Claude Opus 4.8 (1M context) --- README.md | 12 +-- docs/faq.md | 4 +- docs/getting-started.md | 112 ++++++++++++-------- examples/custom_tasks/tasks.jsonl | 7 +- examples/quickstart/arc1_hard/776ffc46.json | 1 + examples/quickstart/arc1_hard/97239e3d.json | 1 + examples/quickstart/arc1_hard/README.md | 28 +++++ examples/quickstart/arc1_hard/d22278a0.json | 1 + scripts/quickstart.sh | 60 +++++++---- tests/test_custom_tasks_example.py | 110 ++----------------- 10 files changed, 158 insertions(+), 178 deletions(-) create mode 100644 examples/quickstart/arc1_hard/776ffc46.json create mode 100644 examples/quickstart/arc1_hard/97239e3d.json create mode 100644 examples/quickstart/arc1_hard/README.md create mode 100644 examples/quickstart/arc1_hard/d22278a0.json diff --git a/README.md b/README.md index 631313d..7b9f85f 100644 --- a/README.md +++ b/README.md @@ -41,12 +41,12 @@ bash scripts/quickstart.sh The script self-bootstraps everything it needs: it synthesizes a provider profile from your key, builds the `ksi-agent:bench` image on first run, -installs the host Node dependencies, then runs three generations over four -deliberately harder tasks in the bundled -[`examples/custom_tasks/`](./examples/custom_tasks/) demo — with the forums on, -so the full execute → forum → distill → seed loop fires. If -anything is missing, `uv run ksi-doctor` prints a ✓/✗ readiness checklist -with the exact command to fix it. +installs the host Node dependencies, then runs three generations over three +hard [ARC-AGI-1](./examples/quickstart/arc1_hard/) tasks (bundled, no download) +— with the forums on, so the full execute → forum → distill → seed loop fires. +ARC tasks are hard for every current model, so they don't all solve on the first +generation and the loop keeps going. If anything is missing, `uv run ksi-doctor` +prints a ✓/✗ readiness checklist with the exact command to fix it. ## Documentation diff --git a/docs/faq.md b/docs/faq.md index 76bc22d..c3cbece 100644 --- a/docs/faq.md +++ b/docs/faq.md @@ -83,8 +83,8 @@ Copy the template for your provider, fill in your key, and pass the path with Every run makes real LLM API calls billed to the key in your provider profile; there is no built-in spending cap. Cost scales with the number of tasks, generations, and the model you choose. To get a feel before committing, -start with the bundled synthetic demo: `bash scripts/quickstart.sh` runs -four harder tasks across three generations with Haiku — the cheapest way to +start with the bundled demo: `bash scripts/quickstart.sh` runs three hard +ARC-AGI-1 tasks across three generations with Haiku — the cheapest way to watch the full loop. Use `DRY_RUN=true` on any experiment wrapper script to print the full CLI command and DB paths without launching anything. diff --git a/docs/getting-started.md b/docs/getting-started.md index 0cd408a..8287857 100644 --- a/docs/getting-started.md +++ b/docs/getting-started.md @@ -1,14 +1,15 @@ # Getting started -Go from a fresh clone to a solved demo task in one command, then learn what just happened. +Go from a fresh clone to the full knowledge loop running in one command, then learn what just happened. ## What you'll do -Run three generations of agents against four bundled, self-contained tasks — pitched hard enough -that they don't all solve on the first try — and watch the full knowledge loop — execute, discuss, -distill, seed — fire end to end. No dataset download, no manual setup. The demo takes several -minutes and leaves you with a working environment ready to run your own tasks or a reference -benchmark. +Run three generations of agents against five bundled **ARC-AGI-1** tasks — hard enough for any +current model that they don't all solve on the first try — and watch the full knowledge loop — +execute, discuss, distill, seed — fire end to end. No dataset download, no manual setup (the ARC-1 +tasks are vendored under `benchmarks/arc1/`). ARC attempts are slow, so a full three-generation +run takes on the order of 15–20 minutes; it leaves you with a working environment ready to run your +own tasks or a reference benchmark. ## Prerequisites @@ -26,11 +27,14 @@ bash scripts/quickstart.sh The script self-bootstraps everything it needs: it synthesizes a provider profile from your key, builds the `ksi-agent:bench` image on first run (this takes a few minutes), installs the host -Node dependencies, then runs three generations over the four bundled tasks under -[`examples/custom_tasks/`](https://github.com/recursive-knowledge/KSI/tree/main/examples/custom_tasks) -(`calc-eval`, `range-queries`, `precise-sum`, `tsp-heuristic`) — each graded on the host by the -task's own eval command, with the per-task and cross-task forums on so every phase of the loop -fires. +Node dependencies, then runs three generations over three hard **ARC-AGI-1** tasks — bundled under +[`examples/quickstart/arc1_hard/`](https://github.com/recursive-knowledge/KSI/tree/main/examples/quickstart/arc1_hard) +(copied from the ARC-1 corpus vendored under `benchmarks/arc1/`, no download) — with the per-task +and cross-task forums on so every phase of the loop fires. Each agent studies an ARC task's +input→output training examples and writes its predicted grids, scored by the `arc_session` +evaluator (exact-match, up to two attempts per test). These three tasks are chosen for being hard +for current models (see the [directory README](https://github.com/recursive-knowledge/KSI/blob/main/examples/quickstart/arc1_hard/README.md) +for their observed pass rates), so they don't all solve on the first generation. For the complete benchmark environment (including benchmark preparation and smoke tests), run `bash scripts/setup_all.sh`. Use `--no-test` when you need @@ -64,14 +68,15 @@ The run logs each attempt and its score as it progresses. When it finishes, resu | Execution traces | `analysis/traces//` | For the quickstart, `` defaults to `quickstart_demo`. The run logs -each task's score as it goes, and each generation ends with a -`completed … solved=N/M` line — three in all, one per generation. These tasks -are meant to be hard, so expect only some to solve on generation 1, with the -solve count and the `tsp-heuristic` score generally improving across the three -generations as distilled knowledge accumulates (the hardest tasks may stay -unsolved — that's fine). **The signal that your environment is set up correctly -is that attempts run and get scored at all** — not that everything solves. -Elapsed times, token counts, and exact solve counts vary by model and run. +each task's score (ARC scoring is exact-match: `1.0` solved, `0.0` not) as it +goes, and ends with a single `completed traces=… tasks=… solved=N/M` summary +line. These ARC tasks are hard for current models, so expect at least one to +remain unsolved after generation 1 (a strong model may solve the rest); the +unsolved tasks carry forward and get re-attempted each generation, now seeded +with what the population distilled. **The signal that your environment is set up +correctly is that attempts run and get scored at all** — not that everything +solves. Whether an unsolved task flips to solved by generation 3 depends on the +model. Elapsed times, token counts, and solve counts vary by model and run. ??? note "A closer look — sample output, optional artifacts, and the knowledge DB" @@ -80,34 +85,49 @@ Elapsed times, token counts, and exact solve counts vary by model and run. run preset, which sets it for you) for a score summary on disk. Traces default to `analysis/traces//` — set `KSI_TRACE_DIR` to change the root. - An illustrative excerpt from a generation's execution phase against - `claude-haiku-4-5-20251001` (the default `configs/ksi/.env.haiku` profile), - timestamps trimmed — each generation logs a block like this, followed by the - forum and distillation phases. Scores are mixed on generation 1 by design - (`tsp-heuristic` is a continuous score, never a clean `1.0`): + An excerpt from a real run (`gpt-5.4-mini`, timestamps trimmed). ARC scores + are binary (exact-match); a strong model solves some tasks on generation 1 + while the hardest carry forward: ```text - INFO ksi.orchestrator.execution_phase: [gen 1] task=range-queries agent=agent-1 done elapsed=31.2s score=1.0000 - INFO ksi.orchestrator.execution_phase: [gen 1] task=tsp-heuristic agent=agent-3 done elapsed=44.7s score=0.7800 - INFO ksi.orchestrator.execution_phase: [gen 1] task=calc-eval agent=agent-0 done elapsed=28.1s score=0.0000 - INFO ksi.orchestrator.execution_phase: [gen 1] task=precise-sum agent=agent-2 done elapsed=26.5s score=0.0000 - INFO ksi.orchestrator.engine: completed traces=4 tasks=4 solved=1/4 (25.0%) + INFO ksi.orchestrator.execution_phase: [gen 1] task=776ffc46 agent=agent-0 done elapsed=281.6s score=1.0000 + INFO ksi.orchestrator.execution_phase: [gen 1] task=97239e3d agent=agent-1 done elapsed=281.6s score=1.0000 + INFO ksi.orchestrator.execution_phase: [gen 1] task=d22278a0 agent=agent-2 done elapsed=287.4s score=0.0000 + INFO ksi.orchestrator.distillation_phase: [ENGINE] distill gen=1: 1 per-task bundle(s), cross_task=0 + INFO ksi.orchestrator.persistence: [gen 2] start agents=1 ``` - **Knowledge DB check** — because the demo now runs the full loop, the - knowledge DB carries rows from every phase, not just execution. Group by - `entry_type` and `source_phase` to see them: + Here two tasks solved and dropped out (`--drop-solved`), while `d22278a0` + stayed unsolved and carried through generations 2 and 3 — each time + re-attempted with freshly distilled guidance seeded in. The run ended with + `completed traces=5 tasks=3 solved=2/3 (66.7%)`: five attempts across three + generations over three unique tasks. + + **Knowledge DB check** — because the demo runs the full loop, the knowledge + DB carries rows from every phase, not just execution. Group by `entry_type` + and `source_phase` to see them (real counts from the run above): ```console $ sqlite3 runtime_state/knowledge/quickstart_demo/quickstart_demo_knowledge.sqlite \ "select entry_type, source_phase, count(*) from knowledge group by entry_type, source_phase order by entry_type, source_phase;" + attempt|execution|5 + distillation|cross_task_distill|1 + distillation|per_task_distill|3 + insight|execution|5 + post|cross_task_forum|1 + ``` + + The `distillation` rows (one per-task bundle per generation, plus a cross-task + bundle) confirm the distill phase ran each generation, and + + ```console + $ sqlite3 runtime_state/knowledge/quickstart_demo/quickstart_demo_knowledge.sqlite \ + "select count(*) from seed_snapshots;" + 2 ``` - You'll see `attempt` and `insight` rows from `execution`, `post` rows from - `per_task_forum` and `cross_task_forum` (the two discussion phases), and - `distillation` rows from `per_task_distill` and `cross_task_distill`. Exact - counts vary with the model and how much each agent posts, and they grow with - each of the three generations. + the two `seed_snapshots` confirm seeding fired between each pair of + generations. Exact counts vary with the model and how much each agent posts. ## What just happened? @@ -119,15 +139,15 @@ KSI runs a knowledge-refinement loop across generations: 4. The system [*distills*](glossary.md#distillation) those discussions into reusable guidance. 5. The next generation is [*seeded*](glossary.md#seeding) with that guidance. -!!! note "Why these tasks are hard on purpose" - Earlier versions of this demo used trivial tasks that every agent solved on - the first attempt — so with `--drop-solved` (on by default) the task pool - emptied after generation 1 and the run **stopped** before the forum, distill, - and seed phases could show their value. These four tasks are pitched beyond a - reliable one-shot solve — a truncation-toward-zero parser trap, a range-query - task that needs a Fenwick tree, numerically-stable summation, and a - continuous-score TSP heuristic that is effectively never "perfect" — so - unsolved tasks carry forward under the default `--drop-solved` and the full +!!! note "Why ARC tasks — and why hard ones" + The demo uses ARC-AGI-1 tasks because they're hard for every current model: + if the tasks were easy, every agent would solve them on the first attempt and + — with `--drop-solved` (on by default) — the task pool would empty after + generation 1, so the run would **stop** before the forum, distill, and seed + phases could show their value. The three bundled tasks + (`97239e3d`, `d22278a0`, `776ffc46`) are chosen for low observed pass rates + (see [`examples/quickstart/arc1_hard/`](https://github.com/recursive-knowledge/KSI/tree/main/examples/quickstart/arc1_hard)), + so unsolved tasks carry forward under the default `--drop-solved` and the full loop runs across all three generations. On your own *easy* tasks, expect the run to stop early once everything is solved; that's the intended behavior. See [experiments.md](experiments.md). diff --git a/examples/custom_tasks/tasks.jsonl b/examples/custom_tasks/tasks.jsonl index 91bffc4..bb38fa5 100644 --- a/examples/custom_tasks/tasks.jsonl +++ b/examples/custom_tasks/tasks.jsonl @@ -1,4 +1,3 @@ -{"task_id": "calc-eval", "prompt": "Create solution.py defining calc(s) that evaluates an arithmetic expression string and returns an int. The expression contains non-negative integer literals, the binary operators + - * /, unary minus (e.g. '-3', '2 - -3', '3*-2'), parentheses, and arbitrary spaces. Standard precedence: * and / bind tighter than + and -, evaluated left-to-right within the same precedence. '/' is integer division whose result is the mathematical quotient rounded TOWARD ZERO (so 7/-2 == -3 and -7/2 == -3, NOT Python floor division, which would give -4). A basic tests.py is provided in repo/ for your own checking, but a hidden grader runs additional tests \u2014 self-verify thoroughly against the full specification (especially division with negative operands), not just the sample tests. Make `python3 tests.py` pass.", "files": {"tests.py": "from solution import calc\nassert calc('3+2*2') == 7\nassert calc('2*3+4') == 10\nassert calc('(2+3)*4') == 20\nassert calc('10/3') == 3\nassert calc('100/(2*5)') == 10\nassert calc('2+3*4') == 14\nassert calc('3/2') == 1\nprint('OK')\n"}, "eval": {"command": "python3 -c \"import solution; c=solution.calc; assert c('3+2*2')==7; assert c(' 3/2 ')==1; assert c(' 3+5 / 2 ')==5; assert c('2*3+4')==10; assert c('2+3*4-6/2')==11; assert c('(2+3)*4')==20; assert c('7/-2')==-3; assert c('-7/2')==-3; assert c('-(3+4)')==-7; assert c('2- -3')==5; assert c('10/3')==3; assert c('-10/3')==-3; assert c('100/(2*5)')==10; assert c('3*-2+10')==4; assert c('((1+2)*(3+4))/5')==4; assert c('0-2*2')==-4; assert c('14-3/2')==13; assert c('2*-3*-4')==24; print('OK')\"", "timeout_sec": 60}} -{"task_id": "range-queries", "prompt": "Create solution.py defining process(n, ops). You manage an integer array `a` of length n, initially all zeros. `ops` is a list of operations applied in order, each a tuple:\n - ('set', i, v): assign a[i] = v (0 <= i < n, v may be negative)\n - ('sum', l, r): query a[l] + a[l+1] + ... + a[r], inclusive (0 <= l <= r < n)\nReturn a list with the answer to each ('sum', ...) query, in the order the queries appear. Both n and the number of operations can be as large as 200,000, and the hidden grader runs under a strict time budget \u2014 an O(n) scan per query (e.g. sum(a[l:r+1])) is far too slow and scores 0. Use a data structure that answers updates and range sums in about O(log n) each (a Fenwick / binary indexed tree or a segment tree). A basic tests.py with a tiny case is provided; make `python3 tests.py` pass, then make sure it scales.", "files": {"tests.py": "from solution import process\nout = process(5, [('set', 0, 3), ('set', 4, 10), ('sum', 0, 4),\n ('set', 2, 5), ('sum', 1, 3), ('sum', 4, 4)])\nassert list(out) == [13, 5, 10], out\nprint('OK')\n"}, "eval": {"command": "python3 -c \"$(cat <<'PYEOF'\nimport random, signal, solution\nrng = random.Random(12345)\nn = 200000\nops = []\nfor _ in range(200000):\n if rng.random() < 0.5:\n ops.append(('set', rng.randrange(n), rng.randrange(-1000, 1000)))\n else:\n l = rng.randrange(n); r = rng.randrange(l, n)\n ops.append(('sum', l, r))\n# expected answers via Fenwick\ntree = [0] * (n + 1); cur = [0] * n\ndef upd(i, d):\n i += 1\n while i <= n:\n tree[i] += d; i += i & (-i)\ndef pre(i):\n s = 0\n while i > 0:\n s += tree[i]; i -= i & (-i)\n return s\nexpected = []\nfor op in ops:\n if op[0] == 'set':\n _, i, v = op; upd(i, v - cur[i]); cur[i] = v\n else:\n _, l, r = op; expected.append(pre(r + 1) - pre(l))\nclass _TO(Exception): pass\ndef _h(s, f): raise _TO()\nsignal.signal(signal.SIGALRM, _h)\nsignal.setitimer(signal.ITIMER_REAL, 5.0)\ntry:\n got = solution.process(n, [tuple(o) for o in ops])\nfinally:\n signal.setitimer(signal.ITIMER_REAL, 0)\nassert list(got) == expected, 'wrong answers'\nprint('OK')\nPYEOF\n)\"", "timeout_sec": 60}} -{"task_id": "precise-sum", "prompt": "Create solution.py defining precise_sum(nums) that returns the sum of a list of Python floats as accurately as possible, minimizing floating-point rounding error. The input may mix very large and very small magnitudes (for example 1e16 alongside many 1.0 and -1e16 values), where a naive left-to-right accumulation silently loses the low-order terms. Implement the summation yourself with a compensated / error-tracking technique (e.g. Kahan-Babuska / Neumaier summation); do NOT call math.fsum or use numpy. A basic tests.py with benign sums is provided; a hidden grader checks accuracy on adversarial inputs against an exact reference, so self-verify against the full specification, not just the sample tests. Make `python3 tests.py` pass.", "files": {"tests.py": "import math\nfrom solution import precise_sum\nassert math.isclose(precise_sum([0.1, 0.2, 0.3]), 0.6, rel_tol=1e-12, abs_tol=1e-9)\nassert math.isclose(precise_sum([1.0, 2.0, 3.0, 4.0]), 10.0, rel_tol=1e-12, abs_tol=1e-9)\nassert math.isclose(precise_sum([1e6, 2e6, 3e6]), 6e6, rel_tol=1e-12, abs_tol=1e-9)\nprint('OK')\n"}, "eval": {"command": "python3 -c \"$(cat <<'PYEOF'\nimport math, random, solution\nrng = random.Random(777)\ncases = []\ncases.append([1e16, 1.0, -1e16] + [1.0] * 100)\ncases.append([1e18, 1.0, -1e18, 1.0])\nbig = []\nfor _ in range(2000):\n big += [1e15, 1.0, -1e15, 0.5, -0.25]\ncases.append(big)\ncases.append([10.0 ** e for e in range(-15, 16)] + [-(10.0 ** e) for e in range(-15, 16)] + [3.14])\nfor c in cases:\n want = math.fsum(c)\n got = solution.precise_sum(list(c))\n assert math.isclose(got, want, rel_tol=1e-12, abs_tol=1e-9), (got, want)\nprint('OK')\nPYEOF\n)\"", "timeout_sec": 60}} -{"task_id": "tsp-heuristic", "prompt": "Create solution.py defining solve(points) for the Euclidean Traveling Salesman Problem. `points` is a list of (x, y) float coordinate tuples. Return a tour: a list containing every index 0..len(points)-1 exactly once, giving the visiting order of a closed loop (it returns to the start after the last point). Your goal is to MINIMIZE the total Euclidean length of the closed tour. You are scored continuously in [0,1] as reference_length / your_length against a strong reference tour, averaged over several hidden point sets of 60-180 points, so a naive in-order tour scores well below 1.0 and better heuristics score higher \u2014 invest in local search (e.g. nearest-neighbor construction followed by 2-opt / or-opt improvement, ideally with multiple restarts) within the time limit. A basic tests.py is provided to check your output is a valid permutation.", "files": {"tests.py": "from solution import solve\npts=[(0,0),(0,1),(1,1),(1,0),(0.5,0.5)]\nt=solve(pts)\nassert sorted(t)==list(range(5)), 'must be a permutation of all indices'\nprint('OK')\n"}, "eval": {"command": "python3 -c \"$(cat <<'PYEOF'\nimport math, random, json, solution\nINST=[(0, 60), (1, 120), (2, 180)]\nREF={0:6195.97445814033,1:8743.12899328959,2:10615.721199788375}\ndef gp(seed,n):\n r=random.Random(1000+seed);return [(r.random()*1000,r.random()*1000) for _ in range(n)]\ndef d(a,b):return math.hypot(a[0]-b[0],a[1]-b[1])\ndef tl(p,t):return sum(d(p[t[i]],p[t[(i+1)%len(t)]]) for i in range(len(t)))\nratios=[]\nfor seed,n in INST:\n p=gp(seed,n)\n tour=list(solution.solve([tuple(x) for x in p]))\n assert sorted(tour)==list(range(n)),'not a valid permutation for seed '+str(seed)\n L=tl(p,tour)\n ratios.append(min(1.0, REF[seed]/L) if L>0 else 0.0)\nscore=sum(ratios)/len(ratios)\njson.dump({'score':score},open('score.json','w'))\nprint('score=%.4f ratios=%s'%(score,[round(r,3) for r in ratios]))\nPYEOF\n)\"", "timeout_sec": 120}} +{"task_id": "fizzbuzz", "prompt": "Create solution.py defining fizzbuzz(n) that returns a list of strings for 1..n using the classic FizzBuzz rules ('Fizz' for multiples of 3, 'Buzz' for 5, 'FizzBuzz' for both, else the number as a string). Make `python3 tests.py` pass.", "files": {"tests.py": "from solution import fizzbuzz\nout = fizzbuzz(15)\nassert out[0] == '1' and out[2] == 'Fizz' and out[4] == 'Buzz' and out[14] == 'FizzBuzz'\nassert len(out) == 15\nprint('OK')\n"}, "eval": {"command": "python3 tests.py", "timeout_sec": 60}} +{"task_id": "reverse-words", "prompt": "Create solution.py defining reverse_words(s) that reverses the order of words in a whitespace-separated string, collapsing runs of whitespace to single spaces. Make `python3 tests.py` pass.", "files": {"tests.py": "from solution import reverse_words\nassert reverse_words('the sky is blue') == 'blue is sky the'\nassert reverse_words(' hello world ') == 'world hello'\nprint('OK')\n"}, "eval": {"command": "python3 tests.py", "timeout_sec": 60}} +{"task_id": "anagram-groups", "prompt": "Create solution.py defining group_anagrams(words) that groups a list of lowercase words into anagram groups, returned as a list of sorted lists, sorted by each group's first word. Make `python3 tests.py` pass.", "files": {"tests.py": "from solution import group_anagrams\nassert group_anagrams(['eat','tea','tan','ate','nat','bat']) == [['ate','eat','tea'], ['bat'], ['nat','tan']]\nprint('OK')\n"}, "eval": {"command": "python3 tests.py", "timeout_sec": 60}} diff --git a/examples/quickstart/arc1_hard/776ffc46.json b/examples/quickstart/arc1_hard/776ffc46.json new file mode 100644 index 0000000..c3befd0 --- /dev/null +++ b/examples/quickstart/arc1_hard/776ffc46.json @@ -0,0 +1 @@ +{"train": [{"input": [[5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0], [5, 0, 0, 2, 0, 0, 5, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0], [5, 0, 2, 2, 2, 0, 5, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0], [5, 0, 0, 2, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], "output": [[5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0], [5, 0, 0, 2, 0, 0, 5, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0], [5, 0, 2, 2, 2, 0, 5, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0], [5, 0, 0, 2, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 2, 2, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 2, 2, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]]}, {"input": [[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 1, 1, 1, 0, 0], [0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], "output": [[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 3, 3, 3, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 3, 3, 3, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 3, 3, 3, 0, 0], [0, 0, 0, 3, 3, 3, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 3, 3, 3, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 3, 3, 3, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]]}, {"input": [[0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5], [0, 0, 2, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5], [0, 0, 2, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 5, 0, 2, 2, 2, 0, 5], [0, 0, 2, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 5, 0, 2, 2, 2, 0, 5], [0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5], [5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 1, 1, 1, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 1, 1, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], "output": [[0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5], [0, 0, 2, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5], [0, 0, 2, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 5, 0, 2, 2, 2, 0, 5], [0, 0, 2, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 5, 0, 2, 2, 2, 0, 5], [0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5], [5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 2, 2, 2, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 2, 2, 2, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 2, 2, 2, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 2, 2, 2, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]]}, {"input": [[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 3, 3, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 3, 3, 0, 0], [0, 5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 5, 0, 3, 3, 3, 3, 0], [0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0], [0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5], [0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0], [0, 5, 5, 5, 5, 5, 5, 5, 0, 0, 1, 1, 1, 0, 0, 0, 1, 1, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 1, 1, 1, 1, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], "output": [[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 3, 3, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 3, 3, 0, 0], [0, 5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 5, 0, 3, 3, 3, 3, 0], [0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0], [0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5], [0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 5, 0, 3, 3, 3, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0], [0, 5, 5, 5, 5, 5, 5, 5, 0, 0, 3, 3, 3, 0, 0, 0, 1, 1, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 3, 3, 3, 0, 0, 1, 1, 1, 1, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 3, 3, 3, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]]}], "test": [{"input": [[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 2, 2, 2, 0, 0], [0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 5, 0, 0, 2, 2, 2, 0, 0], [0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5, 5], [0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 1, 0, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0], [0, 1, 1, 1, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0], [0, 0, 1, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5, 0], [0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0], [0, 0, 2, 2, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 2, 0, 0, 5, 0], [0, 2, 2, 2, 2, 0, 5, 0, 0, 0, 0, 0, 5, 0, 2, 2, 2, 0, 5, 0], [0, 2, 2, 2, 2, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 2, 0, 0, 5, 0], [0, 0, 2, 2, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0], [0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5, 0]], "output": [[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 2, 2, 2, 0, 0], [0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 5, 0, 0, 2, 2, 2, 0, 0], [0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5, 5], [0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 2, 0, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0], [0, 2, 2, 2, 0, 0, 1, 1, 1, 1, 0, 0, 0, 0, 2, 2, 2, 0, 0, 0], [0, 0, 2, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [5, 5, 5, 5, 5, 5, 5, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5, 0], [0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0], [0, 0, 2, 2, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 2, 0, 0, 5, 0], [0, 2, 2, 2, 2, 0, 5, 0, 0, 0, 0, 0, 5, 0, 2, 2, 2, 0, 5, 0], [0, 2, 2, 2, 2, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 2, 0, 0, 5, 0], [0, 0, 2, 2, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 0], [0, 0, 0, 0, 0, 0, 5, 0, 0, 0, 0, 0, 5, 5, 5, 5, 5, 5, 5, 0]]}]} \ No newline at end of file diff --git a/examples/quickstart/arc1_hard/97239e3d.json b/examples/quickstart/arc1_hard/97239e3d.json new file mode 100644 index 0000000..6513de7 --- /dev/null +++ b/examples/quickstart/arc1_hard/97239e3d.json @@ -0,0 +1 @@ +{"train": [{"input": [[2, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], "output": [[2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 0, 0, 0, 0], [2, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 2, 8, 8, 8, 0], [2, 8, 2, 8, 0, 8, 2, 8, 0, 8, 2, 8, 2, 8, 0, 8, 0], [2, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 2, 8, 8, 8, 0], [2, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0], [2, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 2, 8, 8, 8, 0], [2, 8, 2, 8, 0, 8, 2, 8, 0, 8, 2, 8, 2, 8, 0, 8, 0], [2, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 2, 8, 8, 8, 0], [2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]]}, {"input": [[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 6, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 6], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 1, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], "output": [[0, 0, 0, 0, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6], [0, 8, 8, 8, 6, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 6], [0, 8, 0, 8, 6, 8, 6, 8, 0, 8, 6, 8, 0, 8, 6, 8, 6], [0, 8, 8, 8, 6, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 6], [0, 0, 0, 0, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0], [1, 8, 8, 8, 0, 8, 8, 8, 1, 8, 8, 8, 0, 8, 8, 8, 0], [1, 8, 1, 8, 0, 8, 1, 8, 1, 8, 0, 8, 0, 8, 0, 8, 0], [1, 8, 8, 8, 0, 8, 8, 8, 1, 8, 8, 8, 0, 8, 8, 8, 0], [1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0]]}, {"input": [[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 7, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 7], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 3, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 3, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], "output": [[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 7, 7, 7, 7, 7], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 7, 8, 8, 8, 7], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 7, 8, 7, 8, 7], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 7, 8, 8, 8, 7], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 7, 7, 7, 7, 7], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 0, 0, 0, 0], [3, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 3, 8, 8, 8, 0], [3, 8, 3, 8, 0, 8, 3, 8, 0, 8, 3, 8, 3, 8, 0, 8, 0], [3, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 3, 8, 8, 8, 0], [3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]]}], "test": [{"input": [[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 4, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 4, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 2, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 2, 8, 0, 8, 2, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 2, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], "output": [[0, 0, 0, 0, 4, 4, 4, 4, 4, 4, 4, 4, 4, 0, 0, 0, 0], [0, 8, 8, 8, 4, 8, 8, 8, 0, 8, 8, 8, 4, 8, 8, 8, 0], [0, 8, 0, 8, 4, 8, 4, 8, 0, 8, 4, 8, 4, 8, 0, 8, 0], [0, 8, 8, 8, 4, 8, 8, 8, 0, 8, 8, 8, 4, 8, 8, 8, 0], [0, 0, 0, 0, 4, 0, 0, 0, 0, 0, 0, 0, 4, 0, 0, 0, 0], [0, 8, 8, 8, 4, 8, 8, 8, 0, 8, 8, 8, 4, 8, 8, 8, 0], [0, 8, 0, 8, 4, 8, 4, 8, 0, 8, 4, 8, 4, 8, 0, 8, 0], [0, 8, 8, 8, 4, 8, 8, 8, 0, 8, 8, 8, 4, 8, 8, 8, 0], [0, 0, 0, 0, 4, 4, 4, 4, 4, 4, 4, 4, 4, 0, 0, 0, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0], [0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0, 8, 8, 8, 0], [0, 0, 0, 0, 2, 2, 2, 2, 2, 2, 2, 2, 2, 0, 0, 0, 0], [0, 8, 8, 8, 2, 8, 8, 8, 0, 8, 8, 8, 2, 8, 8, 8, 0], [0, 8, 0, 8, 2, 8, 2, 8, 0, 8, 2, 8, 2, 8, 0, 8, 0], [0, 8, 8, 8, 2, 8, 8, 8, 0, 8, 8, 8, 2, 8, 8, 8, 0], [0, 0, 0, 0, 2, 2, 2, 2, 2, 2, 2, 2, 2, 0, 0, 0, 0]]}]} \ No newline at end of file diff --git a/examples/quickstart/arc1_hard/README.md b/examples/quickstart/arc1_hard/README.md new file mode 100644 index 0000000..510f942 --- /dev/null +++ b/examples/quickstart/arc1_hard/README.md @@ -0,0 +1,28 @@ +# Quickstart hard ARC-AGI-1 tasks + +Three ARC-AGI-1 tasks used by `scripts/quickstart.sh` to demonstrate the full +knowledge-refinement loop (execute → forum → distill → seed) across generations. + +They are chosen to be **hard for current models** — none is reliably one-shot, +so with `--drop-solved` (on by default) the unsolved tasks carry forward and the +forum/distill/seed phases actually run across all three generations, instead of +the pool emptying after generation 1. + +| task | ARC-1 split | observed pass rate (single attempt) | +|------|-------------|-------------------------------------| +| `97239e3d` | evaluation | GPT ~0/22, Haiku ~4/14 | +| `d22278a0` | training | GPT ~0/4, Haiku ~1/6 | +| `776ffc46` | training | GPT ~0/4, Haiku ~1/6 | + +Pass rates are approximate, from internal runs; they show these tasks are +solvable *sometimes* (so the knowledge loop has something to learn and transfer) +but rarely on the first try. + +## Provenance + +The task JSONs are copied verbatim from the vendored ARC-AGI-1 corpus under +`benchmarks/arc1/source/data/{evaluation,training}/`, originally from +[`fchollet/ARC-AGI`](https://github.com/fchollet/ARC-AGI) (Apache-2.0). They are +duplicated here only so the quickstart can point at a single `--tasks-path` +directory holding all three (they span two splits), keeping the demo +self-contained with no dataset download. diff --git a/examples/quickstart/arc1_hard/d22278a0.json b/examples/quickstart/arc1_hard/d22278a0.json new file mode 100644 index 0000000..bc84f5b --- /dev/null +++ b/examples/quickstart/arc1_hard/d22278a0.json @@ -0,0 +1 @@ +{"train": [{"input": [[1, 0, 0, 0, 0, 0, 0, 0, 0, 2], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], "output": [[1, 0, 1, 0, 1, 2, 0, 2, 0, 2], [0, 0, 1, 0, 1, 2, 0, 2, 0, 0], [1, 1, 1, 0, 1, 2, 0, 2, 2, 2], [0, 0, 0, 0, 1, 2, 0, 0, 0, 0], [1, 1, 1, 1, 1, 2, 2, 2, 2, 2], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [1, 1, 1, 1, 1, 2, 2, 2, 2, 2], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [1, 1, 1, 1, 1, 2, 2, 2, 2, 2], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0]]}, {"input": [[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 3], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [8, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], "output": [[0, 3, 0, 3, 0, 3, 0, 3, 0, 3, 0, 3], [8, 0, 0, 3, 0, 3, 0, 3, 0, 3, 0, 0], [0, 0, 0, 3, 0, 3, 0, 3, 0, 3, 3, 3], [8, 8, 8, 0, 0, 3, 0, 3, 0, 0, 0, 0], [0, 0, 0, 0, 0, 3, 0, 3, 3, 3, 3, 3], [8, 8, 8, 8, 8, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 3, 3, 3, 3, 3], [8, 8, 8, 8, 8, 0, 8, 0, 0, 0, 0, 0], [0, 0, 0, 0, 8, 0, 8, 0, 0, 3, 3, 3], [8, 8, 8, 0, 8, 0, 8, 0, 8, 0, 0, 0], [0, 0, 8, 0, 8, 0, 8, 0, 8, 0, 0, 3], [8, 0, 8, 0, 8, 0, 8, 0, 8, 0, 8, 0]]}, {"input": [[2, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [4, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]], "output": [[2, 0, 2, 0, 2, 0, 2, 0, 2, 0, 2, 0, 2], [0, 0, 2, 0, 2, 0, 2, 0, 2, 0, 2, 0, 2], [2, 2, 2, 0, 2, 0, 2, 0, 2, 0, 2, 0, 2], [0, 0, 0, 0, 2, 0, 2, 0, 2, 0, 2, 0, 2], [2, 2, 2, 2, 2, 0, 2, 0, 2, 0, 2, 0, 2], [0, 0, 0, 0, 0, 0, 2, 0, 2, 0, 2, 0, 2], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 4, 0, 4, 0, 4, 0, 4], [4, 4, 4, 4, 4, 0, 4, 0, 4, 0, 4, 0, 4], [0, 0, 0, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4], [4, 4, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4], [0, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4], [4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4]]}, {"input": [[1, 0, 0, 0, 0, 0, 2], [0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0], [8, 0, 0, 0, 0, 0, 0]], "output": [[1, 0, 1, 0, 2, 0, 2], [0, 0, 1, 0, 2, 0, 0], [1, 1, 1, 0, 2, 2, 2], [0, 0, 0, 0, 0, 0, 0], [8, 8, 8, 0, 0, 2, 2], [0, 0, 8, 0, 8, 0, 0], [8, 0, 8, 0, 8, 0, 0]]}], "test": [{"input": [[4, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [8, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1]], "output": [[4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 0], [0, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 0], [4, 4, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 0, 1, 1], [0, 0, 0, 0, 4, 0, 4, 0, 4, 0, 4, 0, 4, 0, 0, 0, 0], [4, 4, 4, 4, 4, 0, 4, 0, 4, 0, 4, 0, 0, 1, 1, 1, 1], [0, 0, 0, 0, 0, 0, 4, 0, 4, 0, 4, 0, 0, 0, 0, 0, 0], [4, 4, 4, 4, 4, 4, 4, 0, 4, 0, 0, 1, 1, 1, 1, 1, 1], [0, 0, 0, 0, 0, 0, 0, 0, 4, 0, 0, 0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1], [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], [8, 8, 8, 8, 8, 8, 8, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1], [0, 0, 0, 0, 0, 0, 8, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0], [8, 8, 8, 8, 8, 0, 8, 0, 0, 0, 1, 0, 1, 1, 1, 1, 1], [0, 0, 0, 0, 8, 0, 8, 0, 0, 0, 1, 0, 1, 0, 0, 0, 0], [8, 8, 8, 0, 8, 0, 8, 0, 0, 0, 1, 0, 1, 0, 1, 1, 1], [0, 0, 8, 0, 8, 0, 8, 0, 0, 0, 1, 0, 1, 0, 1, 0, 0], [8, 0, 8, 0, 8, 0, 8, 0, 0, 0, 1, 0, 1, 0, 1, 0, 1]]}]} diff --git a/scripts/quickstart.sh b/scripts/quickstart.sh index df828b1..0da8d9e 100755 --- a/scripts/quickstart.sh +++ b/scripts/quickstart.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# One command, fresh clone -> solved task on the bundled custom-tasks demo. +# One command, fresh clone -> the full knowledge loop on 5 hard ARC-AGI-1 tasks. # # No external dataset download and no prior setup_all.sh run needed. Given # Docker and Node.js 22.16.0 installed (and either uv or a local @@ -10,8 +10,10 @@ # from the environment); # - the ksi-agent:bench Docker image (built on first run if missing); # - the host runtime_runner Node dependencies. -# Then it runs 3 generations with the forums on over four deliberately harder -# tasks (not all solvable in one shot), so the full knowledge loop +# Then it runs 3 generations with the forums on over 3 hard ARC-AGI-1 tasks +# (bundled under examples/quickstart/arc1_hard/, no download). ARC tasks are hard +# for every current model, so they don't all solve on generation 1 — unsolved +# tasks carry forward under the default --drop-solved and the full knowledge loop # (execute -> forum -> distill -> seed) fires end to end across generations. # # Usage: @@ -20,7 +22,10 @@ # PROFILE=configs/ksi/.env.openai bash scripts/quickstart.sh # # Env knobs: -# TASKS_PATH= run your own tasks .jsonl/.json instead of the bundled demo +# TASKS_PATH= run your own custom tasks .jsonl/.json (command evaluator) +# instead of the default ARC-AGI-1 demo +# ARC_DATA_DIR= directory of ARC task JSONs (default: the bundled 3 tasks) +# ARC_TASK_MAP=

optional ARC task-map JSON to filter/pin the selection # EXPERIMENT_NAME=x name the run (default: quickstart_demo) # PROFILE= provider profile to use (default: configs/ksi/.env.haiku) # SKIP_BOOTSTRAP=1 don't build the image / install deps / synthesize a profile @@ -35,10 +40,16 @@ KSI_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)" cd "$KSI_ROOT" PROFILE="${PROFILE:-configs/ksi/.env.haiku}" -TASKS_PATH="${TASKS_PATH:-examples/custom_tasks/tasks.jsonl}" EXPERIMENT_NAME="${EXPERIMENT_NAME:-quickstart_demo}" AGENT_IMAGE="ksi-agent:bench" +# Demo task selection. Default: 3 hard ARC-AGI-1 tasks bundled under +# examples/quickstart/arc1_hard/ (no download). Set TASKS_PATH to a custom +# .jsonl/.json to run your own tasks through the command evaluator instead. +TASKS_PATH="${TASKS_PATH:-}" +ARC_DATA_DIR="${ARC_DATA_DIR:-examples/quickstart/arc1_hard}" +ARC_TASK_MAP="${ARC_TASK_MAP:-}" + # Prefer uv, but fall back to a plain interpreter so a `pip install`d package # (no uv) still works. uv is a convenience here, not a hard requirement. if command -v uv >/dev/null 2>&1; then @@ -132,29 +143,40 @@ if [[ ! -f "$PROFILE" ]]; then exit 1 fi -# Full-loop demo: 3 generations with the per-task and cross-task forums on, over -# four tasks deliberately pitched beyond a single-shot solve (an arithmetic -# parser with a truncation trap, a range-query task that needs a Fenwick tree, a -# numerically-stable summation, and a continuous-score TSP heuristic). Because -# they don't all solve on generation 1, unsolved tasks carry forward under the -# default --drop-solved and every phase fires across generations — -# execute -> forum -> distill -> seed -> next generation. This takes longer than -# a single generation but exercises the whole knowledge-refinement loop rather -# than just the execution phase. +# Full-loop demo: 3 generations with the per-task and cross-task forums on. The +# default task set is 3 hard ARC-AGI-1 tasks — hard enough for any current model +# that they don't all solve on generation 1, so unsolved tasks carry forward +# under the default --drop-solved and every phase fires across generations +# (execute -> forum -> distill -> seed -> next generation). Setting TASKS_PATH +# switches to a custom .jsonl/.json graded by the command evaluator. +if [[ -n "$TASKS_PATH" ]]; then + TASK_FLAGS=(--task-source custom --tasks-path "$TASKS_PATH" --evaluator command) + TASK_BANNER="custom tasks from $TASKS_PATH" + CONCURRENCY=4 +else + TASK_FLAGS=( + --task-source arc + --tasks-path "$ARC_DATA_DIR" + --evaluator arc_session + --arc-max-trials 2 + ) + [[ -n "$ARC_TASK_MAP" ]] && TASK_FLAGS+=(--task-map-path "$ARC_TASK_MAP") + TASK_BANNER="3 hard ARC-AGI-1 tasks" + CONCURRENCY=3 +fi + CMD=( "${PYRUN[@]}" -m ksi.cli - --task-source custom - --tasks-path "$TASKS_PATH" - --evaluator command + "${TASK_FLAGS[@]}" --provider-profile "$PROFILE" --generations 3 --per-task-forum-rounds 1 --cross-task-forum-rounds 1 - --max-concurrent-tasks 4 + --max-concurrent-tasks "$CONCURRENCY" --experiment-name "$EXPERIMENT_NAME" ) -echo "==> Running quickstart demo (4 harder Python tasks: calc-eval, range-queries, precise-sum, tsp-heuristic)" +echo "==> Running quickstart demo ($TASK_BANNER)" echo " ${CMD[*]}" if [[ "${DRY_RUN:-false}" == "true" ]]; then diff --git a/tests/test_custom_tasks_example.py b/tests/test_custom_tasks_example.py index e505456..daeea5b 100644 --- a/tests/test_custom_tasks_example.py +++ b/tests/test_custom_tasks_example.py @@ -7,119 +7,27 @@ def test_example_tasks_load_and_are_well_formed(): tasks = load_custom_tasks(EXAMPLE) - assert len(tasks) == 4 + assert len(tasks) == 3 for t in tasks: assert t.metadata["eval_command"].startswith("python3 ") -# Reference solutions for each demo task. These deliberately-hard tasks are not -# expected to be one-shot for a small model; the references here are the -# known-good solutions used to prove the graders are satisfiable end-to-end. -_REFERENCE_SOLUTIONS = { - # '/' truncates toward zero (int(a/b)), not Python floor division. - "calc-eval": r""" -import re -def calc(s): - toks = re.findall(r"\d+|[-+*/()]", s.replace(" ", "")) - pos = 0 - def peek(): - return toks[pos] if pos < len(toks) else None - def eat(): - nonlocal pos - t = toks[pos]; pos += 1; return t - def atom(): - t = peek() - if t == "(": - eat(); v = expr(); eat(); return v - if t == "-": - eat(); return -atom() - if t == "+": - eat(); return atom() - return int(eat()) - def term(): - v = atom() - while peek() in ("*", "/"): - op = eat(); r = atom(); v = v * r if op == "*" else int(v / r) - return v - def expr(): - v = term() - while peek() in ("+", "-"): - op = eat(); r = term(); v = v + r if op == "+" else v - r - return v - return expr() -""", - # Fenwick / binary indexed tree: O(log n) point-set and range-sum. - "range-queries": r""" -def process(n, ops): - tree = [0] * (n + 1) - cur = [0] * n - def upd(i, d): - i += 1 - while i <= n: - tree[i] += d - i += i & (-i) - def pre(i): - s = 0 - while i > 0: - s += tree[i] - i -= i & (-i) - return s - out = [] - for op in ops: - if op[0] == "set": - _, i, v = op - upd(i, v - cur[i]); cur[i] = v - else: - _, l, r = op - out.append(pre(r + 1) - pre(l)) - return out -""", - # Neumaier / Kahan-Babuska compensated summation. - "precise-sum": r""" -def precise_sum(nums): - s = 0.0 - c = 0.0 - for x in nums: - t = s + x - if abs(s) >= abs(x): - c += (s - t) + x - else: - c += (x - t) + s - s = t - return s + c -""", - # Nearest-neighbor tour: a valid permutation (score is continuous, so the - # grader only requires a valid tour to exit 0). - "tsp-heuristic": r""" -import math -def solve(points): - n = len(points) - if n <= 1: - return list(range(n)) - def d(a, b): - return math.hypot(a[0] - b[0], a[1] - b[1]) - unvisited = set(range(1, n)) - tour = [0] - cur = 0 - while unvisited: - nxt = min(unvisited, key=lambda j: d(points[cur], points[j])) - tour.append(nxt); unvisited.discard(nxt); cur = nxt - return tour -""", -} - - def test_example_evals_fail_on_starter_and_pass_on_reference(tmp_path): # The demo must be gradeable end-to-end without an agent: seed each # task's files, confirm the eval FAILS pre-solution, then write a # reference solution and confirm it PASSES. import subprocess + solutions = { + "fizzbuzz": "def fizzbuzz(n):\n return ['FizzBuzz' if i%15==0 else 'Fizz' if i%3==0 else 'Buzz' if i%5==0 else str(i) for i in range(1, n+1)]\n", + "reverse-words": "def reverse_words(s):\n return ' '.join(reversed(s.split()))\n", + "anagram-groups": "def group_anagrams(words):\n groups = {}\n for w in words:\n groups.setdefault(''.join(sorted(w)), []).append(w)\n return sorted([sorted(g) for g in groups.values()], key=lambda g: g[0])\n", + } for task in load_custom_tasks(EXAMPLE): seed = Path(task.metadata["repo_path"]) cmd = task.metadata["eval_command"] - pre = subprocess.run(cmd, shell=True, cwd=seed, capture_output=True, timeout=120) + pre = subprocess.run(cmd, shell=True, cwd=seed, capture_output=True, timeout=60) assert pre.returncode != 0, f"{task.id}: eval passed with no solution" - (seed / "solution.py").write_text(_REFERENCE_SOLUTIONS[task.id], encoding="utf-8") - post = subprocess.run(cmd, shell=True, cwd=seed, capture_output=True, timeout=120) + (seed / "solution.py").write_text(solutions[task.id], encoding="utf-8") + post = subprocess.run(cmd, shell=True, cwd=seed, capture_output=True, timeout=60) assert post.returncode == 0, f"{task.id}: reference solution failed: {post.stdout} {post.stderr}"