KSI runs a population of disposable agents on your own tasks, each working independently in a sandboxed container. They compare notes on what worked in a structured forum, and the system distills that discussion into reusable guidance that seeds the next generation. Improvement lives in a shared knowledge store — not in any single agent — so it survives across runs.
Point it at any JSON/JSONL file of task records — no benchmark dataset and no loader code required — or at the bundled reference benchmarks (ARC-AGI-1/2, SWE-bench Pro, Polyglot, Terminal-Bench 2).
From a fresh clone to a solved task in one command — no dataset download,
no prior setup step. With Docker and Node.js 22.16.0 installed (and either
uv, or a local editable install via pip install -e .), just provide an
API key:
export ANTHROPIC_API_KEY=sk-ant-... # or: export OPENAI_API_KEY=sk-...
bash scripts/quickstart.shThe script self-bootstraps everything it needs: it synthesizes a provider
profile from your key, builds the ksi-agent:bench image on first run,
installs the host Node dependencies, then runs three generations over three
hard ARC-AGI-1 tasks (bundled, no download)
— with the forums on, so the full execute → forum → distill → seed loop fires.
ARC tasks are hard for every current model, so they don't all solve on the first
generation and the loop keeps going. If anything is missing, uv run ksi-doctor
prints a ✓/✗ readiness checklist with the exact command to fix it.
The docs site is the canonical reference:
- Getting started —
full walkthrough: requirements,
setup_all.sh, provider profiles, first run - Your own tasks —
the task-record schema, the
commandevaluator's scoring contract, and the workspace/repo/layout - Programmatic API —
drive the same runs from Python with
ksi.run(...), no CLI - Architecture — how a generation works: attempts → forum → distillation → seeding
- Extending KSI — add a task source, evaluator, runtime, or improvement strategy
- Benchmarks — dataset preparation and run presets for the reference benchmarks
- FAQ
The same pages are browsable as Markdown under docs/ in this
repo, and benchmark-specific setup lives in
benchmarks/.
For the research behind KSI, see the paper on arXiv — and for the method narrative, results, and interactive knowledge dashboards, the paper page.
If you use KSI in your research, please cite:
@article{wang2026knowledge,
title={Knowledge-Centric Self-Improvement},
author={Wang, Xuefei and Yoon, Lauren Hyoseo and Qu, Chengrui and Wang, Amanda Zichang and Sehgal, Atharva and Mazumdar, Eric and Yue, Yisong},
journal={arXiv preprint arXiv:2607.19592},
year={2026}
}ksi's own code is licensed under Apache-2.0. Task-map manifests
committed under benchmarks/*/task_maps/*.json are KSI-authored under the
same license; the reference-benchmark datasets themselves are third-party
and remain under their own upstream licenses — see
benchmarks/README.md for
sources and attribution.