Skip to content

Repository files navigation

KSI

KSI — Knowledge-centric Self-Improvement

Disposable agents attempt your tasks, share what worked, and distill knowledge that seeds the next attempt.

arXiv alphaXiv Docs Blog X thread License Python


KSI runs a population of disposable agents on your own tasks, each working independently in a sandboxed container. They compare notes on what worked in a structured forum, and the system distills that discussion into reusable guidance that seeds the next generation. Improvement lives in a shared knowledge store — not in any single agent — so it survives across runs.

Point it at any JSON/JSONL file of task records — no benchmark dataset and no loader code required — or at the bundled reference benchmarks (ARC-AGI-1/2, SWE-bench Pro, Polyglot, Terminal-Bench 2).

Quickstart

From a fresh clone to a solved task in one command — no dataset download, no prior setup step. With Docker and Node.js 22.16.0 installed (and either uv, or a local editable install via pip install -e .), just provide an API key:

export ANTHROPIC_API_KEY=sk-ant-...    # or: export OPENAI_API_KEY=sk-...
bash scripts/quickstart.sh

The script self-bootstraps everything it needs: it synthesizes a provider profile from your key, builds the ksi-agent:bench image on first run, installs the host Node dependencies, then runs three generations over three hard ARC-AGI-1 tasks (bundled, no download) — with the forums on, so the full execute → forum → distill → seed loop fires. ARC tasks are hard for every current model, so they don't all solve on the first generation and the loop keeps going. If anything is missing, uv run ksi-doctor prints a ✓/✗ readiness checklist with the exact command to fix it.

Documentation

The docs site is the canonical reference:

  • Getting started — full walkthrough: requirements, setup_all.sh, provider profiles, first run
  • Your own tasks — the task-record schema, the command evaluator's scoring contract, and the workspace/repo/ layout
  • Programmatic API — drive the same runs from Python with ksi.run(...), no CLI
  • Architecture — how a generation works: attempts → forum → distillation → seeding
  • Extending KSI — add a task source, evaluator, runtime, or improvement strategy
  • Benchmarks — dataset preparation and run presets for the reference benchmarks
  • FAQ

The same pages are browsable as Markdown under docs/ in this repo, and benchmark-specific setup lives in benchmarks/.

For the research behind KSI, see the paper on arXiv — and for the method narrative, results, and interactive knowledge dashboards, the paper page.

Citation

If you use KSI in your research, please cite:

@article{wang2026knowledge,
  title={Knowledge-Centric Self-Improvement},
  author={Wang, Xuefei and Yoon, Lauren Hyoseo and Qu, Chengrui and Wang, Amanda Zichang and Sehgal, Atharva and Mazumdar, Eric and Yue, Yisong},
  journal={arXiv preprint arXiv:2607.19592},
  year={2026}
}

Licensing

ksi's own code is licensed under Apache-2.0. Task-map manifests committed under benchmarks/*/task_maps/*.json are KSI-authored under the same license; the reference-benchmark datasets themselves are third-party and remain under their own upstream licenses — see benchmarks/README.md for sources and attribution.

About

No description, website, or topics provided.

Resources

Contributing

Stars

50 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages