Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,11 +6,11 @@
[![License](https://img.shields.io/badge/license-Apache--2.0-blue)](./LICENSE)
[![Python](https://img.shields.io/badge/python-3.12%2B-blue?logo=python&logoColor=white)](./pyproject.toml)

KSI runs a population of disposable agents on **your own tasks**, each
attempting independently inside a sandboxed container. They share what worked
in a structured forum, and the system distills that discussion into reusable
guidance that seeds the next generation. Improvement lives in a shared
knowledge store — not in any single agent — so it survives across runs.
KSI runs a population of disposable agents on **your own tasks**, each working
independently in a sandboxed container. They compare notes on what worked in a
structured forum, and the system distills that discussion into reusable guidance
that seeds the next generation. Improvement lives in a shared knowledge store —
not in any single agent — so it survives across runs.

Point it at any JSON/JSONL file of task records — no benchmark dataset and no
loader code required — or at the bundled reference benchmarks (ARC-AGI-1/2,
Expand Down
47 changes: 25 additions & 22 deletions docs/faq.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,37 +5,40 @@ Common questions about KSI — from first-time setup through research use.
## What is KSI in one sentence?

KSI (Knowledge-centric Self-Improvement) is a benchmark framework that treats
agents as disposable workers and keeps improvement in a persistent shared
knowledge substrate rather than in any individual agent's memory or identity.
Transient agents run in sandboxed Docker containers, write attempts and forum
posts to SQLite-backed knowledge stores, distill reusable guidance, and seed
later generations from those distilled bundles.
agents as disposable workers and keeps improvement in a shared knowledge store
rather than in any single agent's memory.

Under the hood: agents run in sandboxed Docker containers, write their attempts
and forum posts to SQLite knowledge stores, distill what worked into reusable
guidance, and seed later generations from it.

## What problem does it solve — when should I use it?

KSI addresses the question of where improvement should live in an agentic
system. Its core thesis is that improvement should reside in durable shared
knowledge rather than in any individual agent; that an agent's workstream
should be an on-demand capability, not a persistent identity; and that the
resulting knowledge should be model-agnostic and transferable across model
families and new tasks. Use it when you want to study or deploy self-improving
agents on structured benchmarks (coding, reasoning, dialogue) and need a
principled, reproducible substrate for that improvement.
KSI answers a design question: where should an agent system's improvement live?
Its bet is that improvement belongs in durable, shared knowledge — not in any
single agent — so it stays model-agnostic and transfers across model families
and new tasks. Reach for it when you want to study or run self-improving agents
on structured benchmarks (coding, reasoning, dialogue) and need a reproducible
place for that improvement to accumulate.

## When should I NOT use this?

KSI's overhead — Docker sandboxing, a SQLite knowledge substrate, multi-phase
forum/distill/seed generations — pays off when you're studying or running
*multiple generations of self-improvement across a task population*. It's
probably the wrong tool if any of these apply: you need a single agent to
solve one task right now with no multi-run learning loop (a plain agent call
is simpler and faster); you don't have Docker available or can't run
containers in your environment; you need sub-second iteration on prompt/tool
changes (each attempt is a full containerized run); or your benchmark doesn't
fit the task/evaluator model (a `TaskSpec` in, an `EvalResult` out — see
[programmatic_api.md](programmatic_api.md)) and adapting it isn't worth the
investment. If you're unsure, the Quickstart in the README costs one Docker
image build and a few cents of API usage to try.
probably the wrong tool if:

- you just need one agent to solve one task right now, with no learning loop (a
plain agent call is simpler and faster);
- you can't run Docker containers in your environment;
- you need sub-second iteration on prompt/tool changes (each attempt is a full
containerized run); or
- your benchmark doesn't fit the task/evaluator model (a `TaskSpec` in, an
`EvalResult` out — see [programmatic_api.md](programmatic_api.md)) and adapting
it isn't worth the effort.

If you're unsure, the Quickstart in the README costs one Docker image build and
a few cents of API usage to try.

## How is this different from just running an agent in a loop?

Expand Down
18 changes: 9 additions & 9 deletions docs/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -70,11 +70,11 @@ The direct CLI defaults traces to `analysis/traces/<experiment>/`; set
`KSI_TRACE_DIR` before launching if your environment needs a different trace
root.

Here's a real excerpt (trimmed of full timestamps) from an actual quickstart run against
`claude-haiku-4-5-20251001` (the `configs/ksi/.env.haiku` profile the quickstart synthesizes by
default) — task names, `solved=3/3`, and the `[tokens]` line's fields are stable across runs, but
elapsed times and token counts are not, since they depend on the model and how the agent solves
each task:
Here's a real excerpt from a quickstart run against `claude-haiku-4-5-20251001`
(the default `configs/ksi/.env.haiku` profile), with timestamps trimmed. The task
names, `solved=3/3`, and the `[tokens]` fields stay the same across runs; the
elapsed times and token counts won't, since they depend on the model and how the
agent solves each task:

```text
INFO ksi.orchestrator.execution_phase: [gen 1] task=reverse-words agent=agent-1 done elapsed=27.4s score=1.0000
Expand All @@ -85,10 +85,10 @@ INFO ksi.orchestrator.persistence: [tokens] total=418,329 cached_input=346,149 u
```

**Knowledge DB check** — every solved attempt writes an `entry_type='attempt'`
row alongside an `insight` row; with `--per-task-forum-rounds 0
--cross-task-forum-rounds 0` (the quickstart's defaults, for speed) there are
no discussion posts, and with nothing unsolved and no cross-task posts in
this single-generation run, distillation has nothing to write either:
row plus an `insight` row. The quickstart turns off both forums for speed
(`--per-task-forum-rounds 0 --cross-task-forum-rounds 0`), so there are no
discussion posts — and with nothing unsolved in this single-generation run,
distillation has nothing to write either:

```console
$ sqlite3 runtime_state/knowledge/quickstart_demo/quickstart_demo_knowledge.sqlite \
Expand Down
14 changes: 7 additions & 7 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,9 +9,9 @@ hide:
# Knowledge-centric Self-Improvement

<p class="ksi-tagline">
A generational orchestration framework: a population of disposable agents
attempts **your own tasks**, discusses what worked, distills transferable
knowledge, and seeds the next generation with it.
A population of disposable agents attempts **your own tasks**, compares notes
on what worked, distills the lessons into reusable knowledge, and seeds the
next generation with it.
</p>

<div class="ksi-cta" markdown>
Expand All @@ -36,8 +36,7 @@ knowledge, and seeds the next generation with it.

---

Point KSI at any JSON/JSONL file of tasks — the record schema and the
`command` evaluator's scoring contract.
Point KSI at any JSON/JSONL file of tasks — no dataset or loader code needed.

[:octicons-arrow-right-24: Bring your own tasks](your_own_tasks.md)

Expand Down Expand Up @@ -71,8 +70,9 @@ knowledge, and seeds the next generation with it.

---

The four `register_*` extension seams, adding benchmarks and
evaluators, and the [improvement strategies](improvement_strategies.md).
Add a benchmark, evaluator, runtime, or
[improvement strategy](improvement_strategies.md) — one `register_*` call
each, no core edits.

[:octicons-arrow-right-24: Extend KSI](extending.md)

Expand Down
39 changes: 20 additions & 19 deletions docs/your_own_tasks.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,9 @@ file, and each record becomes one attempt.

## Record schema

Only `task_id` and `prompt` are required. Add an `eval` command to score the
attempt, and `files` (or `workspace_dir`) to hand the agent starting files:

```jsonc
{
"task_id": "my-task-1", // required, unique
Expand Down Expand Up @@ -36,25 +39,23 @@ starts the agent from an empty `repo/`.
## The workspace / `repo/` contract

Whichever way you supply starting files, KSI seeds them into the agent's
workspace under a `repo/` directory before the attempt starts (the same seam
the benchmark task sources use). The agent is told in its prompt that
`repo/` holds the task's starting files and that it should create or edit
files there. After the attempt, the `command` evaluator (below) runs in a
post-attempt copy of that same directory.

**Known limitation — workspace capture is capped at 12 files.** By default
KSI wipes the container workspace after each task
(`--wipe-workspace-per-task true`), so grading runs against a *captured* copy
of `repo/` rather than the live container filesystem. For a generic (non
-benchmark) task source, that capture channel reads back at most 12 files
from the workspace, skipping anything named `score.json` or containing
`test` in its filename (plus `.pyc`/`__pycache__` noise and any single file
over 1 MB); a solution spread across more than 12 files silently
loses the extras from this channel. This is fine for small, self-contained
tasks (a script or two). For a larger multi-file solution, either point
`--wipe-workspace-per-task false` at your run so the evaluator sees the live
on-disk workspace instead, or design the eval command to check for what
matters rather than relying on every generated file surviving capture.
workspace under a `repo/` directory before the attempt starts. The agent's
prompt tells it that `repo/` holds the task's starting files and that it should
create or edit files there. After the attempt, the `command` evaluator (below)
runs against a copy of that same directory.

**Known limitation — workspace capture is capped at 12 files.** By default KSI
wipes the container workspace after each task (`--wipe-workspace-per-task true`),
so grading runs against a *captured* copy of `repo/`, not the live container
filesystem. For a custom (non-benchmark) task, that capture reads back at most
**12 files** and skips anything named `score.json`, anything with `test` in its
filename, `.pyc`/`__pycache__` noise, and any single file over 1 MB. A solution
spread across more files silently loses the extras.

This is fine for small, self-contained tasks (a script or two). For a larger
multi-file solution, either pass `--wipe-workspace-per-task false` so the
evaluator grades the live on-disk workspace, or write the eval command to check
what matters rather than relying on every generated file surviving capture.

Note that `--wipe-workspace-per-task false` also bypasses the capture-path
anti-tamper filtering: with the live workspace graded directly, an agent
Expand Down
Loading