From c1f3218ca558ab50f042e5c40e092cc419b876ce Mon Sep 17 00:00:00 2001 From: Lauren Hyoseo Yoon Date: Thu, 16 Jul 2026 10:24:44 -0700 Subject: [PATCH 1/2] docs: tighten new-user descriptions across landing, getting-started, and FAQ Make the first-touch copy more concise and friendlier: - Landing tagline and cards lead with what KSI does, not the framework label; drop wordy card fragments. - README intro matches the landing tagline's phrasing. - Getting started: split two run-on paragraphs (sample-output note, knowledge DB check) into shorter sentences. - FAQ: answer "in one sentence" as one sentence, de-jargon the "what problem" answer, and turn the "when NOT to use" wall of text into a scannable list. --- README.md | 10 ++++----- docs/faq.md | 47 ++++++++++++++++++++++------------------- docs/getting-started.md | 18 ++++++++-------- docs/index.md | 14 ++++++------ 4 files changed, 46 insertions(+), 43 deletions(-) diff --git a/README.md b/README.md index 517365b..1b3d376 100644 --- a/README.md +++ b/README.md @@ -6,11 +6,11 @@ [![License](https://img.shields.io/badge/license-Apache--2.0-blue)](./LICENSE) [![Python](https://img.shields.io/badge/python-3.12%2B-blue?logo=python&logoColor=white)](./pyproject.toml) -KSI runs a population of disposable agents on **your own tasks**, each -attempting independently inside a sandboxed container. They share what worked -in a structured forum, and the system distills that discussion into reusable -guidance that seeds the next generation. Improvement lives in a shared -knowledge store — not in any single agent — so it survives across runs. +KSI runs a population of disposable agents on **your own tasks**, each working +independently in a sandboxed container. They compare notes on what worked in a +structured forum, and the system distills that discussion into reusable guidance +that seeds the next generation. Improvement lives in a shared knowledge store — +not in any single agent — so it survives across runs. Point it at any JSON/JSONL file of task records — no benchmark dataset and no loader code required — or at the bundled reference benchmarks (ARC-AGI-1/2, diff --git a/docs/faq.md b/docs/faq.md index ab56376..3c8787d 100644 --- a/docs/faq.md +++ b/docs/faq.md @@ -5,37 +5,40 @@ Common questions about KSI — from first-time setup through research use. ## What is KSI in one sentence? KSI (Knowledge-centric Self-Improvement) is a benchmark framework that treats -agents as disposable workers and keeps improvement in a persistent shared -knowledge substrate rather than in any individual agent's memory or identity. -Transient agents run in sandboxed Docker containers, write attempts and forum -posts to SQLite-backed knowledge stores, distill reusable guidance, and seed -later generations from those distilled bundles. +agents as disposable workers and keeps improvement in a shared knowledge store +rather than in any single agent's memory. + +Under the hood: agents run in sandboxed Docker containers, write their attempts +and forum posts to SQLite knowledge stores, distill what worked into reusable +guidance, and seed later generations from it. ## What problem does it solve — when should I use it? -KSI addresses the question of where improvement should live in an agentic -system. Its core thesis is that improvement should reside in durable shared -knowledge rather than in any individual agent; that an agent's workstream -should be an on-demand capability, not a persistent identity; and that the -resulting knowledge should be model-agnostic and transferable across model -families and new tasks. Use it when you want to study or deploy self-improving -agents on structured benchmarks (coding, reasoning, dialogue) and need a -principled, reproducible substrate for that improvement. +KSI answers a design question: where should an agent system's improvement live? +Its bet is that improvement belongs in durable, shared knowledge — not in any +single agent — so it stays model-agnostic and transfers across model families +and new tasks. Reach for it when you want to study or run self-improving agents +on structured benchmarks (coding, reasoning, dialogue) and need a reproducible +place for that improvement to accumulate. ## When should I NOT use this? KSI's overhead — Docker sandboxing, a SQLite knowledge substrate, multi-phase forum/distill/seed generations — pays off when you're studying or running *multiple generations of self-improvement across a task population*. It's -probably the wrong tool if any of these apply: you need a single agent to -solve one task right now with no multi-run learning loop (a plain agent call -is simpler and faster); you don't have Docker available or can't run -containers in your environment; you need sub-second iteration on prompt/tool -changes (each attempt is a full containerized run); or your benchmark doesn't -fit the task/evaluator model (a `TaskSpec` in, an `EvalResult` out — see -[programmatic_api.md](programmatic_api.md)) and adapting it isn't worth the -investment. If you're unsure, the Quickstart in the README costs one Docker -image build and a few cents of API usage to try. +probably the wrong tool if: + +- you just need one agent to solve one task right now, with no learning loop (a + plain agent call is simpler and faster); +- you can't run Docker containers in your environment; +- you need sub-second iteration on prompt/tool changes (each attempt is a full + containerized run); or +- your benchmark doesn't fit the task/evaluator model (a `TaskSpec` in, an + `EvalResult` out — see [programmatic_api.md](programmatic_api.md)) and adapting + it isn't worth the effort. + +If you're unsure, the Quickstart in the README costs one Docker image build and +a few cents of API usage to try. ## How is this different from just running an agent in a loop? diff --git a/docs/getting-started.md b/docs/getting-started.md index 438c682..d30a7c1 100644 --- a/docs/getting-started.md +++ b/docs/getting-started.md @@ -70,11 +70,11 @@ The direct CLI defaults traces to `analysis/traces//`; set `KSI_TRACE_DIR` before launching if your environment needs a different trace root. -Here's a real excerpt (trimmed of full timestamps) from an actual quickstart run against -`claude-haiku-4-5-20251001` (the `configs/ksi/.env.haiku` profile the quickstart synthesizes by -default) — task names, `solved=3/3`, and the `[tokens]` line's fields are stable across runs, but -elapsed times and token counts are not, since they depend on the model and how the agent solves -each task: +Here's a real excerpt from a quickstart run against `claude-haiku-4-5-20251001` +(the default `configs/ksi/.env.haiku` profile), with timestamps trimmed. The task +names, `solved=3/3`, and the `[tokens]` fields stay the same across runs; the +elapsed times and token counts won't, since they depend on the model and how the +agent solves each task: ```text INFO ksi.orchestrator.execution_phase: [gen 1] task=reverse-words agent=agent-1 done elapsed=27.4s score=1.0000 @@ -85,10 +85,10 @@ INFO ksi.orchestrator.persistence: [tokens] total=418,329 cached_input=346,149 u ``` **Knowledge DB check** — every solved attempt writes an `entry_type='attempt'` -row alongside an `insight` row; with `--per-task-forum-rounds 0 ---cross-task-forum-rounds 0` (the quickstart's defaults, for speed) there are -no discussion posts, and with nothing unsolved and no cross-task posts in -this single-generation run, distillation has nothing to write either: +row plus an `insight` row. The quickstart turns off both forums for speed +(`--per-task-forum-rounds 0 --cross-task-forum-rounds 0`), so there are no +discussion posts — and with nothing unsolved in this single-generation run, +distillation has nothing to write either: ```console $ sqlite3 runtime_state/knowledge/quickstart_demo/quickstart_demo_knowledge.sqlite \ diff --git a/docs/index.md b/docs/index.md index 7bf7389..8987970 100644 --- a/docs/index.md +++ b/docs/index.md @@ -9,9 +9,9 @@ hide: # Knowledge-centric Self-Improvement

-A generational orchestration framework: a population of disposable agents -attempts **your own tasks**, discusses what worked, distills transferable -knowledge, and seeds the next generation with it. +A population of disposable agents attempts **your own tasks**, compares notes +on what worked, distills the lessons into reusable knowledge, and seeds the +next generation with it.

@@ -36,8 +36,7 @@ knowledge, and seeds the next generation with it. --- - Point KSI at any JSON/JSONL file of tasks — the record schema and the - `command` evaluator's scoring contract. + Point KSI at any JSON/JSONL file of tasks — no dataset or loader code needed. [:octicons-arrow-right-24: Bring your own tasks](your_own_tasks.md) @@ -71,8 +70,9 @@ knowledge, and seeds the next generation with it. --- - The four `register_*` extension seams, adding benchmarks and - evaluators, and the [improvement strategies](improvement_strategies.md). + Add a benchmark, evaluator, runtime, or + [improvement strategy](improvement_strategies.md) — one `register_*` call + each, no core edits. [:octicons-arrow-right-24: Extend KSI](extending.md) From 873be1252cbd6d56a07967140b0e2166e9e9863f Mon Sep 17 00:00:00 2001 From: Lauren Hyoseo Yoon Date: Thu, 16 Jul 2026 12:00:38 -0700 Subject: [PATCH 2/2] docs: smooth the bring-your-own-tasks data-prep read Make custom-task data prep easier to follow: - Record schema: lead with what's required vs. optional so the minimum is clear before the full annotated example. - repo/ contract: drop internal 'seam' jargon and simplify the phrasing. - 12-file capture caveat: split the wall of text into two paragraphs, surface the 12-file number, and list the skipped files cleanly. --- docs/your_own_tasks.md | 39 ++++++++++++++++++++------------------- 1 file changed, 20 insertions(+), 19 deletions(-) diff --git a/docs/your_own_tasks.md b/docs/your_own_tasks.md index 5d57d4d..dbc6760 100644 --- a/docs/your_own_tasks.md +++ b/docs/your_own_tasks.md @@ -7,6 +7,9 @@ file, and each record becomes one attempt. ## Record schema +Only `task_id` and `prompt` are required. Add an `eval` command to score the +attempt, and `files` (or `workspace_dir`) to hand the agent starting files: + ```jsonc { "task_id": "my-task-1", // required, unique @@ -36,25 +39,23 @@ starts the agent from an empty `repo/`. ## The workspace / `repo/` contract Whichever way you supply starting files, KSI seeds them into the agent's -workspace under a `repo/` directory before the attempt starts (the same seam -the benchmark task sources use). The agent is told in its prompt that -`repo/` holds the task's starting files and that it should create or edit -files there. After the attempt, the `command` evaluator (below) runs in a -post-attempt copy of that same directory. - -**Known limitation — workspace capture is capped at 12 files.** By default -KSI wipes the container workspace after each task -(`--wipe-workspace-per-task true`), so grading runs against a *captured* copy -of `repo/` rather than the live container filesystem. For a generic (non --benchmark) task source, that capture channel reads back at most 12 files -from the workspace, skipping anything named `score.json` or containing -`test` in its filename (plus `.pyc`/`__pycache__` noise and any single file -over 1 MB); a solution spread across more than 12 files silently -loses the extras from this channel. This is fine for small, self-contained -tasks (a script or two). For a larger multi-file solution, either point -`--wipe-workspace-per-task false` at your run so the evaluator sees the live -on-disk workspace instead, or design the eval command to check for what -matters rather than relying on every generated file surviving capture. +workspace under a `repo/` directory before the attempt starts. The agent's +prompt tells it that `repo/` holds the task's starting files and that it should +create or edit files there. After the attempt, the `command` evaluator (below) +runs against a copy of that same directory. + +**Known limitation — workspace capture is capped at 12 files.** By default KSI +wipes the container workspace after each task (`--wipe-workspace-per-task true`), +so grading runs against a *captured* copy of `repo/`, not the live container +filesystem. For a custom (non-benchmark) task, that capture reads back at most +**12 files** and skips anything named `score.json`, anything with `test` in its +filename, `.pyc`/`__pycache__` noise, and any single file over 1 MB. A solution +spread across more files silently loses the extras. + +This is fine for small, self-contained tasks (a script or two). For a larger +multi-file solution, either pass `--wipe-workspace-per-task false` so the +evaluator grades the live on-disk workspace, or write the eval command to check +what matters rather than relying on every generated file surviving capture. Note that `--wipe-workspace-per-task false` also bypasses the capture-path anti-tamper filtering: with the live workspace graded directly, an agent