From cd16b17daa09feb8e80f81acc166e12f2b40d499 Mon Sep 17 00:00:00 2001 From: Lauren Hyoseo Yoon Date: Wed, 22 Jul 2026 03:25:20 -0700 Subject: [PATCH 1/2] docs: fix stale file references and missing registry entries Correct documentation drift found against the current code: - adding_a_benchmark.md: built-in task sources register and set their loader/domain-hint in src/ksi/benchmarks/sources.py, not tasks/registry.py. - architecture.md: ARC runs natively (no snapshot mount, no ARC MCP server, no direct-ARC adapter); only forum tasks use the direct Anthropic adapter. - glossary.md: add the `command` evaluator and `custom` task source (both used by the quickstart). - adding_an_evaluator.md: Evaluator.evaluate returns dict[str, Any]. - custom_tasks/README.md: add the profile-creation step and Node.js prereq. - scripts/README.md: dataprep/arc_prep live under benchmarks/scripts/. --- docs/adding_a_benchmark.md | 10 +++++----- docs/adding_an_evaluator.md | 2 +- docs/architecture.md | 13 +++++++------ docs/glossary.md | 4 ++-- examples/custom_tasks/README.md | 10 ++++++++-- scripts/README.md | 2 +- 6 files changed, 24 insertions(+), 17 deletions(-) diff --git a/docs/adding_a_benchmark.md b/docs/adding_a_benchmark.md index e3c752b..57fe2b6 100644 --- a/docs/adding_a_benchmark.md +++ b/docs/adding_a_benchmark.md @@ -24,7 +24,7 @@ rather than comparing name strings. | `execution_prompt_builder` | optional per-source builder; called as `execution_prompt_builder(task, *, has_memory=..., generation=...)` → prompt string. Consulted before the `prompt_kind` fallback chain | | `task_markdown_builder` | optional per-source builder; called as `task_markdown_builder(task)` → `TASK.md` string. Consulted before the `prompt_kind` fallback chain (the `task_md_override` metadata hook still wins over both) | | `distill_domain_hint` | optional per-source distillation domain hint (`src/ksi/distillation/prompts.py::_domain_hint`): the hint **string**, or a zero-arg **callable** returning it. Opt-in — when unset, **no** domain-hint paragraph is injected (the generic hint is reserved for the unresolvable/cross-task case) | -| `loader` | task-loader callable; `load_tasks_for_source` calls `spec.loader(tasks_path, *, task_source=..., evals_path=..., arc_max_trials=...)` (built-in loaders are attached by `src/ksi/tasks/loaders.py` at import time) | +| `loader` | task-loader callable; `load_tasks_for_source` calls `spec.loader(tasks_path, *, task_source=..., evals_path=..., arc_max_trials=...)` (built-in loaders are attached by `src/ksi/benchmarks/loaders.py` at import time) | | `supports_mcp_arc` | marks the source as ARC-native: the container materializes `payload.json` + attempt files and the agent uses native file tools (name retained for back-compat) | | `is_offline` | sealed/offline benchmark; provider-native tools disabled | | `uses_repo_snapshots` | needs SWE-bench-style repo cloning/snapshots | @@ -38,7 +38,7 @@ rather than comparing name strings. ### A real example: how `arc` is registered The shortest of the four built-in registrations is `arc` -(`src/ksi/tasks/registry.py`): +(`src/ksi/benchmarks/sources.py`): ```python register_task_source( @@ -75,7 +75,7 @@ it *doesn't* set (`loader`, `validate_tasks_path`, `uses_repo_snapshots`, `supports_classification`, `needs_eval_records`, `delegates_runtime`, ...) keeps its conservative default — `loader` and `validate_tasks_path` are populated separately by -`src/ksi/tasks/loaders.py` and `src/ksi/tasks/path_validation.py` at import +`src/ksi/benchmarks/loaders.py` and `src/ksi/tasks/path_validation.py` at import time rather than inline in the registration call, which is why a real source can look shorter than the full field table suggests. @@ -87,7 +87,7 @@ different capability profile than ARC's. ## Steps 1. **Register the spec.** Add a `register_task_source(TaskSourceSpec(...))` call - in `src/ksi/tasks/registry.py` (or at runtime via `register_task_source` for a + in `src/ksi/benchmarks/sources.py` (or at runtime via `register_task_source` for a plugin). Set only the flags your benchmark needs; defaults are the conservative generic behavior. @@ -162,7 +162,7 @@ different capability profile than ARC's. for the source (the generic `_GENERIC_DOMAIN_HINT` is reserved for the unresolvable/cross-task case, where there is no single benchmark to key on). The four built-in sources set it on their own specs in - `src/ksi/tasks/registry.py`. + `src/ksi/benchmarks/sources.py`. - Evaluator: register the evaluator with `register_evaluator` (see [adding_an_evaluator.md](./adding_an_evaluator.md)) if new. - CLI `--tasks-path` validation needs no dispatch edit: set `validate_tasks_path` on the spec (as above). A source without it is diff --git a/docs/adding_an_evaluator.md b/docs/adding_an_evaluator.md index f6a4a85..3aaa434 100644 --- a/docs/adding_an_evaluator.md +++ b/docs/adding_an_evaluator.md @@ -9,7 +9,7 @@ An evaluator implements the `Evaluator` protocol (`src/ksi/protocols.py`): ```python class Evaluator(Protocol): - def evaluate(self, *, task: TaskSpec, model_output: str, **kwargs: Any) -> EvalResult: ... + def evaluate(self, *, task: TaskSpec, model_output: str, **kwargs: Any) -> dict[str, Any]: ... ``` `EvalResult` is a `TypedDict` (`src/ksi/models.py`, exported from `ksi`), so diff --git a/docs/architecture.md b/docs/architecture.md index 651c40a..6cba995 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -141,10 +141,10 @@ subdirectory mount, gated to forum-phase task sources only. The MCP server selects the SQLite filename from the payload/environment and uses that file as the authoritative knowledge substrate. -ARC tools are independent from memory tools. `arc_tools` is emitted even when -`--no-memory` or `--disable-memory-mcp` removes agent-facing knowledge tools, -and the TypeScript runner mounts the ARC snapshot file itself (not its -directory) under `/app/memory-db` so `arc_load_task` still works. +ARC runs natively for every provider — it neither mounts an ARC snapshot nor +registers an ARC MCP server. The agent reads `payload.json` from its workspace +and writes attempt files with native file tools; the host synthesizes the +`arc_submit_trial` trace after the container exits. ## 3. Runtime Status And Normalization @@ -394,8 +394,9 @@ Provider configuration is loaded from `configs/ksi/*.env` files and normalized by `src/ksi/providers.py`. - Claude task execution defaults to the Claude Code SDK path in - `agent-runner/src/index.ts`. Scheduled ARC and forum tasks can use direct - Anthropic adapters unless configured back to the Claude Code path. + `agent-runner/src/index.ts`, which also serves ARC natively. Scheduled forum + tasks can use a direct Anthropic adapter unless configured back to the Claude + Code path. - OpenAI task execution uses `@openai/agents` in `agent-runner/src/openai.ts`. - Host-side reflection, lesson extraction, and distillation use `src/ksi/runtime/llm.py`, which wraps the Python Anthropic SDK or OpenAI diff --git a/docs/glossary.md b/docs/glossary.md index b2435b9..c739c02 100644 --- a/docs/glossary.md +++ b/docs/glossary.md @@ -16,11 +16,11 @@ The set of [agents](#agent) working in a given [generation](#generation). Popula ### task source -A registry-backed plugin that supplies tasks for agents to attempt. Maintained task sources: `arc`, `swebench_pro`, `polyglot`, `terminal_bench_2`. Register a new one via `src/ksi/tasks/registry.py`. See [Adding a benchmark](adding_a_benchmark.md) for details. +A registry-backed plugin that supplies tasks for agents to attempt. Maintained task sources: `arc`, `swebench_pro`, `polyglot`, `terminal_bench_2`, `custom`. Register a new one in `src/ksi/benchmarks/sources.py`. See [Adding a benchmark](adding_a_benchmark.md) for details. ### evaluator -A registry-backed plugin that scores an [agent's](#agent) [attempt](#attempt) against a task. Maintained evaluators: `none`, `arc_session`, `swebench_pro`, `polyglot_harness`, `terminal_bench_2`. Register a new one via `src/ksi/eval/registry.py`. +A registry-backed plugin that scores an [agent's](#agent) [attempt](#attempt) against a task. Maintained evaluators: `none`, `command`, `arc_session`, `swebench_pro`, `polyglot_harness`, `terminal_bench_2`. Register a new one via `src/ksi/eval/registry.py`. ### runtime diff --git a/examples/custom_tasks/README.md b/examples/custom_tasks/README.md index f509181..450be56 100644 --- a/examples/custom_tasks/README.md +++ b/examples/custom_tasks/README.md @@ -36,6 +36,12 @@ KSI at. Do not run a custom tasks file you don't trust. ## Run it +First create a provider profile from a template (once): + +```bash +cp configs/ksi/.env.haiku.template configs/ksi/.env.haiku # then set your API key +``` + CLI form: ```bash @@ -48,8 +54,8 @@ Programmatic form: uv run python examples/custom_tasks/run.py ``` -Both need Docker running, the `ksi-agent:bench` image built, and a -provider profile with a real API key (see `configs/ksi/*.template`). +Both need Docker running, Node.js (>=22.16.0 <23), the `ksi-agent:bench` image +built, and a provider profile with a real API key (see `configs/ksi/*.template`). ## Expected output diff --git a/scripts/README.md b/scripts/README.md index 0cb20c3..fb99b96 100644 --- a/scripts/README.md +++ b/scripts/README.md @@ -15,7 +15,7 @@ presets live under [`benchmarks/`](../benchmarks/README.md) instead. | `typecheck_agent_runner.sh` | Type-checks the TypeScript agent runner. | | `dev/` | Developer guardrails such as worktree discipline checks. | -Benchmark dataset preparation (`dataprep/`, `arc_prep/`) and the run presets +Benchmark dataset preparation (`scripts/dataprep/`, `scripts/arc_prep/`) and the run presets (`run_arc.sh`, `run_polyglot.sh`, `run_swebench_pro.sh`, `run_terminal_bench_2.sh`) live under [`benchmarks/`](../benchmarks/README.md); see From 3d2eab0e8aa8b9fed0525a0d4db1ff502210e091 Mon Sep 17 00:00:00 2001 From: xuefei-wang Date: Wed, 22 Jul 2026 08:21:12 -0700 Subject: [PATCH 2/2] docs: correct dataprep/arc_prep paths to benchmarks/scripts/ The paths added in cd16b17 (`scripts/dataprep/`, `scripts/arc_prep/`) resolve relative to the repo root, where no such directories exist. The real locations are `benchmarks/scripts/dataprep/` and `benchmarks/scripts/arc_prep/`, as the commit message intended and as benchmarks/docs/BENCHMARK_PREPARE.md spells out. The bare `scripts/dataprep/` form is correct in benchmarks/README.md because it is relative to benchmarks/; it does not carry over here. Co-Authored-By: Claude Opus 4.8 (1M context) Claude-Session: https://claude.ai/code/session_019FhmqhjcboNTbqzZD39d3c --- scripts/README.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/scripts/README.md b/scripts/README.md index fb99b96..963dba5 100644 --- a/scripts/README.md +++ b/scripts/README.md @@ -15,7 +15,8 @@ presets live under [`benchmarks/`](../benchmarks/README.md) instead. | `typecheck_agent_runner.sh` | Type-checks the TypeScript agent runner. | | `dev/` | Developer guardrails such as worktree discipline checks. | -Benchmark dataset preparation (`scripts/dataprep/`, `scripts/arc_prep/`) and the run presets +Benchmark dataset preparation (`benchmarks/scripts/dataprep/`, +`benchmarks/scripts/arc_prep/`) and the run presets (`run_arc.sh`, `run_polyglot.sh`, `run_swebench_pro.sh`, `run_terminal_bench_2.sh`) live under [`benchmarks/`](../benchmarks/README.md); see