Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions docs/adding_a_benchmark.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ rather than comparing name strings.
| `execution_prompt_builder` | optional per-source builder; called as `execution_prompt_builder(task, *, has_memory=..., generation=...)` → prompt string. Consulted before the `prompt_kind` fallback chain |
| `task_markdown_builder` | optional per-source builder; called as `task_markdown_builder(task)` → `TASK.md` string. Consulted before the `prompt_kind` fallback chain (the `task_md_override` metadata hook still wins over both) |
| `distill_domain_hint` | optional per-source distillation domain hint (`src/ksi/distillation/prompts.py::_domain_hint`): the hint **string**, or a zero-arg **callable** returning it. Opt-in — when unset, **no** domain-hint paragraph is injected (the generic hint is reserved for the unresolvable/cross-task case) |
| `loader` | task-loader callable; `load_tasks_for_source` calls `spec.loader(tasks_path, *, task_source=..., evals_path=..., arc_max_trials=...)` (built-in loaders are attached by `src/ksi/tasks/loaders.py` at import time) |
| `loader` | task-loader callable; `load_tasks_for_source` calls `spec.loader(tasks_path, *, task_source=..., evals_path=..., arc_max_trials=...)` (built-in loaders are attached by `src/ksi/benchmarks/loaders.py` at import time) |
| `supports_mcp_arc` | marks the source as ARC-native: the container materializes `payload.json` + attempt files and the agent uses native file tools (name retained for back-compat) |
| `is_offline` | sealed/offline benchmark; provider-native tools disabled |
| `uses_repo_snapshots` | needs SWE-bench-style repo cloning/snapshots |
Expand All @@ -38,7 +38,7 @@ rather than comparing name strings.
### A real example: how `arc` is registered

The shortest of the four built-in registrations is `arc`
(`src/ksi/tasks/registry.py`):
(`src/ksi/benchmarks/sources.py`):

```python
register_task_source(
Expand Down Expand Up @@ -75,7 +75,7 @@ it *doesn't* set (`loader`,
`validate_tasks_path`, `uses_repo_snapshots`, `supports_classification`,
`needs_eval_records`, `delegates_runtime`, ...) keeps its conservative
default — `loader` and `validate_tasks_path` are populated separately by
`src/ksi/tasks/loaders.py` and `src/ksi/tasks/path_validation.py` at import
`src/ksi/benchmarks/loaders.py` and `src/ksi/tasks/path_validation.py` at import
time rather than inline in the registration call, which is why a real source
can look shorter than the full field table suggests.

Expand All @@ -87,7 +87,7 @@ different capability profile than ARC's.
## Steps

1. **Register the spec.** Add a `register_task_source(TaskSourceSpec(...))` call
in `src/ksi/tasks/registry.py` (or at runtime via `register_task_source` for a
in `src/ksi/benchmarks/sources.py` (or at runtime via `register_task_source` for a
plugin). Set only the flags your benchmark needs; defaults are the
conservative generic behavior.

Expand Down Expand Up @@ -162,7 +162,7 @@ different capability profile than ARC's.
for the source (the generic `_GENERIC_DOMAIN_HINT` is reserved for the
unresolvable/cross-task case, where there is no single benchmark to key
on). The four built-in sources set it on their own specs in
`src/ksi/tasks/registry.py`.
`src/ksi/benchmarks/sources.py`.
- Evaluator: register the evaluator with `register_evaluator` (see [adding_an_evaluator.md](./adding_an_evaluator.md)) if new.
- CLI `--tasks-path` validation needs no dispatch edit: set
`validate_tasks_path` on the spec (as above). A source without it is
Expand Down
2 changes: 1 addition & 1 deletion docs/adding_an_evaluator.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ An evaluator implements the `Evaluator` protocol (`src/ksi/protocols.py`):

```python
class Evaluator(Protocol):
def evaluate(self, *, task: TaskSpec, model_output: str, **kwargs: Any) -> EvalResult: ...
def evaluate(self, *, task: TaskSpec, model_output: str, **kwargs: Any) -> dict[str, Any]: ...
```

`EvalResult` is a `TypedDict` (`src/ksi/models.py`, exported from `ksi`), so
Expand Down
13 changes: 7 additions & 6 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -141,10 +141,10 @@ subdirectory mount, gated to forum-phase task sources only. The MCP server
selects the SQLite filename from the payload/environment and uses that file
as the authoritative knowledge substrate.

ARC tools are independent from memory tools. `arc_tools` is emitted even when
`--no-memory` or `--disable-memory-mcp` removes agent-facing knowledge tools,
and the TypeScript runner mounts the ARC snapshot file itself (not its
directory) under `/app/memory-db` so `arc_load_task` still works.
ARC runs natively for every provider — it neither mounts an ARC snapshot nor
registers an ARC MCP server. The agent reads `payload.json` from its workspace
and writes attempt files with native file tools; the host synthesizes the
`arc_submit_trial` trace after the container exits.

## 3. Runtime Status And Normalization

Expand Down Expand Up @@ -394,8 +394,9 @@ Provider configuration is loaded from `configs/ksi/*.env` files and
normalized by `src/ksi/providers.py`.

- Claude task execution defaults to the Claude Code SDK path in
`agent-runner/src/index.ts`. Scheduled ARC and forum tasks can use direct
Anthropic adapters unless configured back to the Claude Code path.
`agent-runner/src/index.ts`, which also serves ARC natively. Scheduled forum
tasks can use a direct Anthropic adapter unless configured back to the Claude
Code path.
- OpenAI task execution uses `@openai/agents` in `agent-runner/src/openai.ts`.
- Host-side reflection, lesson extraction, and distillation use
`src/ksi/runtime/llm.py`, which wraps the Python Anthropic SDK or OpenAI
Expand Down
4 changes: 2 additions & 2 deletions docs/glossary.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,11 +16,11 @@ The set of [agents](#agent) working in a given [generation](#generation). Popula

### task source

A registry-backed plugin that supplies tasks for agents to attempt. Maintained task sources: `arc`, `swebench_pro`, `polyglot`, `terminal_bench_2`. Register a new one via `src/ksi/tasks/registry.py`. See [Adding a benchmark](adding_a_benchmark.md) for details.
A registry-backed plugin that supplies tasks for agents to attempt. Maintained task sources: `arc`, `swebench_pro`, `polyglot`, `terminal_bench_2`, `custom`. Register a new one in `src/ksi/benchmarks/sources.py`. See [Adding a benchmark](adding_a_benchmark.md) for details.

### evaluator

A registry-backed plugin that scores an [agent's](#agent) [attempt](#attempt) against a task. Maintained evaluators: `none`, `arc_session`, `swebench_pro`, `polyglot_harness`, `terminal_bench_2`. Register a new one via `src/ksi/eval/registry.py`.
A registry-backed plugin that scores an [agent's](#agent) [attempt](#attempt) against a task. Maintained evaluators: `none`, `command`, `arc_session`, `swebench_pro`, `polyglot_harness`, `terminal_bench_2`. Register a new one via `src/ksi/eval/registry.py`.

### runtime

Expand Down
10 changes: 8 additions & 2 deletions examples/custom_tasks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,12 @@ KSI at. Do not run a custom tasks file you don't trust.

## Run it

First create a provider profile from a template (once):

```bash
cp configs/ksi/.env.haiku.template configs/ksi/.env.haiku # then set your API key
```

CLI form:

```bash
Expand All @@ -48,8 +54,8 @@ Programmatic form:
uv run python examples/custom_tasks/run.py
```

Both need Docker running, the `ksi-agent:bench` image built, and a
provider profile with a real API key (see `configs/ksi/*.template`).
Both need Docker running, Node.js (>=22.16.0 <23), the `ksi-agent:bench` image
built, and a provider profile with a real API key (see `configs/ksi/*.template`).

## Expected output

Expand Down
3 changes: 2 additions & 1 deletion scripts/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,8 @@ presets live under [`benchmarks/`](../benchmarks/README.md) instead.
| `typecheck_agent_runner.sh` | Type-checks the TypeScript agent runner. |
| `dev/` | Developer guardrails such as worktree discipline checks. |

Benchmark dataset preparation (`dataprep/`, `arc_prep/`) and the run presets
Benchmark dataset preparation (`benchmarks/scripts/dataprep/`,
`benchmarks/scripts/arc_prep/`) and the run presets
(`run_arc.sh`, `run_polyglot.sh`, `run_swebench_pro.sh`,
`run_terminal_bench_2.sh`) live under
[`benchmarks/`](../benchmarks/README.md); see
Expand Down
Loading