Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 11 additions & 2 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -446,12 +446,21 @@ The allowlist is derived from the provider profile at launch time:
| `openai` | `api.openai.com` + `OPENAI_BASE_URL` host (if set) |

Operator-supplied extras are appended via `KSI_EGRESS_ALLOW=host1,host2`
(comma-separated). Use this for Bedrock, Vertex, or other provider endpoints:
(comma-separated) — intended for Bedrock, Vertex, or other provider endpoints:

```bash
KSI_EGRESS_ALLOW=bedrock-runtime.us-east-1.amazonaws.com bash scripts/run_ksi.sh ...
KSI_EGRESS_ALLOW=bedrock-runtime.us-east-1.amazonaws.com ...
```

!!! note "These variables are read from the runner subprocess environment"
`deriveEgressAllowlist` reads `process.env` of the TypeScript runner. The
Python CLI builds that environment from the provider profile plus a fixed
system-key list (`_build_runner_env` in
`src/ksi/runtime/container_host.py`), and neither `KSI_EGRESS_ALLOW` nor the
`*_BASE_URL` variables are on either list — so a value exported in your
shell does not currently reach the allowlist through a `ksi` /
`scripts/run_ksi.sh` run.

**Escape hatch**: set `KSI_EGRESS=open` to disable isolation and restore
the legacy direct-bridge behavior (no internal network, no proxy). Use only for
debugging — not for production campaigns.
Expand Down
28 changes: 28 additions & 0 deletions docs/faq.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,6 +78,34 @@ Opus) and OpenAI models through provider profiles stored under
Copy the template for your provider, fill in your key, and pass the path with
`--provider-profile configs/ksi/.env.haiku` (or equivalent).

`MODEL_PROVIDER` accepts only `anthropic` and `openai` — there is no `vllm`,
`ollama`, or `openrouter` provider. For open-weight models, see the next
question.

## Can I run an open-source model (Llama, Qwen, DeepSeek)?

Not out of the box today, but the gap is wiring rather than architecture — you
would not need a new provider adapter.

The shortest path is to put a proxy that speaks the Anthropic Messages format
(e.g. [LiteLLM](https://docs.litellm.ai/docs/anthropic_unified/)) in front of
your model and let the existing Claude Agent SDK path talk to it, keeping
`MODEL_PROVIDER=anthropic`. The host-side phases (forum, distillation,
reflection) can already be redirected this way today; the containerized agent
cannot, because the base-URL variables aren't forwarded into the container yet.

!!! warning "Don't half-configure this"
Setting a base URL today does not error — it **splits your traffic**:
knowledge phases hit your local server while every task container still
calls the hosted API.

Expect prompt-cache savings to largely disappear through a proxy, and cost
reporting to read `$0.00` for unrecognized model names. The real risk is
neither — it is whether your model holds up in a long multi-turn tool-calling
loop.

More detail: [Open-source & self-hosted models](./open_source_models.md).

## What does a run cost?

Every run makes real LLM API calls billed to the key in your provider profile;
Expand Down
7 changes: 4 additions & 3 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,9 @@ next generation with it.

---

Fresh clone to a solved demo task in one command — no dataset download.
Fresh clone to a solved demo task in one command — then
[your own tasks](your_own_tasks.md) and
[larger runs](experiments.md).

[:octicons-arrow-right-24: Run the demo](getting-started.md)

Expand All @@ -52,8 +54,7 @@ next generation with it.

---

Drive a run from Python with `ksi.run(...)` — plus
[larger runs](experiments.md) and the
Drive a run from Python with `ksi.run(...)` — plus the auto-generated
[API reference](reference/api.md).

[:octicons-arrow-right-24: API guide](programmatic_api.md)
Expand Down
87 changes: 87 additions & 0 deletions docs/open_source_models.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,87 @@
# Open-source and self-hosted models

Can KSI run against Llama, Qwen, DeepSeek, or a model you host yourself on
vLLM/Ollama? **Not out of the box today**, but the gap is wiring rather than
architecture. This page sketches how it would work, what it touches, and the
problems to expect. Contributions welcome.

## Where things stand

`MODEL_PROVIDER` accepts exactly two values — `anthropic` and `openai`. There is
no `vllm`, `ollama`, or `openrouter` provider, and an unrecognized value is
rejected up front rather than silently falling back.

## How you would wire it

Don't write a new provider adapter. Both SDKs KSI uses are built to be
repointed at a different endpoint, so the shortest path is to put a **proxy that
speaks the Anthropic Messages format** in front of your model —
[LiteLLM](https://docs.litellm.ai/docs/anthropic_unified/) translates that
format to most backends — and let the existing Claude Agent SDK path talk to it.
The SDK honors `ANTHROPIC_BASE_URL` and `ANTHROPIC_AUTH_TOKEN`, so the whole
agent loop (multi-turn, native tools, the MCP memory server, hooks) keeps
working unchanged. `MODEL_PROVIDER` stays `anthropic`; `MODEL` becomes whatever
name your proxy routes.

Pointing at an OpenAI-compatible endpoint instead is also possible, but it is
more work and needs care about which API shape your server implements — many
implement Chat Completions only, and a server that does expose a Responses
endpoint may still not handle the multi-turn history an agent loop replays
through it.

This split is why it half-works today:

- **Host-side phases** — forum, distillation, reflection, task-claiming — already
honor the base-URL environment variables, so they can be redirected right now.
- **The containerized agent** cannot. Task execution runs in a Docker container
whose environment is built from a fixed set of keys, and the base-URL variables
aren't among them. Getting them forwarded is the change a contributor would
need to make.

!!! warning "Don't half-configure this"
Setting a base URL today does not error — it **splits your traffic**:
knowledge phases hit your local server while every task container still calls
the hosted API, billing your key or failing on the egress allowlist. Treat
the current state as useful for experimenting with the knowledge phases only.

## What it affects

- **Egress isolation.** Agent containers reach the network only through an
allowlisting proxy (see
[Architecture § Egress isolation](./architecture.md#10-egress-isolation)), so a
self-hosted endpoint has to be allowlisted. A model server on the Docker host
is its own problem: `localhost` inside the container is the container, and the
internal network has no route back to the host.
- **Prompt caching.** KSI places cache breakpoints on stable prompt prefixes.
Whether those survive a proxy translation varies, so expect cache savings — and
the cache columns in token accounting — to largely disappear.
- **Cost reporting.** Unrecognized model names price at `$0.00`. Reasonable for a
self-hosted model, but it makes cost comparisons against a hosted baseline
misleading.
- **Structured output.** Forum, distillation, and task-claiming ask for
schema-constrained JSON. There is already a fallback for callers that can't
provide it, so this degrades rather than breaks — but whether the looser output
holds up in quality is an open question.
- **Direct adapters.** A couple of paths (forum, ARC) bypass the SDK and call the
Anthropic API directly against a hardcoded URL. They would need the same
treatment, or to be configured back onto the SDK path.

## Problems to expect

The plumbing is the easy part. The real risk is **tool-calling robustness**: the
agent drives a long multi-turn loop with native tools and an MCP server attached,
and smaller open models tend to fail there — dropping tool calls, malforming
arguments, or looping — long before any of the above matters. Reasoning-effort
handling is also currently keyed to specific hosted model families, so
reasoning-capable open models would need that revisited.

Worth validating cheaply before investing: redirect the host-side phases (which
works today) and see whether one forum round and one distillation produce usable
output. If they don't, the container work won't save it.

## If you just want a cheap model

If the goal is lower cost rather than open weights specifically, the supported
route is a small hosted model — `.env.haiku.template` or `.env.openai.template`.
See the [FAQ](./faq.md#what-models-and-providers-can-i-use-do-i-need-an-api-key)
for the full list of bundled profile templates.
10 changes: 6 additions & 4 deletions mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -84,23 +84,25 @@ markdown_extensions:

nav:
- Home: index.md
- Getting started: getting-started.md
- Your own tasks: your_own_tasks.md
- Getting started:
- getting-started.md
- Your own tasks: your_own_tasks.md
- Running larger runs: experiments.md
- Benchmarks: benchmarks.md
- FAQ: faq.md
- Programmatic API:
- programmatic_api.md
- Running larger runs: experiments.md
- Python API reference: reference/api.md
- Architecture: architecture.md
- Extending KSI:
- Extension seams: extending.md
- Add a benchmark / task source: adding_a_benchmark.md
- Add an evaluator / runtime: adding_an_evaluator.md
- Improvement strategies: improvement_strategies.md
- Open-source & self-hosted models: open_source_models.md
- Reference:
- Glossary: glossary.md
- Artifacts & cleanup: artifacts.md
- Runtime startup performance: runtime-startup-performance.md
- Benchmarks: benchmarks.md
- Contributing: CONTRIBUTING.md
- Blog ↗: https://recursive-knowledge.github.io/knowledge-centric-self-improvement/
Loading