diff --git a/docs/architecture.md b/docs/architecture.md index 0269b57..651c40a 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -446,12 +446,21 @@ The allowlist is derived from the provider profile at launch time: | `openai` | `api.openai.com` + `OPENAI_BASE_URL` host (if set) | Operator-supplied extras are appended via `KSI_EGRESS_ALLOW=host1,host2` -(comma-separated). Use this for Bedrock, Vertex, or other provider endpoints: +(comma-separated) — intended for Bedrock, Vertex, or other provider endpoints: ```bash -KSI_EGRESS_ALLOW=bedrock-runtime.us-east-1.amazonaws.com bash scripts/run_ksi.sh ... +KSI_EGRESS_ALLOW=bedrock-runtime.us-east-1.amazonaws.com ... ``` +!!! note "These variables are read from the runner subprocess environment" + `deriveEgressAllowlist` reads `process.env` of the TypeScript runner. The + Python CLI builds that environment from the provider profile plus a fixed + system-key list (`_build_runner_env` in + `src/ksi/runtime/container_host.py`), and neither `KSI_EGRESS_ALLOW` nor the + `*_BASE_URL` variables are on either list — so a value exported in your + shell does not currently reach the allowlist through a `ksi` / + `scripts/run_ksi.sh` run. + **Escape hatch**: set `KSI_EGRESS=open` to disable isolation and restore the legacy direct-bridge behavior (no internal network, no proxy). Use only for debugging — not for production campaigns. diff --git a/docs/faq.md b/docs/faq.md index 51cda64..4939629 100644 --- a/docs/faq.md +++ b/docs/faq.md @@ -78,6 +78,34 @@ Opus) and OpenAI models through provider profiles stored under Copy the template for your provider, fill in your key, and pass the path with `--provider-profile configs/ksi/.env.haiku` (or equivalent). +`MODEL_PROVIDER` accepts only `anthropic` and `openai` — there is no `vllm`, +`ollama`, or `openrouter` provider. For open-weight models, see the next +question. + +## Can I run an open-source model (Llama, Qwen, DeepSeek)? + +Not out of the box today, but the gap is wiring rather than architecture — you +would not need a new provider adapter. + +The shortest path is to put a proxy that speaks the Anthropic Messages format +(e.g. [LiteLLM](https://docs.litellm.ai/docs/anthropic_unified/)) in front of +your model and let the existing Claude Agent SDK path talk to it, keeping +`MODEL_PROVIDER=anthropic`. The host-side phases (forum, distillation, +reflection) can already be redirected this way today; the containerized agent +cannot, because the base-URL variables aren't forwarded into the container yet. + +!!! warning "Don't half-configure this" + Setting a base URL today does not error — it **splits your traffic**: + knowledge phases hit your local server while every task container still + calls the hosted API. + +Expect prompt-cache savings to largely disappear through a proxy, and cost +reporting to read `$0.00` for unrecognized model names. The real risk is +neither — it is whether your model holds up in a long multi-turn tool-calling +loop. + +More detail: [Open-source & self-hosted models](./open_source_models.md). + ## What does a run cost? Every run makes real LLM API calls billed to the key in your provider profile; diff --git a/docs/index.md b/docs/index.md index 5d1f0d1..2381874 100644 --- a/docs/index.md +++ b/docs/index.md @@ -28,7 +28,9 @@ next generation with it. --- - Fresh clone to a solved demo task in one command — no dataset download. + Fresh clone to a solved demo task in one command — then + [your own tasks](your_own_tasks.md) and + [larger runs](experiments.md). [:octicons-arrow-right-24: Run the demo](getting-started.md) @@ -52,8 +54,7 @@ next generation with it. --- - Drive a run from Python with `ksi.run(...)` — plus - [larger runs](experiments.md) and the + Drive a run from Python with `ksi.run(...)` — plus the auto-generated [API reference](reference/api.md). [:octicons-arrow-right-24: API guide](programmatic_api.md) diff --git a/docs/open_source_models.md b/docs/open_source_models.md new file mode 100644 index 0000000..0c5b886 --- /dev/null +++ b/docs/open_source_models.md @@ -0,0 +1,87 @@ +# Open-source and self-hosted models + +Can KSI run against Llama, Qwen, DeepSeek, or a model you host yourself on +vLLM/Ollama? **Not out of the box today**, but the gap is wiring rather than +architecture. This page sketches how it would work, what it touches, and the +problems to expect. Contributions welcome. + +## Where things stand + +`MODEL_PROVIDER` accepts exactly two values — `anthropic` and `openai`. There is +no `vllm`, `ollama`, or `openrouter` provider, and an unrecognized value is +rejected up front rather than silently falling back. + +## How you would wire it + +Don't write a new provider adapter. Both SDKs KSI uses are built to be +repointed at a different endpoint, so the shortest path is to put a **proxy that +speaks the Anthropic Messages format** in front of your model — +[LiteLLM](https://docs.litellm.ai/docs/anthropic_unified/) translates that +format to most backends — and let the existing Claude Agent SDK path talk to it. +The SDK honors `ANTHROPIC_BASE_URL` and `ANTHROPIC_AUTH_TOKEN`, so the whole +agent loop (multi-turn, native tools, the MCP memory server, hooks) keeps +working unchanged. `MODEL_PROVIDER` stays `anthropic`; `MODEL` becomes whatever +name your proxy routes. + +Pointing at an OpenAI-compatible endpoint instead is also possible, but it is +more work and needs care about which API shape your server implements — many +implement Chat Completions only, and a server that does expose a Responses +endpoint may still not handle the multi-turn history an agent loop replays +through it. + +This split is why it half-works today: + +- **Host-side phases** — forum, distillation, reflection, task-claiming — already + honor the base-URL environment variables, so they can be redirected right now. +- **The containerized agent** cannot. Task execution runs in a Docker container + whose environment is built from a fixed set of keys, and the base-URL variables + aren't among them. Getting them forwarded is the change a contributor would + need to make. + +!!! warning "Don't half-configure this" + Setting a base URL today does not error — it **splits your traffic**: + knowledge phases hit your local server while every task container still calls + the hosted API, billing your key or failing on the egress allowlist. Treat + the current state as useful for experimenting with the knowledge phases only. + +## What it affects + +- **Egress isolation.** Agent containers reach the network only through an + allowlisting proxy (see + [Architecture § Egress isolation](./architecture.md#10-egress-isolation)), so a + self-hosted endpoint has to be allowlisted. A model server on the Docker host + is its own problem: `localhost` inside the container is the container, and the + internal network has no route back to the host. +- **Prompt caching.** KSI places cache breakpoints on stable prompt prefixes. + Whether those survive a proxy translation varies, so expect cache savings — and + the cache columns in token accounting — to largely disappear. +- **Cost reporting.** Unrecognized model names price at `$0.00`. Reasonable for a + self-hosted model, but it makes cost comparisons against a hosted baseline + misleading. +- **Structured output.** Forum, distillation, and task-claiming ask for + schema-constrained JSON. There is already a fallback for callers that can't + provide it, so this degrades rather than breaks — but whether the looser output + holds up in quality is an open question. +- **Direct adapters.** A couple of paths (forum, ARC) bypass the SDK and call the + Anthropic API directly against a hardcoded URL. They would need the same + treatment, or to be configured back onto the SDK path. + +## Problems to expect + +The plumbing is the easy part. The real risk is **tool-calling robustness**: the +agent drives a long multi-turn loop with native tools and an MCP server attached, +and smaller open models tend to fail there — dropping tool calls, malforming +arguments, or looping — long before any of the above matters. Reasoning-effort +handling is also currently keyed to specific hosted model families, so +reasoning-capable open models would need that revisited. + +Worth validating cheaply before investing: redirect the host-side phases (which +works today) and see whether one forum round and one distillation produce usable +output. If they don't, the container work won't save it. + +## If you just want a cheap model + +If the goal is lower cost rather than open weights specifically, the supported +route is a small hosted model — `.env.haiku.template` or `.env.openai.template`. +See the [FAQ](./faq.md#what-models-and-providers-can-i-use-do-i-need-an-api-key) +for the full list of bundled profile templates. diff --git a/mkdocs.yml b/mkdocs.yml index 5bbd7a4..89d010d 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -84,12 +84,14 @@ markdown_extensions: nav: - Home: index.md - - Getting started: getting-started.md - - Your own tasks: your_own_tasks.md + - Getting started: + - getting-started.md + - Your own tasks: your_own_tasks.md + - Running larger runs: experiments.md + - Benchmarks: benchmarks.md - FAQ: faq.md - Programmatic API: - programmatic_api.md - - Running larger runs: experiments.md - Python API reference: reference/api.md - Architecture: architecture.md - Extending KSI: @@ -97,10 +99,10 @@ nav: - Add a benchmark / task source: adding_a_benchmark.md - Add an evaluator / runtime: adding_an_evaluator.md - Improvement strategies: improvement_strategies.md + - Open-source & self-hosted models: open_source_models.md - Reference: - Glossary: glossary.md - Artifacts & cleanup: artifacts.md - Runtime startup performance: runtime-startup-performance.md - - Benchmarks: benchmarks.md - Contributing: CONTRIBUTING.md - Blog ↗: https://recursive-knowledge.github.io/knowledge-centric-self-improvement/