From b97477f8945fa839338d81b1caeecc0938afa159 Mon Sep 17 00:00:00 2001 From: xuefei-wang Date: Tue, 21 Jul 2026 21:18:27 -0700 Subject: [PATCH 1/3] docs: correct the egress-allowlist recipe in architecture MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit `KSI_EGRESS_ALLOW` and the `*_BASE_URL` variables are read from the TypeScript runner's `process.env`, which the Python CLI builds from the provider profile plus a fixed system-key list. Neither is on either list, so a value exported in the shell never reaches `deriveEgressAllowlist` — the documented `KSI_EGRESS_ALLOW=... bash scripts/run_ksi.sh` form silently no-ops. Note the constraint rather than leaving a recipe that appears to work. Co-Authored-By: Claude Opus 4.8 (1M context) --- docs/architecture.md | 13 +++++++++++-- 1 file changed, 11 insertions(+), 2 deletions(-) diff --git a/docs/architecture.md b/docs/architecture.md index 0269b57..651c40a 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -446,12 +446,21 @@ The allowlist is derived from the provider profile at launch time: | `openai` | `api.openai.com` + `OPENAI_BASE_URL` host (if set) | Operator-supplied extras are appended via `KSI_EGRESS_ALLOW=host1,host2` -(comma-separated). Use this for Bedrock, Vertex, or other provider endpoints: +(comma-separated) — intended for Bedrock, Vertex, or other provider endpoints: ```bash -KSI_EGRESS_ALLOW=bedrock-runtime.us-east-1.amazonaws.com bash scripts/run_ksi.sh ... +KSI_EGRESS_ALLOW=bedrock-runtime.us-east-1.amazonaws.com ... ``` +!!! note "These variables are read from the runner subprocess environment" + `deriveEgressAllowlist` reads `process.env` of the TypeScript runner. The + Python CLI builds that environment from the provider profile plus a fixed + system-key list (`_build_runner_env` in + `src/ksi/runtime/container_host.py`), and neither `KSI_EGRESS_ALLOW` nor the + `*_BASE_URL` variables are on either list — so a value exported in your + shell does not currently reach the allowlist through a `ksi` / + `scripts/run_ksi.sh` run. + **Escape hatch**: set `KSI_EGRESS=open` to disable isolation and restore the legacy direct-bridge behavior (no internal network, no proxy). Use only for debugging — not for production campaigns. From 4cbdc2e94ceec20accfa068782fad1ac31ab20a0 Mon Sep 17 00:00:00 2001 From: xuefei-wang Date: Tue, 21 Jul 2026 21:18:38 -0700 Subject: [PATCH 2/3] docs: document open-source / self-hosted model support KSI accepts only `anthropic` and `openai` as `MODEL_PROVIDER`, and nothing in the docs said so or explained what running an open-weight model would involve. Adds a page covering the high-level shape: put an Anthropic-Messages-format proxy in front of the model and reuse the existing SDK path rather than writing a provider adapter; host-side phases can be redirected today while the containerized agent cannot, since the base-URL variables aren't forwarded into the container. Notes what it affects (egress allowlist, prompt caching, cost reporting, structured output) and flags multi-turn tool-calling robustness as the real risk. Left as a future contribution rather than an implementation spec. Also warns that setting a base URL today splits traffic between the local server and the hosted API instead of failing loudly. Co-Authored-By: Claude Opus 4.8 (1M context) --- docs/faq.md | 28 ++++++++++++ docs/open_source_models.md | 87 ++++++++++++++++++++++++++++++++++++++ mkdocs.yml | 1 + 3 files changed, 116 insertions(+) create mode 100644 docs/open_source_models.md diff --git a/docs/faq.md b/docs/faq.md index 51cda64..4939629 100644 --- a/docs/faq.md +++ b/docs/faq.md @@ -78,6 +78,34 @@ Opus) and OpenAI models through provider profiles stored under Copy the template for your provider, fill in your key, and pass the path with `--provider-profile configs/ksi/.env.haiku` (or equivalent). +`MODEL_PROVIDER` accepts only `anthropic` and `openai` — there is no `vllm`, +`ollama`, or `openrouter` provider. For open-weight models, see the next +question. + +## Can I run an open-source model (Llama, Qwen, DeepSeek)? + +Not out of the box today, but the gap is wiring rather than architecture — you +would not need a new provider adapter. + +The shortest path is to put a proxy that speaks the Anthropic Messages format +(e.g. [LiteLLM](https://docs.litellm.ai/docs/anthropic_unified/)) in front of +your model and let the existing Claude Agent SDK path talk to it, keeping +`MODEL_PROVIDER=anthropic`. The host-side phases (forum, distillation, +reflection) can already be redirected this way today; the containerized agent +cannot, because the base-URL variables aren't forwarded into the container yet. + +!!! warning "Don't half-configure this" + Setting a base URL today does not error — it **splits your traffic**: + knowledge phases hit your local server while every task container still + calls the hosted API. + +Expect prompt-cache savings to largely disappear through a proxy, and cost +reporting to read `$0.00` for unrecognized model names. The real risk is +neither — it is whether your model holds up in a long multi-turn tool-calling +loop. + +More detail: [Open-source & self-hosted models](./open_source_models.md). + ## What does a run cost? Every run makes real LLM API calls billed to the key in your provider profile; diff --git a/docs/open_source_models.md b/docs/open_source_models.md new file mode 100644 index 0000000..0c5b886 --- /dev/null +++ b/docs/open_source_models.md @@ -0,0 +1,87 @@ +# Open-source and self-hosted models + +Can KSI run against Llama, Qwen, DeepSeek, or a model you host yourself on +vLLM/Ollama? **Not out of the box today**, but the gap is wiring rather than +architecture. This page sketches how it would work, what it touches, and the +problems to expect. Contributions welcome. + +## Where things stand + +`MODEL_PROVIDER` accepts exactly two values — `anthropic` and `openai`. There is +no `vllm`, `ollama`, or `openrouter` provider, and an unrecognized value is +rejected up front rather than silently falling back. + +## How you would wire it + +Don't write a new provider adapter. Both SDKs KSI uses are built to be +repointed at a different endpoint, so the shortest path is to put a **proxy that +speaks the Anthropic Messages format** in front of your model — +[LiteLLM](https://docs.litellm.ai/docs/anthropic_unified/) translates that +format to most backends — and let the existing Claude Agent SDK path talk to it. +The SDK honors `ANTHROPIC_BASE_URL` and `ANTHROPIC_AUTH_TOKEN`, so the whole +agent loop (multi-turn, native tools, the MCP memory server, hooks) keeps +working unchanged. `MODEL_PROVIDER` stays `anthropic`; `MODEL` becomes whatever +name your proxy routes. + +Pointing at an OpenAI-compatible endpoint instead is also possible, but it is +more work and needs care about which API shape your server implements — many +implement Chat Completions only, and a server that does expose a Responses +endpoint may still not handle the multi-turn history an agent loop replays +through it. + +This split is why it half-works today: + +- **Host-side phases** — forum, distillation, reflection, task-claiming — already + honor the base-URL environment variables, so they can be redirected right now. +- **The containerized agent** cannot. Task execution runs in a Docker container + whose environment is built from a fixed set of keys, and the base-URL variables + aren't among them. Getting them forwarded is the change a contributor would + need to make. + +!!! warning "Don't half-configure this" + Setting a base URL today does not error — it **splits your traffic**: + knowledge phases hit your local server while every task container still calls + the hosted API, billing your key or failing on the egress allowlist. Treat + the current state as useful for experimenting with the knowledge phases only. + +## What it affects + +- **Egress isolation.** Agent containers reach the network only through an + allowlisting proxy (see + [Architecture § Egress isolation](./architecture.md#10-egress-isolation)), so a + self-hosted endpoint has to be allowlisted. A model server on the Docker host + is its own problem: `localhost` inside the container is the container, and the + internal network has no route back to the host. +- **Prompt caching.** KSI places cache breakpoints on stable prompt prefixes. + Whether those survive a proxy translation varies, so expect cache savings — and + the cache columns in token accounting — to largely disappear. +- **Cost reporting.** Unrecognized model names price at `$0.00`. Reasonable for a + self-hosted model, but it makes cost comparisons against a hosted baseline + misleading. +- **Structured output.** Forum, distillation, and task-claiming ask for + schema-constrained JSON. There is already a fallback for callers that can't + provide it, so this degrades rather than breaks — but whether the looser output + holds up in quality is an open question. +- **Direct adapters.** A couple of paths (forum, ARC) bypass the SDK and call the + Anthropic API directly against a hardcoded URL. They would need the same + treatment, or to be configured back onto the SDK path. + +## Problems to expect + +The plumbing is the easy part. The real risk is **tool-calling robustness**: the +agent drives a long multi-turn loop with native tools and an MCP server attached, +and smaller open models tend to fail there — dropping tool calls, malforming +arguments, or looping — long before any of the above matters. Reasoning-effort +handling is also currently keyed to specific hosted model families, so +reasoning-capable open models would need that revisited. + +Worth validating cheaply before investing: redirect the host-side phases (which +works today) and see whether one forum round and one distillation produce usable +output. If they don't, the container work won't save it. + +## If you just want a cheap model + +If the goal is lower cost rather than open weights specifically, the supported +route is a small hosted model — `.env.haiku.template` or `.env.openai.template`. +See the [FAQ](./faq.md#what-models-and-providers-can-i-use-do-i-need-an-api-key) +for the full list of bundled profile templates. diff --git a/mkdocs.yml b/mkdocs.yml index 5bbd7a4..f9cf0f0 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -96,6 +96,7 @@ nav: - Extension seams: extending.md - Add a benchmark / task source: adding_a_benchmark.md - Add an evaluator / runtime: adding_an_evaluator.md + - Open-source & self-hosted models: open_source_models.md - Improvement strategies: improvement_strategies.md - Reference: - Glossary: glossary.md From e3c31d895c540fd873317837268b89103ad535da Mon Sep 17 00:00:00 2001 From: xuefei-wang Date: Tue, 21 Jul 2026 21:24:54 -0700 Subject: [PATCH 3/3] docs: regroup nav so pages sit under what they're about MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Four placements didn't match content: - "Running larger runs" sat under "Programmatic API", but it documents CLI flags and explicitly frames `ksi.run(...)` as a layer over the parser it describes. It opens by naming its own place in the progression after the quickstart and your-own-tasks walkthroughs — so it belongs with them. - "Your own tasks" was a standalone tab despite being step two of that same progression. - "Contributing" sat under "Getting started", where a first-time user meets PR conventions before running the demo. It's contributor material, and extending.md already points at it. - "Benchmarks" trailed after "Reference" even though it's a usage guide for running the reference benchmarks yourself. Groups the three-step new-user path under "Getting started", moves Contributing under "Extending KSI", and lifts Benchmarks out of the post-reference slot. Landing-page cards track the change (they were deliberately unified with the nav tabs in 7f99358). Page URLs are derived from file paths, so no links change. Co-Authored-By: Claude Opus 4.8 (1M context) --- docs/index.md | 7 ++++--- mkdocs.yml | 11 ++++++----- 2 files changed, 10 insertions(+), 8 deletions(-) diff --git a/docs/index.md b/docs/index.md index 5d1f0d1..2381874 100644 --- a/docs/index.md +++ b/docs/index.md @@ -28,7 +28,9 @@ next generation with it. --- - Fresh clone to a solved demo task in one command — no dataset download. + Fresh clone to a solved demo task in one command — then + [your own tasks](your_own_tasks.md) and + [larger runs](experiments.md). [:octicons-arrow-right-24: Run the demo](getting-started.md) @@ -52,8 +54,7 @@ next generation with it. --- - Drive a run from Python with `ksi.run(...)` — plus - [larger runs](experiments.md) and the + Drive a run from Python with `ksi.run(...)` — plus the auto-generated [API reference](reference/api.md). [:octicons-arrow-right-24: API guide](programmatic_api.md) diff --git a/mkdocs.yml b/mkdocs.yml index f9cf0f0..89d010d 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -84,24 +84,25 @@ markdown_extensions: nav: - Home: index.md - - Getting started: getting-started.md - - Your own tasks: your_own_tasks.md + - Getting started: + - getting-started.md + - Your own tasks: your_own_tasks.md + - Running larger runs: experiments.md + - Benchmarks: benchmarks.md - FAQ: faq.md - Programmatic API: - programmatic_api.md - - Running larger runs: experiments.md - Python API reference: reference/api.md - Architecture: architecture.md - Extending KSI: - Extension seams: extending.md - Add a benchmark / task source: adding_a_benchmark.md - Add an evaluator / runtime: adding_an_evaluator.md - - Open-source & self-hosted models: open_source_models.md - Improvement strategies: improvement_strategies.md + - Open-source & self-hosted models: open_source_models.md - Reference: - Glossary: glossary.md - Artifacts & cleanup: artifacts.md - Runtime startup performance: runtime-startup-performance.md - - Benchmarks: benchmarks.md - Contributing: CONTRIBUTING.md - Blog ↗: https://recursive-knowledge.github.io/knowledge-centric-self-improvement/