Status: initial version (2026-09-12), mapped against the OWASP Agentic/LLM Top 10 checklist (issue #72). Every item resolves to control (where it lives) / gap (tracked) / N/A (with rationale).
| Asset | Where | Notes |
|---|---|---|
| Provider API key | TOLE_/OPENAI_* env |
Never persisted; scrubbed from wire errors/messages |
| Session logs (JSONL) | sessions dir | Full transcript incl. tool results — local only |
| Job logs | <workspace>/tole-jobs/ |
Child stdout/stderr, attacker-influenceable content |
| Workspace files | --workspace dir (default cwd) |
Readable AND writable by the model via tools |
| Host env of child processes | inherited by spawned commands | See ENV-1 |
| Subprocess execution authority | run_command / job_start / git / gh | The model chooses argv within argv-split + jail rules |
- Provider responses → context. The model's output becomes intent entries; tool args are argv-validated but otherwise untrusted.
- Tool results → context. File contents, job logs, command stdout flow back into the transcript and are sent to the provider. A malicious workspace file or a compromised job subprocess can inject instructions here (OWASP: prompt injection).
- Workspace → filesystem. File tools jail to
--workspace(component-validated, TOCTOU-safe walk, #68);run_command/job_startrun with cwd jailed but FULL process authority — the jail bounds cwd, not capability. - Child env. Spawned commands inherit the host environment by default.
| OWASP item | Status | Where / rationale |
|---|---|---|
| Prompt injection (direct/indirect) | control + documented residual | Tool results on the wire are wrapped in TOOL_RESULT_BEGIN/END fences with inner fence markers neutralized (this fix); secrets scrubbed (E11). Residual: an injected instruction inside data can still steer the model — no framework fully solves this; delimiters + user-visible tool audit lines (what: …) limit blast radius. |
| Insecure output handling | control | Same fence treatment; result sizes capped (read_file 1 MiB cap, run_command 4000/2000 chars, job_poll 8 KiB window). |
| Excessive agency / tool misuse | control (defense-in-depth) + documented residual | Primary control: the per-call approval gate (every non-ReadOnly spawn needs consent). Secondary: check_destructive_argv blocklist refuses headline catastrophic patterns (teardown tools, recursive rm escaping the workspace, wrapper-nested variants) in BOTH run_command and job_start. The blocklist is best-effort by nature — wrapper/interpreter bypasses (sudo apt ..., timeout 5 dd ..., find / -delete, python -c rmtree) are expected and accepted residual risk; a real argv sandbox is a different product and explicitly out of scope. |
| Sensitive data disclosure (env) | gap → fixed here (ENV-1) | Children now inherit a scrubbed env: any var whose name contains API_KEY, _SECRET, _TOKEN, or PASSWORD is removed (GITHUB_TOKEN kept for gh/git auth — documented trade-off). |
| Supply chain (deps) | control | Cargo Audit + Trivy FS Scan required by rulesets on every PR and push. |
| Resource exhaustion | control + fixed here (JOB-2) | Provider 120 s timeout, subprocess 30 s ceiling, step budget 32, loop guard, read caps; job logs are read via bounded windows on poll (never a whole-log load). Since the 2026-09-18 scan triage, a ReadOnly poll no longer rewrites the log file — runaway log disk growth is bounded by the job's lifetime and clearing a runaway log is an operator action. |
| Session/message tampering | control | Append-only JSONL, CAS state transitions, seq monotonicity, torn-line recovery with truncation (#68); compaction never drops entries. |
| Injection via config | control (MCP shipped) | .cora.yaml/session headers are owner-controlled. MCP servers (#74) add an untrusted surface: server-supplied tool metadata is never trusted for risk tier (everything is Write → approval gate), tool results flow through the same fence/scrub path as native results, and the connection env is scrubbed. Residual: a malicious MCP server controls its own tool descriptions (model-visible) — treat server config as operator trust. |
| Identity & authz of sub-agents | N/A | tole v0 is single-agent, single-tenant. MCP servers are external tools, not sub-agents. |
| Human oversight | control | Risk-tiered approval gate (every non-ReadOnly call; Destructive never auto-allowed), #68 closed the replay-without-consent hole. |
GITHUB_TOKENstays in child env — gh/git push auth needs it; the token can therefore be read by any command the model runs. Mitigating context: the model already holds provider credentials for its own API calls by necessity; both are operator-scoped secrets on a single-operator host.- Workspace width is the operator's choice —
--workspace /is technically possible and a terrible idea; the flag validates existence, not sanity. Documented in CONTRIBUTING. - Symlink swap inside
tole-jobs/— requires a local attacker with write access to the operator's workspace, at which point host compromise is already achieved. Jobs dirs are created 0700 to reduce exposure; no further control planned. - Injection residual (above) — fences are a mitigation, not immunity; the model must still treat fenced content as data.
- Per-call Destructive classification heuristics (from RC-1) — needs a design discussion before code.
- MCP (#74, SHIPPED): stdio client live; trust-model extensions applied (metadata never trusted for risk, scrubbed env, fenced results). Remaining MCP surface: HTTP/SSE transport, server auth, resource subscriptions — each re-opens this document.
The session summary stored by the memory loop keeps its content contract
(first prompt + final answer only — never raw tool output), so the
prompt-injection-into-memory surface is unchanged. What #143 adds is a
TYPE: sessions that executed a Write/Destructive tool are stored with
--type decision and a wrote tag. The wrote bit is derived from the
harness's own tool-risk accounting (durable, session-scoped
fact.wrote_this_turn register — set on fresh execution AND
crash-replay, never reset by turn machinery), not from model-asserted
content — the model cannot claim or suppress the decision typing.
--on-turnend gates run owner-controlled scripts at the final-message
boundary. Two surfaces are deliberate and bounded: (1) the deny reason is
owner-authored input (the gate script is trusted the same way
--on-pretool scripts are — argv-executed, env-scrubbed, timed out), and
it enters the transcript as a user-role entry the model will read; (2) the
final-text preview handed to the gate is capped at 8 KiB and never leaves
the host process. Denials are capped at 3 per turn, so a permanently
failing gate cannot livelock the loop — the turn settles durably as
StopGateBlocked, visible to replay.