Run untrusted, AI-generated code on your own hardware. A Go library over Docker and gVisor, small enough to read in an afternoon — no control plane, no database, no scheduler. A sandbox is a container, and the container is the state.
openblox.sh · Docs · Getting started · Production · Architecture · Threat model · Security · Releasing · Discord
Status: pre-release (
0.x). The API will change; breaking changes bump the minor version and are listed in the changelog.
Linux on amd64 or arm64, Docker, and gVisor registered as the runsc
runtime — or a microVM runtime such as Kata, selected explicitly
(how).
The daemon. Your application never touches the Docker socket:
curl -fsSL https://openblox.sh/install.sh -o install.sh
less install.sh # the bytes you are about to run
sh install.shFetched once and run from disk, so what you read is what executes. Piping
curl straight into sh reads one response and runs another.
The script installs one binary and starts nothing. It verifies the published
checksum, and the Sigstore build attestation too when the gh CLI is present.
Its source is www/install.sh — that URL serves this file.
Pin a version in production, and pick your own target if you want one:
OPENBLOX_VERSION=v0.8.1 OPENBLOX_BIN_DIR=~/.local/bin sh install.shA dedicated host. On a fresh Debian 13 or Ubuntu 24.04 machine, one script installs Docker, gVisor, the daemon and its service, a firewall and, if you want remote callers, an mTLS listener. Then it proves the host with a sandbox:
curl -fsSL https://openblox.sh/setup.sh -o setup.sh && sh setup.shRunning it again is safe, and it writes openblox-uninstall.sh, which removes
exactly what it added. Sizing, remote callers and upgrades:
A dedicated host.
Benchmark your hardware (optional). Once the daemon runs, bench.sh times what a
caller feels: creating a sandbox, an exec, a Python job next to the same job
on the host. --mode stress steps concurrency up until the host tops out, and
names the level where it does:
curl -fsSL https://openblox.sh/bench.sh | sudo sh
curl -fsSL https://openblox.sh/bench.sh | sudo sh -s -- --mode stressWhat the numbers mean: Benchmarking.
The library. For a single process that may hold the Docker socket:
go get github.com/blox-eng/openbloxbackend, err := docker.New()
if err != nil {
return err
}
defer backend.Close()
// No options: no network, non-root, read-only rootfs, capped CPU/memory/PIDs,
// gVisor runtime, reaped when idle.
sb, err := backend.Create(ctx, "session-1",
sandbox.WithImage("ghcr.io/blox-eng/openblox-sandbox:latest"))
if err != nil {
return err
}
res, err := sb.Exec(ctx, sandbox.Command{
Argv: []string{"python3", "-c", "print(6 * 7)"},
})
fmt.Println(string(res.Stdout)) // 42That image is the reference sandbox userland; any image that
meets the contract works. :latest is fine here —
pin a digest anywhere it matters.
The example above imports the library, so your process holds the Docker socket,
which is root-equivalent on the host. In production, run the openbloxd daemon
on the host instead. It owns the socket, and your application talks to it over a
Unix socket with pkg/brokerclient — the same Backend interface, no Docker
access:
application ──unix socket──► openbloxd ──Docker API──► Docker + gVisor ──► sandbox
(no Docker access) (policy per profile,
not settable by requests)
Deployment, verification, compatibility, upgrades and troubleshooting:
Running in production. For a
machine of its own, setup.sh does all of it:
A dedicated host.
openblox is the layer below a sandbox platform, not a smaller one.
your scheduler, your tenancy, your API ← yours to build, if you ever need it
──────────────────────────────────────
openblox ← isolation, done correctly
──────────────────────────────────────
Docker + gVisor ← the boundary itself
One rule decides what belongs here:
How a sandbox is isolated is openblox's problem. Which sandbox runs where is yours.
Egress, capabilities, filesystem, resource caps, lifetime, runtime: openblox's. Placement, queueing, tenancy, metering, snapshots: not openblox's, and not planned. Build those on top when something actually asks for them. That is what a lower layer is for, and it is why there is no control plane to adopt first.
The comparison is libvirt, not OpenStack.
The zero value of every option is the most restrictive one. A sandbox created with no options gets:
| Isolation | gVisor (runsc) — syscalls handled in user space, not by the host kernel |
| Network | no external interface, so no egress and no DNS side channel |
| Filesystem | read-only root, non-root user (root is refused), noexec scratch |
| Resources | bounded CPU, memory (no swap), disk, process count, and captured output |
| Privileges | all capabilities dropped, no-new-privileges |
| Lifetime | commands killed at their timeout; sandboxes reaped when idle and at max age |
Relaxing anything is explicit and greppable at the call site. If the host cannot
provide the runtime asked for — gVisor, unless you chose a microVM runtime such as
Kata — Create fails with ErrRuntimeUnavailable; it never falls back to a
weaker boundary.
These are isolation measures, not a guarantee: the boundary is the runtime's —
gVisor's by default — and
THREAT_MODEL.md lists what is defended, the test behind each
claim, and what is not defended. Under Kata two of them differ: /dev/shm is
writable and executable, because Kata discards its mount options, and crash
recovery is unmeasured (details).
Two levels of the same guarantee. In the library, your code chooses: the
defaults are safe, and every relaxation is explicit and greppable at the call
site. Through openbloxd
the choice stops being the caller's at all — profiles live in the daemon's
config file and no request can reach them. A caller names a profile. It cannot
name an image, a runtime, a user, an egress policy, or a resource cap.
That is the difference between weakening being visible and weakening being unreachable, and it is the whole reason the daemon exists.
The claims in THREAT_MODEL.md are backed by pkg/conformance, the same
adversarial suite pkg/docker runs against itself, expressed against the
sandbox.Backend interface rather than against one implementation. Point it at
your own backend to find out how it scores — it does not skip:
func TestConformance(t *testing.T) {
cfg := conformance.Config{
Name: "mine",
New: func() (sandbox.Backend, error) { return mypkg.New() },
}
conformance.Run(t, cfg)
conformance.RunHostLocal(t, cfg)
}New takes no *testing.T on purpose: the suite, not the implementation,
decides what a construction failure means, and there is nothing to skip with.
It has only ever run against pkg/docker: under gVisor on every pull request,
and under Kata on amd64 in a separate workflow that does not gate merges, where
21 of 23 properties pass. See the
package doc for what each tier covers.
| Exec | run a command with a per-call timeout, get stdout, stderr, exit code |
| Files | read and write inside the sandbox without a shell round-trip |
| Processes | start a detached background command, idempotently |
| Preview links | HMAC-signed reverse proxy to a port inside the sandbox |
| Reaping | idle timeout and max age, enforced without a scheduler |
openbloxd |
a policy broker so callers never touch Docker |
Often you should. openblox adds no isolation of its own: the boundary is gVisor's.
- If your app is the only caller, and you trust it to pass the same dozen
docker run --runtime=runscflags on every call, you get the same sandbox at the same speed. - We measured that side by side. Create was 1154 ms with openblox and 1136 ms by
hand; exec
truewas 39 ms and 44 ms. Every operation was within noise.
What openblox adds is narrow:
- Those flags are the defaults. Without them, a gVisor container reaches the internet, has a writable root, keeps a capability bounding set, and has no memory, process or disk cap.
- The caller cannot weaken them.
openbloxdowns the policy, and a caller only names a profile. The caller never holds the Docker socket, which is root on the host, and the sandboxes can live on a host that holds none of your app's secrets. This is the main reason openblox exists. - The claims are tested. An adversarial suite probes a running sandbox from inside on every pull request.
- The lifecycle chores are done. Exec timeouts, output caps, idle and max-age reaping, and signed preview links.
If none of that matters to you, use gVisor directly. Method and numbers.
We measured openblox against Docker (runc and gVisor), Kata Containers, microsandbox, OpenSandbox and llm-sandbox on one 4-core, 3.8 GiB host, with the same image, workload and limits. openblox is not the fastest:
| openblox | the others | |
|---|---|---|
| Cold start | 1154 ms create, 39 ms exec | microsandbox is faster: 636 ms create, 20 ms exec. Kata took 8040 ms to create on this CPU. |
| Throughput, CPU-bound | 8.2 jobs/s at saturation | runc 9.4, microsandbox 9.5 (gVisor's syscall cost); Kata 5.8 |
| Host memory per idle sandbox | 41 MiB | microsandbox 74 MiB, Kata 205 MiB |
| Defaults | no network, read-only root, no capabilities, non-root, memory/process/disk caps | each needs options to get there, or cannot. OpenSandbox on gVisor cannot turn the network off. microsandbox's SDK "no network" is a deny policy with the device still attached. llm-sandbox runs as root by default. |
| What the caller holds | a profile name, over a unix socket or mTLS | the Docker socket (Docker, Kata, llm-sandbox); microsandbox's VM monitor, unconfined, inside the caller's own process tree; or an API key to a server that publishes each sandbox's command API on all interfaces by default (OpenSandbox) |
| Claims tested | adversarial suite on every pull request | none found that probes a running sandbox from inside |
Pick something else when:
- Cold start matters most. Pick microsandbox, and confine its VM monitor yourself: it runs with no seccomp, namespaces or cgroup.
- You want a separate guest kernel per sandbox. Pick Kata, but budget a slow create on older CPUs and about 200 MiB per VM.
- You only want a Python API over Docker. Pick llm-sandbox.
- You want a hosted-style platform. Pick E2B or Daytona, on hardware bigger than 4 GB.
Under gVisor, a sandbox that goes over its memory or process limit is killed as a whole, not just the offending process. E2B self-hosted, Daytona, CubeSandbox and Kubernetes agent-sandbox did not fit a 4 GB host, or are no longer self-hostable as open source. Full results, method and the scripts to reproduce them.
- You need tenants isolated from each other at the API. Every caller of one
openbloxdcan reach every sandbox; tenancy is yours to enforce in front of it. - You need a fleet. One host, one daemon. No scheduling, no fairness.
- You need protection from side channels between co-resident sandboxes. A microVM runtime such as Kata removes the shared kernel, not the shared CPU: microarchitectural side channels remain.
- You need snapshots, fork, pause/resume, or sub-second cold starts.
- You cannot run Linux with gVisor or a microVM runtime, or cannot keep that runtime patched.
Most of these are placement rather than isolation, which the rule above puts on your side of the line; the rest are trades made deliberately. None are gaps waiting to be filled. They are the boundary that keeps openblox small enough to be worth reading, and requests to cross it get declined on that basis. See ARCHITECTURE.md for the reasoning.
Linux on amd64 and arm64, both tested natively in CI against a real gVisor
runtime. Docker Engine with runsc registered, or a microVM runtime such as Kata
selected explicitly; Kata is measured on amd64 in a separate workflow that does not
gate merges. Go 1.25+ for library users. The
compatibility matrix covers
openbloxd, clients, images, Docker and the runtime.
The badges above are measured from the source on every deploy, not typed here.
Three direct dependencies (docker/docker, containerd/errdefs, yaml.v3).
CI runs lint, race-enabled tests, CodeQL and govulncheck (gating on newly
reachable vulnerabilities), plus the integration and conformance suites against
a real gVisor runtime on amd64 and arm64 — including attacks on the network,
filesystem, privileges, resource caps and timeouts.
Every release is cut by CI from a verified commit. The openbloxd binaries are
reproducible and ship with SBOMs; binaries and the sandbox image carry Sigstore-signed
build provenance. RELEASING.md shows how to
verify them.
Written for Blox, where it is the only sandbox backend and
replaced a hosted platform. Its own production rollout is gated on migrating its
callers off the Docker socket and onto openbloxd. No support SLA.
Report vulnerabilities privately — see SECURITY.md. Please do not open a public issue.
Issues and PRs welcome. See CONTRIBUTING.md. Commits follow Conventional Commits; CI enforces it. For a question that is not an issue, there is a Discord.
MIT — see LICENSE.
The openbloxd binaries are statically linked, so they also carry the code of
their dependencies. Every release attaches THIRD_PARTY_LICENSES.txt with the
full licence text of each linked module, generated from what is actually in the
binary. Build it yourself with make licenses.