# Zeroboot: sub-millisecond VM sandboxes for AI agents, built on copy-on-write forks

> Zeroboot forks a pre-booted Firecracker snapshot into a fresh KVM virtual machine in under a millisecond, so an agent can get a real hardware-isolated sandbox per code execution. It is a working prototype with a self-host path and a managed API, and its own README lists the sharp edges.

**zerobootdev/zeroboot** — Sub-millisecond VM sandboxes for AI agents via copy-on-write forking

- Repository: https://github.com/zerobootdev/zeroboot
- Website: https://zeroboot.dev
- Stars: 2,452 · Forks: 110
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/zerobootdev-zeroboot

## The problem Zeroboot targets: per-execution isolation without container latency

An agent that writes and runs code needs somewhere to run it. Containers share a kernel, so a kernel-level escape from one sandbox is an escape from the host. Full virtual machines give hardware-enforced memory isolation, but booting one per execution is far too slow for a tool call. Zeroboot's answer is to boot once and fork many times: the README describes each fork as a separate KVM virtual machine with hardware-enforced memory isolation, created from a snapshot instead of from a boot sequence.

The audience is narrow and specific. You need Linux with KVM, and you need a workload where a sandbox lives for one execution or a short burst. A long-running service that wants persistent state, package installs at runtime, or a network connection inside the sandbox is a poor fit, and the README says so. The project is a Rust binary, edition 2021, licensed Apache-2.0, with Python and TypeScript client SDKs under sdk/python and sdk/node.

## How copy-on-write forking turns a snapshot into a live VM

The mechanism is three steps, and the README states them directly. Template creation is a one-time cost: Firecracker boots a VM, pre-loads your runtime, and snapshots memory plus CPU state. Forking then creates a new KVM VM, maps the snapshot memory with mmap(MAP_PRIVATE) so pages are copy-on-write, and restores all CPU state. The third step is the property that matters for untrusted code: each fork is its own KVM VM, so memory isolation is enforced by hardware rather than by namespaces or a shared kernel.

The dependency list in Cargo.toml matches that design. kvm-ioctls and kvm-bindings drive the hypervisor directly, vmm-sys-util and libc supply the low-level pieces, nix covers fs, mman and ioctl, and axum with tokio serves the HTTP API. A guest/ directory holds the init program, built by the Makefile with musl-gcc into a static binary, which is what runs inside the template.

The README's benchmark table reports spawn latency at 0.79ms p50 and 1.74ms p99, memory per sandbox at roughly 265KB, fork plus exec of Python at about 8ms, and 1000 concurrent forks at 815ms. Those are the project's own published numbers; the README does not describe the hardware or the methodology behind them, so where they matter to you they are claims to check against your own host rather than settled figures.

## Installing Zeroboot and running your first sandbox

The README points to docs/DEPLOYMENT.md for deployment and describes self-hosting as "on any Linux box with KVM". The Makefile is the entry point for a local build. After cloning the repository, build the Rust binary and the guest init program:

```bash
make build
make guest
```

The first target runs cargo build --release, producing target/release/zeroboot. The second compiles guest/init.c with musl-gcc into a static guest/init binary. That guest binary is what the template boots, so it has to exist before you snapshot anything. If musl-gcc is missing, the guest target is the step that fails.

With both artifacts in place, the Makefile's serve target starts the API against a working directory:

```bash
make serve
```

That expands to ./target/release/zeroboot serve workdir, so the server reads and writes its state under workdir. The README does not document the port the server listens on, and it does not describe a configuration file; docs/API.md and docs/DEPLOYMENT.md are the places to check for the request shape and the bind address.

If you would rather not run infrastructure yet, the README's quickest path is the managed API, which needs no local KVM. This is the exact request from the README, with the demo key it publishes:

```bash
curl -X POST https://api.zeroboot.dev/v1/exec \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer zb_demo_hn2026' \
  -d '{"code":"import numpy as np; print(np.random.rand(3))"}'
```

The response body is not shown in the README, so the exact JSON shape is something to read off the live call or from docs/API.md. The Python SDK is the shorter route for a real integration:

```python
from zeroboot import Sandbox
sb = Sandbox("zb_live_your_key")
result = sb.run("print(1 + 1)")
```

The TypeScript SDK under sdk/node follows the same shape, constructing a Sandbox with an API key and awaiting run. Note that the demo key in the README is a shared credential for trying the endpoint, not a key for your own workloads.

## The limitations the README admits: randomness, one vCPU, no network

The most interesting limitation is the one that follows from forking itself. Forks inherit the CSPRNG state of the snapshot, so every sandbox created from the same template starts with the same generator state. The README says kernel entropy is reseeded via RNDADDENTROPY, but that userspace PRNGs such as numpy and OpenSSL need explicit reseeding per fork, and it links to Firecracker's guidance on randomness for clones. For code that generates keys, tokens, or nonces, that is not a footnote; it is a correctness and security requirement you have to handle in the guest. It is also the clearest sign that this is a prototype rather than a hardened service.

The other three limits are structural. Each fork gets a single vCPU; multi-vCPU is described as architecturally possible but not implemented. There is no networking inside forks, so sandboxes communicate via serial I/O only, which rules out pip installs, HTTP calls, and anything that expects a socket. And template updates require a full re-snapshot of roughly 15 seconds, with no incremental patching, so a change to your runtime image is a stop-and-rebuild event rather than a rolling update.

Taken together, these constraints define the tool's shape: short, deterministic, CPU-bound snippets with no network and no cryptographic randomness. Anything else needs a different sandbox.

## Zeroboot and Firecracker's own snapshotting compared with container sandboxes

The closest alternative is Firecracker used directly with its snapshot and clone support. Zeroboot is built on Firecracker, and the README links to Firecracker's documentation on randomness for clones. The difference is what each gives you out of the box. Firecracker provides the VMM and the snapshot primitive; you still write the fork orchestration, the API, the SDKs, and the guest init. Zeroboot packages that layer: a Rust server with an axum HTTP API, Python and TypeScript clients, a Makefile that builds the guest, and a template workflow. If you already run Firecracker and have written your own clone path, Zeroboot duplicates work you have done. If you have not, it is the difference between a hypervisor and a service.

Against container-based sandboxes such as the ones the README's benchmark table names, the difference is the isolation boundary. A container shares the host kernel, so its isolation is a kernel feature; a Zeroboot fork is a KVM VM, so its isolation is hardware-enforced memory separation. That is a stronger boundary, and it comes with the costs listed above: no networking, one vCPU, and a snapshot-based lifecycle. The README's table also reports spawn latency and memory per sandbox for those other tools, but it does not state where those figures came from, so the only way to use them is to put them beside numbers you obtain yourself.

## Maintenance, licensing, and what a template change costs

The last push to the repository was on 2026-03-21, and the repository is not archived. The README labels the project a working prototype: the fork primitive, benchmarks, and API are described as real, but not production-hardened. There are no retrieved releases, and Cargo.toml still carries version 0.1.0, so expect to track main rather than pin a tagged version. The README invites issues if you are interested in the project, and points to a Tally form for early access to the managed service, which suggests the hosted offering is not generally available yet.

The licence is Apache-2.0, declared in Cargo.toml and shipped as LICENSE. That is a permissive licence with an explicit patent grant, which matters if you embed the fork logic in a commercial product. It does not settle the questions you would ask about the managed API, such as where code runs and what happens to it; the README says nothing about data handling or retention, so those are questions for the early-access process rather than for the repository.

The operational cost is dominated by template updates. A full re-snapshot takes roughly 15 seconds and there is no incremental patching, so every change to the pre-loaded runtime is a rebuild of the template, not a patch. Plan for a small number of templates that change rarely, and treat the README's latency table as something to reproduce on your own hardware before you size anything against it.

## Conclusion

Adopt Zeroboot if you run untrusted, short-lived code for agents and you have Linux hosts with KVM, or if you want to try the managed API before standing up infrastructure. Do not adopt it if your workloads need networking inside the sandbox, multiple vCPUs, or reproducible randomness without per-fork reseeding, because the README lists all three as open limitations. Verify first that your host exposes /dev/kvm, that you can run make build and make guest, and that you accept a full re-snapshot of roughly 15 seconds for every template change.

## FAQ

### What is a Zeroboot VM sandbox?

It is a KVM virtual machine created by forking a pre-booted Firecracker snapshot, with snapshot memory mapped copy-on-write and CPU state restored. The README describes each fork as a separate KVM VM with hardware-enforced memory isolation, created in under a millisecond.

### How do I install and self-host Zeroboot?

The README says to self-host on any Linux box with KVM. The Makefile provides make build for the Rust binary, make guest to compile guest/init.c into a static guest binary, and make serve to run ./target/release/zeroboot serve workdir. docs/DEPLOYMENT.md is the referenced guide.

### Can Zeroboot sandboxes use the network or multiple vCPUs?

No. The README lists both as known limitations: there is no networking inside forks, so sandboxes communicate via serial I/O only, and each fork has a single vCPU, with multi-vCPU described as architecturally possible but not implemented.

## Sources

- [Issues](https://github.com/zerobootdev/zeroboot/issues)
- [License: Apache-2.0](https://github.com/zerobootdev/zeroboot/blob/main/LICENSE)
- [Project website](https://zeroboot.dev)
- [README](https://github.com/zerobootdev/zeroboot/blob/main/README.md)
- [zerobootdev/zeroboot on GitHub](https://github.com/zerobootdev/zeroboot)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zerobootdev-zeroboot
