forkd: forking live microVMs for AI agent fan-out
Fork() for AI agent microVMs. Spawn 100 children in ~100ms from a warm parent; BRANCH a live VM in ~150ms. KVM-isolated, snapshot CoW.
At a glance
- What is it?
- forkd is a Rust microVM runtime built on Firecracker that spawns children from a warmed parent snapshot instead of cold-booting a kernel. The promise is per-child KVM isolation at fork(2)-like cost, and the documentation is candid about where that promise gets expensive.
- Who is it for?
- forkd fits teams running many short-lived, mutually isolated agent sandboxes on Linux hosts they control, and who can accept a vendored Firecracker fork and a per-link cost on deep snapshot chains. It is the wrong tool if you need macOS or Windows hosts, unprivileged operation, or a single stable upstream VMM.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The cold-boot tax that forkd exists to remove
Agent fan-out turns one task into many parallel sandboxes. The usual way to get those sandboxes is to boot a VM per child, and booting a kernel, mounting a rootfs and importing a Python interpreter is the dominant cost when the actual work is a few hundred milliseconds of tool calls. forkd targets that gap directly. The README frames the project as "a microVM sandbox runtime for AI agent fan-out", where children fork from a warmed parent snapshot and inherit its address space copy-on-write rather than booting their own kernel.
The intended user is someone running agent workloads that need real isolation per child, not threads or containers sharing a kernel. The parent VM boots once, imports the runtime (the README lists Python plus dependencies, a JIT-warmed JVM, an already-loaded ML model) and is paused to disk. Everything after that is a fork. The headline claim in the README is 100 microVMs in 101 ms, and a BRANCH of a live VM in 56 ms p50 on a 1.5 GiB source. Those are the project's own numbers from its bench directory, not an independent measurement.
How the copy-on-write fork actually works
forkd is built on Firecracker, and the mechanism is the interesting part. Each child is a separate Firecracker process that mmaps the parent's memory image with MAP_PRIVATE. The kernel then implements copy-on-write at the page level, so children share the parent's resident memory until they write to it and diverge. That is what buys per-child KVM isolation at a spawn cost closer to fork(2) than to a cold boot.
The second mechanism is BRANCH, which snapshots a running sandbox rather than a paused one. The README describes v0.4 live BRANCH as collapsing the source-pause window from roughly 200 ms (the older Diff mode) to 56 ms p50 and 64 ms p90, and states that the pause is disk-independent because the memory copy runs after resume rather than during it. That is a real architectural distinction: a slow disk hurts Diff mode and does not hurt Live mode. The prerequisite is that the source booted with live_fork=True, which gives it memfd-backed RAM so UFFD_WP can observe writes from the running parent. Miss that flag and the live path is not available to you.
The third mechanism arrived in v0.5: diff-snapshot chains. Each layer records a parent_tag and a content-hash edge to the layer below, and the daemon walks the chain at spawn time to assemble the memory image in one pass. This is aimed at the case where an agent caches pip install numpy, pip install pandas and pip install scikit-learn as separate snapshots and you do not want three copies of the same 1.5 GiB base.
Installing forkd and taking a first fork
The README does not give a cargo install line or a release tarball walkthrough, so the honest starting point is the environment check. forkd requires Linux 5.7 or newer, the sysctl vm.unprivileged_userfaultfd=1 (or CAP_SYS_PTRACE), and the vendored Firecracker fork from deeplethe/firecracker on the branch forkd-v0.4-mem-backend-shared-v1.12. The README states that forkd doctor probes both the kernel setting and the Firecracker build, so run it before anything else.
sudo -E forkd doctorIf doctor passes, the CLI fork path spawns live-fork-capable children locally. The README notes that this path and the daemon-tracked path do not compose yet, and points at issue #209 for the status of daemon-side spawn from the CLI.
sudo -E forkd fork --tag pyagent -n 1 --per-child-netns --live-forkSnapshotting a running sandbox uses the live mode with the wait disabled, which returns before the background memory copy finishes.
sudo -E forkd snapshot --from-sandbox <sb-id> --live --no-waitFrom Python, the same flow goes through the Controller class. Note the comment in the README's own example: the parent must boot with live_fork=True, and after a wait=False branch the call returns with status="writing", so you poll list_snapshots until status="ready".
from forkd import Controller
c = Controller()
parent = c.spawn_sandboxes("pyagent", n=1, live_fork=True)[0]
branch = c.branch_sandbox(parent["id"], mode="live", wait=False)For the chain workflow, snapshot-diff builds layers and fork consumes the chain head. The daemon walks the edges and verifies parent content hashes, so the caller sees a single POST /v1/sandboxes round-trip.
forkd snapshot-diff --from py-base --tag py-numpy --exec "pip install numpy==2.0.2"
forkd snapshot-diff --from py-numpy --tag py-pandas --exec "pip install pandas==2.2.3"
forkd fork --tag py-pandas -n 1The Dockerfile in the repository builds only the controller daemon, and it is explicit that the image does not contain Firecracker or KVM tooling. The comment states the container is expected to run with --privileged --network=host --pid=host, or equivalent capabilities, because the daemon manages netns and cgroups.
Where the fork model gets expensive
The chain feature is also where the cost curve turns ugly, and the project publishes the numbers. In the v0.5 Phase 5 bench on a 512 MiB base, ext4, i7-12700, a flat base spawns at 59 ms p50. Add one numpy layer and the depth-1 head spawns at 751 ms, a per-link tax of +692 ms. Depth 2 lands at 1222 ms, depth 3 at 1668 ms. The README attributes the per-link tax to SHA-256 of the base, around 460 ms, which means the tax is proportional to the size of the base image rather than to the layer you added. A flat-equivalent snapshot holding all three packages in one diff spawns at 746 ms, which is roughly the depth-1 number.
That comparison is the real design tension. Stacking layers saves disk and lets you reuse a layer across chains, but a deep chain is slower to spawn than a single flattened diff at the same content. The project acknowledges this by shipping forkd snapshot-compact, which flattens a chain from one tag to another. If your agent spawns thousands of children from a chain head, the per-link tax is paid on every spawn, and compaction is not optional.
The second limitation is the environment. This is a Linux-only, KVM-only runtime. The README's requirement list includes a specific vendored Firecracker branch, which means you are tracking a fork rather than upstream Firecracker, and the repository carries a patch file at the top level named for the MAP_SHARED memory backend option. The Dockerfile's privileged invocation is another boundary: the daemon needs CAP_NET_ADMIN and CAP_SYS_ADMIN for netns and cgroup work, so this is not something you drop into a locked-down multi-tenant cluster without thinking about what the daemon can reach. The README also does not document rollback for a failed snapshot or unpack, so plan your own recovery story rather than expecting one in the docs.
forkd against Firecracker snapshot restore and gVisor
The closest comparison is plain Firecracker with its own snapshot and restore support. Firecracker can already snapshot a microVM and restore it, and that gets you a warm start. The difference in approach is what happens when you restore many children from one parent. forkd's children mmap the parent's memory image with MAP_PRIVATE and share resident pages until they diverge, so the marginal memory cost of the hundredth child is the pages it actually touches. A naive restore-per-child model materializes each guest's memory separately. The README's framing of the result is "per-child KVM isolation, and a spawn cost that's closer to fork(2)", and that combination is the specific thing forkd adds on top of stock Firecracker.
The other axis is gVisor, which takes a different route entirely: it reimplements the Linux system call surface in userspace and runs the workload inside that, rather than giving each workload its own KVM-backed kernel. gVisor starts faster than a cold VM and needs no KVM, which makes it a better fit on hosts where you cannot get /dev/kvm or where you do not want hardware virtualization in the trust boundary. What you give up is kernel fidelity: syscalls are emulated, and anything relying on unusual kernel behaviour may not work. forkd goes the other way, keeping a real kernel per child and paying for it with a KVM requirement and a Firecracker fork to track.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-02, so the project is under current work. The release cadence visible in the release list is v0.5.1 on 2026-06-05, v0.5.2 on 2026-06-08 and v0.5.3 on 2026-07-22, with the workspace version in Cargo.toml pinned at 0.5.3. The default branch is dev, not main, which is worth knowing if you script against the repository.
The upgrade cost is unusual for a Rust project because it is not just a Cargo bump. The README ties the live BRANCH path to a named Firecracker fork branch, forkd-v0.4-mem-backend-shared-v1.12, and the top-level patch file exists to carry the MAP_SHARED memory backend option. The FIRECRACKER-UPSTREAM-PROPOSAL.md file at the repository root suggests the intent is to upstream that change, but as of the README the project still directs users to the fork. Until that lands, upgrading forkd can mean rebuilding Firecracker as well.
The licence is Apache-2.0, declared in both the LICENSE file and the workspace package metadata. Apache-2.0 is permissive and includes an explicit patent grant, and the repository also carries a NOTICE file, which is the file Apache-2.0 expects downstream redistributors to preserve. If you ship a product that embeds the controller binary, read the NOTICE and keep it with your distribution. That is a description of what the licence requires, not legal advice; get counsel for your own case.
Editorial conclusion
forkd fits teams running many short-lived, mutually isolated agent sandboxes on Linux hosts they control, and who can accept a vendored Firecracker fork and a per-link cost on deep snapshot chains. It is the wrong tool if you need macOS or Windows hosts, unprivileged operation, or a single stable upstream VMM. Before adopting it, run forkd doctor on the target host, confirm vm.unprivileged_userfaultfd=1, and build the Firecracker branch forkd-v0.4-mem-backend-shared-v1.12, since the README states the stock binary will not do.
Frequently asked questions
What does forking mean in forkd?
In forkd, forking means spawning a child microVM from a warmed parent snapshot so the child inherits the parent's address space copy-on-write instead of cold-booting its own kernel. Each child is a separate Firecracker process that mmaps the parent's memory image with MAP_PRIVATE.
What is the difference between forking and spooning in forkd?
The documentation does not describe a spooning mode or any counterpart to forking, so there is nothing to compare. forkd documents two related operations: fork, which spawns children from a parent snapshot, and BRANCH, which snapshots a running sandbox and resumes it.
Does forkd work on macOS or Windows?
No. The README requires Linux 5.7 or newer, KVM, vm.unprivileged_userfaultfd=1 or CAP_SYS_PTRACE, and the vendored Firecracker fork, all of which are Linux and KVM specific.
Why is spawning from a deep forkd snapshot chain slower?
The v0.5 bench shows a per-link tax on each chain layer, which the README attributes to SHA-256 of the base image at roughly 460 ms. On a 512 MiB base, spawn p50 goes from 59 ms at depth 0 to 751 ms at depth 1 and 1668 ms at depth 3, so forkd ships snapshot-compact to flatten a chain.
Can forkd run in Docker?
The repository Dockerfile builds the forkd-controller daemon, and its comment states the image does not contain Firecracker or KVM tooling because those must be available on the host kernel. The documented run command uses --privileged --network=host --pid=host so the daemon can manage netns and cgroups.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/deeplethe-forkd)