Model or dataset
akitaonrails/ai-jail avatar
akitaonrails/ai-jail

ai-jail: an OS sandbox for AI coding agents on Linux and macOS

Multi-OS sandbox to run AI agents with better constraints (it is not 100% secure, but enough)

1,308 stars120 forksRustGPL-3.0

At a glance

What is it?
ai-jail wraps a coding agent in bubblewrap, Landlock and seccomp on Linux, and sandbox-exec on macOS, so the agent sees a project directory and little else. It is a containment layer, not a disposable VM.
Who is it for?
Adopt ai-jail if you already run an agent against a repository and want the project directory writable while network, GPU, display, X11, host shared memory and credential state stay off unless you ask for them. Do not adopt it as a substitute for a disposable VM when the code being executed is hostile, and do not use it on Windows, which the README says is unsupported; WSL2 with the Linux backend is the stated path.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem ai-jail targets: an agent that can read your whole home directory

A coding agent runs as your user. It can read your SSH keys, your shell history, your cloud credentials and every repository on the machine, and with a network connection it can send any of that somewhere else. The usual mitigation is a container, but a container built for development is normally a convenience wrapper around the same filesystem, not a boundary.

ai-jail takes the opposite default. The project directory is writable; host capabilities are not. Private home is on by default, so the agent gets a fresh tmpfs $HOME instead of your real one. Agent credential state is not mounted unless you pass --agent-state. Network, GPU, display, X11, host /dev/shm, terminal passthrough, Docker, SSH, Tailscale and the systemd user bus are all off by default. The intended user is a developer who wants to hand an agent a checkout without handing it the machine.

The README is direct about the ceiling: this is a useful layer, not a replacement for a disposable VM when running hostile code. That sentence should govern how you deploy it.

How the sandbox is built: bubblewrap, Landlock, seccomp, sandbox-exec

On Linux the mechanism is bubblewrap plus Landlock and seccomp. bubblewrap constructs the mount and namespace view the agent sees; Landlock restricts filesystem access from inside; seccomp filters syscalls. The Rust crate dependencies match that split: landlock and seccompiler are pulled in only for cfg(target_os = "linux"), while nix supplies signal, process, poll, term and resource handling on Unix generally. macOS goes through Apple's deprecated /usr/bin/sandbox-exec interface, which the README names as deprecated rather than hiding.

The interesting part is how flags map to mounts rather than to policy language. --display mounts only the validated Wayland socket and explicitly never all of XDG_RUNTIME_DIR; X11 is a separate flag, and the README states that X11 access permits keylogging and screenshots. --terminal-passthrough forwards raw terminal data, whereas the default filters output through a VT parser, which the README says reduces exposure of the terminal clipboard, query and parser surface. --worktree makes the per-worktree git dir and the shared common dir writable so the agent can commit, and --lockdown keeps both read-only. Each flag is documented with its consequence, not just its name.

The environment is also a boundary. The default is a minimal allowlist; --inherit-env passes the full parent environment, secrets included. That is the flag most likely to undo the rest of the configuration by accident.

Installing ai-jail and running an agent for the first time

The README lists four install paths. Homebrew on macOS, an Arch package, crates.io, and a Nix flake that sets BWRAP_BIN automatically. Pick one and stop; there is no reason to mix them.

Configuration files, fail-closed behaviour and the .ai-jail directory

Persistent settings live in a global TOML file at ~/.ai-jail, keyed by command. The README shows agent_state being enabled for claude there, which means you can stop typing the flag on every launch:

toml
# ~/.ai-jail
[commands.claude]
agent_state = true

The behaviour worth knowing is failure mode. Existing unreadable or invalid project or global configuration fails closed rather than launching with a weakened policy. A typo in ~/.ai-jail therefore stops the run instead of silently dropping a restriction. Bootstrap output is always mode 0600. The first ordinary run may create a .ai-jail directory in the project, and --dry-run never writes it, which makes --dry-run the safe way to inspect what a configuration would do before it does it.

The README does not document a rollback or an uninstall path for the state a run creates, so treat .ai-jail as something you add to .gitignore deliberately rather than something the tool cleans up for you.

Where ai-jail is the wrong tool

Three cases stand out. First, hostile code: the README itself says this is not a replacement for a disposable VM when running hostile code, so if the agent is executing untrusted input, the sandbox is a layer and not the answer. Second, --docker. The README states that --docker mounts an actual Unix Docker socket and is effectively host-root, because the daemon can create host-mounted containers. Enabling it discards most of what the rest of the flags bought.

Third, --allow-tcp-port. It remains accepted for backward compatibility, but launch fails closed because UDP cannot be securely constrained through this option. Anyone upgrading from an older configuration that used it will find runs refusing to start rather than silently narrowing. The correct replacement, per the README, is --network, and only when unrestricted network access is explicitly what you want, since --network permits full network exfiltration of any readable data.

Windows is not supported at all. The README points to WSL2 and the Linux backend inside it.

The BWRAP_BIN check and what it says about the threat model

ai-jail does not simply trust a path handed to it through BWRAP_BIN. The README describes the acceptance rule in detail: the variable is accepted only when it canonically resolves to a root-owned executable that is not group- or world-writable, or to an executable with no write bits under a /nix/store whose own owner is root, is not world-writable, and carries the sticky bit if it is group-writable, the standard multi-user store layout at mode 1775. A group-writable store without the sticky bit is refused, because a group member could replace the binary. A single-user store owned by the invoking user does not qualify either.

That is a narrow, specific rule, and it tells you the project is defending against a swapped bubblewrap binary rather than only against the agent. It also creates a practical constraint: if you build bubblewrap yourself into a writable prefix, ai-jail will refuse it, and the Nix flake exists partly to sidestep that by setting BWRAP_BIN to a store path that satisfies the rule.

How ai-jail compares with a development container

Docker or Podman is the obvious alternative, and the difference is not speed. A development container isolates by giving the workload a different filesystem image and, typically, mounting the source tree into it. The agent still runs inside a full userland with a package manager, and the boundary is the image plus whatever mounts you configured. Getting the agent to authenticate usually means passing credentials into the container.

ai-jail isolates by subtraction. It runs the agent binary you already have, on your host, but with a constructed view: tmpfs home, explicit mounts, capabilities off by default, and a minimal environment allowlist. There is no image to build and no second toolchain. The trade-off is that the isolation primitives are OS-level and platform-specific, so the Linux and macOS paths are genuinely different implementations, and macOS relies on an interface Apple has deprecated. If you need one reproducible environment across a team, a container is the better fit. If you want your existing checkout and your existing agent with fewer capabilities, ai-jail is aimed at exactly that.

Editorial conclusion

Adopt ai-jail if you already run an agent against a repository and want the project directory writable while network, GPU, display, X11, host shared memory and credential state stay off unless you ask for them. Do not adopt it as a substitute for a disposable VM when the code being executed is hostile, and do not use it on Windows, which the README says is unsupported; WSL2 with the Linux backend is the stated path. Before trusting a run, check three things: that --dry-run shows the mounts you expect, that your ~/.ai-jail parses, since invalid configuration fails closed, and whether you actually need --agent-state, because mounting Claude's ~/.claude and ~/.claude.json exposes that login material to everything inside the sandbox.

Frequently asked questions

What is ai-jail?

ai-jail is a multi-OS sandbox that runs AI coding agents with tighter constraints, using bubblewrap plus Landlock and seccomp on Linux and sandbox-exec on macOS. The README describes it as a useful layer rather than a replacement for a disposable VM when running hostile code.

How do I install ai-jail on Linux or macOS?

The README lists Homebrew via the akitaonrails/tap tap, the Arch packages ai-jail-bin and ai-jail, cargo install --locked ai-jail, a Nix flake, and signed GitHub release archives. Linux additionally requires bwrap from the bubblewrap package.

Does ai-jail need the bubblewrap binary on Linux?

Yes. The README says Linux requires bwrap, installed through pacman, apt or dnf. BWRAP_BIN is accepted only when it canonically resolves to a root-owned executable that is not group- or world-writable, or to a qualifying /nix/store path.

Can ai-jail run on Windows?

No. The README states that Windows is not supported and directs users to WSL2 with the Linux backend inside it.

Does ai-jail give the agent network access by default?

No. Network is off by default, and --network enables unrestricted access, which the README notes permits full network exfiltration of any readable data. The status bar update check is also off by default and is the only documented outbound request.

Official sources

  1. akitaonrails/ai-jail on GitHub
  2. Issues
  3. License: GPL-3.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/akitaonrails-ai-jail.svg)](https://hysenlabs.com/projects/akitaonrails-ai-jail)