# NVIDIA OpenShell: a sandboxed runtime for autonomous AI agents

> OpenShell puts each coding agent in its own container and routes every outbound connection through a YAML policy engine. Here is how the install works, what the policy model actually enforces, and where the alpha status shows.

**NVIDIA/OpenShell** — OpenShell is the safe, private runtime for autonomous AI agents.

- Repository: https://github.com/NVIDIA/OpenShell
- Website: https://docs.nvidia.com/openshell/latest/
- Stars: 10,338 · Forks: 1,400
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-openshell

## The problem OpenShell targets: agents with shell access and no egress boundary

An autonomous coding agent is a process that reads your repository, calls a model API, and writes files. Give it a shell and it can also read anything else your user account can read, and post anything it finds to any host that answers. OpenShell's answer is to run the agent inside a container whose outbound traffic is intercepted by a policy engine, with filesystem and process constraints applied underneath. The README frames the project as "the safe, private runtime for autonomous AI agents" and lists what it protects: data, credentials, and infrastructure.

The audience is narrower than the tagline suggests. OpenShell is for teams already running agents such as claude, opencode, codex, or copilot on developer machines or in a cluster, and who want the network boundary to be a file in the repository rather than a paragraph in a system prompt. It is not a model, a harness, or an agent framework. It is the box the agent runs in, plus the control plane that manages boxes.

## Inside the architecture: gateway, sandbox, policy engine, provider access

Four components appear in the README's table. The gateway is the control-plane API that coordinates sandbox lifecycle and acts as the auth boundary. The sandbox is the isolated runtime with container supervision and policy-enforced egress routing. The policy engine enforces filesystem, network, and process constraints "from application layer down to kernel." Provider access handles profile-defined endpoints, binary policy, and endpoint-bound credential injection for model APIs and other services.

The data flow matters more than the component list. Every outbound connection from a sandbox is intercepted by the policy engine, which does one of three things according to the README: allows the request when the destination and binary match a policy block, binds credentials to endpoints by injecting provider credentials only after policy admits a request to a profile-authorized endpoint, or denies the request and logs it. That middle case is the interesting design choice. The agent never holds the API key. The proxy holds it and attaches it after the request has already passed policy, so a prompt-injected agent cannot exfiltrate a credential it never possessed.

Sandbox lifecycle runs through a configured compute driver. Supported compute platforms listed in the README are Docker, Podman, MicroVM, and Kubernetes. The gateway is a Rust workspace under crates/, with protobuf definitions in proto/ and a Python SDK in python/ that speaks gRPC, which is why pyproject.toml depends on grpcio and protobuf.

## Four protection layers, and why only two of them reload

The README splits policy into four domains and is explicit about when each applies. Filesystem policy prevents reads and writes outside allowed paths and is locked at sandbox creation. Process policy blocks privilege escalation and dangerous syscalls, also locked at creation. Network policy blocks unauthorized outbound connections and is hot-reloadable at runtime. Provider attachments grant endpoint-bound credentials and network access, also hot-reloadable.

That asymmetry is the part worth internalizing before you design a policy. You can tighten or widen egress while an agent session is live, which the quickstart demonstrates by applying a policy and reconnecting without recreating the sandbox. You cannot do the same for the filesystem or process rules. If your agent needs a path that was not in the creation-time policy, the fix is a new sandbox, not a policy edit. Treat the static layers as the shape of the box and the dynamic layers as the dial you turn during a session.

The network layer is enforced at HTTP method and path granularity. The README's example shows a GET to api.github.com succeeding while a POST to a specific repository path returns a policy_denied error naming the method and path. That is L7 enforcement rather than a host allowlist, and it is the reason a single allowed domain does not automatically mean write access to it.

## Installing OpenShell and running a first sandbox

Prerequisites come first. The README lists a supported host (Linux, macOS on Apple Silicon, or Windows with WSL 2 marked experimental) and a local runtime: Docker, Podman, or host virtualization enabled for MicroVM-backed sandboxes. There is no native Windows install path documented; the Windows route is WSL 2.

The recommended install is the binary installer script, which pulls the latest stable release unless OPENSHELL_VERSION is set:

```bash
curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh
```

A dev release that tracks the latest commit on main is also published. Note that the openshell package on PyPI is the Python SDK only and does not install the CLI, so adding it to a project is a separate step:

```bash
uv add openshell
```

For a cluster, the README publishes an OCI Helm chart to GHCR and flags that path as experimental, with breaking changes expected:

```bash
helm install openshell oci://ghcr.io/nvidia/openshell/helm-chart
```

With the CLI in place, create a sandbox and name the agent to run inside it. The README shows claude, opencode, codex, and copilot as options:

```bash
openshell sandbox create -- claude
```

The base sandbox image ships python 3.14, node 22, gh, git, vim and nano, plus ping, dig, nslookup, nc, traceroute, and netstat. Then watch policy do something. A fresh sandbox starts with minimal outbound access, so a curl inside it fails at the proxy:

```bash
sandbox$ curl -sS https://api.github.com/zen
curl: (56) Received HTTP code 403 from proxy after CONNECT
```

Exit, apply the quickstart policy from the host, and reconnect. The README's walkthrough applies examples/sandbox-policy-quickstart/policy.yaml with a wait flag:

```bash
openshell policy set demo --policy examples/sandbox-policy-quickstart/policy.yaml --wait
openshell sandbox connect demo
```

After reconnecting, the GET to api.github.com returns a zen quote while a POST to the same host is rejected with a policy_denied error. The repository also ships an automated demo of the same flow:

```bash
bash examples/sandbox-policy-quickstart/demo.sh
```

## Where OpenShell is the wrong tool

The project labels itself alpha, and the README repeats that on the Kubernetes path: "Expect rough edges and breaking changes." If you need a stable API surface for a production agent platform, the versioning alone should give you pause. The most recent tagged release documented is v0.0.116, and a development build is published alongside it, which tells you the interface is still moving.

There is also a real cost to the isolation model. Every sandbox is its own container with a gateway coordinating lifecycle, so you are running a control plane to run an agent. For a single developer running one agent against one repository on a laptop, that overhead buys you a boundary you could approximate with a container and a firewall rule. OpenShell earns its place when agents run unattended, when they hold provider credentials, or when several people need the same egress rules enforced identically.

The static policy layers are a limitation in practice, not just in documentation. A workflow that discovers mid-session that it needs a new filesystem path cannot be fixed by editing YAML and reloading. And the README does not document rollback for a policy that breaks a running session, so plan to test policies against a throwaway sandbox before applying them to one you care about.

## How OpenShell differs from a plain Docker sandbox or a generic sandbox runtime

The obvious alternative is hand-rolling it: docker run with a mounted workspace, plus iptables or a proxy container for egress. The difference is where the decision lives. In the hand-rolled setup, egress rules are host configuration that drifts from the repository and is invisible in code review. OpenShell makes the rule a YAML file that the proxy enforces at HTTP method and path level, applied with a CLI command and hot-reloadable. The policy is an artifact you can diff.

The second difference is credential handling. A container with an API key in an environment variable hands the agent the key. OpenShell's provider access injects credentials only after policy admits a request to a profile-authorized endpoint, which the README describes as binding credentials to endpoints. That is a structural difference, not a configuration one, and it is the feature that most distinguishes OpenShell from a container plus a firewall.

Compared with a general-purpose sandbox runtime, OpenShell is opinionated toward agents: the base image ships four agent CLIs, and the repository includes agent skills for operating OpenShell itself under skills/. If your workload is not an agent, that opinionation is dead weight.

## Licence, maintenance, and what upgrading costs

OpenShell is Apache-2.0, and the repository carries the standard files that licence implies: LICENSE, THIRD-PARTY-NOTICES, and a SECURITY.md with a vulnerability reporting path. Apache-2.0 permits commercial use and modification and includes an explicit patent grant, which matters for a runtime that sits between an agent and a model provider's API. The repository also includes a DCO file and a GOVERNANCE.md, so contributions are signed rather than assigned. None of this is legal advice; if you are embedding OpenShell in a product, have counsel read THIRD-PARTY-NOTICES, since the Rust workspace pulls in TLS, HTTP, and gRPC dependencies.

On maintenance, the last push to main was on 2026-09-10, and the repository is not archived. The most recent tagged release documented is v0.0.116 from 2026-08-28, with a VM runtime capability release on 2026-09-05 and a dev build tracking main. That cadence is the upgrade cost: a dev channel that follows main plus frequent tagged releases means pinning matters. The installer defaults to the latest stable release and accepts OPENSHELL_VERSION, so pinning is a supported move rather than a workaround. The Python SDK requires Python 3.11 or newer, and the Rust workspace declares rust-version 1.94, so both toolchains need to be current before you build from source.

## Conclusion

Adopt OpenShell if you run coding agents that touch a real filesystem and real credentials, and you want egress decisions expressed in version-controlled YAML rather than in a prompt. Skip it if you need a stable interface: the project labels itself alpha, the Kubernetes path is experimental, and Windows support runs through WSL 2 only. First verify that your host runtime (Docker, Podman, or virtualization for MicroVM) is one the gateway can drive, then check whether the static filesystem and process policy you need is expressible, since those sections lock at sandbox creation and cannot be reloaded the way network policy can.

## FAQ

### What is NVIDIA OpenShell?

It is a runtime that runs autonomous AI agents inside sandboxed execution environments, with declarative YAML policies that restrict file access, network egress, and process behaviour. The README describes it as the safe, private runtime for autonomous AI agents.

### What does OpenShell do?

It creates isolated containers for agents, intercepts every outbound connection through a policy engine that allows, denies, or attaches endpoint-bound credentials, and coordinates sandbox lifecycle through a gateway control plane. Supported compute platforms are Docker, Podman, MicroVM, and Kubernetes.

### Is OpenShell safe to use?

The project applies defense in depth across filesystem, network, process, and provider policy, and the README notes it is alpha software with an experimental Kubernetes path subject to breaking changes. Whether it is safe enough depends on your threat model; the design goal is to keep agents from reading files or reaching hosts outside policy.

### How do I install OpenShell on Windows?

There is no native Windows install documented. The README lists Windows with WSL 2 as an experimental supported host, so the Windows route is to install inside WSL 2 and then use the binary installer script.

### How do I install OpenShell?

Run the binary installer script from the README, which installs the latest stable release unless OPENSHELL_VERSION is set. The openshell package on PyPI is the Python SDK only and does not install the CLI.

### How do I use OpenShell with an agent?

Create a sandbox with openshell sandbox create -- followed by an agent such as claude, opencode, codex, or copilot. The sandbox starts with minimal outbound access, so you apply a policy with openshell policy set and then reconnect to the sandbox.

## Sources

- [License: Apache-2.0](https://github.com/NVIDIA/OpenShell/blob/main/LICENSE)
- [NVIDIA/OpenShell on GitHub](https://github.com/NVIDIA/OpenShell)
- [Project website](https://docs.nvidia.com/openshell/latest/)
- [README](https://github.com/NVIDIA/OpenShell/blob/main/README.md)
- [Releases](https://github.com/NVIDIA/OpenShell/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-openshell
