# CodeRunner sandboxes your agents, and leaves the network open by default

> CodeRunner runs an AI agent's code inside an Apple-container VM on macOS and exposes it over MCP on port 8222, with persistent Jupyter kernels, Playwright scraping and named per-agent sessions. The design is careful about isolation and careless about the network: unrestricted egress is the default, the lock is an environment variable, and the setting cannot be changed after the container exists.

**instavm/coderunner** — A local sandbox for your AI agents

- Repository: https://github.com/instavm/coderunner
- Website: https://instavm.io
- Stars: 893 · Forks: 41
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/instavm-coderunner

## The network is open unless you ask for it to be shut, and you cannot change your mind

This is the first thing to know about the sandbox, and it is stated plainly: by default, code running in the sandbox has unrestricted network access.

Isolation from your host filesystem is the default. Isolation from the internet is opt-in, with one environment variable:

```bash
CODERUNNER_NETWORK=none ./install.sh
```

That puts the sandbox on a host-only network with no internet, and in this mode the MCP server is at `http://127.0.0.1:8222/mcp` rather than on a hostname.

The constraint that makes this a design decision rather than a switch is the next sentence: the setting is fixed when the container is created, and the installer refuses to resume a container with a different network mode. So tightening or loosening the network after the fact is not a configuration change, it is a delete and reinstall.

Combined with the fact that the default is open, the sequence matters. A sandbox created the quick way will have egress, and the code an agent writes in it will have egress, for as long as the container exists.

There is a second thing the page does not resolve, which is who may talk to the sandbox. No authentication, token or allowlist appears anywhere in the security section or the integration guides.

## Three addresses for one MCP server, and no auth described on any of them

The same server is given three different addresses in three different sections, and the difference is not cosmetic.

In the network-isolation mode, the endpoint is `http://127.0.0.1:8222/mcp`, which is loopback and therefore only reachable from the machine. In the general integration section the endpoint is written as `http://coderunner.local:8222/mcp`, a hostname, which has to resolve somewhere and is not loopback. And the compose file in the repository publishes port 8222 to the host and sets `FASTMCP_HOST=0.0.0.0`, which is a bind to every interface.

So the sandbox exposes an HTTP MCP endpoint that runs Python on request, over plain HTTP, with the server bound to all interfaces in the Docker path, and nothing in the documentation mentions a credential.

The Kiro configuration makes the consequence concrete. That example lists an `autoApprove` array containing `execute_python_code`, which means the client is instructed not to ask a human before running agent-written code. Combined with an open network and a published port, that is the arrangement to think hardest about before putting anything sensitive in the container.

The tool surface behind that endpoint is not small: Python execution in a persistent kernel, session management, Playwright page scraping, and skill file reads.

## Apple silicon only, and two install commands that disagree about sudo

The prerequisite line is short and restrictive: a Mac with macOS on Apple silicon, named as M1, M2, M3 or M4, and Python 3.10 or newer on the host.

That is not an arbitrary choice. The security section quotes the technical overview of Apple's container project to justify VM-level isolation, and that project is a macOS virtualisation facility. So the host is macOS-only, even though the container image is Debian based, which is the normal arrangement for a VM: a Linux guest on a macOS host.

The lifecycle interface is the `container` command that ships with it. Stop the sandbox when you are done, start it again to resume with uploads, kernels and installed packages preserved, and delete the container and rerun the installer to start clean.

That last point has a consequence for how you read the word sandbox. A resumed container keeps its kernels and its installed packages, so state survives across sessions by design rather than being thrown away.

The install commands on the page are inconsistent about privilege. The quick start clones, makes the script executable and runs `./install.sh`. The Claude Code CLI section, further down the same page, runs `sudo ./install.sh`. One of those needs root to create the VM and the other does not, or one of them is out of date.

## The image installs systemd, sudo and an SSH server

The container image is a single Dockerfile, and the system package list is long enough to be worth reading as a security surface.

The base is a Debian-based Python 3.13.3 image. The first apt-get layer installs systemd and sudo, openssh-client and openssh-server, curl, jq, kmod, iproute2, procps, build-essential and cargo, then xvfb, ffmpeg, wkhtmltopdf, poppler-utils, a default JRE, unzip, p7zip, ripgrep, fd-find, sqlite3, bc, and around twenty shared libraries for a headless browser.

What that buys you is real: a shell you can SSH into, a headless display for rendering, PDF tooling, a Rust and C toolchain for compiling packages, and the browser libraries Playwright needs. What it costs is a container with a running init system, a login server and a compiler inside it, which is a much larger attack surface than a minimal runtime image would present.

The skills split is visible in the same file. Public skills are copied into the image at build time, so they change only when the image is rebuilt, and a user skills directory is created as an empty mount point for whatever you add. The compose file mounts the host's `uploads` directory at `/app/uploads`, which is where user skills and any uploaded files both land.

## One exact version pin in a forty package requirements file

The dependency file is the other file worth auditing, because the image is built by installing it directly.

Out of roughly forty-five entries, three carry any version constraint at all. `playwright==1.53.0` is the only exact pin, `mcp[cli]>=1.26,<2` is the only entry with an upper bound, and `requests>=2.33.0` is the only one with a lower bound. Everything else, from jupyter-server and fastapi to pandas, scikit-learn, python-pptx, reportlab, sympy and beautifulsoup4, is declared bare.

That distribution is not random. The two constrained packages are the two that break: a browser automation library is tied to a specific browser build, and the MCP SDK is on a fast-moving protocol where a major version means a different transport. Everything else is a library whose API has been stable, so leaving it floating costs nothing on a working build and saves maintenance.

There is no lockfile in the tree, so a rebuild six months from now resolves those forty packages afresh. One detail in the Dockerfile is worth noting for anybody who wants to switch package managers: the bash kernel for Jupyter is installed with its own installer module, with a comment saying it does not work with uv.

## Six configuration files, one empty block and a sibling repository for the plugin

The integration list is long, and it is mostly configuration files, which is the point: an MCP server is only useful once something is told to call it.

Claude Desktop is the first route. You copy an example configuration, replace the placeholder `/path/to/your/python` with the full path to the interpreter in `~/.coderunner/venv`, which the installer creates for that proxy, and replace `/path/to/coderunner` with your clone. Restart the app afterwards. The other Python examples are told to install their dependencies in your own virtualenv with `pip install -r examples/requirements.txt`, so the installer-managed environment is for the desktop proxy alone.

The Claude Code route does not use the MCP server in this repository at all. It installs a plugin from a separate project, `github.com/instavm/coderunner-plugin`, with `claude plugin marketplace add` and `claude plugin install instavm-coderunner`, then reconnects with `/mcp`. So the Claude Code integration is a different codebase, and this repository is where the server it talks to lives.

The rest are config files: OpenCode with a remote MCP entry pointing at `coderunner.local`, a Python OpenAI agents example that needs `OPENAI_API_KEY` exported and runs `openai_client.py`, Gemini CLI with an `httpUrl` entry, and Kiro with a command and args plus the auto-approve list already mentioned.

The seventh entry is the odd one. Coderunner-UI is described as the project's own offline AI workspace for full privacy and local processing, linked, and then given an empty details block. Whatever it is, the page does not say.

## Five named kernels, and stopping one discards its state

The tool list is the clearest statement of what the sandbox is for, and it is worth reading as an isolation model.

Code execution happens in a persistent Jupyter kernel. Alongside it there are three session tools: start a Python session to reserve an isolated kernel for a named session, list the active named sessions, and stop a session, which discards its kernel state. The pattern is that the return value of a session start is a session id, which you pass to code execution so state stays isolated between agents, and up to five named sessions can run concurrently.

That is a real boundary between concurrent agents. One agent's variables are not another's, and stopping a session is how you destroy whatever that agent accumulated.

The plugins expose this through the MCP tool names, including the scraping tool built on Playwright, and three skill tools for listing skills, reading skill documentation and reading skill files and examples. The skills themselves are the ones the image bakes in, covering formats such as docx, xlsx, pptx, pdf and image processing.

The repository treats this as something that needs testing and an incident plan. Alongside the install, entrypoint and cleanup scripts there are `test-build.sh`, `test-e2e.sh` and `test-sessions.py` in the root, plus SECURITY.md, CONTRIBUTING.md, third-party notices and an INCIDENT_RESPONSE.md. There are no GitHub releases, and the last push was on 2026-08-13.

## Conclusion

Take CodeRunner if you are on Apple silicon and want your agent's code execution to leave your host files alone, with the option of cutting the network entirely before the container is ever created. Leave it if you need Linux hosts, if you need to change the network policy after setup without rebuilding, or if you need per-call confirmation on every execution the agent requests. Three things to check before pointing an agent at it: whether FASTMCP_HOST is bound to all interfaces on a port you publish, since the compose file does exactly that, whether your client is configured to auto-approve the execute tool, as the Kiro example does, and which MCP endpoint you are using, since the page gives three different addresses for the same server.

## FAQ

### What is CodeRunner?

A local sandbox for AI agents. Code runs in an isolated container with VM-level isolation using Apple's container project, and an MCP server on port 8222 lets an assistant execute Python in a persistent Jupyter kernel, scrape pages with Playwright and read skill files.

### Does the CodeRunner sandbox have internet access?

By default it does: the README states that code running in the sandbox has unrestricted network access. Installing with `CODERUNNER_NETWORK=none ./install.sh` puts it on a host-only network with no internet, and in that mode the MCP endpoint is `http://127.0.0.1:8222/mcp`. The mode is fixed when the container is created.

### What are the system requirements for CodeRunner?

A Mac running macOS on Apple silicon, listed as M1, M2, M3 or M4, with Python 3.10 or newer on the host. The container image itself is Debian based and built on Python 3.13.3, so the Linux side is the guest rather than the host.

### Which AI assistants can connect to CodeRunner?

Six documented configurations: Claude Desktop, the Claude Code CLI through a plugin installed from a separate repository, OpenCode, a Python OpenAI agents example, Gemini CLI and Amazon's Kiro. A seventh, Coderunner-UI, is named and linked with no configuration shown.

### Can CodeRunner keep state between sessions?

Yes. Starting the container again preserves uploads, kernels and installed packages, so a session is resumable rather than disposable. Inside the sandbox, code runs in a persistent Jupyter kernel, and named sessions give each agent its own kernel, with up to five running concurrently.

### How do I reset a CodeRunner sandbox or change its network mode?

Delete the container and rerun the installer with `container delete coderunner && ./install.sh`. Because the network mode is fixed at creation and the installer refuses to resume a container created with a different mode, changing it means a full reinstall rather than a configuration edit.

## Sources

- [instavm/coderunner on GitHub](https://github.com/instavm/coderunner)
- [Issues](https://github.com/instavm/coderunner/issues)
- [License: Apache-2.0](https://github.com/instavm/coderunner/blob/main/LICENSE)
- [Project website](https://instavm.io)
- [README](https://github.com/instavm/coderunner/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/instavm-coderunner
