vndee/llm-sandbox: a Python library for running LLM-generated code in containers
Lightweight and portable LLM sandbox runtime (code interpreter) Python library.
At a glance
- What is it?
- llm-sandbox wraps Docker, Kubernetes or Podman behind one Python session object so generated code runs isolated from the host. Here is what the install path looks like, where the design gets thin, and who should pick something else.
- Who is it for?
- Adopt llm-sandbox if your agent already produces Python, JavaScript, Java, C++, Go or R snippets and you want container isolation without writing the Docker plumbing yourself; the Docker backend is the shortest path and the docs give a working session in a few lines.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap llm-sandbox fills between a model output and a running process
A model returns a string that looks like Python. Something has to turn that string into a process, and that process will import packages, write files and possibly open sockets. Doing this with subprocess on the host is the default mistake: the generated code inherits your filesystem, your environment variables and your network. llm-sandbox exists to make the container step shorter than the subprocess step. Its stated purpose is to run LLM-generated code "in a safe and isolated mode", and the unit of work is a session object rather than a raw docker run invocation.
The audience is narrow and identifiable. You are writing an agent, a code interpreter feature, or a data-analysis assistant, and you already accept that containers are the isolation boundary. The library does not add a new isolation primitive; it wraps Docker, Kubernetes or Podman so that language selection, dependency installation, file copying and artifact capture happen through one API. If you have never containerised anything, this library will not teach you that, and the Docker daemon still has to be running.
SandboxSession, backends and the container lifecycle
The central object is SandboxSession, imported from llm_sandbox. It is a context manager. Entering it creates or attaches to a container for the requested language; leaving it tears the container down. Inside the block, session.run() takes a code string and returns an object with stdout, so the data flow is: code string in, container executes it, captured output back to your Python process.
The backend is selected by what you install and how the session is configured. Docker is described as the most popular and widely supported option, Kubernetes as the orchestration path for scalable deployments, and Podman as the rootless alternative. Language support is handled by container images, and the R example shows the escape hatch: you pass image="ghcr.io/vndee/sandbox-r-451-bullseye" explicitly when the default image for a language is not what you want. Dependency installation is a parameter, not a shell command you write: libraries=["numpy"] triggers the install inside the sandbox before the code runs. That is the whole mechanism. There is no parser, no AST rewriting and no syscall filtering visible in the README; the isolation is the container's, and the library is the orchestration layer on top.
Installing llm-sandbox and running a first snippet
The base package is installed from PyPI, but the base install does not bring a container client. The README separates them, so the Docker path is the one most people want:
pip install 'llm-sandbox[docker]'Kubernetes and Podman get their own extras, and they can be combined. The optional dependency groups in pyproject.toml confirm the split: docker pulls docker>=7.1.0, k8s pulls kubernetes>=32.0.1, podman pulls both docker and podman. Installing the bare package and then wondering why no backend is available is a self-inflicted error, not a bug.
Once installed, the smallest useful program is a session with a language and a run call. The README's basic example prints from inside the container:
from llm_sandbox import SandboxSession
with SandboxSession(lang="python") as session:
result = session.run("""
print("Hello from LLM Sandbox!")
print("I'm running in a secure container.")
""")
print(result.stdout)What you should see is the two printed lines arriving on result.stdout in your host process. The code never ran on the host.
The second thing worth trying early is dependency installation, because that is where the container startup cost becomes visible. Passing libraries as a list makes the session install them before executing:
from llm_sandbox import SandboxSession
with SandboxSession(lang="python") as session:
result = session.run("""
import numpy as np
arr = np.array([1, 2, 3, 4, 5])
print(f"Mean: {np.mean(arr)}")
""", libraries=["numpy"])
print(result.stdout)That install happens inside the container, so it is paid on every fresh session unless you reuse containers. The README addresses this with a Fast Production Mode that skips environment setup, and with container pooling for pre-warmed containers.
Where the defaults may not match your threat model
The README's security claims are broad: isolated execution, security policies, CPU, memory and time limits, and network isolation. Those are container capabilities the library exposes, and the honest reading is that llm-sandbox is a convenient way to reach them, not a hardened sandbox of its own design. If your threat model includes a model actively trying to escape, container configuration is the thing that matters, and the README does not walk through a reference hardening profile.
The repository is more informative than the prose here. examples/container_user_privilege.py and examples/k8s_readonly_file_system.py exist precisely because user privilege and read-only filesystems are settings you have to think about, and their presence in the examples directory implies they are not the default you get for free. The related search phrase "llm sandbox escape" is the right question to be asking, and the answer is not in the README. A second limitation is operational: every session depends on a running Docker daemon or a reachable cluster, so the library is useless in a serverless environment without one. A third is startup latency. Installing libraries per session is slow, which is why pooling and fast production mode exist; if your workload is one short snippet per request, the container overhead can exceed the execution time.
llm-sandbox against E2B, Pyodide and plain Docker
The nearest managed alternative is a hosted code-interpreter service such as E2B, where the sandbox runs on someone else's infrastructure and you call an API instead of holding a Docker socket. The difference is where the trust boundary sits and who operates it. llm-sandbox runs in your process and your infrastructure, which means your credentials, your cluster and your patching schedule; a hosted service removes that operational load and adds a network dependency and a vendor.
A second alternative is an in-process interpreter such as Pyodide or a restricted Python evaluator. Those are faster to start and need no daemon, but they only cover Python and the isolation is a language-level restriction rather than a kernel boundary. llm-sandbox's advantage is the opposite: it accepts Python, JavaScript, Java, C++, Go and R through the same session API, at the cost of a container per session.
The third comparison is the one that matters most. Plain Docker with the docker SDK is what llm-sandbox wraps. If you only ever run Python and you already have image-building and cleanup code you trust, the library saves you a few hundred lines of lifecycle management and adds a dependency you now have to track. That is a real trade, not a clear win.
Maintenance cadence, versioning and the MIT licence
The repository is not archived, and the last push was on 2026-09-10. Releases have moved quickly: 0.3.42, 0.3.43 and 0.3.44 all landed within a day of each other in early August 2026, and the pyproject.toml in the repository still declares version 0.3.13. That mismatch between the file and the published releases is worth noting before you pin anything. The practical consequence is that you should pin an exact version in your own dependency file rather than a range, because patch releases arrive frequently and the changelog is not the primary documentation channel.
Upgrade cost is dominated by the backend clients, not the library. The extras pin minimums for docker, kubernetes and podman, and the MCP extras pull mcp>=1.28.1,<3. If you use the MCP server, the executable is llm-sandbox-mcp, installed via the mcp-docker, mcp-podman or mcp-k8s extras. Development uses uv dependency groups, and make install runs uv sync plus pre-commit install, so contributing has a defined path.
The licence is MIT, which permits commercial use and modification with the copyright notice retained. That is a permissive licence with no copyleft obligation on your own code. Nothing here is legal advice, and the usual caveat applies: your organisation's policy on container images you pull is a separate question from the library's licence, and the R example pulls an image from ghcr.io rather than building locally.
Editorial conclusion
Adopt llm-sandbox if your agent already produces Python, JavaScript, Java, C++, Go or R snippets and you want container isolation without writing the Docker plumbing yourself; the Docker backend is the shortest path and the docs give a working session in a few lines. Skip it if you need a strong isolation boundary against adversarial code, since the README's own framing is about isolation from the host, and the examples directory is where the real behaviour lives rather than the prose docs. Before committing, read examples/container_user_privilege.py and examples/k8s_readonly_file_system.py, then verify that the container user and filesystem settings match your threat model rather than assuming the defaults do.
Frequently asked questions
What is llm-sandbox?
It is a Python library that runs LLM-generated code inside an isolated container instead of on the host. It supports Docker, Kubernetes and Podman as backends and exposes them through a single SandboxSession context manager.
What is the llm-sandbox alternative if I do not want to run containers myself?
A hosted code-interpreter service such as E2B is the closest alternative, since the sandbox runs on the provider's infrastructure and you call an API rather than holding a Docker socket. An in-process interpreter like Pyodide is another option, but it covers Python only and provides a language-level restriction rather than a kernel boundary.
Is llm-sandbox the same as a VM?
No. llm-sandbox uses containers, which share the host kernel, so the isolation boundary is weaker than a virtual machine's. The README describes isolated containers with resource limits and network isolation, and the container configuration is what determines how strong that boundary actually is.
What is sandbox vs Docker?
In llm-sandbox the two are not opposed: Docker is one of the container backends the library drives, alongside Kubernetes and Podman. The sandbox is the session you create, and Docker is the runtime that provides the isolation underneath it.
What is the purpose of a sandbox in this project?
The README states the purpose as running LLM-generated code in a safe and isolated mode, with no access to the host system. It also lists CPU, memory and execution time limits and network isolation as the controls available inside a session.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vndee-llm-sandbox)