Self-hosted service
pydantic/monty avatar
pydantic/monty

pydantic/monty: a sandboxed Python interpreter in Rust for AI-written code

A minimal, secure Python interpreter written in Rust for use by AI

8,361 stars425 forksRustMIT

At a glance

What is it?
Monty runs model-generated Python without Docker, a VM or a sandboxing service. It is a VM in Rust with host-granted functions, enforced resource limits and serialisable snapshots, and it is not a general-purpose Python.
Who is it for?
Monty fits teams running model-generated Python inside an agent loop, where a per-session sandbox must start in milliseconds and the host must control every side effect through explicitly passed functions. It does not fit workloads that need the full CPython standard library, arbitrary third-party packages, or code that expects a filesystem, environment variables or a network socket, because the README states none of those exist inside the sandbox.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Monty actually solves for AI-generated Python

The problem is not running Python. It is running Python that a language model wrote, in the same process that holds your API keys, your database connection and your filesystem. The usual answer is a container per execution or a hosted sandboxing service. Both work, and both put a network hop or an image pull between the model's output and the result.

Monty takes the other route. It is a Python interpreter implemented in Rust that executes inside your process, so there is no container, VM or sandboxing service in the loop. The README states that creating a sandbox and running ten commands takes 5 ms, against 900 ms for Docker and 1900 ms for a sandboxing service. Those are the project's own figures from its introduction page, not independent measurements, but the architectural reason for the gap is visible in the repository: the workspace contains crates/monty, crates/monty-runtime, crates/monty-pool and crates/monty-proto, which is what an in-process interpreter with a worker pool looks like.

The intended user is an agent framework author or an application developer who lets a model write code as a tool call. Pydantic AI runs what the docs call Code Mode on Monty, and the repository ships examples for a SQL playground, a web scraper and an expense analysis, which is the shape of the target workload: short scripts that transform data the host already holds.

How the interpreter, the host boundary and the pool fit together

The mechanism is a VM with a deliberately closed world. According to the README, filesystem, environment variables and network do not exist inside the sandbox. The only path out is the functions and mounts you pass in. That is not a permission system layered on top of Python; it is the absence of the capabilities in the first place, which removes a whole class of escape routes before they can be considered.

The example in the README shows the boundary in practice. The model writes code that calls nutrition('chocolate bar'), the host supplies a lambda for nutrition, and the sandbox sees only the return value. The function body runs on the host, in your process, with your privileges. So the sandbox protects you from the model's code, not from the function you wrote to serve it.

Three further pieces are visible. Resource limits for memory, time and recursion are enforced by the VM itself rather than by an external supervisor. A paused interpreter serialises to bytes and can be resumed later, which the docs cover under snapshots. And the runtime is pooled: the Python API checks a session out of a Monty pool, which is where the per-execution startup cost is amortised. The workspace also lists crates/monty-wasm-runtime, crates/monty-js and crates/monty-python, so the same core is exposed to browsers, Node and CPython, with crates/monty-proto presumably carrying the wire format between them.

Installing pydantic-monty and running a first script

The README gives three install lines, one per language. Pick the one matching your host. For a Python application, the package name on PyPI is pydantic-monty even though the import is pydantic_monty.

bash
uv add pydantic-monty        # Python
npm install @pydantic/monty  # JavaScript / TypeScript
cargo add monty              # Rust

The repository's Makefile shows the developer path if you are building from source rather than installing a wheel: make install runs uv sync --all-packages --only-dev --inexact, then cargo check --workspace, and make dev-py builds the Python packages with maturin develop against crates/monty-runtime/Cargo.toml and crates/monty-python/Cargo.toml. That is the path to take when you want to patch the interpreter, not the one to take to evaluate it.

The first real use is the example from the README. A pool is opened as a context manager, a session is checked out from it, and feed_run is given the code, an inputs mapping and an external_lookup mapping.

python
from pydantic_monty import Monty

code = """
kcal = nutrition('chocolate bar')['kcal']
hours = kcal * 4184 / (bulb_watts * 3600)
print(f'a chocolate bar powers a {bulb_watts} W bulb for {hours:.1f} hours')
"""

with Monty() as pool:
    with pool.checkout() as session:
        session.feed_run(
            code,
            inputs={'bulb_watts': 10},
            external_lookup={'nutrition': lambda food: {'kcal': 230}},
        )

The README records the output as a chocolate bar powers a 10 W bulb for 26.7 hours. Two things are worth noting about the call shape. Inputs are passed as a mapping rather than injected as globals, and external functions are looked up by name at call time, so the set of host functions a script can reach is fixed by the caller and not by the script.

The Python subset is the real constraint, not the sandbox

The repository carries a top-level entry literally named limitations, and the docs site has a page for the Python subset. That is the honest signal here. A from-scratch interpreter in Rust does not implement everything CPython does, and the project does not claim it does.

The practical consequence is that code which imports a third-party package, reads a file, opens a socket or relies on a CPython-specific module will not run. The README's own framing supports this: the target is code written by a model, which tends to be short, self-contained and computational. If your workload is a user-supplied script that imports pandas, Monty is the wrong tool, and the comparison page the README links to is where the project itself draws the line against Docker, Pyodide, WASI and sandboxing services.

There is a second, subtler cost. Because the sandbox has no filesystem, every piece of data the script needs must arrive through the host boundary, and every result must leave the same way. For a large dataset that means serialising it across the boundary rather than mounting a volume. The README mentions mounts as a host-provided mechanism, but the details live in the security model documentation, not in the README.

Finally, the version number. The workspace version is 0.0.23, and the three most recent releases are 0.0.21, 0.0.22 and 0.0.23, all dated within about four weeks of each other. The README also notes that Hack Monty Round 3 is the last round before Monty V1. Pre-1.0 software with that release cadence means API surface can move under you.

Monty against Docker, Pyodide and WASI

The README links a comparison page that names Docker, Pyodide, WASI and sandboxing services as the alternatives, so the project has already picked its opponents. The differences are architectural rather than cosmetic.

Docker gives you a real Linux userland with a real CPython, which means any package, any file and any syscall. You pay for it in startup: the README's figure is 900 ms to create a sandbox and run ten commands, against 5 ms for Monty. If your agent executes code once per conversation turn, that difference is the difference between a snappy loop and a visible pause.

Pyodide compiles CPython to WebAssembly and runs it in a browser or a WASM runtime. You keep much more of the standard library, but you inherit a WASM runtime and a much larger artefact, and the sandbox boundary is the WASM instance rather than a purpose-built capability model. Monty's own WASM crate, crates/monty-wasm-runtime, suggests the project sees that deployment target too, with a smaller interpreter inside it.

A hosted sandboxing service removes the operational work entirely and is the slowest option in the README's numbers at 1900 ms. It is the right choice when you need a full OS and cannot run containers yourself. Monty is the right choice when you control the host process and the code is computational. These are not competing on the same axis; they are competing on how much of a real machine the model's code is allowed to see.

Licence, maintenance and what an upgrade costs

Monty is MIT licensed, stated in the README badge, in the LICENSE file at the repository root and in the workspace package metadata in Cargo.toml. MIT is permissive: you can embed the interpreter in a closed product, and the obligation is essentially attribution and keeping the licence text. The commercial Full Monty offering, which the README describes as running the same workers as a container image, is a separate product with its own terms. Nothing here is legal advice; if you are redistributing a modified interpreter, read the LICENSE file itself.

The repository is not archived, and the last push was on 2026-09-20. The release history shows three versions in roughly a month, and the workspace version in Cargo.toml is kept in lockstep across the internal crates, with a comment stating that the versions must be bumped together. That lockstep matters for upgrade cost: a version bump touches every internal crate at once.

Upgrading means re-testing the Python your model produces, because the subset can change between releases. The repository gives you the material to do that: crates/monty/test_cases/ holds files whose names encode the behaviour under test, including cases for duplicate function parameters, module-level nonlocal errors and traceback scoping. Those files are the closest thing to a compatibility contract you can run against your own corpus.

Editorial conclusion

Monty fits teams running model-generated Python inside an agent loop, where a per-session sandbox must start in milliseconds and the host must control every side effect through explicitly passed functions. It does not fit workloads that need the full CPython standard library, arbitrary third-party packages, or code that expects a filesystem, environment variables or a network socket, because the README states none of those exist inside the sandbox. Before adopting it, read the Python subset page at pydantic.dev/docs/monty/limitations/ and check your target code against the test cases under crates/monty/test_cases/, then pin the version: the workspace is at 0.0.23 and the release history shows 0.0.21, 0.0.22 and 0.0.23 landing within about a month.

Frequently asked questions

Does pydantic/monty require Docker or a virtual machine?

No. The README states that Monty runs Python with no container, VM or sandboxing service in the loop, and that the sandbox reaches the host only through the functions and mounts you pass in.

Which languages can I call pydantic/monty from?

The README gives install lines for Python (pydantic-monty), JavaScript and TypeScript (@pydantic/monty) and Rust (monty). Community bindings for Go and Dart are also listed.

Can pydantic/monty read files or make network requests?

No. The README states that filesystem, environment variables and network do not exist inside the sandbox, so anything the code needs must come through a host function you supply.

What happens to a pydantic/monty session that is paused mid-run?

The README states that a paused interpreter serialises to bytes you can resume later, and the documentation site covers this under snapshots.

What licence does pydantic/monty use?

MIT, according to the README badge, the LICENSE file at the repository root and the workspace package metadata in Cargo.toml.

Official sources

  1. License: MIT
  2. Project website
  3. pydantic/monty on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pydantic-monty.svg)](https://hysenlabs.com/projects/pydantic-monty)