# alexzhang13/rlm: Recursive Language Models as a Drop-In Completion Call

> The rlms package turns an ordinary llm.completion(prompt, model) call into rlm.completion(prompt, model), offloading context into a REPL the model programs against. Here is what the repository documents, where the sandbox boundary actually sits, and who should stay away.

**alexzhang13/rlm** — General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.

- Repository: https://github.com/alexzhang13/rlm
- Website: https://arxiv.org/abs/2512.24601
- Stars: 5,650 · Forks: 905
- Language: Python
- License: MIT
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/alexzhang13-rlm

## The problem rlms solves: context that does not fit, and a model that can only ask once

The README frames Recursive Language Models as a task-agnostic inference paradigm for near-infinite length contexts. The mechanism is not a bigger attention window. It is delegation: the model is allowed to programmatically examine, decompose, and recursively call itself over its input. The target user is an engineer who already has an API-based or local LLM and keeps hitting the point where the prompt is too long, the retrieval is too coarse, or the answer requires reading many parts of a document and combining them. The library's own framing is a bet on future language model design, which is worth taking literally. The authors argue for a CodeAct-style harness where every language model has a code environment and sub-calls are functions in code. They state they want to move away from the JSON tool-calling standard for both sub-agents and generic tool calls. That is a design position, not a neutral utility, and it shapes everything below.

## How the recursion works: context as a REPL variable, sub-calls as functions

The canonical call is llm.completion(prompt, model). rlms replaces it with rlm.completion(prompt, model), so the object presents itself as a language model with the same shape of interface. Underneath, the context is offloaded as a variable inside a REPL environment that the model can interact with, and the model can launch sub-LM calls from inside that environment. The README describes the default client as running on the host process through Python exec calls, using the same virtual environment as the host, with some limitations on available global modules. Depth is a real parameter in this design: release v0.1.1a added Depth > 1 support and compaction with offloaded history, and the examples directory contains depth_metadata_example.py and compaction_example.py, so recursion depth and history compaction are first-class concerns rather than afterthoughts. A visualizer directory sits at the top level, and the README notes the quickstart generates a log you can use with the visualizer to explore trajectories. That matters because a recursive run is not readable as a single transcript.

## Installing rlms and running a first completion against a real model

The README requires Python 3.11 or later. The fastest path is PyPI:

```bash
pip install rlms
```

With that installed, the README's example constructs an RLM bound to the OpenAI backend and asks for output. The verbose flag prints to the console with rich and is disabled by default.

```python
from rlm import RLM

rlm = RLM(
    backend="openai",
    backend_kwargs={"model_name": "gpt-5-nano"},
    verbose=True,
)

print(rlm.completion("Print me the first 100 powers of two, each on a newline.").response)
```

The result object carries a response attribute, which is what the print statement reads. If you prefer a source checkout, the README documents a uv-based manual setup and a Makefile with make install, make check, and a make quickstart target that runs examples.quickstart using your OPENAI_API_KEY environment variable. That target also produces the log the visualizer consumes. Pick a sandbox before you point this at anything untrusted: environment accepts "local", "ipython", "docker", "modal", "prime", "daytona", or "e2b", with environment_kwargs for per-sandbox settings.

## The sandbox boundary is the real deployment decision

rlms splits environments into isolated and non-isolated, and the default is the non-isolated one. LocalREPL runs in the same process as the RLM with specified global and local namespaces for minimal security. The README says using it is generally safe but should not be used for production settings, and that it shares the same virtual environment as the host process. Read that as: code the model writes can touch your dependencies. The README is also direct that non-isolated execution is problematic when prompts or tool calls can interact with malicious users. On the isolated side, DockerREPL launches the REPL as a Docker container, defaults to the python:3.11-slim image, and the README states the container runs fully isolated from the host with a lightweight host-side proxy bridging LM access back into the container. DockerREPL is documented as supporting the full feature set of the local environment, including llm_query, llm_query_batched, rlm_query and rlm_query_batched with parallel batching. IPythonREPL is a middle option: in-process by default, or a separate ipykernel subprocess where subprocess mode adds hard cell_timeout enforcement and full namespace isolation. The optional extras are modal, e2b, daytona, prime and ipython, each pulling its own SDK plus dill.

## Where rlms is the wrong tool

Three cases stand out. First, if your task is a single well-scoped question over a short prompt, the recursion machinery is pure overhead: you pay for a REPL, a code-generation step, and possibly sub-calls where one completion would do. Second, if your prompts can be influenced by untrusted users and you cannot stand up Docker, Modal, E2B, Daytona or Prime, the default local environment is the wrong place to run this. The README says so itself. Third, if your existing stack is built around JSON tool calling, rlms is pointedly moving the other way, and the README states that intent explicitly. Mixing a code-execution harness into a tool-calling pipeline means maintaining two mental models of how the model acts. There is also a maturity signal worth weighing: pyproject.toml classifies the project as Development Status 4 - Beta, and the version is 0.1.3. The last push to the repository was on 2026-08-26, and the most recent release, v0.1.3, was on 2026-06-26 with Docker REPL fixes.

## How rlms differs from a code-executing agent framework

The obvious alternative is a general agent framework that gives a model a Python tool and lets it loop. The difference in approach is where the context lives. In a typical tool-calling loop, the context stays in the message history and the model copies the parts it needs into each tool call, so the history grows with every step. In rlms, the context is a variable in the REPL and the model writes code that reads it, which is why the README can call the paradigm task-agnostic and aim at near-infinite length contexts. The second difference is recursion. rlm_query and rlm_query_batched let the model spawn sub-RLM calls from inside the environment, and the release notes for v0.1.1a add Depth > 1 plus compaction with offloaded history. A plain agent loop has no built-in notion of a sub-agent sharing the same context object. There is also rlm-minimal, linked from the README, which is the smaller reference implementation if you want to read the idea before adopting the full engine. The repository additionally ships a verifiers training environment under training/, built on Prime Intellect's prime-rl, so the same harness can be used for training rather than only inference.

## Licence, maintenance and what upgrading costs

The project is MIT licensed, declared both in pyproject.toml and in the LICENSE file at the repository root. MIT is permissive, so the practical implication is that you can vendor or modify the code and ship it, provided you keep the copyright notice; that is a description of the licence text, not legal advice for your situation. On maintenance: the repository is not archived, the last push was on 2026-08-26, and releases have appeared at a steady clip through 2026, from v0.1.1a in February to v0.1.3 in June. The README states the repository is maintained by the authors of the paper from the MIT OASYS lab. Upgrade cost is the part to plan for. The dependency floor is Python 3.11, and the core dependencies pin minimum versions of anthropic, google-genai, openai, portkey-ai and rich, so a major bump in any provider SDK can reach you. Sandbox extras are separate: modal, e2b, daytona, prime and ipython each bring their own SDK and dill, so a change in one sandbox vendor only affects installations that opted into it. The v0.1.2 release changed the final answer function, which is the kind of change that can silently alter how a run terminates. Pin the version and read the release notes before moving.

## Conclusion

Adopt rlms if you need a model to inspect and recurse over context far larger than a prompt window and you are willing to run its generated code somewhere you control. Skip it if you need a stable API, a production-hardened sandbox, or a tool-calling interface rather than a code-execution one; the package is classified Beta and the local REPL is explicitly not for production. Before committing, read the IPython and Docker environment docs for the cell_timeout and isolation semantics, and check the examples folder for a sandbox matching your deployment.

## FAQ

### What does RLM stand for in the context of AI?

In this project it stands for Recursive Language Model. The README describes RLMs as a task-agnostic inference paradigm in which a language model programmatically examines, decomposes, and recursively calls itself over its input.

### What is MIT RLM?

The README states the repository is maintained by the authors of the RLM paper from the MIT OASYS lab. The package itself is named rlms and is distributed under the MIT licence.

### What is the Recursive Language Model (RLM) and how does it work?

The README says RLMs replace the canonical llm.completion(prompt, model) call with rlm.completion(prompt, model), offloading the context as a variable in a REPL environment that the LM can interact with and launch sub-LM calls inside of.

## Sources

- [alexzhang13/rlm on GitHub](https://github.com/alexzhang13/rlm)
- [License: MIT](https://github.com/alexzhang13/rlm/blob/main/LICENSE)
- [Project website](https://arxiv.org/abs/2512.24601)
- [README](https://github.com/alexzhang13/rlm/blob/main/README.md)
- [Releases](https://github.com/alexzhang13/rlm/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/alexzhang13-rlm
