Library / SDK
huggingface/OpenEnv avatar
huggingface/OpenEnv

OpenEnv: a Gymnasium-style interface for agentic RL environments

An interface library for RL post training with environments.

2,606 stars460 forksPythonBSD-3-Clause

At a glance

What is it?
OpenEnv defines step(), reset() and state() over WebSocket to Docker-hosted FastAPI environments, plus a CLI for packaging and deploying them. It is early-stage software aimed at RL researchers and environment authors, not at production serving.
Who is it for?
Adopt OpenEnv if you are writing RL post-training loops and want one client shape for many sandboxed environments, or if you are an environment author who wants HTTP and Docker packaging handled for you. Do not adopt it as a production serving layer: the README states it is experimental and that APIs may change.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OpenEnv fixes for RL post-training loops

Reinforcement learning against an agentic environment usually means writing glue: one harness for a browser sandbox, another for a code interpreter, a third for a game. Each has its own reset semantics, its own action encoding, its own way of reporting a reward. OpenEnv's answer is a single interface shape borrowed from Gymnasium. The README describes it as a standard for interacting with agentic execution environments via "simple Gymnasium style APIs - `step()`, `reset()`, `state()`".

That shape is the whole product thesis. If every environment exposes the same three calls and returns the same result object, an RL framework writer can target one client instead of N integrations, and a researcher can swap environments without rewriting the training loop. The README frames the audience in two halves: "users of agentic execution environments" who drive them during training loops, and "environment creators" who need isolation, packaging and a deploy path. The second half is where the framework does more work than a bare protocol would.

WebSocket to FastAPI inside Docker: the actual data flow

The architecture diagram in the README is specific. A client application holds concrete client classes, shown as `EchoEnv` and `CodingEnv`, both described as `EnvClient`. Those clients speak WebSocket to the server, carrying the `reset`, `step` and `state` calls. On the other side sit Docker containers, each running a FastAPI server that hosts an environment class derived from an `Environment` base, with `EchoEnvironment` and `PythonCodeActEnv` given as the two examples.

So the boundary is a container. The environment does not run in your training process; it runs in its own image and is reached over a socket. That is what makes isolation possible, and it is also the main cost: every step is a network round trip, and any state you care about lives on the far side of it. The result objects returned by `step()` carry an `observation` and a `reward`, which is the minimum an RL loop needs to compute a loss.

The README also documents a built-in web interface for interactive exploration and debugging, with a two-pane layout (human or agent interaction on the left, state observation on the right), WebSocket-based live updates, action forms generated automatically from the environment's Action types, and a log of actions and results. It is conditionally enabled: disabled by default for local development, and switched on with `ENABLE_WEB_INTERFACE=true`. That default is a reasonable choice, since the Gradio dependency it pulls in is not something every training run wants loaded.

Installing OpenEnv and running a first step against the Echo environment

The core package installs from PyPI. The README gives a single command, and `pyproject.toml` confirms the project name and the Python floor of 3.10.

bash
pip install openenv

The core install does not include any environment client. Clients are distributed separately, and the README's example installs one from a Hugging Face Space rather than from PyPI:

bash
pip install git+https://huggingface.co/spaces/openenv/echo_env

With the client installed, the README's asynchronous example connects to a running Space, resets the environment, and sends one action. Note that the client is an async context manager, and that the action is a typed object with a tool name and an arguments dictionary.

python
import asyncio
from echo_env import CallToolAction, EchoEnv

async def main():
    async with EchoEnv(base_url="https://openenv-echo-env.hf.space") as client:
        result = await client.reset()
        print(result.observation.echoed_message)
        result = await client.step(
            CallToolAction(
                tool_name="echo_message",
                arguments={"message": "Hello, World!"},
            )
        )
        print(result.observation.result)
        print(result.reward)

asyncio.run(main())

The README states that `reset()` yields the observation `"Echo environment ready!"` and that the step returns `"Hello, World!"` in `result.observation.result`. If you prefer not to write async code, the README documents a `.sync()` wrapper that turns the client into a synchronous context manager, with the same `reset()` and `step()` calls in the body. For a fuller walkthrough the README points at the getting-started page in the docs and at a Colab tutorial notebook under `examples/`.

The environment author's side: CLI, Docker, and a 23-backend sandbox list

The half of OpenEnv that is less visible from the quick start is the authoring path. The README says the `openenv` CLI provides commands to initialize new environments and deploy them to Hugging Face Spaces, and that creators can package environments "using canonical technologies like docker". The repository layout backs this up with `envs/`, `scripts/`, and a `src/` tree alongside a `tests/` directory.

The optional dependency groups in `pyproject.toml` are where the sandbox story lives. Separate extras exist for `daytona`, `aca` (Azure Container Apps sandbox, pinned with an upper bound because it is a preview SDK), `modal`, `novita`, `inspect`, and `harbor`. The comment on the `harbor` extra is unusually candid: every sandbox backend is included, and a bare `harbor` install "gives one that lists all 23 backends and can instantiate none of them", because each raises `MissingExtraError` from its constructor. That is a deliberate packaging stance, and it tells you the project expects you to install only the backend you actually run.

The design implication is that OpenEnv does not implement isolation itself. It defines the protocol and the packaging, and delegates the actual sandboxing to Daytona, Modal, Azure Container Apps, Novita, or whichever backend you select. If your threat model depends on a specific isolation primitive, that primitive belongs to the provider, not to OpenEnv.

Where OpenEnv is the wrong tool

The README carries an explicit early development warning: OpenEnv is "currently in an experimental stage", and you "should expect bugs, incomplete features, and APIs that may change in future versions". The release cadence is consistent with that. Three releases landed within about two weeks in September 2026, and `pyproject.toml` on `main` reads `0.6.1.dev0`, so the version you pin is the version you should expect to move.

The project also asks contributors to discuss significant changes before implementing them, and points to a technical committee. That is a governance signal worth reading literally: if your fork needs an API that the interface specification does not cover, you are negotiating with a committee, not merging a patch.

The third constraint is structural. Because environments live in containers reached over WebSocket, OpenEnv is a poor fit for anything latency-sensitive or for environments whose state cannot be serialized across a process boundary. A grid-world or a card game works fine. A training loop that needs microsecond stepping, or an environment that holds a GPU-resident simulator, will spend more time in transport than in learning. There is also no client in the core install, so a fresh `pip install openenv` gives you the protocol and the CLI but nothing to point them at.

How OpenEnv differs from Gymnasium and from Inspect

Gymnasium is the obvious comparison, and the relationship is closer to inheritance than rivalry. OpenEnv borrows the `reset()` and `step()` vocabulary directly and adds `state()`, then puts a network boundary and a container under it. Classic Gymnasium environments are in-process Python objects; you construct one and call it. OpenEnv environments are services. That difference is what buys isolation and reproducibility, and it is also why the two are not interchangeable: a Gymnasium environment you cannot containerize is not an OpenEnv environment.

Inspect is the more interesting contrast, and the repository treats it as adjacent rather than competing. There is an `inspect` extra in `pyproject.toml` and an `examples/evaluation_inspect.ipynb` in the tree, so the two are meant to be used together. The distinction is purpose. Inspect is built around evaluating model behaviour against a task suite, with scoring and logs as first-class outputs. OpenEnv is built around producing the interaction stream an RL update consumes. Evaluation asks whether a model is good; post-training asks what gradient to apply. The reward field on the step result is the tell: it exists to be differentiated, not merely reported.

The third point of comparison is the harness ecosystem the examples point at. The `examples/` directory includes BrowserGym harnesses and evaluations, a torchforge GRPO BlackJack example, Daytona terminal-bench runs, and a CARLA environment. Those are integrations with existing sandbox and benchmark projects, not replacements for them.

Maintenance, licensing and what an upgrade actually costs you

The repository is not archived, and the last push was on 2026-09-24, with v0.6.0 tagged the same day. The licence is BSD-3-Clause, declared in both the README badge and `pyproject.toml` via `license = "BSD-3-Clause"` and `license-files = ["LICENSE"]`. That is a permissive licence with no copyleft obligation on your own code, which matters if you are wrapping a proprietary simulator in an OpenEnv server. It says nothing about the licences of the environments you install, and those are separate packages with their own terms.

Upgrade cost concentrates in three places. The first is the interface itself: with RFC 001 still an open proposal for the baseline API and interface specifications, a `step()` signature change is a breaking change for every client you have written. The second is the sandbox extras. The `aca` extra is pinned `>=0.1.0b2,<0.2.0` with a comment explaining that the provider talks to it through a private adapter precisely because the SDK churns, so an Azure Container Apps upgrade is a re-validation, not a version bump. The third is environment clients installed from Git URLs, as the Echo example does; those track whatever the Space's default branch holds, so pinning them means pinning a commit rather than a version.

A practical hedge is to keep your own environment logic behind the `Environment` base class and keep the client thin, so that a protocol change touches the adapter rather than the simulator.

Editorial conclusion

Adopt OpenEnv if you are writing RL post-training loops and want one client shape for many sandboxed environments, or if you are an environment author who wants HTTP and Docker packaging handled for you. Do not adopt it as a production serving layer: the README states it is experimental and that APIs may change. Before committing, verify that the environment you need already ships a client package, since the core install does not include one, and read RFC 001, which is where the baseline API and interface specifications are still being settled.

Frequently asked questions

What is OpenEnv?

It is an end-to-end framework for creating, deploying and using isolated execution environments for agentic RL training, built on Gymnasium style APIs. Environments expose `step()`, `reset()` and `state()` and run as FastAPI servers inside Docker containers, reached by a client over WebSocket.

How do I install openenv?

The README gives `pip install openenv` for the core package, which requires Python 3.10 or later. Environment clients install separately; the README's example installs the Echo client from a Hugging Face Space with `pip install git+https://huggingface.co/spaces/openenv/echo_env`.

What is an OpenEnv environment?

It is a class derived from the `Environment` base that runs inside a Docker container behind a FastAPI server, with the README naming `EchoEnvironment` and `PythonCodeActEnv` as examples. A matching `EnvClient` subclass on the client side connects to it over WebSocket and exposes `reset`, `step` and `state`.

Official sources

  1. huggingface/OpenEnv on GitHub
  2. License: BSD-3-Clause
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huggingface-openenv.svg)](https://hysenlabs.com/projects/huggingface-openenv)