NeMo Gym: NVIDIA's Environment Layer for Agent Evaluation and RL Rollouts
Evaluate and improve models and agents using environments
At a glance
- What is it?
- NeMo Gym is an Apache-2.0 Python library that packages tasks, agent harnesses, verifiers and per-task state into runnable environments for evaluation and reinforcement learning. It is worth adopting when you need scale or training; a plain script still beats it for stateless scoring.
- Who is it for?
- Adopt NeMo Gym if you need stateful, repeatable evaluation that feeds straight into RL training, or if you want to reuse its environment hub instead of writing sandboxes and verifiers yourself. Do not adopt it if a stateless check against a fixed dataset answers your question, because the README says a script is probably sufficient there, and do not adopt it if you cannot run Python 3.13.14 or newer.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem NeMo Gym solves: stateful agent tasks, not single-shot scoring
Most evaluation code assumes one prompt, one output, one score. Agents break that assumption. An agent that calls tools, executes code or works inside a sandbox produces a trajectory, and the score depends on what happened at each step, not just the final string. NeMo Gym is built around that distinction. The README defines an environment as four things: a dataset of tasks, an agent harness describing how the model interacts with the world, a verifier that scores task completion, and state, described as per-task execution context.
The intended audience is narrow and worth stating plainly. The README lists four signals: you need stateful environments such as code execution, tool calling or sandboxes; you want reproducible evaluation across teams using shared environments and verifiers; you need scale, meaning multiple repeats per task or thousands of concurrent requests for training; or you want to move between evaluation, agent optimization and training without rewriting the harness. The README also says the opposite case out loud: if you are scoring model outputs with a stateless check and do not need scale or training, a script is probably sufficient. That sentence is the most useful one in the document, because it tells you when to close the tab.
How the environment abstraction maps onto the repository layout
The four-part environment definition is not just documentation. It shows up in the directory structure. The repository has separate top-level directories for environments/, resources_servers/, responses_api_agents/ and responses_api_models/. That split suggests a request flow: an agent harness in responses_api_agents/ drives a model served through responses_api_models/, while resources_servers/ provides the tooling and sandbox side that gives the agent something to act on. Verifiers and datasets sit alongside the environment definitions themselves.
The consequence is that environments are meant to be shared artifacts rather than one-off scripts. The README describes an environment hub of popular benchmarks and training environments, and lists integrations with other environment libraries including Aviary, Harbor, OpenEnv, Reasoning Gym and Verifiers. Agent harnesses listed out of the box include OpenHands, Mini SWE Agent and LangGraph, with Codex CLI, KiloCode, RemoteAgent and anyswe_agent added in v0.5.0. Training integrations listed are NeMo RL, Unsloth and VeRL.
Two design choices are worth calling out. First, environments are coupled to a training loop by design, which is why the same object can serve evaluation and RL. Second, the optional OpenTelemetry tracing spans the agent, model and resources servers, so a rollout can be inspected end to end. Release v0.6.0 adds a standardized ng_trajectory schema and health checks for validating rollouts, which matters because a broken harness looks identical to a failing model until you can see the trace.
Installing nemo-gym and running a first evaluation
The README points to PyPI, and pyproject.toml declares the distribution name as nemo-gym with a Python floor of 3.13.14. That floor is unusually high and is the first thing to check on your machine, because it will rule out many existing images and CI runners. The repository ships a uv.lock, so uv is the natural installer.
python --version
uv pip install nemo-gymThe first command tells you whether you clear the 3.13.14 requirement at all. If it prints something lower, stop and upgrade the interpreter rather than fighting the resolver. The second installs the library from PyPI.
The README's news entries reference a unified gym CLI introduced in v0.4.0, and v0.5.0 documents a subcommand for recomputing rewards from stored rollouts without re-running inference.
gym eval reverifyThat command is the one to remember if you are iterating on a verifier: it lets you change scoring logic and re-score rollouts you already collected, instead of paying for inference again. The README does not document the full flag surface of the CLI, so treat the subcommand name as the confirmed part and check the docs site for arguments.
The examples/ directory is the best starting point for a real run. It contains environment_manifest.yaml plus several Slurm and vLLM launch files, including slurm_vllm_1IN.yaml and a multi-node Ray variant.
ls examples/You should see the manifest and the Slurm YAML files listed above. The manifest is the configuration object that ties a dataset, harness and verifier together; the Slurm files are launch configurations for scaling vLLM evaluation jobs across GPUs or nodes, which v0.6.0 lists as a highlight. Start from the manifest, not the Slurm files, unless you already have a cluster.
Where NeMo Gym gets in your way
The README carries a warning block that is unusually direct for a corporate project page: NeMo Gym is in early development, and you should expect evolving APIs, incomplete documentation and occasional bugs. It asks contributors to open an issue before making changes. Take that at face value. pyproject.toml classifies the package as Development Status 4 - Beta. If you are building a long-lived internal evaluation platform, the interfaces you code against are the part most likely to move.
The Python requirement is a second, harder constraint. Requires-python is >=3.13.14, and the README repeats 3.13.14 or higher. That is not a soft floor you can ignore with an older interpreter; it excludes most LTS distributions and many managed notebook environments. The repository also has no Windows-native path: the README lists Windows via WSL2, so a Windows team is really adopting Linux.
The third constraint is conceptual. The abstraction is heavier than the problem it solves for simple cases, and the README admits it. If your evaluation is a stateless check over a fixed dataset, you are paying for a dataset/harness/verifier/state split, a server topology and a CLI without using any of it. The same applies if your agents do not touch external state: no sandbox, no tools, no multi-turn execution means the stateful machinery is dead weight.
Finally, the documentation is uneven. The README documents the CLI subcommand names and the release highlights, but not the full argument surface, and it does not describe rollback or downgrade paths between releases. If you need a supported upgrade story before you commit, that is not something the README answers.
NeMo Gym compared with running your own harness on top of an RL framework
The obvious alternative is to keep using a training framework directly and write your own evaluation harness. NeMo RL, VeRL and Unsloth are all listed as training integrations rather than competitors, and that framing is accurate: those projects handle the training loop, while NeMo Gym supplies the environments and rollout collection that feed it. If you already have a working harness and only need a trainer, adding NeMo Gym inserts a layer you may not want.
The real difference is where the verifier and the per-task state live. In a hand-rolled setup they live in your training script, which means evaluation and training drift apart as the script grows. In NeMo Gym they live in a named environment that both paths import. That is the trade: you accept an opinionated four-part abstraction and a beta API surface, and in return the same environment that scores a model in evaluation can generate rollouts for training without a second implementation.
A second alternative is to use one of the environment libraries the README lists for interoperability, such as Verifiers or OpenEnv, alongside NeMo Gym. The README explicitly supports combining environments and benchmarks from other libraries, so this is not an either/or decision at the environment level. The choice is really whether you want NeMo Gym's rollout, tracing and scaling infrastructure around those environments, or whether you would rather call them directly from your own code.
Licence, releases and what an upgrade actually costs
NeMo Gym is Apache-2.0. The LICENSE file is referenced from pyproject.toml, the README carries the Apache 2.0 badge, and the source files carry SPDX headers naming NVIDIA Corporation and Apache-2.0. For most teams that is a permissive licence with no copyleft obligation on your own code, but it is worth confirming with your own legal review rather than treating a licence identifier as advice.
The release cadence visible here is fast: v0.5.0 on 2026-08-07, v0.5.1 on 2026-09-03, v0.6.0 on 2026-09-09, with the last push to the repository on 2026-09-10. Three releases in roughly five weeks, on a project that labels itself beta and warns about evolving APIs. The practical upgrade cost is not the install command; it is re-validating your environment definitions and verifiers against a moving interface, and re-checking that stored rollouts still reverify to the same scores after a verifier change. The gym eval reverify subcommand is what makes that check cheap, since it re-scores stored rollouts without new inference.
There is also a contributor-side cost worth knowing about. The Makefile defines Fern documentation targets, and its comments state that make docs-login must run before make docs, otherwise the autodoc step fails with HTTP 403 because the account is not provisioned. If you plan to contribute documentation, that is a setup step, not a bug.
Editorial conclusion
Adopt NeMo Gym if you need stateful, repeatable evaluation that feeds straight into RL training, or if you want to reuse its environment hub instead of writing sandboxes and verifiers yourself. Do not adopt it if a stateless check against a fixed dataset answers your question, because the README says a script is probably sufficient there, and do not adopt it if you cannot run Python 3.13.14 or newer. Before committing, install nemo-gym with uv, run the example environment manifest under examples/, and confirm that the environment you actually need exists in the hub rather than assuming the catalogue covers it. The project is in early development by its own warning, so the interfaces you build against today may move.
Frequently asked questions
What is NeMo Gym used for?
The README describes it as a library for evaluating and improving models and agents using environments, where an environment bundles a dataset, an agent harness, a verifier and per-task state. It is used both for reproducible evaluation and for generating rollouts that feed RL training through frameworks such as NeMo RL, Unsloth and VeRL.
Is NeMo Gym free to use?
The repository is licensed under Apache-2.0, with the LICENSE file referenced from pyproject.toml and SPDX headers in the source files. The README does not describe any paid tier for the library itself.
How do I install NeMo Gym?
The README links to the PyPI package nemo-gym, and pyproject.toml sets the distribution name to nemo-gym with requires-python of 3.13.14 or higher. The repository ships a uv.lock, so installing with uv is the path the layout suggests.
Does NeMo Gym need a GPU?
The README's requirements table states that a GPU is not required for NeMo Gym library operation, though a GPU may be needed for specific resources servers or model inference. The table lists Linux, macOS and Windows via WSL2 as supported operating systems.
Can I re-score rollouts without running inference again with NeMo Gym?
Yes. The v0.5.0 release notes list a gym eval reverify command that recomputes rewards from stored rollouts without re-running inference. The README does not document the full argument surface of that subcommand.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-nemo-gym)