Library / SDK
google-deepmind/mujoco_playground avatar
google-deepmind/mujoco_playground

MuJoCo Playground: GPU environments for robot learning

An open-source library for GPU-accelerated robot learning and sim-to-real transfer.

2,232 stars363 forksPythonApache-2.0

At a glance

What is it?
An open-source library of GPU-accelerated environments for reinforcement learning and sim-to-real work, built on MuJoCo MJX with a second physics path through MuJoCo Warp. What the package gives you, how the two implementations differ, and where the precision traps are
Who is it for?
Start with the Cartpole environment, because it is the cheapest way to find out whether your GPU, your JAX build and your driver stack agree with each other. Run train-jax-ppo with no flags first, then repeat it with the warp implementation and compare, and set JAX_DEFAULT_MATMUL_PRECISION=highest in your shell before you treat any number as a result rather than an anecdote.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 18 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A library of environments, not a simulator

The repository description calls this an open-source library for GPU-accelerated robot learning and sim-to-real transfer, and that wording is doing real work. The physics engine is not here. The README links MuJoCo MJX, which lives in the separate google-deepmind/mujoco repository, and the project is a suite of environments built on top of it. So when you install this package you are getting task definitions, observations, rewards, resets and training entry points, not a new simulator. That distinction matters for how you evaluate the project, because the environments are where the research contribution sits: someone has already decided what a cartpole balance task should look like, what a G1 joystick walking task should reward, and how a dexterous pick should be scored. The rest is plumbing to run those tasks thousands of times in parallel on a GPU. The README also points at a project website, playground.mujoco.org, which is where the fuller task descriptions live, and the repository records 2,232 stars, 363 forks, an Apache-2.0 license, Python as its primary language, and a last push on 2026-09-19.

What the environment list actually covers

The feature list in the README is short and worth reading item by item, because the range is wider than the name suggests. First come classic control environments from dm_control, which is the fastest way in and the sanity check most people should run before anything heavier. Then quadruped and bipedal locomotion environments, which is the core of the sim-to-real story. Then non-prehensile and dexterous manipulation environments, covering pushing, picking and grasping tasks that involve articulated hands. Vision-based support sits on top of all of this through the MJWarp Batch Renderer, which is what makes pixel observations possible at training scale rather than only for rendering a demo. The repository tree reflects that structure without much decoration: mujoco_playground for the library, learning for training code and notebooks, assets for imagery including the banner and the rough terrain texture, plus a CHANGELOG, a CITATION.cff, a pylintrc, a pre-commit configuration and a uv.lock. One naming detail catches people out early, since the distribution on PyPI is called playground rather than mujoco_playground.

Two physics implementations behind one flag

The most consequential change in this project over the past year is that it can train under two different physics implementations. A note near the top of the README says both the MuJoCo MJX JAX implementation and the MuJoCo Warp implementation at HEAD are supported, and it links a discussion post for the longer explanation. The Warp implementation lives in its own repository, google-deepmind/mujoco_warp, so this project is consuming it rather than containing it. The v0.1.0 release notes describe the first pass-through, which forwarded the implementation to MJX through a config override of the form registry.load with an impl key set to warp. The v0.2.0 notes then say the default MuJoCo implementation for all environments is now MuJoCo Warp, which is a significant default flip for anyone with existing scripts. On the command line the choice is a single flag:

bash
train-jax-ppo --env_name CartpoleBalance --impl warp

So a run that worked in January may behave differently in March without anything in your own code changing. If you care about comparability across time, pin the implementation explicitly rather than inheriting the default.

PyPI gets you moving, source gets you current

Installation has a short path and a longer path, and the README is explicit that the longer one is preferable. The short path is a single command, because the package is published to PyPI:

sh
pip install playground

The README then marks the source route as important, on the grounds that it gets the latest features and bug fixes from MuJoCo itself, which matters more here than in most libraries since the physics backend moves underneath. The source route requires Python 3.10 or later and is built around uv rather than pip. It clones the repository, creates a virtual environment on a chosen interpreter, installs a CUDA 12 build of JAX, and verifies the GPU backend before installing the library with all extras. Two details in that sequence repay attention. The verification step imports jax and prints the default backend, which is the only reliable way to know you have a GPU build rather than a CPU one that silently trains at a fraction of the speed. And Menagerie assets download automatically the first time you load a locomotion or manipulation environment, so a long pause on first run is expected rather than a failure.

Launching training and watching it happen

Once installed, the CLI entry points do the work. A cartpole run is one flag, and the same command works through uv when you have followed the source install:

bash
uv --no-config run train-jax-ppo --env_name CartpoleBalance --impl warp
uv --no-config run train-rsl-ppo --env_name CartpoleBalance --impl warp

The pair of lines is there because there are two training front ends, one from the JAX PPO script and one from the RSL implementation, and picking either with the same environment name is the way to compare them. For interactive inspection the README points at rscope, a separate tool that streams trajectories during training rather than leaving you with a final checkpoint and no idea what happened in between. It needs a second terminal, since the viewer is a long-running process of its own:

bash
python learning/train_jax_ppo.py --env_name PandaPickCube --rscope_envs 16 --run_evals=False --deterministic_rscope=True
# In a separate terminal
python -m rscope

Four Colab notebooks cover the same ground for people who would rather not set up CUDA: an introduction through the DM Control Suite, plus dedicated ones for locomotion, manipulation and vision.

Precision differences that will cost you an afternoon

The README has an FAQ section, and the reproducibility entry in it is the most practically useful text in the file. Users on NVIDIA Ampere GPUs, with RTX 30 and 40 series given as examples, can hit reproducibility problems because JAX defaults to TF32 for matrix multiplications, and that lower precision can hurt reinforcement learning training stability. The suggested fix is to export JAX_DEFAULT_MATMUL_PRECISION=highest before starting experiments, or append it to your shell profile, which restores the full float32 behaviour that older Turing systems gave you by default. A second reproducibility wrinkle is documented alongside it: to match the paper exactly you need the brax training script at a specific pinned commit, and results differ slightly if you use the learning/train_jax_ppo.py script instead. The FAQ also notes that bug reports belong in the issue tracker and that contributions from developers with robotics experience are welcome through the contribution guidelines. With 112 open issues, this is a project with real traffic and real unanswered questions, which is worth factoring into how you plan to spend time on it.

Three releases and a lot of renames

The changelog reads like an API stability report, which is exactly what you want from a library that sits between two fast-moving dependencies. v0.0.5 from 2025-06-23 renamed light_directional to light_type to follow a MuJoCo API change between versions 3.3.2 and 3.3.3, fixed a bug in get_qpos_ids, and implemented render in a wrapper. v0.1.0 from 2026-01-08 added the Warp pass-through, switched environments to contact sensors and removed collision.py, replaced mjx_env.init with mjx_env.make_data because make_data now takes an MjModel rather than an mjx.Model, added a device argument to make_data, and changed AutoResetWrapper so that resets on done are full resets, which also opens a path to curriculum learning through a done count in state info. That release also moved dependencies to mujoco 3.4 or newer and warp-lang 1.11 or newer. v0.2.0 from 2026-03-16 added vision-based PPO configs for CartpoleBalance and PandaPickCubeCartesian, moved the notebooks to MuJoCo Warp, made Warp the default implementation everywhere, and renamed the deprecated nconmax and nccdmax config fields to naconmax and naccdmax across all environments to match the updated MJX API. If you maintain scripts against older releases, those four config field names are the ones to check first.

Editorial conclusion

Start with the Cartpole environment, because it is the cheapest way to find out whether your GPU, your JAX build and your driver stack agree with each other. Run train-jax-ppo with no flags first, then repeat it with the warp implementation and compare, and set JAX_DEFAULT_MATMUL_PRECISION=highest in your shell before you treat any number as a result rather than an anecdote. Once that runs clean, load a locomotion environment such as G1JoystickFlatTerrain and let the Menagerie assets download, since that first fetch is the step most newcomers mistake for a hang.

Frequently asked questions

What is MuJoCo?

MuJoCo is a physics engine for robotics simulation, and this project is not it. MuJoCo Playground is a library of GPU-accelerated environments built on top of MuJoCo MJX, which is the JAX implementation living in the separate google-deepmind/mujoco repository. So the engine, the renderer and the physics backend are upstream dependencies, while the task definitions, rewards and training entry points are what you install from playground.

Is MuJoCo by Google?

MuJoCo is developed by Google DeepMind, and MuJoCo Playground sits in the same organisation under the google-deepmind GitHub account, with a citation file, an Apache-2.0 license and a project website at playground.mujoco.org. Playground also depends on a second Google DeepMind repository for its Warp physics path. Being Google-owned has not meant a closed platform, since the library is installed from PyPI and the source is public.

What does the impl warp flag change?

It selects MuJoCo Warp instead of the MJX JAX implementation as the physics backend for that run. Support for passing Warp through to MJX arrived in v0.1.0 through a config override, and v0.2.0 made Warp the default implementation for all environments. Since the default changed, pin the implementation explicitly when you need runs to be comparable across time.

Why does the first locomotion run take a long time?

Menagerie assets download automatically the first time you load a locomotion or manipulation environment, so the delay is an asset fetch rather than a hang. The README suggests triggering it explicitly with a one-line Python command that loads the G1JoystickFlatTerrain environment, which is a reasonable way to warm the cache before starting a real training run. Verify your GPU backend at the same time, since a CPU JAX build produces the same silence for a different reason.

Official sources

  1. google-deepmind/mujoco_playground on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-deepmind-mujoco-playground.svg)](https://hysenlabs.com/projects/google-deepmind-mujoco-playground)