Library / SDK
learnsyslab/gym-pybullet-drones avatar
learnsyslab/gym-pybullet-drones

gym-pybullet-drones: quadcopter control as a Gymnasium environment, with the model version as the target

PyBullet Gymnasium environments for single and multi-agent reinforcement learning of quadcopter control

2,146 stars570 forksPythonMIT

At a glance

What is it?
A Python physics simulation package that wraps quadcopter tracking and hover tasks in the Gymnasium interface, aimed at reinforcement learning research rather than flight software.
Who is it for?
gym-pybullet-drones earns its place when a research question needs a flying vehicle as the controlled object and the surrounding physics has to be free. It ships the tracking and hover environments, the example policies, and the tests that keep them behaving, and it targets a specific stack: Python 3.12, Gymnasium, Stable Baselines3 2.0 and PyBullet.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 32 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A physics backend with a Gymnasium wrapper around it

The package is small in concept and large in files. `gym_pybullet_drones/` is the whole library, `tests/` holds the test suite, and the metadata lives in `pyproject.toml` under a poetry-core build. The name says what it is: PyBullet computes the rigid body dynamics of one or more quadcopters, and a wrapper class exposes observation spaces, action spaces, reset and step so that a reinforcement learning library can drive it.

That distinction matters more than the marketing wording does. The simulation is not trying to be an accurate model of any particular airframe. It is a physics engine plus a controller-friendly task definition, which is exactly what you want when the research question is about the learning algorithm and not about aerodynamics. The original work behind the package is the IROS 2021 paper, and the repository keeps that publication attached to the project rather than leaving it as an unlinked citation. The homepage field for the repository points at the arXiv PDF of that paper rather than at separate documentation, which tells you where the design rationale actually lives.

Two environments shipped, plus the citation trail back to IROS 2021

A tree walk of the repository shows a deliberately narrow surface. There is no docs directory, no examples folder at the top level, and no tutorial site in the file list, so the README carries the burden of explaining the whole interface. What it does cover well is the range of tasks, and the shipped examples make the scope concrete.

The README names three controllers that run without training: a PID position controller, a PID velocity controller, and MRAC, minimum required attitude control. If you want to see what the observation space looks like or how a policy is scored, running `mrac.py` is the fastest route, because it flies a formation without any checkpoint to load. The RL examples cover single agent hover at a target height and a two-drone hover at two different heights, which is the minimum useful case for multi-agent work.

sh
cd gym_pybullet_drones/examples/
python3 pid.py
python3 pid_velocity.py
python3 mrac.py

The citation block in the README points at the IROS 2021 proceedings entry and at a `paper` branch holding the original codebase. There is also a `master` branch for the same purpose, and the README tells you to check one of them out if you need the pre-refactoring version. That detail is worth keeping in mind when reproducing published numbers, since the current `main` branch has been reorganized around Gymnasium compatibility.

Installing on Python 3.12, including the part where PyBullet needs compiling

The install instructions assume a conda environment on Python 3.12, and the README is explicit that it was tested on Intel x64 with Ubuntu 24.04 and on Apple Silicon with macOS 26. The clone and environment creation steps are unremarkable:

sh
git clone https://github.com/learnsyslab/gym-pybullet-drones.git
cd gym-pybullet-drones/
conda create -n drones python=3.12
conda activate drones

The interesting part comes next. PyBullet stopped shipping pre-built wheels after Python 3.10, so on any newer interpreter pip has to build it from source. On Ubuntu that means installing build tools first, and on macOS it means passing a compiler flag that works around a file descriptor rename in the C sources:

sh
sudo apt install build-essential
sh
CFLAGS="-Dfdopen=fdopen" pip install pybullet --no-cache-dir

Once the source build is out of the way, an editable install pulls the rest from pyproject.toml. That dependency list is where the project's real targeting shows up: torch, numpy, scipy, transforms3d, matplotlib for the visuals, plus `gymnasium ^1.3`, `stable-baselines3 ^2.9`, `control ^0.10.2` and `pybullet ^3.2.7`. Notably absent is any hardware specific dependency, so the same install covers a CPU-only workstation. One contrast with many research packages: `control` is a hard dependency rather than something you install separately, which is a small sign that the MATLAB and Python control libraries are treated as equally first class.

Training with Stable Baselines3 PPO and replaying the best checkpoint

Training and evaluation are two separate scripts, and the README shows both, including the shell trick for picking up the newest run directory. That pattern matters in practice, because the output layout is one timestamped folder per training run under `results/`.

sh
cd gym_pybullet_drones/examples/

# single agent, task: single drone hover at z == 1.0
python learn.py
LATEST_MODEL=$(ls -t results | head -n 1) && python play.py --model_path "results/${LATEST_MODEL}/best_model.zip"

The multi-agent variant flips a flag rather than selecting a different script:

sh
python learn.py --multiagent true
LATEST_MODEL=$(ls -t results | head -n 1) && python play.py --multiagent true --model_path "results/${LATEST_MODEL}/best_model.zip"

Using Proximal Policy Optimization as the only demonstrated algorithm is a deliberate choice about what this project is for. It removes the algorithm shopping decision so the environment is the variable under study, which is the right trade if you are writing a paper about the task rather than about the optimizer. It also means the repository says nothing about whether PPO is a good choice for your problem; you would find that out by running it, not by reading. One example goes beyond hover and is worth knowing about: `downwash.py` demonstrates the ground effect problem where one vehicle's rotor wash perturbs a neighbour, which is a real and annoying failure mode for stacked formations.

Betaflight SITL wiring, and the Ubuntu-only caveat

The most unusual feature is a bridge to real flight controller firmware running in software. Betaflight SITL means building a firmware executable per simulated vehicle and letting the PyBullet physics step the same control loop the real board would run. It is the closest this repository comes to hardware, and it is also the most fragile part to set up.

sh
# one-time setup: from the repo's top folder, build one SITL executable per drone (e.g. 2), if needed, `apt install curl`
cd gym-pybullet-drones/
./gym_pybullet_drones/assets/clone_bfs.sh 2

# run the example
cd gym_pybullet_drones/examples/
python3 beta.py --num_drones 2
# --num_drones must be <= the number passed to clone_bfs.sh

The constraint on the final line is the one to remember. The number you pass to `clone_bfs.sh` is how many firmware builds exist, and `--num_drones` has to stay at or below it, so scaling up means rebuilding. The setup also assumes curl is available and is labelled Ubuntu only. For a research group comparing a learned controller against Betaflight's own mixer logic, this is the feature that justifies the package; for everyone else it is one script you can ignore.

Where the model ends, and the siblings that pick up the other half

The most useful thing in the README is the tip block at the very top, which routes readers to three sibling projects depending on what they actually need. For research involving symbolic dynamics and explicit constraints, the recommendation is `safe-control-gym`. For differentiable, GPU accelerated simulation, it is `crazyflow`, built on JAX. For deployment with PX4 or ArduPilot on ROS2 and JetPack hardware, it is `aerial-autonomy-stack`. Read that block as an admission about scope: this package is the physics-grounded Gymnasium interface, and the pieces around it are maintained separately by related groups.

The package version in pyproject.toml is `2.2.0`, while the newest GitHub release is `v1.0.0` dated 2021-03-09. Both are current facts about the repository. The reading that makes sense is that release tagging stopped being the versioning mechanism when the project moved to the poetry-core build with a `main` branch, so the version that pip installs is the one in pyproject.toml and the tags are historical. Pin your install if reproducibility matters to you, because the tags alone will not tell you which code you got.

The open question count is also worth reading honestly: 111 open issues against roughly two thousand stars is a busy tracker for a project this size, and it is the normal state of a research codebase that has become infrastructure for several groups. The last push was 2026-09-06, and the README's commented-out TODO block, which lists motor delay modelling and a switch from roll-pitch-yaw to quaternion state, is still there in the same commented state. That combination of recent commits and a frozen TODO list is what a maintained but slowly evolving research dependency looks like.

Editorial conclusion

gym-pybullet-drones earns its place when a research question needs a flying vehicle as the controlled object and the surrounding physics has to be free. It ships the tracking and hover environments, the example policies, and the tests that keep them behaving, and it targets a specific stack: Python 3.12, Gymnasium, Stable Baselines3 2.0 and PyBullet. What it does not ship is anything about real flight, wind realism, or the model identification work that separates a policy that hovers in simulation from a controller that lands a vehicle. The README is candid about that boundary, since it points at safe-control-gym for symbolic constraints, at crazyflow for GPU and differentiable simulation, and at aerial-autonomy-stack for PX4 and ArduPilot deployment. Start with the control examples rather than the learning examples, because the controllers run without any training and show you what the environment actually rewards. The one caveat to carry through everything is version drift: the package version in pyproject.toml and the newest release tag are two years apart, so pin your install rather than trusting either one alone.

Frequently asked questions

Is gym-pybullet-drones still being developed?

Yes, at a measured pace. The last commit was on 2026-09-06 and the project was featured in GitHub's Maintainer Spotlight for 2026, but the open TODO list in the README, covering motor delay and a move from Euler angles to quaternions, is still commented out rather than scheduled.

What is gym-pybullet-drones used for?

It provides Gymnasium environments for single-agent and multi-agent reinforcement learning of quadcopter control, wrapping PyBullet physics so that hover, velocity tracking and formation tasks can be trained with Stable Baselines3 rather than a custom simulation harness.

Which version should I install, the one in pyproject.toml or the newest release tag?

Install from the repository. The package version in pyproject.toml is 2.2.0 while the newest GitHub release is v1.0.0 from March 2021, because tagging stopped when the project moved to a poetry-core build on the main branch. The tags are history, not a version pointer.

How do I reproduce results from the IROS 2021 paper?

The README tells you to check out the paper or master branch for the original codebase. The current main branch has been refactored for Gymnasium and Stable Baselines3 2.0 compatibility, so it is not a drop-in substitute for the code that produced the published numbers.

Does it support Betaflight SITL?

Yes, on Ubuntu only. You build one SITL executable per drone with the bundled clone_bfs.sh script, and the --num_drones flag passed to beta.py has to stay at or below the number you built.

What should I use instead if I need differentiable simulation?

The README recommends crazyflow, a GPU accelerated, JAX based simulator from the same lab. For symbolic dynamics and explicit constraint handling it points to safe-control-gym instead.

Official sources

  1. learnsyslab/gym-pybullet-drones on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/learnsyslab-gym-pybullet-drones.svg)](https://hysenlabs.com/projects/learnsyslab-gym-pybullet-drones)