sumo-rl, whose readme is stitched from fragments and whose version is scraped
Reinforcement Learning environments for Traffic Signal Control with SUMO. Compatible with Gymnasium, PettingZoo, and popular RL libraries.
At a glance
- What is it?
- Reinforcement learning environments for traffic signal control, wrapping a traffic simulator behind Gymnasium and PettingZoo interfaces with customisable observations and rewards. The newest release is more than two years old, the newest commit is about seven months old, the simulator's only documented install path is an Ubuntu package archive, and the one performance switch on the page costs you the graphical interface.
- Who is it for?
- sumo-rl fits a research project that wants a traffic simulation in a standard reinforcement learning interface rather than a bespoke one, and its extension points are the reason: observation and reward are both replaceable, and the reward example is three lines long. Three things to check before you build on it.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The newest release is two years old and the newest commit is seven
The dates do not line up. The three most recent releases are 1.4.3 in June 2023, 1.4.4 in March 2024, and 1.4.5 in May 2024, while the last push on the repository is dated 2026-03-08. So the release line stopped more than two years before the last commit, and the last commit is itself around seven months old. Nothing on the page claims otherwise, and the arrangement is a normal one for a research dependency: a DOI is published at the top of the readme so a specific version of the source is citable and fixed. The practical consequence is which of two install lines you use. The first gives you the stable release, which is the tag. The second clones the repository and installs it in editable mode, which gives you everything committed since the tag. The page presents the second as the way to get the latest unreleased version, and after a gap this long that is the line most people will want.
The simulator's only install path is an Ubuntu package archive
Installing the library is one pip command. Installing the simulator underneath it is four, and every one of them assumes a Debian-family machine:
sudo add-apt-repository ppa:sumo/stable
sudo apt-get update
sudo apt-get install sumo sumo-tools sumo-docThen the path has to be exported, which the page does by appending a line to a shell profile and sourcing it, with the default install location named as the value. There is no macOS path, no Windows path, and no conda path anywhere on the page, even though the library itself declares a floor of Python 3.9 and says nothing else about the host it needs. So a researcher working on a laptop that is not running Linux reads a Linux installation and has to find the simulator elsewhere. The three packages installed are the simulator, its tools, and its documentation, which is a generous default if you want the manuals locally.
The eight-times speed switch costs you the interface and the parallelism
One environment variable carries a performance claim and a compatibility warning at the same time:
export LIBSUMO_AS_TRACI=1Setting it swaps the control interface for the compiled library, and the page puts the gain at roughly eight times. The warning underneath matters more than the number does. With it active you cannot use the graphical interface, and you cannot run more than one simulation at a time. Both usage examples on the page ask for the graphical interface, and the optional rendering extra installs a virtual display so that a machine without a screen can still draw one. So the two choices pull against each other. A benchmarking run takes the speed and gives up the parallelism; watching what a policy is actually doing takes the interface and gives up the speed. The trade is stated rather than buried. The eight times figure belongs to the simulator's own documentation and is not a measurement made in this repository.
The registered environment is still called v0 while the library is at 1.4
The single-agent example builds its environment through the standard factory using the identifier sumo-rl-v0, passing a network file, a route file, an output path, a request for the interface, and a simulation length in seconds. That identifier still carries a zero version suffix while the library's own releases sit in the 1.4 range, so the name a user types has not moved in years. The multi-agent example takes a different route and calls a parallel environment constructor directly, pointing at a bundled benchmark network and its route file, then resetting and stepping per agent, with a comment marking the exact line where a policy would be inserted. The two return different shapes, correctly for their specifications: the single-agent reset yields an observation and an info dictionary, while the parallel reset yields a per-agent mapping, and the step call comes back with per-agent observations, rewards, terminations, truncations, and infos.
The yellow phase is a constraint the policy never chooses
The action space is discrete, and every agent picks the next green phase configuration once per decision interval. The detail that changes how a result should be read is the sentence after the example: every phase change is preceded by a yellow phase lasting a configured number of seconds. The agent therefore never chooses whether to transition, and never chooses a yellow state; it chooses among green configurations and the simulator inserts the transition for it. In the two-way single intersection example that leaves four discrete actions, and the count is a property of the phase set, not of the environment. So a result reported on that intersection is a result over a four-way green choice with the inter-green handled for you, which is a different problem from choosing freely among four phases that can be held or advanced as the policy likes.
The version is scraped from a source line by a function naming another library
The build configuration declares the version dynamic instead of writing the number down, and a companion build script works it out at build time: it opens the package's init file, walks the lines, and returns the value from the one that starts with the version variable name, raising an error if it cannot find it. That is a workable arrangement for keeping one number in one place rather than two. The function's own docstring, however, says that it gets the version of a different reinforcement learning environment library entirely. It is a copy that was never updated to name this project. Nothing breaks, and the number it returns is correct. Anyone editing the packaging layer should know that the comment sitting next to the version logic is describing something else entirely.
The readme is assembled from fragments and outputs are versioned
Two structural details explain the shape of the page. Each section is wrapped in a pair of marker comments, an opening tag and a closing one, around the intro, the install steps, the observation, the action, the reward, and each of the two API examples. That is the signature of shared includes, and the lint configuration lists a documentation scripts directory among its source paths, so the same text is very likely assembled into the online documentation as well. That is why the page reads as a summary with pointers rather than as a manual: it is a fragment set, not the whole book. The other detail sits at the tree level, where an outputs directory stands beside an experiments directory and a tests directory. In a research repository, a versioned outputs directory usually means generated results are committed alongside the code that produced them, which makes the history readable but the clone heavier than the code alone would suggest.
Editorial conclusion
sumo-rl fits a research project that wants a traffic simulation in a standard reinforcement learning interface rather than a bespoke one, and its extension points are the reason: observation and reward are both replaceable, and the reward example is three lines long. Three things to check before you build on it. The release line stopped more than two years before the last commit, so a plain pip install gets you the tag, not the current tree, and the page's own alternative is an editable install from a clone. Installing the simulator underneath it is documented for one operating system only. And the speed switch that matters is mutually exclusive with the interface and with running more than one simulation, so decide whether you are benchmarking or watching before you set it.
Frequently asked questions
What is sumo-rl and what does it wrap?
It provides reinforcement learning environments for traffic signal control built on the SUMO traffic simulator. The main class behaves like a standard Gymnasium environment when instantiated as single-agent, and multi-agent versions follow the PettingZoo APIs, with one class responsible for retrieving information and actuating on traffic lights through the simulator's control interface.
How do I install sumo-rl and the simulator it needs?
The library installs with pip install sumo-rl for the stable release, or from a clone with an editable install for the latest. The simulator underneath is documented only for a Debian-family machine, by adding its package repository, installing the simulator, its tools and its documentation, then exporting the SUMO_HOME variable to the default install path in a shell profile.
What is the default reward function in sumo-rl?
It is the change in cumulative vehicle delay, meaning how much the total delay, the sum of waiting times of all approaching vehicles, changed relative to the previous time step. You can substitute one of the implementations in the traffic signal class with the reward_fn parameter, or write your own and pass it to the environment constructor.
What does an agent observe in sumo-rl by default?
A vector made of a one-hot encoding of the current active green phase, a binary flag for whether the minimum green time has already elapsed in the current phase, and then per incoming lane both its density and its queue length, each divided by the lane's total capacity. Queued means a speed below 0.1 m/s. You can replace it by inheriting from the observation function class and passing your own to the environment constructor.
What does the LIBSUMO_AS_TRACI variable do in sumo-rl?
It switches the control interface to the compiled library for a performance gain the page puts at roughly eight times. The cost is stated in the same place: with it active you cannot run the graphical interface and you cannot run multiple simulations in parallel. The number comes from the simulator's own documentation rather than from a benchmark in this repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lucasalegre-sumo-rl)