Library / SDK
agi-brain/xuance avatar
agi-brain/xuance

XuanCe advertises three deep learning backends and hard-requires one of them

XuanCe: A Comprehensive and Unified Deep Reinforcement Learning Library

1,087 stars162 forksPythonMIT

At a glance

What is it?
A deep reinforcement learning library with a long algorithm list, three backends in its badges and keywords, and a dependency manifest that requires torch and mentions neither of the other two. The packaging tells its own story too: the version is written in two files, the Python floor is declared below what its own dependencies support, and one environment ships three prebuilt native libraries of which the macOS one is marked Intel only.
Who is it for?
XuanCe suits someone who wants reference implementations of a long list of single-agent and multi-agent algorithms under one API, with parallel environments, multi-GPU training and automatic hyperparameter tuning already wired in. It is a poor fit if you need TensorFlow or MindSpore, since neither appears in the dependency list while the badges and keywords advertise them.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Three backends are advertised and one is an unconditional requirement

The page opens with install badges for three deep learning frameworks, and the package keywords repeat all three names. The introduction states an expectation of compatibility with multiple backends, naming PyTorch, TensorFlow and MindSpore, and hopes the project can become a zoo of algorithms. The dependency list tells a narrower story. `torch>=2.0.0, <3.0.0` is unconditional, and so is `torchvision`, and neither TensorFlow nor MindSpore appears anywhere in the requirements. So the multi-backend framing is a stated goal and a keyword, not an install-time choice: nothing in the manifest lets you pick a different framework, and no extra marker switches it. That gap matters for anyone whose organisation standardises on one of the other two, because the cost of finding out is an install and an import rather than a line in the project metadata. It is also consistent with the rest of the positioning, which is about breadth of algorithms rather than breadth of frameworks.

Version 1.4.4 is written in two files and both have to agree

The repository carries both a `pyproject.toml` and a `setup.py`, and they duplicate each other rather than one deferring to the other. The version 1.4.4 appears in both, the description string appears in both nearly word for word, and the keyword list appears in both. So a release is a two-file edit, and nothing in the build enforces that the two agree. The author field diverges as well: the declarative metadata lists one person with an address, while `setup.py` records the same name followed by an abbreviation for others. The package data story is entirely in `setup.py`, because the declarative file has no setuptools tool section at all, which means the native libraries and the YAML config globs are declared in the older file while the build backend is declared in the newer one. The two also disagree about how strict the license declaration is, one using a text field and the other a bare string, and both agree on MIT. The classifier set adds a topic of software development build tools, which is a slightly odd description for a reinforcement learning library.

The declared Python floor sits below what its own dependencies support

`requires-python` is set to 3.6 or newer, and the classifiers run from 3.6 through 3.12, listing every minor release in between. The dependency floors contradict the bottom of that range: torch 2.0 and gymnasium 0.28 both dropped the older interpreters, so a Python 3.6 or 3.7 interpreter cannot satisfy the manifest it is being asked to install from. The top of the range is better behaved, since the classifiers stop at 3.12 and torch is capped below 3.0. Four dependencies carry no constraint at all and instead have a suggested version in a trailing comment: numpy, scipy, PyYAML and pygame each appear bare, with the suggestion sitting in the source rather than in the requirement. The numpy comment explains the choice, saying it does not need pinning because each backend resolves a compatible version through its own dependency stack. That is a defensible decision, but it means the effective version floor for three of them is decided by whatever else is installed rather than by this package.

Two dependencies are pinned with equality while the rest are suggestions

The dependency list uses three different styles of strictness in one block. Two entries are hard-pinned: `pyglet==1.5.15`, whose comment even repeats the same version as its suggestion, and `moviepy==1.0.3`. Several are bounded ranges, including torch below its next major and gymnasium below 1.3.0, which is an unusual upper bound for a library to place on a widely used environment suite. And the rest are lower bounds or nothing at all. Two of those are not optional in practice: tensorboard and wandb are both unconditional requirements, while the features section describes logging as being with tensorboard or wandb, as if choosing one were enough. Installing the package gets both. `pettingzoo` is likewise unconditional and exists for the multi-agent side, so a user who only wants single-agent algorithms still pulls a multi-agent environment suite. The result is a dependency set where two exact pins will eventually block an upgrade and two logging backends arrive whether wanted or not.

One environment ships three native libraries and the macOS one is Intel only

The package data declaration is where this project stops being pure Python. Three native libraries are listed by exact filename for a single environment: a shared object for Linux, a DLL for Windows, and a dylib for macOS. Each carries a comment naming its platform, and the macOS entry adds a qualifier that the other two do not have, marking it as being for Intel processors. So on an Apple Silicon machine the bundled binary for that environment does not match the architecture, and there is no fourth entry to cover it. The same declaration also carries the configuration tree, using three glob depths of YAML files, which is what lets a library ship hundreds of experiment configurations without a data-file plugin. The pattern is a reasonable one for a project that has to run on three operating systems, and the per-platform comments show the packaging was done deliberately rather than by accident. The gap is the one platform nobody listed.

Two algorithm rows share one paper and one reference implementation

The algorithm list is long, with twenty-one single-agent entries before the model-based and multi-agent sections begin, and it is worth reading for how the citations are wired. Two differently named PPO variants, one with a clipped objective and one with a KL divergence, both cite the same paper and both link the same course homework as their reference code. Two parameterised-action variants, one split and one not, cite the same paper. So four names rest on three documents, which is defensible for variants but means the list cannot be used to tell those implementations apart from their citations. One row has a formatting fault: the TD3 entry runs its paper link straight into the code link without a separator, so that code link does not render as a link while every neighbouring row does. The list also mixes families without grouping them, running value-based methods, policy gradients, actor-critic methods and parameterised-action methods in one sequence, and the earliest entries carry papers only while later ones add code links.

The examples are filed by task, including drones and offline reinforcement learning

The example tree has eight directories, and the names are more informative than the feature list. There is a starting directory, a single-agent directory, a multi-agent directory, and one for offline reinforcement learning, which is a task family the algorithm list does not cover. There is a drones directory, which is the only place a concrete hardware application is named anywhere in the repository. There is a hyperparameter tuning directory, matching the automatic tuning feature, and two directories that read as contribution guides rather than samples, one for adding an algorithm and one for adding an environment. That structure is the practical answer to the project's own stated motivation, which is that reinforcement learning algorithms are sensitive to hyperparameter tuning, vary with different tricks, and suffer unstable training, so implementations tend to be elusive. The examples are organised so that each of those problems has somewhere to look. The documentation itself is built separately, with a Read the Docs configuration at the root and both English and Chinese contributing guides.

Two patch releases went out two minutes apart, and master is ahead of every tag

The release history has a detail worth noting if you pin by version. Two patch releases were published on the same morning two minutes apart, one at 05:49 and the next at 05:51, which is the shape of a packaging fix followed immediately by a correction to the fix. The next release came more than two months later. Since then the branch has been pushed to repeatedly, most recently on 2026-09-22, so the working tree sits well ahead of the newest tag with no release behind it. The declared version in both packaging files matches the newest tag exactly, which means the manifests are not simply stale; they are correct for the last release and simply not bumped for the commits since. Two smaller conventions show up in the same pass. The licence file carries a text extension, so tooling that looks for the unadorned filename will not find it, and the README exists in English and Chinese alongside matching contributing guides in both languages.

Editorial conclusion

XuanCe suits someone who wants reference implementations of a long list of single-agent and multi-agent algorithms under one API, with parallel environments, multi-GPU training and automatic hyperparameter tuning already wired in. It is a poor fit if you need TensorFlow or MindSpore, since neither appears in the dependency list while the badges and keywords advertise them. Before you depend on it, check three things: the Python floor of 3.6 in the metadata is below what torch 2.0 and gymnasium 0.28 support, the version lives in both pyproject.toml and setup.py so a release has to change two files, and two dependencies are hard-pinned with equality while the rest carry only suggested versions in comments.

Frequently asked questions

Which deep learning backends does XuanCe support?

The page states an expectation of compatibility with PyTorch, TensorFlow and MindSpore, and its badges and package keywords name all three. The dependency list, however, requires torch between 2.0 and 3.0 plus torchvision unconditionally, and neither of the other two appears there.

Does XuanCe cover multi-agent reinforcement learning?

Yes. Multi-agent reinforcement learning is one of the two task types the features list names, it has its own algorithm section starting with independent Q-learning, VDN, QMIX, WQMIX and QTRAN, `pettingzoo` is a direct dependency for it, and there is a dedicated MARL examples directory.

What does the magent2 environment need on my machine?

The package ships three prebuilt native libraries for it: a shared object for Linux, a DLL for Windows, and a dylib for macOS that the packaging comment marks as being for Intel processors. No Apple Silicon build is listed.

Does XuanCe include hyperparameter tuning?

Automatic hyperparameter tuning is one of the listed features, and the examples tree has a directory dedicated to it. Parallel environments and distributed multi-GPU training are named as separate features alongside it.

Which Python versions does XuanCe declare support for?

The packaging metadata declares 3.6 or newer with classifiers running through 3.12, and the torch requirement is capped below version 3. The lower end of that declared range cannot be satisfied by the manifest's own dependency floors, since torch 2.0 and gymnasium 0.28 no longer support those interpreters.

Official sources

  1. agi-brain/xuance on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/agi-brain-xuance.svg)](https://hysenlabs.com/projects/agi-brain-xuance)