RLMatrix puts reinforcement learning in C# with source generation, and its platform caveats are on the roadmap
Deep Reinforcement Learning in C#
At a glance
- What is it?
- RLMatrix is a reinforcement learning framework for C# built on TorchSharp, published to NuGet, with PPO, the DQN family up to C51 and Rainbow, and GAIL. Its pitch is a workflow argument rather than an algorithm argument: C# source generation removes the API plumbing so you write domain code. Its platform story is weaker, because testing non-Windows and non-CUDA setups is still listed as roadmap work.
- Who is it for?
- RLMatrix fits a C# or Unity team whose agents need to live in the same language and runtime as the game, since that is the case the framework is built for and where the integration claim has content. It does not fit someone who wants a measured speed comparison, because the performance claims in the readme carry no numbers or method, and it does not fit a Linux or Mac deployment without testing first.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 161 days ago.
- What is it written in?
- Mainly C#, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The performance claim is made twice and quantified nowhere
Two separate bullets in the feature list make performance claims, and neither carries a number.
The opening line says the framework provides performance exceeding Python alternatives, and a later bullet says it is faster and more stable than three named projects: the Python stable-baselines, the Unity ML-Agents toolkit, and the Godot reinforcement learning agents.
Naming three specific competitors is unusual and worth taking seriously. It also means the claim is falsifiable, and a reader has no way to falsify it from the documentation: no benchmark suite, no task list, no hardware description, no throughput or wall-clock figures, and no statement of what the comparison controlled for.
In C# versus Python, that last point is the substantive one. A native AOT-compiled loop and an interpreted one are not comparable on equal terms, so a fair number would depend entirely on which side of the JIT boundary the measurement sits.
The practical response is to treat these as product positioning rather than as evidence, and to build your own evaluation on your own environment if speed is the reason you are choosing this over a Python stack.
Testing non-Windows and non-CUDA platforms is still roadmap work
One roadmap item states the platform limitation plainly: test and optimize support for non-Windows non-CUDA platforms.
That line is worth reading next to the feature list rather than separately. The framework is built on TorchSharp, which is the .NET binding for the PyTorch native library, so it inherits the requirement for a working native backend and accelerator stack. On Windows with a supported GPU that path is the well-trodden one.
Elsewhere the readme is confident about engines rather than platforms. Game engine ready is stated as battle-tested in Unity and Godot, and the roadmap separately commits to further enhancing those integrations, which is consistent with the engines being the primary targets.
So the practical support matrix implied by the documentation is: Windows first, Unity and Godot as the intended hosts, and CPU-only or non-Windows deployments to be verified by the user.
The repository shape supports that reading. It is a Visual Studio solution with a single solution file at the top level, an editor config, and four example directories rather than any indication of per-platform packaging.
Source generation is the argument, not the algorithm list
Most reinforcement learning libraries are described by what they can train. RLMatrix describes itself by what you have to write.
The feature is called a DRL workflow and the mechanism is C# source generation inside the toolkit: you focus on your domain problem instead of wrestling with complex API requirements. In practice that means the repetitive part of a reinforcement learning setup, the observation handling, the action loop, the training plumbing, is generated from your declarations rather than hand-written.
That is a real advantage in a language with a strong type system, and it is the reason a C# team would pick this over a Python one: the compile-time errors you would get from a mis-shaped observation are the same class of error source generation removes.
It also explains the multi-head feature sitting next to it. Continuous, discrete, and mixed action spaces can be handled simultaneously in a single agent, which is impractical when each head means hand-written code.
Taken together, the two features describe the same bet: the value is in the ergonomics of the loop, not in a novel algorithm. If you are evaluating on algorithm coverage instead, the list is short.
The shipped algorithms are PPO, the DQN family, and GAIL
The algorithm list is explicit about what is included and where the ceiling is.
PPO is there. DQN is there with its popular modifications, described as going up to C51 and DQN Rainbow, which is the distributional and multi-step end of that family. GAIL is there. Everything else is described as on the way.
Two entries in the roadmap qualify that list. One commits to expanding the library with additional state-of-the-art methods. The other commits to adding more tools for imitation learning and GAIL, which is a signal that the GAIL implementation is present but not yet accompanied by the surrounding tooling an imitation learning workflow needs.
Recurrent support is a toggle rather than a separate algorithm. RNN integration is enabled with a simple option, and it exists for sequential or partial observability problems, which is the case where a memoryless policy cannot work because the state is not in the current observation.
For training infrastructure, the list covers multi-environment training across parallel environments that are optionally networked, a built-in dashboard for real-time metrics visualisation, and what is called industrial-grade distributed training with a fault-tolerant networked architecture.
Those three together are the difference between a research loop on one machine and a fleet, and they are the parts most likely to matter if you are training at scale.
Four example directories, and the readme sends you to another repository
The examples are organised as four directories, and the names tell you what each is for.
CartPole-v1 is the classic sanity check. LunarLander is a harder control problem with continuous actions. Godot is the engine integration. Networked is the distributed case, and its presence is a hint about how the networked training path is meant to be used.
There is a warning attached to all of them. The examples in this repository are described as being updated, and readers wanting more current examples are pointed to a separate testing repository dedicated to a CartPole example.
That note is a little awkward for a project whose argument is about developer ergonomics, since examples are the main teaching surface for a source-generation workflow. A separate repository for current examples means the answer to what the API looks like today lives somewhere other than where the framework lives.
The documentation itself is hosted separately as well, with the getting started guide at the project's own documentation site rather than in the repository, and the readme links there instead of reproducing a quick start.
Distribution is a NuGet package, and there are no repository releases
There are no GitHub releases for the project, and the package lives on NuGet under the name RLMatrix.
That is the whole distribution story: a .NET developer references the package, and the source repository is where the code and the examples live. Nothing else is documented as an install path.
Support is arranged in three channels rather than one. A Discord community server is offered for discussion, support, and updates. GitHub issues are the route for bugs and feature requests. And a business email address is given separately for commercial enquiries, which is a reasonable split between community support and anything contractual.
The licence is MIT, stored in the repository as a Markdown file rather than a plain licence file, which is a small detail that will show up as an automated scanner warning in some dependency tooling.
The last push to the default branch is dated 2026-04-23, so the tree has moved recently relative to the current date while carrying no tagged releases to indicate which version a given commit corresponds to.
Editorial conclusion
RLMatrix fits a C# or Unity team whose agents need to live in the same language and runtime as the game, since that is the case the framework is built for and where the integration claim has content. It does not fit someone who wants a measured speed comparison, because the performance claims in the readme carry no numbers or method, and it does not fit a Linux or Mac deployment without testing first. Before adopting it, confirm your target platform is one the project considers supported, plan to write your own evaluation rather than trusting the comparison against Python baselines, and read the examples note, which points to a separate repository for current code.
Frequently asked questions
What is an RL framework?
It is a library that supplies the training loop for reinforcement learning rather than the model itself: observation handling, action selection, the update step, and the surrounding infrastructure for running many environments at once. RLMatrix is one of these, written for C# on top of TorchSharp, and adds C# source generation so that plumbing is generated rather than hand-written.
Which algorithms does RLMatrix include?
PPO, DQN with its popular modifications up to C51 and DQN Rainbow, and GAIL, with more described as on the way. The roadmap commits both to expanding the library with further state-of-the-art methods and to adding more tooling for imitation learning and GAIL.
Does RLMatrix work on Linux or macOS?
That is listed as unfinished work. The roadmap includes test and optimize support for non-Windows non-CUDA platforms, and the framework is built on TorchSharp, so it depends on a working native backend and accelerator stack. The documented targets are Windows with Unity and Godot.
How do I install RLMatrix and where are the guides?
It is published on NuGet as RLMatrix and has no GitHub releases, so the package is the install path. The getting started guide lives on the project's own documentation site rather than in the repository, and the repository carries examples for CartPole, LunarLander, Godot, and networked training.
Are the RLMatrix examples up to date?
The readme warns that the examples in the repository are being updated and points readers to a separate testing repository for more current examples. The examples directory does contain CartPole-v1, Godot, LunarLander, and Networked projects, but the guidance is to treat them as work in progress.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/asieradzk-rl-matrix)