Model or dataset
inclusionAI/AReno avatar
inclusionAI/AReno

AReno, a single-node post-training stack that refuses to fall back

An easy-to-use, fast toolkit to scale up RL post-training on a single node.

322 stars127 forksPythonApache-2.0

At a glance

What is it?
AReno is a Python toolkit for reinforcement learning post-training that packs training, serving, and agentic loops into one process on one machine, with no external backend in the loop. The design decisions worth knowing are the strict platform split between CUDA and MLX, a build that needs a manual environment variable to skip its own CUDA extension, and a Dockerfile that pins a library version the packaging metadata contradicts.
Who is it for?
AReno suits a researcher who wants to run a full reinforcement learning loop on one GPU or one Apple Silicon machine without assembling a training framework, an inference server, and a kernel library. Check four things before you start.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 9 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The image pins one transformers version and the metadata asks for another

Two build files in the same repository disagree about a dependency, and it is the dependency everything else loads through.

The package metadata requires transformers at 5.15 or newer and below 6. The Dockerfile declares a build argument defaulting that version to 4.56.2 and installs it with an exact pin.

The image installs its pinned version first, then runs an editable install of the repository with build isolation turned off. That second step resolves the project's own dependency list, which is where the requirement of 5.15 and up lives. So the pin in the image is overridden a few lines later, or the editable install fails, depending on how the resolver treats it.

Either way the two files describe different environments, and the Dockerfile's argument is the more specific of the two claims while the metadata is the one that gets enforced. Anyone building a pinned image from this and then auditing it will find a version in the layers that does not match the constraint in the project.

Two other pins do agree across both files, which makes the transformers mismatch stand out rather than blending into a pattern. The dataset library is pinned to one exact version in each, and the safetensors dependency has the same floor in both.

The build script writes errors as cause and next step

The setup file is longer than a setup file usually is, and the reason is a build that has to fail informatively.

The first decision is that metadata-only commands never touch the CUDA path. If the invocation is egg_info, dist_info, or sdist, the extension list comes back empty before anything is imported, which is what lets a documentation build or a metadata read succeed on a machine with no GPU.

Past that, an environment variable controls the extension. Set to auto, the build proceeds only on Linux with torch already importable, and skips silently otherwise. Set to zero or one of its other false-like spellings, it skips unconditionally, with the documentation noting that this is for docs and metadata installs only.

What follows is a chain of named checks: platform, torch presence, a minimum version of 2.6, and whether the installed torch is a CUDA build. Each has its own failure.

The most instructive failure is a missing memory library. The message explains that the CUDA extensions are built without isolation, that PyTorch's extension builder imports that library while sizing parallel compile jobs, and that the fix is to install it and retry with the no-isolation flag. Three facts, in that order: what broke, why, and what to type. Most build failures do not manage the third one.

The CLI will not substitute one backend for the other

Platform support is stated as a hard split rather than a best effort, and the sentence ruling out fallback is explicit.

Linux on x86_64 or aarch64 with an NVIDIA GPU and a CUDA-enabled PyTorch 2.6 or newer is one path. Apple Silicon macOS through the MLX stack is the other. Windows is supported only by way of a Linux subsystem, and only for the CUDA path. The CLI picks between them by platform and, in the project's words, does not silently fall back between backends.

That is a decision with teeth. A user on an unsupported combination gets a failure rather than a slow emulation, and the diagnostic command exists to tell them which combination they actually have.

The dependency markers in the package metadata are what make the split work. Torch and the attention library are marked for Linux only, while three MLX packages are marked for macOS on arm64 only, with the architecture in the marker rather than assumed. A Mac install therefore does not pull a Linux-only torch or CUDA package, which is the failure mode that makes most cross-platform deep learning installs miserable.

One library is pinned exactly: the attention implementation is fixed to a single version on Linux, while the MLX side has floors. That is a reasonable split, since the pinned one is the piece most sensitive to a torch upgrade.

areno check names six ways a machine can be wrong

There is a command whose only job is to tell you whether the host can run the toolkit, and it enumerates the failure modes rather than reporting one aggregate error.

Running it prints OK, WARN, or FAIL statuses with a concrete next step for each. The named problems are missing CUDA, a PyTorch install that is CPU-only, an unset CUDA home variable, an unavailable compiler, missing optional runtime dependencies, and a missing acceleration extension.

The documented container run uses the same command as its smoke test, which means the diagnostic is also how you confirm a pulled image works on your hardware:

bash
# Please make sure the host has an NVIDIA driver and NVIDIA Container Toolkit.
docker run --gpus all --rm -it \
  ghcr.io/inclusionai/areno:v0.0.8 \
  areno check

That list is a good map of what actually goes wrong on a fresh machine. The CPU-only torch case is the one that catches people most often, because a plain install succeeds and everything looks fine until a kernel launch fails. The acceleration extension is the other half of that story, since it is the optional piece that makes the training fast.

For issue reports there is a second command that emits a machine-readable environment dump. It covers the toolkit version, the interpreter, the platform, the PyTorch and CUDA situation, the GPU, the compiler, the import status of each dependency, and the relevant environment variables.

Both commands are cheap, and running the machine-readable one first would save a lot of back and forth in a bug report.

The operations agent points at your own endpoint, not theirs

There is a natural language front end to the toolkit, and how it is configured tells you what it is.

The agent command is described as a local operations assistant for train and serve tasks. It uses an OpenAI-compatible model to inspect the current checkout, read command help, run diagnostics, and produce or execute AReno commands for the machine it is running on. You configure it once with a base URL, a model name, and an API key.

The example configuration points at a local address on port 8000 with a small fast model and a key read from an environment variable. That is the shape to notice: there is no hosted service and no bundled key. Whoever runs the toolkit supplies the model.

The interaction model is careful in two ways worth copying. It asks follow-up questions through the terminal when a required parameter is missing, rather than guessing one, and it streams command output while it works. A natural language front end that silently invents a parameter is worse than no front end.

There is also a wrapper script at the repository root so the same agent can be run from a source checkout without installing anything. The example request in the README asks for a complete command to run the math demo with a given sample count, sized to the current GPU.

One optional CUDA extension is the whole performance story

The toolkit ships a native extension, and the README is careful to call it optional while the diagnostics treat it as important.

The build file refers to it by module and attribute path, and the diagnostic command has a dedicated status for it being absent alongside missing runtime dependencies. That means the toolkit works without it and simply runs slower, and that the two conditions are reported separately so you can tell a missing library from a missing kernel.

The extension is what the reference build profile is tuned for. The release profile enables fat link-time optimisation, sets the code generation unit count to one, and strips symbols. For a project whose argument is performance on a single node, building the native pieces with no cross-unit boundaries is the whole point of the profile rather than a default.

The example set matches the feature list one for one, which is a sign the features were built from use cases rather than the other way round. There are example directories for agentic work, maths, multimodal, supervised fine-tuning, and vision-language, matching the five headline capabilities of agentic loops, maths rewards, image and video content, SFT style training, and adapters.

The maths example is the one the quick start imports from, and it shows the reward plumbing: a reward record type, sampling parameters, a sequence type, a group advantages function, and a loss function named for a group sequence policy objective.

The repository root holds meeting notes and a blog

The top level tells you what kind of project this is, and two directories are not what you would expect.

Alongside the expected entries, the toolkit package, the tests, the documentation, the scripts, the examples, the continuous integration configuration, and a dashboard directory, there is a directory named for meetings and another named for a blog. There are also four markdown files at the root that are process rather than code: an agent map, a contributors guide, a release guide, and an agent instruction file.

For a research-adjacent project maintained by a team inside a larger organisation, that is normal and arguably a good sign. The project started with engineers from a reinforcement learning team at a large payments firm, is described as initiated by that team and maintained by a community, and its documentation lives on a team site rather than a personal domain.

The version tells the other half of the story. The releases are all in the zero series, and the newest is 0.0.8, tagged on 4 September 2026. The last push to main is 24 September 2026, so a few weeks of work sit past the tag without a release, and the package metadata classifies the project as alpha.

The Docker image is versioned to match, with the tag used in the documented run command pointing at the current release rather than a moving latest.

Editorial conclusion

AReno suits a researcher who wants to run a full reinforcement learning loop on one GPU or one Apple Silicon machine without assembling a training framework, an inference server, and a kernel library. Check four things before you start. Which backend your host actually has, because the CLI will not quietly substitute one for the other. Whether you can build the optional CUDA extension, which needs a toolkit on the machine rather than just a CUDA-enabled wheel. Whether the container image you pull matches your checkout, since its pinned library version differs from what the packaging metadata asks for. And which loss function you want, since the example imports one by name rather than treating it as a default.

Frequently asked questions

What is AReno?

It is a local LLM post-training toolkit for reinforcement learning, SFT and DPO style training, serving, and agentic RL. It stands for ASystem Reinforcement Learning Nano, was originally developed by engineers from the ASystem team at Ant Group, and is designed to run the whole loop on a single node with no external training or inference backend.

Which platforms does AReno support?

Linux on x86_64 or aarch64 with an NVIDIA GPU and CUDA-enabled PyTorch 2.6 or newer, plus Apple Silicon macOS through MLX. Windows is supported through a Linux subsystem for the CUDA path only, and the CLI selects CUDA or MLX by platform without silently falling back between them.

How do I install AReno from source?

On Linux, clone the repository and run the install script from the scripts directory. On Apple Silicon, clone, create a virtual environment, activate it, upgrade pip, and install the project in editable mode. Platform dependency markers install the MLX stack on Apple Silicon without pulling Linux-only Torch or CUDA packages.

What does AReno's reinforcement learning loop look like?

Five steps: construct a Trainer and initialise it to load the tokenizer and start workers, generate on-policy completions inside a rollout session, score each completion and turn the rewards into advantages with your own reward function, pack the rollout into sequence objects and run one optimizer step, then repeat until done and close the trainer.

Official sources

  1. inclusionAI/AReno on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/inclusionai-areno.svg)](https://hysenlabs.com/projects/inclusionai-areno)