AxisRL is the system layer between SGLang and Megatron, and it declares one dependency
AxisRL is an agentic RL post-training framework built on SGLang rollout, Megatron training, and real-world agent workflows.
At a glance
- What is it?
- AxisRL is a post-training framework for agentic reinforcement learning: it coordinates multi-turn rollout, tool calls, verifiers, reward collection, and weight synchronization around a SGLang serving stack and a Megatron training stack. The interesting parts are the ones about silence. Its own argument is that small differences in tokenization, routing, or weight sync only surface later as loss spikes, and its tooling is built to reproduce them. The packaging tells a different story, with one declared dependency and a container image owned by one person's account.
- Who is it for?
- AxisRL is for a team already running SGLang and Megatron at scale with a multi-turn agent workload, where the gap between the rollout and training paths is where the bugs actually are. Check three things before you start.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 63 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The recommended environment is one person's image, pinned in a tag and a filename
The installation section does not tell you to build anything. It tells you to pull a prebuilt image:
docker pull leejunjie/sglang-mcore:cu130-sgl0.5.14-mcore0.18-magiwhich carries SGLang, Megatron Core, MagiAttention, Ray, the CUDA dependencies, and the Python packages the current recipes use. Then the package itself goes in as an editable install:
pip install -e .Two things are worth pausing on. The image lives in an individual's account rather than an organisation namespace, so the environment this project recommends is a container somebody else published. And the versions of the whole stack are encoded in the tag and repeated in the Dockerfile path, a CUDA version, an SGLang version, a Megatron Core version, and the attention implementation, so upgrading means editing a filename as much as editing a file. That is honest version pinning, and it is also a place where the setup can drift out of step with the code if someone pulls a newer image without renaming anything.
setup.py lists one dependency, and it is the cloud sandbox SDK
The package metadata is three fields wide: a name, a version of 0.1.0, and a single requirement, a cloud sandbox SDK at a version floor of 2.34.0 or newer. Everything the framework is actually built on, the serving engine, the training engine, the attention implementation, the cluster runtime, arrives through the container image instead. That is a defensible choice for a framework whose whole premise is a specific prepared environment, and it has a practical consequence: pip will happily install this package on a laptop and hand you an import that fails at the first reference to the serving stack, with the dependency list offering no hint that anything is missing. Two other details belong here. The package finder is told to include the recipe directory as well as the library, so the training scripts are part of the distribution rather than examples beside it. And the Python floor is 3.12, which is why the type checking configuration pins the same version.
Type checking is configured against a path that only exists in the container
Both major Python type checkers are configured here, and both configurations assume the prepared environment. The editor checker sets its Python version to 3.12 and adds an absolute path outside the repository to its search path, a workspace directory for the serving engine's Python package, which will not resolve on a checkout that has not been built into the image. It also points both the checker and the other tool at a stub directory in the tree. Then the muting starts. Six separate categories of report are switched off in the editor checker, including unknown argument types, unknown member types, and general type issues, and the command line checker disables three error codes and tells itself to ignore missing imports. So a green run in this configuration is weak evidence. If you rely on a type gate in CI, the settings to look at are these two blocks, not the pipeline.
The linter selects every rule, then turns off the deserialization one
The linter configuration starts from the strictest possible position, selecting all rules, and then earns it back with a long ignore list. Three of those entries are worth reading as policy rather than taste. One suppresses the rule about unsafe deserialization, and the comment next to it gives the reason: pickle and the modules that wrap it can be unsafe when used to deserialize untrusted data. In a framework whose inputs arrive from rollouts, harnesses, and tool results, that is exactly the rule you would want on. Two others suppress the rule against commented out code and the rule about implicit string concatenation, which are style choices. The numeric limits are generous rather than strict: line length one hundred and fifty, complexity twenty, fifteen arguments per function, twenty branches, fifteen return statements. A comment on the complexity setting describes it as a maximum function name length, which belongs to a different tool entirely.
The black box recipe tunnels out of the training host and is a work in progress
One recipe runs an agent harness AxisRL cannot see into, and the page is upfront about its status. It says the recipe is still a work in progress, demonstrating the integration path, and that the configuration, launch scripts, and proxy interfaces may change. Mechanically, the harness runs inside cloud sandboxes, calls AxisRL through an OpenAI compatible proxy, and AxisRL captures the model inputs, outputs, metadata, and rewards it needs for training. Three prerequisites are listed, and one of them is a network daemon: a tunnel client installed on the training host, because the default path uses a tunnel. The other two are a sandbox API key in the environment or a dotenv file, and a sandbox template with a specific name, which you build once:
cd axis_recipe/blackbox_rl/e2b_template
e2b template build --name axrl-openhands
cd -So the integration you would reach for first is the one flagged as unstable, and it asks for outbound tunnel access from the machine that holds your training job.
Config is overridable from the command line as a dotted path
The recipes are the entry points, and every one of them is a shell script you can adjust without editing. The mechanism is a dotted path and an equals sign, passed as a flag, which means a config field can be set from the shell for one run without touching the file that defines the default. Two of the quick starts differ almost entirely in the recipe directory name, one grouped reinforcement learning objective and one proximal policy optimization, both with the same update count of four. The retrieval recipe is the interesting one, because it needs a second process: it starts a search server on a port, runs a script to build its configuration, and then launches training against that generated file, with a second form that collapses the first three steps into a single script invocation. The output location is passed as an environment variable rather than a flag, which is a small inconsistency in an otherwise uniform interface.
The framework's own premise is that the dangerous failures are the silent ones
The clearest statement of what this project is for comes in the section on why it exists. Post-training workloads are moving past single turn question answering, and in an agentic setting a model interacts with a long lived environment, calls tools, reads the results, updates its context, and receives a reward only after several turns. The framework therefore has to coordinate multi turn rollout, environment state, tool calls, verifiers, reward collection, sample construction, and weight synchronization, which is a list of responsibilities rather than a feature. The sharper claim is about observability: small differences in tokenization, chat templates, log probabilities, routing, packing, or weight versions can appear much later as loss spikes, reward instability, or a mismatch between the rollout and training paths. That is why the debug surface is the largest part of the design, with mismatch analysis, routing replay checks, spike replay, and context packing all named as capabilities, and why the driver is kept thin, moving heavy payloads by handle so the trainer reads them on demand.
Editorial conclusion
AxisRL is for a team already running SGLang and Megatron at scale with a multi-turn agent workload, where the gap between the rollout and training paths is where the bugs actually are. Check three things before you start. There is no release history and the package declares version 0.1.0, so pin the commit. The type checking configuration resolves against a path that only exists inside the prepared container. And the black box recipe needs a tunnel daemon on the training host, so decide whether you are comfortable opening your training box before you enable it.
Frequently asked questions
What is AxisRL?
It is an agentic reinforcement learning post-training framework that sits between SGLang for high throughput rollout and Megatron for distributed training. Its job is the system layer: multi turn rollout, environment state, tool calls, verifiers, reward collection, training sample construction, and weight synchronization.
Which policy optimization objectives does AxisRL support?
The page names PPO, GRPO and GRPO2, GSPO, TOPR, TIS, and related variants, all configurable. Off policy stabilization tools named alongside them are truncated importance sampling, sequence masking, and one called Icepop.
How do I install AxisRL?
Pull the prebuilt image the project recommends, then install the package inside the container with an editable pip install. An optional script downloads the models and datasets the current recipes and tests reference, and the page warns that this bulk download includes multi billion parameter models.
Can AxisRL train an agent running in someone else's harness?
Yes, through black box harness capture over an OpenAI compatible proxy, where AxisRL captures model inputs, outputs, metadata, and rewards. The recipe demonstrating it with OpenHands in cloud sandboxes is described as a work in progress whose configuration, launch scripts, and proxy interfaces may change.
Does AxisRL publish releases?
The repository has no GitHub releases, and the package metadata declares version 0.1.0 with Python 3.12 or newer. The stack versions are pinned in the container image tag and in the name of the Dockerfile under the docker directory instead.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/xyz-ai-lab-axrl)