CLI tool
radixark/miles avatar
radixark/miles

Miles: an RL post-training framework for trillion-parameter models

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

2,991 stars518 forksPythonApache-2.0

At a glance

What is it?
Miles pairs SGLang rollout with Megatron-LM training and adds the precision, routing and fault-tolerance pieces that large RL runs need. It is a fork of slime, and it assumes you already have a cluster.
Who is it for?
Adopt Miles if you are already running Megatron-LM or SGLang at multi-node scale and need async rollout, MoE routing replay or in-place engine recovery. Do not adopt it for single-GPU fine-tuning or for a first RL experiment on one workstation; the FSDP2 backend exists but the README states the recipes, the parallelism and the largest models all live on Megatron-LM.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Miles is for, and who it is aimed at

Miles is a reinforcement learning framework for post-training large language and vision-language models. The README describes it as "enterprise-ready" and aimed at "large-scale model post-training", which is a narrower audience than the phrase suggests. The target user is a team that already has a multi-node GPU cluster and wants to run GRPO, GSPO, PPO, REINFORCE++, SFT or on-policy distillation against a frontier-scale checkpoint.

The design assumption is visible in the dependency list. requirements.txt pins ray[default]>=2.56.0, sglang-router>=0.2.3, tensorboard, wandb, kubernetes_asyncio and nvidia-resiliency-ext, and the README states that rollout runs on SGLang while training runs on Megatron-LM. That is a distributed-systems stack, not a laptop stack. A PyTorch FSDP2 backend exists for runs that would rather train the HuggingFace implementation as-is, but the README is explicit that the recipes, the parallelism and the largest models all live on Megatron-LM. Treat FSDP2 as the on-ramp, not the destination.

If your problem is "I have one A100 and a dataset and I want to try RLHF", Miles is the wrong shape of tool. If your problem is "I have a 128-GPU allocation, a MoE model that keeps diverging, and rollouts that stall when one engine dies", the feature list reads like it was written for you.

The SGLang rollout plus Megatron trainer split

The architecture is a two-plane system. Generation happens on SGLang engines sitting behind a router; training happens in Megatron-LM. The router spreads requests across engines, preserves per-request metadata and health-checks the fleet, which is what makes multi-turn agentic workloads viable: a conversation that spans several turns has to land on an engine that still holds the right context, and per-request metadata is how that survives routing.

Rollout and training workers are decoupled, with configurable on-policy and off-policy schedules. That decoupling is the source of the framework's main selling point and its main complexity. Because the two planes run independently, weights have to move from trainer to engines while the run is live. Miles does this in-loop, and the README claims new weights reach the engines in seconds even on a trillion-parameter model such as Kimi-K2.6, with P2P RDMA as the fast path for disaggregated setups. The disk-delta weight sync path is visible in requirements.txt through the pins blake3==1.0.9 and xxhash==3.7.1 for checksums and zstandard==0.25.0 as the codec, which tells you there is a file-based fallback when RDMA is not available.

Two correctness mechanisms sit on top of this split. Token-in-token-out keeps the token stream intact between rollout and training with no detokenize/retokenize round-trip, and the README states it is supported for every model and every black-box harness. Rollout Routing Replay records expert routing during rollout and replays it in the trainer's forward pass, which addresses the MoE routing mismatch that the README says destabilizes large runs. Both are attempts to make the two planes agree on what actually happened, and both add bookkeeping you will have to budget for.

Installing Miles and running a first training job

The README does not inline installation steps. It points to https://miles.radixark.com/docs/getting-started/installation and to a per-GPU container image table, and it lists the supported hardware as NVIDIA GB300, GB200, B300, B200, H200, H100 and A100, plus AMD MI300X, MI325, MI350 and MI355X via ROCm. The installation page is where per-GPU status and the container image for each accelerator live, so read it before you clone anything.

The repository is packaged with setuptools, and setup.py declares the distribution name miles at version 0.1.0, pulling install_requires from requirements.txt. The extras_require table registers fsdp, mlflow and dashboard, and the dashboard extra is described in setup.py as standalone offline serving with fastapi, uvicorn and prometheus_client. Entry points sit at the repository root: train.py, train_async.py and train_multi_lora_async.py. The README does not print an install command, so the command below is what the packaging in setup.py and pyproject.toml supports, not a line quoted from the README:

bash
pip install -e .

For a first real run, the README directs you to the Quick Start page and to the launch script walkthrough under the user guide. The repository ships worked configurations rather than a single default config, including examples/geo3k_vlm/ for a VLM run, examples/lora/ and examples/multi_lora/ for adapter training, examples/on_policy_distillation/, examples/ppo/ and examples/swe-agent-harbor-docker/ for an agentic coding setup. The async entry point is the file the README's fully async description maps to:

bash
python train_async.py

Run it without a recipe attached and you will get argument errors, not a training run. The launch script walkthrough is the page that explains which flags the shipped examples set.

Where the design costs you: hardware, backends and recovery claims

The first real limitation is hardware. The README's hardware list is datacenter accelerators only. There is no consumer-GPU path documented, and the dependency list reinforces that: ray[default]>=2.56.0, kubernetes_asyncio, nvidia-resiliency-ext~=0.6.0 on Linux, ring_flash_attn on Linux, and torchft-nightly==2026.4.3 pinned specifically for Linux on x86_64. A macOS or Windows workstation cannot run the distributed path described here.

The second is the backend split. FSDP2 is offered, but the README says the recipes, the parallelism and the largest models all live on Megatron-LM. That means the path of least resistance for a small team is also the path with the least documentation and the least coverage in the shipped examples. Choosing FSDP2 to avoid learning Megatron's parallelism configuration buys you a smaller feature surface.

The third is the fault tolerance claim. The README states that when an SGLang engine dies, Miles recovers it and resumes the run in place with no restart and no pause. That is a strong statement and it is the kind of behaviour that depends on your cluster's health-checking and on how the router was configured. The README does not describe what happens when the failure is in the trainer plane rather than the engine plane, nor does it document rollback semantics for a partially completed step. If you are evaluating Miles for a long unattended run, that asymmetry is worth probing on your own infrastructure before you trust it.

Finally, the agentic connectors. Harbor, HUD, NeMo Gym, OpenEnv and Verifiers are listed, with sandboxes on AgentENV, Daytona, E2B or Modal. Each connector is a separate integration with its own failure modes, and the repository keeps the ones that need extra dependencies in extras rather than the base install.

Miles against slime, and against the FSDP2 route

Miles was forked from slime and the README says the two co-evolve. That matters for anyone already running slime: the divergence is deliberate, and the features Miles adds on top are the enterprise-facing ones. Low-precision training with MXFP8 and NVFP4, INT4 QAT, LoRA and multi-LoRA, Rollout Routing Replay for MoE, P2P RDMA weight transfer, and the in-place engine recovery all sit in the Miles column. If you are on slime and none of those matter to you, the fork gives you a second upstream to track without an obvious payoff.

The second comparison is internal rather than external. Megatron-LM versus FSDP2 is the real choice a new user makes. Megatron-LM gives you the recipes, the parallelism configurations and the largest supported models, and it brings a heavy dependency surface: requirements.txt notes that onnxscript is an import-time dependency of transformer_engine 2.17 and that without it any Megatron-backend run dies at import megatron. FSDP2 trains the HuggingFace implementation as-is, which is far easier to reason about, and it is the backend the README implicitly treats as secondary. Pick FSDP2 if your model fits comfortably and you value legibility; pick Megatron-LM if you are chasing the models the README actually names.

A third option worth naming is not using a framework at all. If your post-training is SFT plus a small preference-tuning run, the orchestration Miles provides (routers, weight transfer, engine health checks) is overhead you are paying for and not using.

Licence, releases and the cost of tracking a fork

Miles is Apache-2.0. That is a permissive licence, and it is the same family as the upstream slime project, which keeps the fork legally uncomplicated for commercial use. Apache-2.0 includes an explicit patent grant and requires that you preserve notices and state significant changes. If you modify Miles and redistribute it, the licence's change-notice obligation applies to your fork. Nothing here is legal advice, and if you are shipping a modified Miles inside a product, have counsel read the NOTICE and modification clauses rather than the summary.

The release history is thin: v0.1.0, dated 2026-08-18, is the only release listed, and it is also the date of the last push to main. The README's news section shows a steady cadence of feature landings through 2026, including day-0 support for several frontier models, so development is clearly ongoing even though tagged releases are rare. The practical consequence is that you should expect to track main rather than pin to tags, and that means reading commits.

The upgrade cost is the fork relationship. Miles co-evolves with slime, so upstream changes in slime may or may not arrive in Miles, and Miles-specific changes will not flow back. Budget for a periodic diff between the two trees if you depend on slime behaviour that Miles has modified. The dependency pins are also tight in places: transformers==5.12.1 is pinned with a comment explaining that 5.13 collides with SGLang's qwen3_asr, and safetensors>=0.8.0 is pinned because the samples-reply wire codec's malformed-payload exception contract was validated on that version. Those are the pins most likely to fight you when you try to share a virtual environment with another project.

Editorial conclusion

Adopt Miles if you are already running Megatron-LM or SGLang at multi-node scale and need async rollout, MoE routing replay or in-place engine recovery. Do not adopt it for single-GPU fine-tuning or for a first RL experiment on one workstation; the FSDP2 backend exists but the README states the recipes, the parallelism and the largest models all live on Megatron-LM. Before committing, verify the installation page lists a container image for your exact GPU, check that your model appears on the supported models page, and confirm whether the v0.1.0 release notes describe an upgrade path from the slime commit you are currently on.

Frequently asked questions

What is Miles for?

Miles is a reinforcement learning framework for post-training large language and vision-language models at scale. It pairs SGLang for rollout with Megatron-LM for training, and supports GRPO, GSPO, PPO, REINFORCE++, SFT and on-policy distillation.

What is the meaning of Miles?

The README does not define the name. It closes its About section with the line "A journey of a thousand miles begins with a single rollout", which is the only explanation the documentation offers.

How do I install Miles?

The README does not inline install steps; it points to the Installation page at miles.radixark.com/docs/getting-started/installation, which lists per-GPU status and the container image for each supported accelerator. The repository itself is packaged as a Python distribution named miles at version 0.1.0 through setup.py.

Which GPUs does Miles support?

The README lists NVIDIA GB300, GB200, B300, B200, H200, H100 and A100, plus AMD MI300X, MI325, MI350 and MI355X through ROCm. Per-GPU status and container images are on the Installation page.

Does Miles support LoRA training?

Yes. The README describes low-rank adapters that train frontier-scale models on a fraction of the GPUs, with the same adapters loading straight into SGLang for rollout. The repository ships examples/lora/ and examples/multi_lora/, and there is a train_multi_lora_async.py entry point.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/radixark-miles.svg)](https://hysenlabs.com/projects/radixark-miles)