InternLM/xtuner: a training engine aimed at ultra-large MoE models
A Next-Generation Training Engine Built for Ultra-Large MoE Models
At a glance
- What is it?
- XTuner V1 targets dropless MoE training at 200B to 1T parameters, with a documented focus on Ascend NPU optimization. Here is what the repository actually shows, and where the documentation stops short.
- Who is it for?
- Adopt XTuner V1 if you are training or fine-tuning MoE models in the 200B to 1T parameter range and you either run Ascend A3 Supernodes or want dropless training without a large expert-parallelism dimension. Do not adopt it if you need a stable API surface for small dense-model fine-tuning, if you are on vLLM or SGLang for inference, or if you need checkpoint conversion and rollback documented before you commit.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What XTuner V1 is for, and who it is not for
The README describes XTuner V1 as "a next-generation LLM training engine specifically designed for ultra-large-scale MoE models." The qualifier matters. This is not a general fine-tuning toolkit that happens to work on Mixture-of-Experts architectures; the design decisions in the README all point at the scale problem that appears once a model has hundreds of billions of parameters and a sparse expert layer.
The stated target is dropless training. In a dropless setup no tokens are discarded when experts are unevenly loaded, which is the normal condition during long-sequence training. The README claims 200B-scale MoE models can be trained without expert parallelism at all, and that 600B models need only intra-node expert parallelism. That is a claim about the parallelism layout, not a benchmark, and the README does not publish the configuration behind it.
Who it is for: teams with multi-node clusters who are already committed to MoE pretraining or supervised fine-tuning and who have hit a wall with a conventional 3D parallel layout. The roadmap table lists Intern S1, Intern VL, Qwen3 Dense, Qwen3 MoE, GPT OSS, Deepseek V3 and KIMI K2, each with GPU FP8, GPU BF16 and NPU BF16 columns. Several NPU cells are marked as in progress, so the supported matrix is genuinely narrower than the model list suggests.
Who it is not for: anyone fine-tuning a 7B dense model on a single GPU. The dependency set in pyproject.toml pins transformers==5.14.1, torch>=2.6.0, tilelang==0.1.11 and mmengine==0.11.0rc2, and requires Python 3.10 or newer. That is a heavy, tightly constrained stack to carry for a job that a simpler trainer would handle.
Dropless MoE training and the parallelism layout
The mechanism the README puts forward is a reduced expert-parallelism dimension. Traditional 3D parallelism splits a model across tensor, pipeline and data dimensions, and MoE training typically adds expert parallelism on top. XTuner V1 argues that for the MoE shapes common in current research, that extra dimension can be shrunk or dropped, which removes the cross-node communication that expert parallelism forces.
The long-sequence side is handled by memory optimization rather than by sequence parallelism. The README states that a 200B MoE model can be trained at 64k sequence length without sequence parallelism, and that DeepSpeed Ulysses sequence parallelism is supported when you do need it, with maximum sequence length scaling linearly. It also claims stability under expert load imbalance during long-sequence training, which is the failure mode dropless training is most exposed to: if tokens keep routing to a hot expert, memory and compute skew rather than being dropped.
The repository layout is consistent with a config-driven trainer. There is a recipe/ directory for training recipes, an xtuner/ package, examples/v1/ for the current version, and a patch/ directory. The Dockerfile builds on nvcr.io/nvidia/pytorch:25.03-py3, sets TORCH_CUDA_ARCH_LIST to "9.0 10.0", and installs flash-attn, grouped GEMM, DeepEP and causal-conv1d from source directories under /tmp. That tells you the engine depends on compiled kernels, not only on stock PyTorch operators.
What the README does not give is a data flow diagram or an explanation of how a recipe maps onto the parallelism strategy. A reader has to infer the layout from the feature list and the Dockerfile.
Installing xtuner and running a first recipe
The package is published on PyPI as xtuner, and pyproject.toml declares a hatchling build with a version read from xtuner/version.py. A plain pip install will pull the full runtime dependency list, including the pinned transformers and tilelang versions, so install into a fresh environment rather than an existing one. The README itself does not print an install command; it points readers at the documentation site at xtuner.readthedocs.io and at the PyPI package, and the dependency list in pyproject.toml is what a resolver will act on.
The base image line in the Dockerfile is the documented starting point for a container build:
ARG BASE_IMAGE=nvcr.io/nvidia/pytorch:25.03-py3
FROM ${BASE_IMAGE} AS setup_env
ENV TORCH_CUDA_ARCH_LIST="9.0 10.0"Those three lines set the build argument, derive the build stage from it, and fix the CUDA architecture list. Everything else in that file installs flash-attn, grouped GEMM, DeepEP and causal-conv1d from source directories, which is where most of the build time goes. That is also why the container path exists: the compiled kernels are not part of the PyPI wheel.
For a real run, the repository ships recipes under recipe/ and examples under examples/v1/. The README does not document a launch command, so the config format has to be read from those files directly.
There is also a reinforcement learning extra. pyproject.toml declares an optional dependency group named rl containing ray[default], httpx, fastapi, uvicorn, mathruler, pylatexenc and checkpoint-engine[p2p].
rl = [
"ray[default]",
"httpx",
"fastapi",
"uvicorn",
"mathruler",
"pylatexenc",
"checkpoint-engine[p2p]"
]That group is what the GRPO implementation depends on. The README lists GRPO as implemented, with MPO, DAPO and multi-turn agentic RL marked as coming soon, so the RL surface is the least settled part of the project.
Where the documentation leaves you on your own
The largest gap is checkpoint handling. The README does not document conversion between XTuner V1 checkpoints and the Hugging Face format, and it does not document rollback if a long run diverges. For a trainer aimed at multi-day runs on 200B models, that is not a small omission. The acknowledgements point at Torchtitan, DeepSpeed, MindSpeed and Megatron as inspirations, but inspiration is not compatibility, and nothing in the README claims format-level interoperability with any of them.
The second gap is the speed benchmark. The README includes a benchmark image but the extracted text carries no numbers, no hardware configuration and no baseline. The claim that FSDP throughput surpasses traditional 3D parallel schemes above 200B scale, and the claim that Ascend A3 Supernode efficiency exceeds NVIDIA H800, both appear as prose without a reproducible setup in the README. Treat them as direction, not as figures you can plan a cluster budget around.
The third gap is the inference integration list. LMDeploy is checked. vLLM and SGLang are unchecked. If your serving stack is vLLM, the training-to-serving path is not documented as complete.
There is also a maturity signal in pyproject.toml itself: the classifier reads "Development Status :: 4 - Beta". The 1.0.1 tag does not change that classifier.
How it differs from Megatron-LM and Torchtitan
Megatron-LM is the reference implementation for 3D parallelism at scale, and it is the approach XTuner V1 explicitly positions itself against. The difference is where the parallelism dimension goes. Megatron-style MoE training typically adds expert parallelism as a first-class dimension, which means expert routing traffic crosses node boundaries. XTuner V1's pitch is that this dimension can be reduced or removed for the model sizes in question, keeping the expensive communication intra-node.
That is a real architectural difference, not a packaging difference. It also has a cost: a layout that assumes expert parallelism stays intra-node is sensitive to the ratio of experts to nodes. If your expert count forces a split across nodes anyway, the advantage the README describes does not apply to you.
Torchtitan is the other natural comparison, and XTuner V1 credits it as an inspiration. Torchtitan is a PyTorch-native platform for generative model training. XTuner V1 adds a recipe layer, a reinforcement learning path with GRPO, and an explicit Ascend NPU target that Torchtitan does not carry. If you are on NVIDIA hardware only and want the thinnest possible layer over PyTorch, Torchtitan is closer to that. If you want recipes plus an RL path plus NPU support, XTuner V1 covers more ground at the cost of a heavier dependency set.
Neither comparison is settled by the README alone, because the README does not publish a head-to-head configuration.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-09, which is recent. The release cadence visible in the repository is uneven: v0.2.0 in July 2025, v1.0.0rc0 in November 2025, and v1.0.1 in May 2026. The V1 announcement in the README is dated 2025/09, which sits between the v0.2.0 and v1.0.0rc0 tags. Anyone tracking this project should expect the V1 line to move in larger steps than a patch-per-month cadence.
Upgrade cost is dominated by the pinned dependencies rather than by XTuner's own API. transformers==5.14.1 and mmengine==0.11.0rc2 are exact pins, and bitsandbytes==0.45.0 and tilelang==0.1.11 are pinned as well. If another package in your environment needs a different transformers version, you will be resolving that conflict by hand. The Dockerfile is the intended escape hatch, since it builds the whole stack from a known base image.
The licence is Apache-2.0, declared both in the LICENSE file and in pyproject.toml as {text = "Apache License 2.0"}. That is a permissive licence with an explicit patent grant, which is generally what enterprises want for a training dependency. This is a description of the licence text, not legal advice; if you are redistributing a built container image or a modified fork, read the NOTICE and attribution requirements yourself.
Editorial conclusion
Adopt XTuner V1 if you are training or fine-tuning MoE models in the 200B to 1T parameter range and you either run Ascend A3 Supernodes or want dropless training without a large expert-parallelism dimension. Do not adopt it if you need a stable API surface for small dense-model fine-tuning, if you are on vLLM or SGLang for inference, or if you need checkpoint conversion and rollback documented before you commit. Verify first that your hardware matches one of the checked rows in the roadmap table, that your installed torch and transformers versions satisfy the pinned constraints in pyproject.toml, and that the recipe/ directory contains a config close enough to your model that you are editing rather than authoring.
Frequently asked questions
What is InternLM/xtuner?
It is a Python training engine for large language models, described in its README as a next-generation training engine built for ultra-large MoE models. XTuner V1 emphasizes dropless training, long-sequence support and Ascend NPU optimization.
How do I install InternLM/xtuner?
The package is published on PyPI as xtuner, and the README points readers at the documentation site for setup. The repository also ships a Dockerfile based on nvcr.io/nvidia/pytorch:25.03-py3 for the full stack with compiled kernels.
Which models does InternLM/xtuner support?
The roadmap table lists Intern S1, Intern VL, Qwen3 Dense, Qwen3 MoE, GPT OSS, Deepseek V3 and KIMI K2, each with GPU FP8, GPU BF16 and NPU BF16 columns. Several NPU cells for GPT OSS, Deepseek V3 and KIMI K2 are marked as in progress.
What licence does InternLM/xtuner use?
The repository declares Apache-2.0 in both the LICENSE file and pyproject.toml. That is a permissive licence with a patent grant, though the README does not discuss redistribution requirements for built images.
Does InternLM/xtuner support reinforcement learning?
The README lists GRPO as implemented, and pyproject.toml declares an optional dependency group named rl containing ray[default], httpx, fastapi, uvicorn and related packages. MPO, DAPO and multi-turn agentic RL are listed as coming soon.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/internlm-xtuner)