Tencent YOLO-Master: MoE Inside a YOLO Detector, and What It Costs You
[CVPR2026]🚀🚀🚀Official code for the paper "YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection." *(YOLO = You Only Look Once)* 🔥🔥🔥
At a glance
- What is it?
- YOLO-Master is Tencent Youtu Lab's CVPR 2026 detector that swaps uniform computation for sparse Mixture-of-Experts routing. It ships as a full Ultralytics-style pipeline under AGPL-3.0, which decides most of the adoption question before you benchmark anything.
- Who is it for?
- Adopt YOLO-Master if you are already inside an Ultralytics-style training and export workflow and want to evaluate whether instance-conditional routing helps on your own dense scenes; the ES-MoE blocks, the diagnose_model and prune_moe_model tools and the ONNX/TensorRT paths are the parts worth testing. Do not adopt it if your pipeline is closed source, since the repository is AGPL-3.0, or if you need a long-stable API with no churn.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The mismatch YOLO-Master targets: dense compute on sparse scenes
A standard YOLO detector spends the same budget on a frame that is mostly empty sky as it does on a crowded intersection. The paper's abstract frames this as misallocation: over-allocating representational capacity on trivial inputs while under-serving complex ones. YOLO-Master's answer is instance-conditional adaptive computation, where the amount of work per input depends on how hard the input looks. The audience is not someone training a first detector. It is an engineer who already has a YOLO baseline in production, has hit a latency or accuracy ceiling, and wants to know whether routing buys anything on their data distribution. The README is explicit that gains are most pronounced on challenging dense scenes and that efficiency is preserved on typical inputs, so the value proposition is conditional on your scene mix rather than universal.
How ES-MoE and dynamic routing actually change the forward pass
The mechanism is an Efficient Sparse Mixture-of-Experts block inserted into the YOLO-style backbone or neck, paired with a lightweight dynamic routing network. During training, a diversity-enhancing objective guides expert specialization so the experts become complementary rather than redundant. At inference, the router activates only the experts it judges relevant, which is where the FLOPs savings come from. The README describes this as moving from static dense computation to compute-on-demand. Two consequences follow. First, the model's effective depth varies per image, so latency is a distribution rather than a constant, and a p99 number matters more than a mean. Second, the router is a learned component with its own failure surface: the repository ships a file literally named debug-router-nan-activations.md, which tells you the maintainers have seen the router misbehave. The paper reports 42.4% AP at 1.62ms on MS COCO for YOLO-Master-N, described as +0.8% mAP over YOLOv13-N while being 17.8% faster. Those are the authors' numbers on their hardware, not a guarantee for yours.
Installing YOLO-Master and running a first validation
The repository keeps the Ultralytics package layout: pyproject.toml declares the distribution name as ultralytics, and requirements.txt lists the runtime dependencies. The declared Python floor is 3.8. The pyproject comments give the install commands directly, including the editable form for development.
pip install ultralyticsFor working against the repository itself, the same file documents the editable install, which lets you modify code without reinstalling.
pip install -e .Dependencies come from requirements.txt, which pins numpy>=1.23.0, torch>=1.8.0, torchvision>=0.9.0, opencv-python>=4.6.0 and ultralytics-thop>=2.0.18. Note the LoRA extra, peft>=0.18.0,<0.20.0, which is only needed for parameter-efficient fine-tuning. After install, the README points to a documented end-to-end workflow covering validation, training, inference and deployment to ONNX and TensorRT. There are also hosted entry points: a Hugging Face Space demo and a Colab notebook linked from the top of the README, which is the cheapest way to see routing behaviour before touching your own hardware. What you should expect from a first run is not a single latency figure but a spread, because the router decides per image how many experts fire.
Where YOLO-Master is the wrong tool
The licence is the first hard boundary. The repository is AGPL-3.0, and pyproject.toml carries the same identifier. If your product embeds the detector in a closed-source service, AGPL-3.0 obligations around network use are a real constraint, and that is a decision for your legal team rather than a technical one. Second, the project is young and moving: the release list shows v26.02 in February 2026 and v26.08 in August 2026, with two differently named August tags. Anyone who needs a frozen API for a two-year deployment should treat that cadence as a cost, not a feature. Third, the MoE design only pays off when scene complexity varies. If your inputs are uniformly simple, the router adds parameters, a training-time diversity objective and a new failure mode (NaN activations in the routing path, per the repository's own debugging note) in exchange for savings you will not see. A plain YOLO baseline is the better engineering choice there. Finally, the README does not document rollback behaviour or version-pinning policy, so downgrade paths are something you would have to establish yourself.
Compared with a plain Ultralytics YOLO pipeline
The honest comparison is not against another MoE detector, because the README claims YOLO-Master is the first deep integration of MoE into the YOLO architecture on general-purpose datasets. The comparison that matters is against the dense YOLO pipeline this repository is built on top of, down to the package name. A plain Ultralytics install gives you one deterministic computation graph per input size: latency is predictable, export is well-trodden, and there is no router to debug. YOLO-Master trades that predictability for adaptivity. The repository also carries utilities a stock YOLO install does not: CW-NMS (cluster-weighted non-maximum suppression), a sparse SAHI inference mode for slicing large images, Mixture-of-Attention and Mixture-of-Transformers support, and the diagnose_model and prune_moe_model tools for inspecting and shrinking the expert set. The pruning tool is the interesting one, because it implies you can recover a smaller, denser-ish model after training, which is the usual escape hatch from MoE inference complexity. Whether that trade is worth it depends on whether your accuracy gain on dense scenes exceeds the operational cost of a variable-latency model.
Maintenance, upgrades and what AGPL-3.0 means for a deployer
The last push to the default branch was on 2026-08-08, and the repository is not archived, so it is being worked on rather than frozen. The release history is uneven in naming: v26.08 and YOLO-Master-v26.08 both point at the August 2026 work, while YOLO-Master-v26.02 sits in February 2026. Anyone scripting an upgrade against tag names should read the release notes rather than assume a naming convention. Upgrade cost is dominated by the Ultralytics lineage: because the distribution is named ultralytics, installing this alongside a stock Ultralytics install is a conflict to plan for, not something to discover during a deploy. On licensing, AGPL-3.0 is a copyleft licence with a network-use clause. The practical implication for a hosted detector service is that source disclosure obligations may attach. I am not giving legal advice; the point is that this is the first question to settle, before any benchmark, because a favourable latency number does not change the answer.
Editorial conclusion
Adopt YOLO-Master if you are already inside an Ultralytics-style training and export workflow and want to evaluate whether instance-conditional routing helps on your own dense scenes; the ES-MoE blocks, the diagnose_model and prune_moe_model tools and the ONNX/TensorRT paths are the parts worth testing. Do not adopt it if your pipeline is closed source, since the repository is AGPL-3.0, or if you need a long-stable API with no churn. Verify first that your inference runtime supports the exported graph, then reproduce the 42.4% AP / 1.62ms COCO figure on your own hardware before committing a product roadmap to it.
Frequently asked questions
What exactly is YOLO?
YOLO stands for You Only Look Once, a family of single-stage object detectors. YOLO-Master is a YOLO-style framework for real-time object detection that adds Efficient Sparse Mixture-of-Experts blocks and dynamic routing on top of that architecture.
Is YOLO free to use?
YOLO-Master is released under AGPL-3.0, as stated in the repository and in pyproject.toml. That is a copyleft licence with a network-use clause, so it is free to use but not free of obligations for closed-source deployments.
How do I install YOLO-Master?
The pyproject.toml comments give the install command as pip install ultralytics, or pip install -e . for an editable development install. Runtime dependencies are listed in requirements.txt, which requires Python 3.8 or newer.
What accuracy and latency does YOLO-Master report?
The README states that on MS COCO, YOLO-Master-N reaches 42.4% AP at 1.62ms latency, described as a +0.8% mAP gain over YOLOv13-N while being 17.8% faster. These are the authors' reported figures, and the README notes the gains are most pronounced on challenging dense scenes.
What tools does YOLO-Master add beyond a standard YOLO pipeline?
The README lists MoE pruning, the diagnose_model and prune_moe_model diagnostic utilities, CW-NMS, sparse SAHI inference modes, and Mixture-of-Attention and Mixture-of-Transformers support. It also provides an end-to-end workflow covering installation, validation, training, inference and deployment to ONNX and TensorRT.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tencent-yolo-master)