Qwen-VLA: one vision-language-action generalist for manipulation, navigation, and every arm
The official repository of Qwen-VLA
At a glance
- What is it?
- The Qwen team unified vision-language-action modeling across tasks, environments, and robot embodiments with a Qwen3.5-4B backbone and a 1.15B DiT action decoder, and a generalist that matches or beats per-task specialists.
- Who is it for?
- Qwen-VLA makes the strongest published case yet that a single vision-language-action generalist can hold its own against per-task specialists, with the ALOHA out-of-distribution gap, 76.9 versus 41.5 percent success, as the headline evidence.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 123 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The generalist claim, stated precisely
Qwen-VLA is the Qwen team's official entry into robotics: a unified vision-language-action generalist built on a Qwen3.5-4B vision-language backbone with a 1.15B DiT flow-matching action decoder. The claim the README makes is the one that matters in this field: a single model trained across all tasks and embodiments matches or outperforms task-specific specialists fine-tuned independently per benchmark, across both simulation and real-world evaluations.
That claim cuts against the dominant pattern in robot learning, where each platform and each task gets its own policy. Qwen-VLA casts manipulation, navigation, egocentric action modeling, and trajectory prediction into one shared action-and-trajectory prediction framework, so heterogeneous embodied data from different robots contributes to the same weights. The stated direction of travel is from skill specialists toward generalist actors, and the benchmark tables are structured to test exactly that.
Embodiment-aware prompt conditioning
The mechanism that lets one set of weights serve multiple robots is embodiment-aware prompt conditioning. Switching platforms requires changing a text prompt rather than swapping output heads or fine-tuning a new policy. There are no per-platform output heads at all, which is the architectural commitment that makes the generalist training possible.
The practical consequence is data efficiency at the fleet level. In the old regime, a new robot arm starts from zero unless someone hand-engineers transfer. Here, everything the model learned from bimanual ALOHA data, from navigation trajectories, from widow-style manipulation, is in the same weights the new platform conditions into. The README also documents a progressive training recipe: large-scale action pretraining, multimodal continued pretraining, supervised fine-tuning, then reinforcement learning, bridging discrete vision-language tokens and continuous action trajectories stage by stage.
Simulation benchmarks against specialists
The simulation results cover manipulation and navigation side by side. Qwen-VLA-Instruct scores 97.9 on LIBERO, 56.7 on RoboCasa-GR1, 73.7 on Simpler-WidowX, 86.1 and 87.2 on RoboTwin Easy and Hard, and 69.0, 57.5, and 59.6 on the R2R and RxR navigation metrics, every number achieved by one model evaluated across all platforms without per-benchmark adaptation.
The out-of-distribution table is arguably more important than the in-domain one. Fine-tuned solely on simple pick-and-place and evaluated on unseen spatial and visual tasks, Qwen-VLA-Instruct reaches 32.0 percent success on SimplerEnv-OOD. On DOMINO, zero-shot dynamic manipulation with moving objects using no dynamic training data, it records 26.6 success and 39.5 on the MS metric. Generalization under distribution shift is where specialist policies historically collapse, and these numbers are the README's main evidence that large-scale embodied pretraining changes that.
Real ALOHA results against GR00T and pi0.5
The real-world comparison sets the unified model against two per-task specialists, GR00T N1.6 and pi0.5, both fine-tuned independently per task, on an ALOHA bimanual platform. In-domain across six tasks, Qwen-VLA with pretraining averages 83.6 percent success against 71.6 for pi0.5 and 28.6 for GR00T N1.6, taking the best score in five of six categories including 96.2 on pick and place and 98.7 on bowl stacking.
The out-of-distribution real-world table widens the gap. Across color, instance, position, background, and instruction shifts, Qwen-VLA with pretraining averages 76.9 percent success against 41.5 for pi0.5 and 25.4 for GR00T. The without-pretraining rows in both tables do extra work: they show the same architecture learns far worse from scratch, which isolates the contribution of the large-scale embodied pretraining rather than the architecture alone.
What the repo contains
This repository is the official information hub rather than a full training release. It carries the overview diagram, demo videos, the benchmark tables, an Issues channel for questions, and links out to the technical report on arXiv and the team blog. The technical report is available at arXiv eprint 2605.30280, and the README links the BibTeX reference for researchers who need to cite the work.
For teams tracking the VLA field, the repo is worth watching for follow-up releases of weights or training details, since the benchmark tables and the architecture description define the evaluation standard the team is holding itself to publicly.
Why the unified framing matters beyond the numbers
The deeper significance of Qwen-VLA is what a unified action space enables downstream. Data flywheels become possible: footage from any embodiment feeds the same model instead of fragmenting across per-robot fine-tunes. Evaluation becomes comparable: one policy measured across platforms is a fairer picture than a grid of specialists each tuned on their own benchmark. And deployment simplifies: a lab or warehouse with mixed hardware conditions one model per capability rather than one per robot-task pair.
The training recipe also matters for the field's direction. Treating action generation as flow matching with a DiT decoder, conditioned on a strong vision-language backbone, continues the convergence of robot learning with the multimodal model stack, and the progressive schedule from action pretraining through reinforcement learning is a documented path others can attempt to reproduce. Whether generalist actors keep beating specialists as task diversity grows is now an empirical question with a published baseline to test against.
Editorial conclusion
Qwen-VLA makes the strongest published case yet that a single vision-language-action generalist can hold its own against per-task specialists, with the ALOHA out-of-distribution gap, 76.9 versus 41.5 percent success, as the headline evidence. The combination of a Qwen3.5-4B backbone, a 1.15B DiT flow-matching action decoder, and embodiment-aware prompt conditioning turns platform switching into a prompt change, and the progressive training recipe documents how the gap between language tokens and robot actions gets closed.
Frequently asked questions
What is the best VLA model?
There is no single measured best, but Qwen-VLA is the strongest recent published claim: its unified generalist matches or outperforms per-benchmark specialists on LIBERO, RoboCasa, SimplerEnv, RoboTwin, and navigation suites, and on a real ALOHA platform it averages 83.6 percent in-domain success against 71.6 for the pi0.5 specialist with 76.9 versus 41.5 out of distribution.
What does Qwen.vl do?
Qwen-VL is the Qwen team's vision-language model family for image and video understanding. Qwen-VLA extends that line into robotics: it builds on a Qwen3.5-4B vision-language backbone and adds a 1.15B DiT flow-matching action decoder so the model outputs robot actions and trajectories, not just text.
How does Qwen-VLA switch between different robots?
Through embodiment-aware prompt conditioning. One set of weights serves multiple platforms and switching embodiments requires only changing a text prompt, with no per-platform output heads or separate fine-tuned policies, which is what lets heterogeneous robot data train a single generalist model.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/qwenlm-qwen-vla)
Community notes