MolmoAct2: Ai2 opens its action reasoning models for real-world robot control
Official Repository for MolmoAct2
At a glance
- What is it?
- Ai2 pairs the Molmo2-ER embodied reasoning backbone with a flow-matching continuous action expert, ships base and finetuned checkpoints for DROID, bimanual YAM, SO-100/101, and LIBERO, and plugs the whole stack into LeRobot and Hugging Face.
- Who is it for?
- MolmoAct2 is a full pipeline rather than a single model: an embodied reasoning backbone, staged base checkpoints with optional depth-token reasoning, finetuned policies for Franka, bimanual YAM, SO-100/101, and LIBERO, and the datasets behind all of it.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 38 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
An open stack for action reasoning
MolmoAct2 is the Allen Institute for AI's open family of action reasoning models for robot control and real-world deployment. The architecture stacks three pieces: the Molmo2-ER embodied-reasoning vision-language backbone, a robot state and action modeling layer on top of it, and a flow-matching continuous action expert that closes the loop for manipulation. In plain terms, the model looks at camera input, reasons about the scene and its own state, and produces smooth continuous control outputs for a robot arm rather than just describing what it sees.
The release is deliberately complete. It includes base checkpoints for continued training, fine-tuned robot policies ready for evaluation and deployment, and the datasets used to build both MolmoAct2 and Molmo2-ER. Everything is documented on the project blog and in the accompanying paper on arXiv, with model weights and dataset collections hosted on Hugging Face. For a field where many action models exist only as paper artifacts, a fully open stack from backbone to policy is the differentiator.
Base checkpoints for every training stage
Ai2 publishes base checkpoints at every stage of the training pipeline rather than a single finished model. MolmoAct2 itself is the post-trained model with the continuous flow-matching action expert attached, and it serves as the default foundation checkpoint for adapting to a new robot embodiment or benchmark. MolmoAct2-Think adds depth-token reasoning, for downstream policies that should reason over compact depth predictions before acting.
Two earlier-stage checkpoints round out the set. MolmoAct2-Pretrain is the pre-trained discrete autoregressive VLA backbone captured before the continuous action expert is attached, meant for continuing the training stages rather than for direct continuous-control inference. Molmo2-ER is the embodied-reasoning VLM backbone that served as the starting point for the whole action model line. Having each stage published separately means a lab can enter the pipeline at any point, whether that is full pre-training on their own data or a targeted fine-tune of the finished policy.
Finetuned policies for common robot platforms
For teams that want to deploy rather than train, the finetuned checkpoints cover the platforms the open-source robotics community actually uses. MolmoAct2-DROID is fine-tuned on the filtered DROID Franka mixture with absolute joint-pose control, intended for DROID-style policy inference or further fine-tuning. MolmoAct2-BimanualYAM covers the bimanual YAM mixture with absolute joint-pose control and annotated language instructions, and MolmoAct2-SO100_101 targets SO-100 and SO-101 datasets, the low-cost arms that dominate hobbyist and academic benchtop setups. A LIBERO variant is fine-tuned on the full LIBERO training mixture for benchmark work.
The README is honest about what a finetuned policy guarantees: performance depends on hardware, cameras, calibration, action conventions, and the language and task distribution. These checkpoints are starting points for closely related robots or direct policies for their training platform, not universal manipulation skills, and treating them that way is the difference between a working deployment and a disappointing one.
Ecosystem integration, release by release
The update log shows a project built to live inside an existing ecosystem rather than beside it. MolmoAct2 launched in early May 2026, and within nine days a LeRobot workflow for fine-tuning and inference appeared on a dedicated branch. By the end of that month the model was fully integrated into the official Hugging Face LeRobot repository documentation, so the standard LeRobot training and evaluation tooling works with it out of the box.
The following weeks added the operational pieces: FastAPI inference servers for DROID and YAM setups contributed by Jie Wang, evaluation rollouts of MolmoAct2-Cortex on bimanual YAM rigs packaged for failure annotation and reward model training, zero-shot evaluation on ManiSkill simulation in June, and then pre-training and post-training code with full experimental details. A maintenance update in August fixed finetuning issues in LeRobot's own MolmoAct2 training scripts. The pattern is consistent: each release closes a gap between research checkpoint and something an engineer can run on real hardware.
Evaluation before deployment
The project puts visible effort into evaluation infrastructure, which is where most open robotics releases are thinnest. A badge on the repository tracks MolmoAct2's standing as the first VLA on the MolmoSpace leaderboard, and the sim_eval directory provides zero-shot evaluation for the DROID and Bimanual YAM policies on ManiSkill simulation, so teams can sanity-check behavior without booking robot time.
The released evaluation rollouts serve a second purpose that matters for anyone training successor models: the MolmoAct2-Cortex rollouts on bimanual YAM setups were published specifically because they are useful for failure annotation and reward model training. Instead of only showing successful trajectories, the release gives the community labeled failure data, which is what reward models and better verification loops are actually built from. That choice signals how Ai2 expects the ecosystem around these models to develop.
Hardware reach, including Intel GPUs
Inference has been validated on Intel XPU, meaning Intel GPUs, and it runs without any code changes to the repository. The requirement is a PyTorch build with Intel XPU support and selecting the xpu device at runtime, which opens the model up to hardware beyond the usual NVIDIA-centric robotics lab.
Combined with the LeRobot integration, the SO-100/101 checkpoints, and the FastAPI inference servers, the practical picture is a stack a small team can adopt end to end: pick a finetuned policy for a supported arm, serve it behind the bundled inference server, evaluate in ManiSkill first, and fine-tune through LeRobot when the target platform drifts from the training distribution. The related searches around the project, from LeRobot integration to the SO-100/101 datasets, mirror exactly that adoption path, and the repository is structured so each step has its own documented entry point.
Editorial conclusion
MolmoAct2 is a full pipeline rather than a single model: an embodied reasoning backbone, staged base checkpoints with optional depth-token reasoning, finetuned policies for Franka, bimanual YAM, SO-100/101, and LIBERO, and the datasets behind all of it. With official LeRobot integration, ManiSkill evaluation, published failure rollouts, and Intel GPU inference with no code changes, it is one of the most complete open action reasoning releases available, and the fastest starting point for any team that wants real robot policies they can inspect, retrain, and deploy.
Frequently asked questions
What is MolmoAct2?
MolmoAct2 is Ai2's open family of action reasoning models for robot control. It combines the Molmo2-ER embodied-reasoning vision-language backbone with robot state and action modeling and a flow-matching continuous action expert for closed-loop manipulation, and it ships with base checkpoints, finetuned robot policies, and the training datasets on Hugging Face.
Which robots does MolmoAct2 support?
The finetuned checkpoints cover DROID-style Franka arms, the bimanual YAM platform, SO-100 and SO-101 low-cost arms, and the LIBERO benchmark mixture, all with absolute joint-pose control and annotated language instructions. Because every base checkpoint is also published, other embodiments can be supported through fine-tuning, and performance depends on hardware, cameras, calibration, and task distribution.
How is MolmoAct2 different from a standard vision-language model?
A standard VLM produces text about an image. MolmoAct2 adds robot state and action modeling on top of the vision-language backbone and connects it to a flow-matching continuous action expert, so the model outputs smooth continuous control signals for closed-loop manipulation instead of only descriptions, with an optional depth-token reasoning variant for policies that reason before acting.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/allenai-molmoact2)