# microsoft/Magma: one multimodal model aimed at both screen clicks and robot arms

> A Microsoft Research foundation model for multimodal agents, released as an 8B checkpoint with inference code, training code, and two families of visual annotation datasets.

**microsoft/Magma** — [CVPR 2025] Magma: A Foundation Model for Multimodal AI Agents

- Repository: https://github.com/microsoft/Magma
- Stars: 1,947 · Forks: 165
- Language: Python
- License: MIT
- Published: 2026-10-08 · Updated: 2026-10-08 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-magma

## What the authors claim, and what is checkable

The repository is Microsoft Research's Magma, described in the title line as a foundation model for multimodal AI agents and accepted at CVPR 2025. Thirteen authors are listed across five affiliations: Microsoft Research, the University of Maryland, the University of Wisconsin-Madison, KAIST and the University of Washington. Jianwei Yang, Reuben Tan and Qianhui Wu are marked as project leads and first authors, and Jianfeng Gao is listed under leadership.

The highlights section makes four claims. That Magma is the first foundation model for multimodal agents handling both virtual and real environments. That a single model does generic image and video understanding and also generates goal-driven visual plans and actions. That it reaches state-of-the-art results on UI navigation, robotic manipulation and generic image and video understanding, with spatial reasoning called out specifically. And that the pretraining strategy scales from unlabeled video in the wild alongside existing agentic data.

The first claim is the authors' own framing and should be read that way. A single model spanning screen interaction and physical manipulation is the unusual part, and it is the part the demo selection supports: a UI agent, a gaming agent and a robot visual planner all appear in the news timeline, which is a wider range of embodiments than most multimodal releases demonstrate.

The project page is worth reading on its own terms. There are no GitHub releases, no topics and no homepage field, and the license is MIT with the language given as Python. The paper and the hosted project page are linked from the README instead.

## Set-of-Mark and Trace-of-Mark, the two annotation families

Magma's training data comes with two kinds of visual annotation, and the naming is easy to get wrong. SoM appears in the README as SoM prompting annotations attached to web and mobile interaction datasets, and ToM appears as visual traces attached to video and robot demonstration data.

The datasets are published on Hugging Face under the MagmaAI organisation, and the news timeline dates each one. On 2025.04.06 the Open X-Embodiment pretraining data with visual traces went up as Magma-OXE-ToM. On 2025.04.12 the pretraining videos with visual traces went up as Magma-Video-ToM. On 2025.04.29 the interaction datasets landed as Magma-Mind2Web-SoM and Magma-AITW-SoM, and the README says those two were used for the downstream finetuning results reported in the paper.

So the split is by data source rather than by task: web and mobile interaction traces get one treatment, embodiment and instructional video get the other. Mind2Web and AITW are the two benchmarks that show up in agent papers, so the fact that the annotated versions are released is more useful than it first appears, because it lets you inspect the supervision rather than take the numbers on faith.

The README also links a generation script for SoM and ToM on instructional videos, describing it as Algorithm 2 in the paper, with a news entry dated 2025.03.16. That is the piece you would need to reproduce the annotations on video you have rather than only consume the published sets.

## The repository layout tells you what is runnable

The tree is a research repository with a clear separation between model code, training code and demos:

```bash
agents/
data_configs/
tools/
server/
train.py
trainer/
```

`magma/` holds the model package, and `pyproject.toml` at the root means it installs as a Python project. `train.py` and `trainer/` are the training entry point and its supporting code, `data_configs/` holds the dataset configurations the README's training sections refer to, and `data/` plus `scripts/` cover preprocessing. `server/` is the API server, which matters if you intend to run inference against something rather than import the package.

`agents/` is the most interesting directory for a reader, because it holds the three demos the news timeline names: the UI agent, the gaming agent and the robot visual planner. The robot one is started with a single command from the README:

```bash
python agents/robot_traj/app.py
```

That news entry is dated 2025.03.06 and describes it as a demo for robot planning capabilities. The UI and gaming demos from 2025.02.28 were published as Gradio spaces on Hugging Face, and the README still carries a commented-out block linking both spaces, which suggests the author left the section in place after moving the links elsewhere. That is the kind of small residue that tells you the front page has not been rewritten in a while.

## A release timeline compressed into six weeks

The news timeline is the most useful part of the README, because it turns a research release into a sequence you can reconstruct. The arXiv paper appeared on 2025.02.18. Inference code followed on 2025.02.23. CVPR 2025 acceptance was announced on 2025.02.26, the day before the model landed on Hugging Face and Azure AI Foundry as Magma-8B on 2025.02.25. The README also notes the project reached the front page of Hacker News on 2025.02.20.

Training code and an example for training Magma-8B on the Magma-820K dataset shipped on 2025.03.09, so the gap between weights and a training recipe was two weeks. Datasets followed in April, as described above.

Against that, two facts deserve equal weight. The GitHub releases list is empty, so there is no tagged version to pin and no release notes to diff. And the last push was on 2026-03-03, while the newest news entry is dated 2025.04.29. Whatever moved in the repository after April 2025 is not described on the front page, which makes the timeline the reliable part and everything after it guesswork.

That matters for anyone reading the checklist. Two items remain unchecked: SeeClick and Vision2UI pretraining data with SoM, and a video finetune script. A third, the UI and LIBERO finetuning script, is also unchecked in the list as published. Everything marked done is inference code, the two agent demos, the checkpoint, the training code and both trace datasets.

## Reading the outline as a map of the documentation

The README carries a full table of contents, and it is worth reading as a statement of intent rather than navigation. After the What is Magma section it moves to pretraining, installation, data preprocessing with SoM and ToM generation, and then a Model Training section split into pretraining on Open X without SoM or ToM and finetuning on Magma-820K.

The usage half is more granular still. Inference is documented four ways: with Hugging Face Transformers, with local code from the repository, with bitsandbytes for quantized loading, and a benchmarking entry. Separate sections cover evaluation with lmms-eval, evaluation with SimplerEnv, multi-image or video input, and the API server. The three agent demos then get their own subsections.

Two absences are visible from the front page itself. There is no installation snippet above the fold, even though the outline points to an Installation section further down, so the quick start lives in a part of the page you have to scroll to. And the project carries no topics at all, which for a repository of this profile is unusual and means discovery on GitHub rests on the star count rather than on tags.

For the model itself, the README links Magma-8B on Hugging Face under the microsoft organisation and a versioned listing on Azure AI Foundry. Both are worth checking before assuming the checkpoint is the same, since the Foundry entry carries an explicit version number and a registry path.

## Who this is for, and where a reader should start

Magma is a research release with an MIT license and a permissive model distribution, which makes it usable, but the intended reader is fairly specific. The people who will get the most out of it are those working on GUI agents, where Mind2Web and AITW are the standard benchmarks and the SoM-annotated versions of both are published, and on embodied work, where the Open X-Embodiment traces give you annotated robot demonstration data rather than raw episodes.

The spatial reasoning emphasis in the highlights is the differentiator to test first. If your problem involves locating an element on a screen or a target in a scene, that is the capability the paper claims to improve and the one the released annotation sets let you examine. If your problem is general image understanding, the choice among multimodal backbones is less clear-cut and the case for this one rests on the paper's benchmarks rather than on anything the repository demonstrates.

The practical path is short. Read the arXiv paper for the architecture and the pretraining recipe, clone the repository for the model code in `magma/` and the training loop in `trainer/`, and pull the annotated datasets before downloading any weights, since the annotation sets are what tell you whether the supervision style matches what you want to train. The model is 8B, so the quantization section on bitsandbytes is the practical entry point for inference on a single machine.

Two caveats to carry into that evaluation. The training data for SeeClick and Vision2UI with SoM is listed as not yet released, so part of the pretraining mixture described in the paper is not available to reproduce. And with no GitHub releases and a last push on 2026-03-03, pinning to a commit rather than a tag is the only way to be sure of what you are running.

## Conclusion

Magma is interesting less as a model you would deploy today than as a research artifact released unusually completely: inference code, training code, an 8B checkpoint on Hugging Face and Azure AI Foundry, and the annotated datasets the paper's claims rest on, all under MIT. The two unchecked boxes in the project checklist are the ones that would matter to a practitioner, a UI and LIBERO finetuning script and a video finetune script, so anyone planning to adapt the model to their own tasks is waiting on those. Start with the arXiv paper at 2502.13130 for the method, the `magma/` package for the model code, and the Mind2Web-SoM and AITW-SoM annotation sets if you want to see what the supervision actually looks like.

## FAQ

### What is the Magma model from Microsoft Research?

Magma is an 8B multimodal foundation model aimed at agents rather than single tasks, released with inference code, training code, a checkpoint on Hugging Face and Azure AI Foundry, and SoM and ToM annotated datasets. The paper appeared on arXiv as 2502.13130 and was accepted at CVPR 2025.

### Is Magma open source and can I use it commercially?

The repository is MIT licensed, and the model checkpoint is distributed through Hugging Face and Azure AI Foundry. The repository carries no GitHub releases, so there is no tagged version to pin, and the last push was on 2026-03-03.

### What datasets did Magma release for training?

Four datasets on Hugging Face under the MagmaAI organisation: Magma-OXE-ToM and Magma-Video-ToM carrying visual traces for pretraining, and Magma-Mind2Web-SoM and Magma-AITW-SoM carrying SoM prompting annotations used for the reported downstream finetuning. SeeClick and Vision2UI data with SoM is listed as not yet released.

### Can I finetune Magma for my own task?

Training code shipped on 2025-03-09 along with an example for training Magma-8B on the Magma-820K dataset, and `train.py` with the `trainer/` directory is the entry point. The UI and LIBERO finetuning script and the video finetune script are both still unchecked in the project list, so those two recipes are not there yet.

## Sources

- [Issues](https://github.com/microsoft/Magma/issues)
- [License: MIT](https://github.com/microsoft/Magma/blob/main/LICENSE)
- [microsoft/Magma on GitHub](https://github.com/microsoft/Magma)
- [README](https://github.com/microsoft/Magma/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-magma
