# microsoft/unilm: Microsoft's Monorepo for Foundation Model Research and Pre-trained Models

> microsoft/unilm is a large research repository from Microsoft that hosts pre-trained models and training code across language, vision, speech, and multimodal domains, organized as a collection of separate sub-projects each with its own subdirectory, paper, and weights.

**microsoft/unilm** — Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities

- Repository: https://github.com/microsoft/unilm
- Website: https://aka.ms/GeneralAI
- Stars: 22,227 · Forks: 2,705
- Language: Python
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-unilm

## What microsoft/unilm Is and Who It Is For

microsoft/unilm is a research monorepo maintained by the Foundation Models team at Microsoft Research. It is not a single model or library; it is a collection of individual research projects, each living in its own subdirectory with its own README, training code, and links to pre-trained model weights hosted on Hugging Face or other model repositories.

The intended users are machine learning researchers and engineers who want to use or reproduce specific Microsoft Research models. Each sub-project targets a different problem domain: document AI, text understanding, multilingual NLP, vision pre-training, speech representation learning, and multimodal understanding. A practitioner building a document extraction pipeline would typically clone the repository and navigate to `layoutlmv3/` or `layoutlm/`, not use the monorepo as a whole.

The repository homepage at aka.ms/GeneralAI links to the team's broader research agenda. The team also maintains a separate TorchScale library (github.com/microsoft/torchscale) for its foundation architecture work.

## Repository Organization and Major Sub-projects

The top-level directory contains over 50 subdirectories, each representing a distinct research project or model family. The README groups them into categories:

Foundation architectures and the TorchScale library cover DeepNet (scaling Transformers to 1,000 layers), Foundation Transformers (Magneto), Length-Extrapolatable Transformer, and X-MoE (sparse Mixture-of-Experts). These represent the structural research that underpins many of the downstream models.

Language and multilingual models include the original UniLM (unified pre-training for understanding and generation), InfoXLM and XLM-E (100+ language models), DeltaLM and mT6 (encoder-decoder for multilingual generation), MiniLM (small and fast models), E5 (text embeddings), and MiniLLM (knowledge distillation for large language models).

Vision models include BEiT and BEiT-2 (generative self-supervised pre-training for images), DiT (document image transformers), and TextDiffuser and TextDiffuser-2 (diffusion models for text rendering in images).

Speech models include WavLM (full-stack speech tasks) and VALL-E (neural codec language model for text-to-speech).

Multimodal models include the LayoutLM family (document understanding combining text, layout, and image), Kosmos-1, Kosmos-2, and Kosmos-2.5 (multimodal LLMs), VLMo, VL-BEiT, and BEiT-3 (vision-language pre-training).

Recent additions visible in the top-level directory include Diff-Transformer, LatentLM, PFPO, ReSA, and YOCO.

## LayoutLM: The Repository's Most Widely Used Sub-project

LayoutLM and its successors (LayoutLMv2, LayoutLMv3, LayoutXLM) are among the most practically adopted models in the repository. They address document AI: understanding scanned documents, PDFs, and form images by combining text, bounding-box layout information, and visual features. LayoutLM is Microsoft's document foundation model, and the models are used for tasks like form key-value extraction, receipt understanding, and document classification.

LayoutLMv3, in the `layoutlmv3/` subdirectory, is a unified architecture that processes text, layout, and image through a single Transformer, removing the need for a separate image encoder used in v2. LayoutXLM extends the family to multilingual document understanding.

The models are hosted on the Hugging Face Hub under the `microsoft/` organization and are accessible through the `transformers` library, making them straightforward to use without cloning the repository. The repository code provides the original training scripts and reproduction instructions.

## BEiT, WavLM, and E5: Notable Other Sub-projects

BEiT (BERT Pre-Training of Image Transformers) applies masked image modeling to vision pre-training, analogous to masked language modeling in BERT. BEiT-2 extends this with a vector-quantized teacher. BEiT-3 is described as a general-purpose multimodal foundation model and the culmination of the Big Convergence agenda of pre-training across tasks, languages, and modalities.

WavLM is the speech pre-training model in the repository, targeting full-stack speech tasks: automatic speech recognition, speaker verification, speech separation, and others. It is hosted on Hugging Face and accessible through `transformers`.

E5 is a family of text embedding models trained for similarity matching tasks. The `e5/` subdirectory provides the training code and links to models on Hugging Face. E5 is relevant to engineers building retrieval-augmented generation pipelines or semantic search.

TrOCR, in `trocr/`, is the transformer-based OCR model that combines a visual encoder with a language model decoder for text recognition in images. It is available on Hugging Face as several variants tuned for different document types.

## Using Models from the Repository

Most models in the repository are published to the Hugging Face Hub under the `microsoft/` namespace. Practitioners using the `transformers` library can load them directly without cloning the repository. For example, LayoutLMv3 and E5 variants are loadable via `AutoModel.from_pretrained()` with the appropriate model name from the Hub.

The repository itself is needed when reproducing training runs, running the fine-tuning scripts, or working with models not yet pushed to the Hub. Each subdirectory has its own `requirements.txt` or setup instructions, and dependency versions vary across sub-projects. The `s2s-ft` subdirectory has a published PyPI package (`s2s-ft`) with its own version history (the v0.3 release from 2020 is listed in GitHub releases, the most recent for that specific sub-package).

The repository does not have a unified install command or a top-level `requirements.txt`. Researchers navigate to the sub-project directory of interest and follow its own documentation.

## Limitations and Research vs. Production Considerations

The repository is primarily a research artifact, not a production library. Each sub-project was created to accompany a paper and demonstrate reproducibility. The code quality, documentation depth, and maintenance frequency vary substantially between subdirectories. Older entries like UniLM-v1 (`unilm-v1/`) received their last updates years before the more recent additions.

Some subdirectories contain models that are research previews rather than production-ready implementations. Diff-Transformer, YOCO, and LatentLM are listed in the top-level directory but without detailed descriptions in the README, suggesting they are newer entries whose documentation may be in the subdirectory README rather than the top-level summary.

The repository does not provide a compatibility matrix indicating which Python, PyTorch, or CUDA versions each sub-project requires. Engineers reproducing results from an older paper may encounter dependency conflicts that require environment isolation (conda or virtual environments) per sub-project.

The `gitmodules` file and the `.github/` directory are present but the monorepo does not use submodules for the individual model directories; each directory contains its own code directly.

## Maintenance, Hiring, and License

The last push was on 2026-09-21, confirming the repository continues to receive updates. New sub-projects are added as research results are published.

The repository README includes an explicit hiring notice: the Foundation Models team at Microsoft Research is hiring FTE researchers and interns in the areas of Foundation Models, NLP, MT, Speech, Document AI, and Multimodal AI, with a contact email provided.

The repository is licensed under the MIT license. Individual sub-projects may have additional license terms for their model weights, particularly where those weights incorporate third-party data or components. The s2s-ft package releases (listed in GitHub releases) are MIT licensed. Engineers deploying models in products should check the Hugging Face model card for each model's specific license terms alongside the repository MIT license.

## Conclusion

microsoft/unilm is worth bookmarking for ML engineers who track foundation model research across text, vision, speech, and document understanding. The repository is the authoritative source for reproducing experiments from the associated papers and for obtaining pre-trained weights not yet available elsewhere. Teams building products should evaluate each sub-project individually: LayoutLM and LayoutLMv3 are production-grade document AI models with extensive Hugging Face support, while newer entries like Diff-Transformer and LatentLM are research previews. The last push was on 2026-09-21, confirming the repository remains updated.

## FAQ

### What is the microsoft/unilm repository?

microsoft/unilm is a research monorepo from Microsoft Research that hosts pre-trained models and training code across language, vision, speech, and document AI domains. Each sub-project lives in its own subdirectory with its own paper, weights, and documentation.

### What are pre-trained language models in the context of unilm?

In this repository, pre-trained language models include the original UniLM (for understanding and generation), MiniLM (compact models for inference efficiency), E5 (text embeddings), and InfoXLM and XLM-E (multilingual models covering 100+ languages). Each has its own subdirectory with training code and links to weights on Hugging Face.

### What NLP models are in microsoft/unilm?

The repository includes UniLM, MiniLM, E5 (text embeddings), InfoXLM, XLM-E, DeltaLM, AdaLM, SimLM, and MiniLLM among its language models, along with multimodal models like LayoutLM (document AI) and Kosmos-1, Kosmos-2, and Kosmos-2.5 (multimodal LLMs). Practically adopted models like LayoutLMv3 and E5 are available through Hugging Face.

## Sources

- [License: MIT](https://github.com/microsoft/unilm/blob/master/LICENSE)
- [microsoft/unilm on GitHub](https://github.com/microsoft/unilm)
- [Project website](https://aka.ms/GeneralAI)
- [README](https://github.com/microsoft/unilm/blob/master/README.md)
- [Releases](https://github.com/microsoft/unilm/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-unilm
