Open-source project
microsoft/XPretrain avatar
microsoft/XPretrain

XPretrain: Microsoft Research's Collection of Multimodal Pre-Training Models

Multi-modality pre-training

512 stars35 forksPythonNOASSERTION

At a glance

What is it?
XPretrain is a repository from Microsoft Research's Multimedia Search and Mining group that collects training code and datasets for several multimodal pre-training models, covering both video-language and image-language tasks. Each model is in its own subdirectory with its own setup, and the repository has no unified install or single entry point.
Who is it for?
Researchers reproducing video-language or image-language pre-training results from Microsoft Research's MSM group can find training code for HD-VILA, LF-VILA, CLIP-ViP, and VisualParsing in this repository. Each model lives in its own subdirectory with its own environment requirements, so cloning the repository is the beginning of the setup process, not the end.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 33 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What XPretrain Collects and Who It Serves

XPretrain is not a library you install and import. It is a collection point for research code from Microsoft Research's MSM group, where each project has its own subdirectory, its own dependencies, and its own training and evaluation instructions. The repository groups the work into two tracks: video and language, and image and language.

The primary audience is machine learning researchers who want to reproduce published results, start a fine-tuning run from one of these checkpoints, or study the architectural decisions in the training pipelines. There is no single model you download and run; you clone the repository, navigate to the subdirectory for the model you care about, and follow that subdirectory's documentation.

The repository was last pushed on 2026-08-28. The most recent model entry is CLIP-ViP, which was accepted at ICLR 2023 and whose code was released in March 2023. The top-level repository entries visible in this review are CLIP-ViP/, LF-VILA/, hd-vila-100m/, hd-vila/, and visualparsing/, along with standard Microsoft open-source support files.

Video-Language Models: HD-VILA, LF-VILA, and CLIP-ViP

Three video-language models are included. HD-VILA, accepted at CVPR 2022, is a high-resolution and diversified video-language pre-training model. Its accompanying dataset, HD-VILA-100M, is a high-resolution diversified video-language dataset released publicly in March 2022; the dataset lives in the hd-vila-100m/ subdirectory.

LF-VILA, accepted at NeurIPS 2022, focuses on long-form video-language pre-training. Long-form video presents a different challenge than clip-level pre-training because the text description must align with content that spans many minutes rather than a few seconds.

CLIP-ViP, accepted at ICLR 2023, adapts an image-language pre-training model to the video-language domain. The README notes that the code for both CLIP-ViP and LF-VILA was released in March 2023. CLIP-ViP is the most recent of the three and represents the group's current direction on video-language alignment.

All three models are in their own subdirectories. A researcher wanting to run any one of them must read that subdirectory's README for environment requirements, checkpoint download instructions, and evaluation scripts.

Image-Language Models: Pixel-BERT, SOHO, and VisualParsing

Three image-language models are listed. Pixel-BERT is described as an end-to-end image and language pre-training model, with a link to an arXiv paper (arxiv.org/pdf/2004.00849.pdf) rather than a subdirectory in this repository. It may have been included before the code was hosted here, or the code was never added to XPretrain.

SOHO, accepted as an oral at CVPR 2021, is described as an improved end-to-end image and language pre-training model using quantized visual tokens. The README links to a separate repository (github.com/researchmm/soho) rather than a subdirectory in XPretrain, so its code is not in this repository.

VisualParsing, accepted at NeurIPS 2021, is a Transformer-based end-to-end image and language pre-training model. Unlike Pixel-BERT and SOHO, VisualParsing has a subdirectory here (visualparsing/) and is one of the five main subdirectories visible in the repository.

The result is that XPretrain as a single repository only contains code for four of the six listed models: HD-VILA, LF-VILA, CLIP-ViP, and VisualParsing. For Pixel-BERT and SOHO, the README points outward.

How to Navigate the Repository

There is no unified install command, no shared requirements file, and no common entry point for the collection. Each subdirectory has its own setup instructions. The workflow is:

1. Clone the repository. 2. Navigate to the subdirectory for the model you need (for example, CLIP-ViP/ or hd-vila/). 3. Follow that subdirectory's README for Python environment setup, checkpoint downloads, and training or evaluation commands.

The main README serves as a table of contents. It lists each model with a brief description, a link to the subdirectory or external repository, and the venue and year of the associated paper. The News section lists release dates for code and datasets, which is useful for understanding what is available versus what might still be pending release.

Practically, if you only need one model, there is no reason to clone the full repository with all subdirectories. A sparse checkout of just the CLIP-ViP/ subdirectory would reduce the download size and avoid pulling in assets from other models.

License and What NOASSERTION Means

GitHub shows the license as NOASSERTION, which means the repository contains a LICENSE file but GitHub's automated scanner could not match it to a recognized SPDX identifier. This is different from having no license at all: a LICENSE file is present, but its exact terms are not automatically surfaced.

For researchers reproducing published experiments in an academic context, this is unlikely to cause problems. For teams considering commercial use, code redistribution, or incorporation into a product, the terms of the LICENSE file should be reviewed directly.

Microsoft's standard open-source repositories typically use MIT or Apache-2.0, and many other Microsoft Research repositories use those licenses. But XPretrain's NOASSERTION status means that assumption cannot be made here without reading the file. The SUPPORT.md and CODE_OF_CONDUCT.md follow standard Microsoft open-source policies.

How XPretrain Differs from Single-Model Pre-Training Repositories

OpenAI's CLIP repository (openai/CLIP) is a single model with a single install path, a consistent API, and pretrained weights that can be loaded in three lines of Python. A researcher wanting to extract image features from a frozen backbone can have CLIP running in minutes.

XPretrain is organized differently. It is a research archive rather than a deployable library. Each model entry reflects the state of the code at the time the paper was published. There are no shared abstractions across the models, no common feature extraction API, and no plan visible in the repository for updating earlier models with newer training techniques.

The practical consequence is that XPretrain is appropriate for reproducing specific published results, studying particular architectural choices, or fine-tuning from a specific checkpoint. It is less appropriate for building a production pipeline that needs to swap between models with a consistent interface. For that use case, Hugging Face's model hub, which hosts many of these model families under a uniform transformers API, provides a more maintainable foundation.

Maintenance and Contribution

The most recent push was on 2026-08-28. The contributions page points to Microsoft's Contributor License Agreement process at cla.opensource.microsoft.com. Contributors must sign a CLA once to contribute across all repositories using Microsoft's CLA infrastructure.

The repository has no GitHub releases. There is no release history that would allow a downstream project to pin to a specific version by tag. The repository also has no issues or pull request template visible in the top-level entries, though a SECURITY.md file exists for reporting security concerns.

Contact for issues with the pre-trained models goes to the GitHub issue tracker. For other communications, the README lists Bei Liu ([email protected]) and Jianlong Fu ([email protected]) from Microsoft Research as the points of contact.

Editorial conclusion

Researchers reproducing video-language or image-language pre-training results from Microsoft Research's MSM group can find training code for HD-VILA, LF-VILA, CLIP-ViP, and VisualParsing in this repository. Each model lives in its own subdirectory with its own environment requirements, so cloning the repository is the beginning of the setup process, not the end. The license is listed as NOASSERTION on GitHub; read the LICENSE file before using this code in a commercial context. Teams that need a consistent inference API across models are better served by loading the model checkpoints through the Hugging Face transformers library rather than working directly from these training codebases.

Frequently asked questions

What does pre-training mean in the context of XPretrain?

In XPretrain's context, pre-training refers to training a model on a large, general-purpose dataset of paired data (video and text, or image and text) before any task-specific fine-tuning. The goal is to learn representations that transfer well to downstream tasks such as video retrieval or visual question answering.

Does XPretrain contain a unified install command for all models?

No. There is no shared requirements file or central entry point. Each model lives in its own subdirectory with its own setup instructions, and researchers must follow the individual subdirectory README to set up the environment for a specific model.

What is the HD-VILA-100M dataset in XPretrain?

HD-VILA-100M is a high-resolution, diversified video-language dataset from Microsoft Research, released publicly in March 2022 and living in the hd-vila-100m/ subdirectory. It was built to accompany the HD-VILA pre-training model accepted at CVPR 2022.

Official sources

  1. Issues
  2. microsoft/XPretrain on GitHub
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-xpretrain.svg)](https://hysenlabs.com/projects/microsoft-xpretrain)