# MMF: FAIR's Modular Framework for Vision-Language Multimodal Research

> MMF is a Python framework from Facebook AI Research that provides modular, composable building blocks for training and evaluating vision-language models. It powered multiple FAIR research projects and served as the starter codebase for several academic challenges including Hateful Memes, TextVQA, and VQA.

**facebookresearch/mmf** — A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)

- Repository: https://github.com/facebookresearch/mmf
- Website: https://mmf.sh/
- Stars: 5,633 · Forks: 938
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/facebookresearch-mmf

## What MMF provides for vision-language research teams

Vision-language research involves coordinating image encoders, text encoders, fusion modules, task-specific heads, and dataset loaders in a consistent way across experiments. Wiring these components together from scratch for each new paper is time-consuming and makes results harder to reproduce. MMF was built to solve that coordination problem. According to the README, MMF contains reference implementations of state-of-the-art vision and language models and has powered multiple research projects at Facebook AI Research. The README describes it as un-opinionated, scalable, and fast, and states that it supports distributed training via PyTorch.

The description of the repository as a starter codebase for challenges around vision and language datasets is notable. The Hateful Memes, TextVQA, TextCaps, and VQA challenges are all named explicitly. Using MMF as a competition baseline means getting a working training loop, dataset integration, and evaluation pipeline without building them from scratch. That focus on challenge participation shapes the design: MMF prioritises reproducible baselines over flexibility for arbitrary new architectures.

## The modular design and the Pythia lineage

The README states that MMF was formerly known as Pythia. Pythia was an earlier FAIR project for VQA research; MMF represents a generalized and modularized evolution of that codebase to cover a broader set of vision-language tasks. The framework is described as built on PyTorch and uses omegaconf for configuration management, as visible in both the pyproject.toml and the requirements.txt.

The repository layout reflects the modular architecture: the mmf/ directory contains the core library and the mmf_cli/ directory contains the command-line interface. A separate projects/ directory holds implementations of specific papers built using the framework. The docs/ and website/ directories contain documentation and the project website source respectively. Tests live in tests/. This separation means the core library, the CLI, the per-project implementations, and the documentation are maintained independently, which is consistent with a research codebase that accumulates new models over time without refactoring the shared infrastructure.

## Installation: where the documentation lives and what to expect

The README directs users to the documentation at mmf.sh/docs/ for installation instructions, without reproducing the steps in the README itself. The README says 'Follow installation instructions in the documentation.' The project is distributed as a Python package; setup.py is present at the repository root. However, the most recent formal GitHub release is v0.3.1, dated 2019-08-26. That gap between the 2019 release and the continued repository activity through 2026 means that the PyPI package version and the current repository state may differ significantly, and teams installing from PyPI rather than from the repository directly may not get the current codebase.

The requirements.txt at the root specifies the full dependency set with exact or bounded version pins. A representative excerpt shows the core stack:

```python
torch==1.11.0
torchaudio==0.11.0
torchvision==0.12.0
transformers>=3.4.0, <=4.10.1
pytorch-lightning==1.6.0
```

The file also includes fasttext, nltk, editdistance, lmdb, iopath, datasets, and pycocotools alongside the deep learning stack. The diversity of dependencies reflects the range of tasks MMF covers: multimodal datasets require evaluation scripts from pycocotools for captioning metrics and nltk for text processing.

## The torch 1.11.0 pin and why it matters for adoption in 2026

The requirements.txt specifies torch==1.11.0, torchaudio==0.11.0, and torchvision==0.12.0. PyTorch 1.11 was released in 2022. By 2026, PyTorch has advanced through the 2.x series. These pinned versions create a compatibility challenge on several fronts: GPU drivers and CUDA versions that work with PyTorch 1.11 may not be the versions shipping with current hardware, and operating systems that support PyTorch 1.11 prebuilt wheels may have moved past the versions that PyPI still hosts.

The transformers version range in requirements.txt is >=3.4.0,<=4.10.1. Hugging Face Transformers 4.10 dates to 2021. Many model architectures introduced since then (including those relevant to vision-language work) are not available in that range. pytorch-lightning==1.6.0 is also pinned to a 2022 release. These pins are likely the last tested combination, not a deliberate restriction; but engineers who want to use current model weights or current training infrastructure will need to resolve those conflicts themselves, without guidance from the project.

## The research challenge datasets MMF was built for

The README identifies four specific challenges where MMF was used as a starter codebase: the Hateful Memes Challenge, TextVQA, TextCaps, and the VQA Challenge. Each of these tasks combines visual and textual reasoning. TextVQA requires reading text visible in an image to answer questions. TextCaps generates captions for images that contain text. VQA (Visual Question Answering) answers natural-language questions about general images. The Hateful Memes Challenge involves classifying whether a meme combining an image and text is hateful.

The projects/ directory in the repository layout is where the specific research papers are implemented. Using MMF for one of these four tasks means starting from a codebase that already includes the dataset loading, evaluation metrics, and baseline model architectures relevant to that challenge, rather than assembling them from disparate sources. Outside those specific challenges, the documentation at mmf.sh/docs/ provides the list of supported datasets and models.

## How MMF compares to PyTorch Lightning for training vision-language models

PyTorch Lightning is the most common comparison for MMF among research engineers. Both abstract the training loop and distributed execution on top of PyTorch. The key difference is scope. PyTorch Lightning is general-purpose: it abstracts the training loop for any PyTorch model, provides callbacks, checkpointing, and multi-GPU training, but does not include vision-language dataset pipelines, evaluation metrics for VQA or captioning, or reference implementations of specific multimodal models.

MMF goes further in the vision-language direction: it provides the dataset integration, task-specific evaluation code, and research model implementations as part of the framework. The trade-off is that MMF's scope is narrower and its dependency set is more specific to a particular era of research. An engineer training a custom vision-language model from scratch in 2026 who does not need the specific baselines MMF provides would find PyTorch Lightning a more current and dependency-compatible starting point. An engineer reproducing a FAIR paper or entering a competition that used MMF as a reference implementation is in the intended use case.

## Conclusion

MMF is best suited to researchers who need a reference implementation of a FAIR-published model or who are participating in one of the competitions MMF was designed for. For teams starting new vision-language work in 2026, the dependency version pins deserve careful scrutiny before adoption: requirements.txt specifies torch==1.11.0, transformers between 3.4.0 and 4.10.1, and pytorch-lightning==1.6.0, all of which are considerably older than current releases. The last push was on 2026-07-07 and the repository is not archived, but there have been no formal GitHub releases since v0.3.1 in 2019. The README states the licence is BSD; confirm the exact terms in the LICENSE file before redistribution.

## FAQ

### What is MMF and what is it used for?

MMF (Modular Framework) is a Python framework from Facebook AI Research for building and training vision-language multimodal models. It provides reference implementations of FAIR research models and was the starter codebase for the Hateful Memes, TextVQA, TextCaps, and VQA challenges.

### What version of PyTorch does MMF require?

The requirements.txt file in the repository pins torch==1.11.0, torchaudio==0.11.0, and torchvision==0.12.0. These are 2022-era PyTorch releases and may require specific CUDA and hardware configurations that differ from current setups.

### Is MMF still maintained?

The repository is not archived and the last push was on 2026-07-07. However, no formal GitHub releases have been made since v0.3.1 in August 2019. Activity continues in the repository, but the gap between commits and formal releases is significant.

## Sources

- [facebookresearch/mmf on GitHub](https://github.com/facebookresearch/mmf)
- [Issues](https://github.com/facebookresearch/mmf/issues)
- [Project website](https://mmf.sh/)
- [README](https://github.com/facebookresearch/mmf/blob/main/README.md)
- [Releases](https://github.com/facebookresearch/mmf/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/facebookresearch-mmf
