Open-source project
open-mmlab/mmaction2 avatar
open-mmlab/mmaction2

MMAction2: a PyTorch toolbox for five video understanding tasks

OpenMMLab's Next Generation Video Understanding Toolbox and Benchmark

5,160 stars1,359 forksPythonApache-2.0

At a glance

What is it?
MMAction2 packages action recognition, temporal localization, spatio-temporal detection, skeleton action detection and video retrieval into one config-driven framework. It is broad, and that breadth is also the main thing to weigh before adopting it.
Who is it for?
MMAction2 fits teams that need several video understanding tasks behind one config system and are willing to pin a PyTorch and MMEngine stack to get it. It is the wrong choice if you only need to classify short clips and want a single lightweight dependency, or if you need a release cadence faster than the last tagged version, v1.2.0 from 2023-10-12.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem MMAction2 solves: five video tasks behind one config system

Video understanding is not one problem. Recognizing an action in a trimmed clip, finding where an action starts and ends in a long recording, drawing a box around a person and labeling what they are doing, classifying motion from pose keypoints, and retrieving a video from a text query are five different pipelines with different inputs, losses and evaluation metrics. Most research codebases pick one. MMAction2's stated goal is to hold all five: the README lists action recognition, action localization, spatio-temporal action detection, skeleton-based action detection and video retrieval as the tasks implemented in the repository.

The intended user is not someone who wants a one-line classifier. It is a team that already works in PyTorch, has a dataset in one of these five shapes, and wants to swap backbones and heads without rewriting the training loop. The modular design claim in the README is the whole pitch: a video understanding framework is decomposed into components, and you combine modules to build a customized pipeline. If your problem is "I have 200 labeled clips and need a baseline by Friday", the framework's config layer is overhead you will feel before you feel any benefit.

How the config, registry and tooling fit together

The repository layout tells you most of the architecture. Top-level entries include configs/, mmaction/, tools/, demo/, projects/ and tests/. The Python package lives in mmaction/, and the configs/ tree holds the model and dataset definitions. This is the OpenMMLab pattern: a model is described by a config file rather than constructed in code, and training, testing and inference are all driven by tools/ scripts that read that config.

The registry is the mechanism behind the modularity. Components such as backbones, necks and heads are registered under names, and a config refers to them by name, so replacing a backbone is a config edit rather than a code edit. The README also points to a model zoo and an API reference in the documentation, which is where the individual config names and checkpoints are listed.

Two structural details are worth noting because they shape day-to-day work. First, the default branch is main, and the README states it was switched from master (the 0.x line); the migration guide is the document that covers the differences. Second, the repository carries a projects/ directory alongside configs/, which is where work that is not yet part of the main config tree tends to live. If a model you read about is not in configs/, it may still be in projects/.

Installing MMAction2 and running a first demo

The README links to an installation page in the documentation rather than inlining the steps, and the repository's requirements.txt is a thin aggregator that pulls in three other files: requirements/build.txt, requirements/optional.txt and requirements/tests.txt. That means the dependency list is split by purpose, and the installation page is the place where the version constraints for the underlying OpenMMLab packages are stated. Read it before you pip install anything, because the versions of the framework's dependencies are the usual source of import errors.

There is a PyPI package, mmaction2, referenced by the README badge. A minimal install from PyPI looks like this:

Where MMAction2 gets in the way

The dependency chain is the first real cost. MMAction2 sits on top of shared OpenMMLab libraries, and the installation page is the authority on which versions pair with which PyTorch and CUDA builds. When those drift, the failure is an import error at startup, not a warning, and the fix is usually to reinstall the whole stack rather than the one package that complained.

The second limitation is release cadence. The most recent tagged release listed for this repository is v1.2.0 from 2023-10-12, and the last push to the repository was on 2026-03-18. Development activity between those points does not necessarily appear in a tagged version, but anyone who needs a versioned, reproducible artifact is working from a release that is well over a year old. If your compliance process requires a tagged release for every dependency, that is a real constraint.

The third is scope mismatch. The framework is built for the five listed tasks and their standard datasets. A narrow problem, such as classifying a fixed set of gestures from a webcam stream at low latency, will spend most of its complexity budget on machinery you do not use. The webcam demo scripts in demo/ show the streaming path exists, but the config system, the registry and the dataset abstractions remain in the way.

How it compares with a single-task video model library

The closest alternative in spirit is a general video model library that exposes pretrained models as callable objects and leaves training loops to you. The difference is where the abstraction sits. In that style of library you get a model, you feed it a tensor, you get logits, and any dataset, sampler or metric work is yours to write. In MMAction2 the dataset, the pipeline of transforms, the sampler, the head and the metric are all declared in the same config file, and tools/ scripts execute that declaration end to end.

That trade favors reproducibility over flexibility. Reproducing a paper's training setup is mostly a matter of using the config that ships with the repository. Writing a training loop for a task the framework does not model is harder, because you are working against the registry rather than with plain PyTorch. There is also a middle option visible in the repository itself: the demo scripts and the inferencer interface in demo/demo_inferencer.py expose inference without requiring you to adopt the training side at all. If you only need predictions, that is the shallow end of the pool and it is worth using before committing to the full stack.

Licence, maintenance and what an upgrade costs

MMAction2 is released under Apache-2.0, which is a permissive licence that allows commercial use and modification, with the usual requirements around preserving notices. The repository also ships a CITATION.cff and a citation section in the README, so published work built on it is expected to cite the project. This is a description of the licence terms, not legal advice; check the LICENSE file and your own obligations.

The upgrade cost is concentrated in the 0.x to 1.x transition. The README states plainly that the default branch moved from master (0.x) to main (1.x) and that users are encouraged to migrate, with a migration guide in the documentation. Configs written for 0.x do not carry over unchanged. Anyone adopting the project now starts on the 1.x line, so that particular migration is behind them, but the same pattern applies to future major versions: configs are the interface, and the interface changes.

On maintenance, the honest statement is that the last push was on 2026-03-18 and the newest tagged release is v1.2.0 from 2023-10-12. The repository is not archived. Whether that gap between commits and tags matters depends on whether you consume the main branch or a released version.

Editorial conclusion

MMAction2 fits teams that need several video understanding tasks behind one config system and are willing to pin a PyTorch and MMEngine stack to get it. It is the wrong choice if you only need to classify short clips and want a single lightweight dependency, or if you need a release cadence faster than the last tagged version, v1.2.0 from 2023-10-12. Before committing, verify that the model you want has a config under configs/ and a checkpoint in the model zoo, and check the installation page for the exact mmcv and MMEngine versions your CUDA build requires.

Frequently asked questions

How do I install MMAction2?

The README links to an installation page in the documentation rather than listing steps inline, and the repository's requirements.txt aggregates requirements/build.txt, requirements/optional.txt and requirements/tests.txt. There is also a PyPI package named mmaction2, and the OpenMMLab mim installer is the recommended path because it resolves matching versions of the shared libraries.

What is MMAction2 used for?

It is a PyTorch toolbox for video understanding that implements five task families: action recognition, action localization, spatio-temporal action detection, skeleton-based action detection and video retrieval. The README describes it as a modular framework where components are combined to build a customized video understanding pipeline.

Does MMAction2 support skeleton-based action recognition?

Yes. Skeleton-based action detection is one of the five tasks listed in the README, and the repository ships a demo_skeleton.py script with a demo_skeleton.mp4 sample video. The README's overview images also show skeleton-based results on NTU-RGB+D-120.

Which version of MMAction2 should I start with?

Start on the 1.x line, which is the main branch. The README states that the default branch was switched from master (0.x) to main (1.x) and points to a migration guide for the differences, so configs written for 0.x do not carry over unchanged.

Official sources

  1. License: Apache-2.0
  2. open-mmlab/mmaction2 on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/open-mmlab-mmaction2.svg)](https://hysenlabs.com/projects/open-mmlab-mmaction2)