EasyMocap: markerless motion capture from RGB video, and where it stops being the right tool
Make human motion capture easier.
At a glance
- What is it?
- EasyMocap fits SMPL, SMPL+H, SMPL-X and MANO body models to ordinary RGB footage, from one internet video up to 23 calibrated cameras. The fit is the product, and the setup cost is real.
- Who is it for?
- Adopt EasyMocap if you already have RGB footage or a calibrated multi-camera rig and you need SMPL-family parameters rather than a skeleton overlay: the emc command and the config-driven pipeline match that workflow. Do not adopt it if you need a maintained release cadence, a documented rollback path, or a phone-only capture tool; the repository's only tagged release is v0.1 from 2021-03-29 and setup.py declares version 0.2.1, so the two disagree.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What EasyMocap actually produces, and who needs that output
EasyMocap is a Python toolbox for markerless human motion capture and novel view synthesis from RGB videos. The README frames it as a collection of demos in different settings rather than a single product, and that framing matters: the repository ships several pipelines, each aimed at a different capture situation.
The output is not a stick figure. The core feature is fitting SMPL, SMPL+H, SMPL-X or MANO to the footage, which means the result is a parametric body model with pose and shape parameters. That is the right artefact if you plan to retarget motion, drive an avatar, or feed a downstream renderer. It is the wrong artefact if all you want is a 2D skeleton drawn over a video, because you would be paying for a body-model fit you never use.
The intended audience is visible in the demo list. Multi-view single-person capture is illustrated with ZJU-MoCap footage from 23 calibrated and synchronized cameras, and a separate MANO demo was captured with 8 cameras. Internet video capture is a lower-setup path that fits SMPL using 2D keypoint estimation and CNN initialization. Multi-view multi-person capture is demonstrated with 8 consumer cameras. So the project spans a research rig and a single downloaded clip, and the amount of work between those two ends is large.
The pipeline: 2D keypoints in, SMPL-family parameters out
The mechanism differs by pipeline, but the internet-video path shows the general shape. 2D keypoints are estimated first, then a CNN initialization proposes a starting body, and then SMPL is fitted. That ordering is deliberate: the fit is a non-convex optimization, so the initialization decides whether it converges to something plausible. The README credits the 2D estimation to Cao et al. and HRNet, and the CNN initialization to Kolotouros et al.
The multi-view path replaces the single-view prior with geometry. With calibrated and synchronized cameras, the fit is constrained by agreement across views, which is why the README can show body, hand and face poses from the same capture. The repository layout reflects this split: easymocap/ holds the package modules, including config, dataset, smplmodel, pyfitting, estimator and annotator, while apps/ holds the runnable entry points and config/ holds the YAML configuration. The console script registered in setup.py points at apps.mocap.run:main_entrypoint, so the command-line surface is a thin wrapper over those modules.
Two supporting pieces sit outside the fit itself. Camera calibration has its own app and README under apps/calibration, and a 3D realtime visualization module is documented separately, with demos for a 25-joint skeleton, a full skeleton, multi-person skeletons, and SMPL, SMPL-X and MANO meshes. Novel view synthesis is not implemented here; the README points at NeuralBody and the SIGGRAPH 2022 work on human interaction from sparse multi-view videos.
Installing EasyMocap and running a first capture
The README does not carry an install section; it links to a Quick Start in the public documentation, and the repository supplies requirements.txt and setup.py. The requirements file pins several packages tightly, including mediapipe==0.10.0, pytorch-lightning==1.5.0, tensorboard==2.8.0 and setuptools==59.5.0, and it pulls chumpy directly from its GitHub repository. Those pins are the first thing to check against your environment, because a newer PyTorch Lightning or setuptools will not satisfy them.
A conventional install from a clone looks like this. The setup.py declares the package name easymocap and registers the emc console script, so after installation the emc command is the entry point.
pip install -r requirements.txt
pip install -e .Once installed, emc is the entry point registered by setup.py, and its implementation is the main_entrypoint function in apps/mocap/run.py. The README's quickstart badge links to the documentation for the arguments to pass rather than listing them here.
emcThe configuration is YAML under config/, and the runnable apps live under apps/, so a real run means choosing an app, pointing it at a config, and pointing that config at your data. The README directs readers to the Quick Start page for the internet-video path, which is the lowest-setup route: one video in, SMPL parameters out. For the multi-view path you need calibrated cameras first, and camera calibration has its own app and its own README under apps/calibration. Expect the calibration step to dominate the setup time, not the fit.
Where EasyMocap breaks down, and what the documentation does not cover
The initialization dependency is the sharpest limitation. On the internet-video path, the fit is seeded by a CNN and by 2D keypoints, so a clip with heavy occlusion, unusual clothing or a camera angle far from the training distribution gives the optimizer a bad starting point. The README does not describe a fallback when the initialization fails, and it does not document rollback or recovery for a bad fit. You either re-run with different settings or you discard the sequence.
Multi-view capture trades that risk for a different one. The 23-camera ZJU-MoCap example is a calibrated, synchronized rig, and the README does not claim the same quality from an uncalibrated set of phones. The 8-camera examples are the realistic middle ground, and even there the word calibrated is doing the work. If your cameras drift or your sync is approximate, the cross-view constraint degrades.
The release history is a separate concern. The repository's only tagged release is v0.1 from 2021-03-29, while setup.py declares version 0.2.1 and the README announces an EasyMocap v0.2 release. A user pinning by tag gets the 2021 code, not the version the README describes. The last push to the repository was on 2026-09-15, so the code has moved since that tag, but the tag itself has not.
Finally, several features are explicitly unfinished. Multiple internet videos with a specific action is marked Coming soon and links to doc/todo.md. Novel view synthesis for challenging motion is marked coming soon in the same figure caption. The README also notes that the ZJU-MoCap dataset requires signing an agreement and emailing it to the listed contacts before you get a download link, so the data is not a pip install away.
EasyMocap compared with a pose-estimation library
The natural alternative is a 2D or 3D pose-estimation library, and the difference is what comes out. A pose estimator gives you joint locations in image or camera coordinates. EasyMocap fits a body model, so you get pose and shape parameters for SMPL, SMPL+H, SMPL-X or MANO, which can be retargeted to a rig or rendered as a mesh. The README's realtime visualization demos show both outputs side by side, skeletons and meshes, which is a fair picture of the gap.
The cost of that extra output is the fit. A pose estimator runs a network and returns. EasyMocap runs an optimization that needs an initialization and, on the multi-view path, needs calibration. If your downstream task only consumes joint angles and you have no use for shape, the body-model fit is overhead. If you need a mesh, a shape parameter, or hand and face detail from the same capture, the estimation library cannot give you that and EasyMocap is the shorter path.
There is a middle option inside the project itself. The internet-video path uses 2D keypoint estimation and CNN initialization as components, so you can treat EasyMocap as a fitting layer on top of an estimator you already trust, provided the interfaces line up. The README does not document swapping in a different 2D estimator, so that would be your own integration work.
Licence, maintenance and the upgrade cost you should price in
The repository's licence is reported as NOASSERTION, which means the licence file could not be matched to a standard identifier. The practical consequence is that you should read LICENSE yourself rather than assume an OSI licence applies. This is not legal advice, and the licence text is the only authority.
The dependency pins are the real upgrade cost. requirements.txt fixes mediapipe at 0.10.0, pytorch-lightning at 1.5.0, tensorboard at 2.8.0 and setuptools at 59.5.0, and installs chumpy from a git URL. Moving any of those forward means testing the fitting code against a different API, and the README does not document a supported upgrade path or a compatibility matrix. If you vendor EasyMocap into a larger environment, those four pins constrain the rest of it.
On maintenance, the last push was on 2026-09-15, so the repository is not dormant, but the only tagged release remains v0.1 from 2021-03-29. Anyone who pins to a tag is pinning to 2021 code. If you need reproducibility, pin a commit rather than the tag, and record the commit hash alongside your config.
Editorial conclusion
Adopt EasyMocap if you already have RGB footage or a calibrated multi-camera rig and you need SMPL-family parameters rather than a skeleton overlay: the emc command and the config-driven pipeline match that workflow. Do not adopt it if you need a maintained release cadence, a documented rollback path, or a phone-only capture tool; the repository's only tagged release is v0.1 from 2021-03-29 and setup.py declares version 0.2.1, so the two disagree. Before committing, verify that the SMPL model files you are entitled to use can be downloaded, that your Python and PyTorch versions satisfy mediapipe 0.10.0 and pytorch-lightning 1.5.0, and that the NOASSERTION licence text in the repository covers your intended use.
Frequently asked questions
Is motion capture still used?
The README does not argue for motion capture as a field. It presents EasyMocap as an open-source toolbox for markerless human motion capture from RGB videos, with demos in several capture settings, and points at published work built on the ZJU-MoCap dataset such as HumanNeRF, KeypointNeRF and Texel-Aligned Features.
Is there any free motion capture software available?
EasyMocap itself is an open-source toolbox for markerless human motion capture from RGB videos, and the README describes free demos across several settings, including a Colab notebook linked from the multi-view single-person section. The licence is reported as NOASSERTION, so read the LICENSE file before assuming the terms. The ZJU-MoCap dataset is separate and requires signing an agreement and emailing it to the listed contacts.
Can you do motion capture with a phone?
The README does not describe a phone-based capture path. Its lowest-setup route is internet video, which fits SMPL using 2D keypoint estimation and CNN initialization, and its multi-view demos use calibrated and synchronized camera rigs, from 8 consumer cameras up to 23 cameras for the ZJU-MoCap footage.
Which AI tool is best for motion capture?
The README does not rank tools. It positions EasyMocap as a toolbox that fits SMPL, SMPL+H, SMPL-X or MANO to RGB video, and for novel view synthesis it points at NeuralBody and the SIGGRAPH 2022 work on human interactions from sparse multi-view videos rather than claiming that ground itself.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zju3dv-easymocap)