# OpenPose keeps body keypoints cheap and makes you pay for faces

> OpenPose is a C++ library from the CMU Perceptual Computing Lab that estimates 135 body, hand, face and foot keypoints on a single image, and ships a command-line demo plus C++ and Python APIs. Two things decide whether it fits: the constant runtime claim covers the body model only, and the last tagged release is v1.7.0 from November 2020.

**CMU-Perceptual-Computing-Lab/openpose** — OpenPose: Real-time multi-person keypoint detection library for body, face, hands, and foot estimation

- Repository: https://github.com/CMU-Perceptual-Computing-Lab/openpose
- Website: https://cmu-perceptual-computing-lab.github.io/openpose
- Stars: 34,467 · Forks: 8,041
- Language: C++
- License: NOASSERTION
- Published: 2026-08-17 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/cmu-perceptual-computing-lab-openpose

## Body runtime is fixed, but the face and hand passes bill you per person

The runtime promise and the runtime caveat sit in the same feature list, and the caveat is the one that matters in a crowd. Body and foot estimation comes in 15, 18 or 25 keypoint variants, including 6 foot keypoints, and its runtime is invariant to the number of detected people. That is the property the project builds its speed argument on. Hand estimation is 2x21 keypoints, and its runtime depends on the number of detected people. Face estimation is 70 keypoints, and its runtime also depends on the number of detected people. So the shape of your bill is set by which passes you switch on and how full the frame is: a busy street scene costs roughly the same for bodies whether one person or ten are in it, and then every face and every pair of hands you enable adds to it. The published inference time comparison against Alpha-Pose and Mask R-CNN, on the same hardware and under the same conditions, is about that body-side invariance, since the two alternatives grow linearly with the number of people. The named alternative for the per-person cost is OpenPose Training, a separate repository.

## The 3D module triangulates one person at a time

The whole-body 135 keypoint claim is a 2D claim, and the 3D side of this library is explicitly narrower. What is offered is 3D real-time single-person keypoint detection, built on 3D triangulation from multiple single views. A rig of cameras observing one person reconstructs that person in three dimensions. A rig observing two people reconstructs one of them. That is the single-word limit in the feature list doing all the work, and it is easy to miss because the rest of the library is built around multi-person detection. Two supporting pieces come with it. Synchronization of Flir cameras is handled for you, and the supported hardware is Flir and Point Grey cameras. Alongside that sits a calibration toolbox that estimates distortion, intrinsic and extrinsic camera parameters, which is a separate step you run before the reconstruction means anything. If your plan is an interactive multi-camera volume with several moving people in it, the 3D module is not the piece that delivers it, and the fix is code of your own rather than a flag in the demo.

## Three ways in, and only one of them lets you change the pipeline

There are three entry points and they are not equivalent. The first needs no installation and no code: download the latest Windows portable version of OpenPose. The second is compiling and running from source. The third is the API surface, with a C++ API and a Python API for custom functionality such as your own inputs, pre-processing, post-processing and output steps. The quick start shows the first two shapes of invocation, one per platform:

```
# Ubuntu
./build/examples/openpose/openpose.bin
```

```
:: Windows - Portable Demo
bin\OpenPoseDemo.exe --video examples\media\video.avi
```

Flags can be added in any order, so the demo is a thin shell over the same estimator the API exposes. The cost of the portable route is that it is a binary and a fixed feature set; the cost of the source route is the build. The cost of the API route is that you now own the loop. A related Unity plugin lives in its own repository rather than in this tree, so engine integration is a separate dependency decision. Pick the entry point by how much of the pipeline you intend to change, not by which demo runs first.

## 3rdparty/ and .gitmodules mean the third-party code is not in your clone

The top level of this repository carries a .gitmodules file next to a 3rdparty/ directory, which means the external code is pulled in as submodules rather than committed into the tree. That is a small fact with an outsized effect on a first build, because a clone that does not bring the submodules along leaves you with a source tree that has no dependencies in it and a CMakeLists.txt that has nothing to configure against. The rest of the layout is conventional: cmake/, src/, include/, python/, scripts/, models/, doc/ and examples/. The examples tree is where the shape of the project shows. tutorial_api_cpp/ and tutorial_api_python/ are the two API walkthroughs, user_code/ is where a custom input or output is meant to live, calibration/ covers the calibration toolbox, tests/ and media/ hold fixtures, and deprecated/ is still sitting in the tree. That last directory is the one to watch: an example you find by browsing may be one the project has already retired, so check the path before you copy from it. Continuous integration runs through GitHub Actions for Linux and macOS and through AppVeyor for Windows.

## Ubuntu 14 and the Nvidia TX2 are still named as targets

The supported operating system list reads Ubuntu 20, 18, 16 and 14, Windows 10 and 8, Mac OSX, and Nvidia TX2. Two of those entries are useful signals. Ubuntu 14 and Windows 8 place the snapshot in an era, which tells you the build scripts, the CUDA expectations and the compiler flags inside this tree were written against a specific set of toolchains rather than against whatever your distribution ships now. The Nvidia TX2 sits in the same list as the desktop platforms, which is a hint at the intended range: a Jetson TX2 developer kit is a named deployment target, not an accident. On the hardware side there are three paths: CUDA for Nvidia GPUs, OpenCL for AMD GPUs, and a non-GPU, CPU-only version. That third path means the library is not GPU-locked, but the published runtime comparison is a comparison between three pose libraries on shared hardware, not a per-accelerator table, so the CPU-only number for your particular scene is something you have to establish yourself. The list is a snapshot of a supported set, not a compatibility guarantee.

## The newest tag is v1.7.0 from 17 November 2020

Two dates define the lifecycle here. The last push to the master branch was on 3 August 2024. The last tagged release is v1.7.0, published on 17 November 2020, ahead of v1.6.0 from 27 April 2020 and v1.5.1 from 4 September 2019. The repository is not archived, so the history is intact and the tree is still reachable, but there is no release after v1.7.0, and the gap between that tag and the last commit runs to nearly four years. That gap is the practical fact for anyone planning a deployment. Anything you build from master is newer than every binary the project ever tagged, which is fine for research work and awkward for a product that needs to name and pin a version. The release trail that would normally answer questions about behaviour lives in doc/08_release_notes.md, with the feature history in doc/07_major_released_features.md, and the newest entry in that trail corresponds to the 2020 tag. Questions about how the code behaves today have to be answered from the source and the issue tracker instead.

## Built-in inputs and outputs are fixed points, and the API is how you replace them

The input side accepts an image, a video, a webcam, Flir or Point Grey hardware, and an IP camera, and the project states that you can add your own custom input source, naming a depth camera as the example. The output side does the same: a basic image plus keypoint display and saving in formats such as PNG, JPG and AVI, keypoint saving in JSON, XML and YML, keypoints as an array class, and the ability to add your own custom output code. Read those two lists as the built-in surface, and read the trailing sentence on each as the real answer for anything not on it. Three keypoint serialisations and a handful of media containers is a short list, so if your consumer is a database, a renderer, a spreadsheet or a network service, none of it is built in and the documented route is a custom implementation on the API. There is one optimisation in the feature list worth knowing about: single-person tracking, offered for further speedup or visual smoothing, and it is single-person by name. The project also asks that its IEEE TPAMI and CVPR papers be cited if OpenPose assists your research.

## Conclusion

OpenPose remains a defensible choice for a research pipeline, a multi-camera calibration rig, or an integration that needs body, foot, hand and face keypoints in one pass and can build from master. It is a poor choice for a shipped product that needs a tagged release to pin, for multi-person 3D capture, or for a crowded scene where per-person face and hand cost is exactly what you are trying to avoid. Before you commit, read the runtime note attached to each keypoint type, confirm your accelerator is one of CUDA, OpenCL or CPU-only, and read the LICENSE file in the repository yourself.

## FAQ

### What does OpenPose do?

It is a real-time multi-person keypoint detection library for body, face, hand and foot estimation, described as the first real-time multi-person system to jointly detect all four, 135 keypoints in total, on single images. Authored at the CMU Perceptual Computing Lab and built on the CMU Panoptic Studio dataset.

### Is OpenPose free to use?

The repository ships a LICENSE file at the top level and the README carries its own License section, but the terms are not spelled out on the front page, so read that file before you ship a product. Separately, the project asks that its IEEE TPAMI and CVPR papers be cited in your publications if OpenPose helps your research.

### Is there anything better than OpenPose?

The project does not argue that case, and it does not name a successor. What it publishes is an inference time comparison against Alpha-Pose and Mask R-CNN on the same hardware and conditions, where its runtime stays constant while the other two grow linearly with the number of people.

### how to use openpose

Run the demo and add flags in any order. The quick start runs the webcam demo with ./build/examples/openpose/openpose.bin on Ubuntu, and bin\OpenPoseDemo.exe --video examples\media\video.avi from the Windows portable demo. A flag selects the input source, --video takes a path, and --face switches the 70-keypoint face pass on.

### how to use openpose in python

Through the Python API, which is one of two routes for custom functionality alongside the C++ API. Both let you add custom inputs, pre-processing, post-processing and output steps, and the examples tree carries a tutorial_api_python directory next to tutorial_api_cpp.

## Sources

- [Official documentation](https://cmu-perceptual-computing-lab.github.io/openpose)
- [Official README](https://github.com/CMU-Perceptual-Computing-Lab/openpose#readme)
- [Project repository](https://github.com/CMU-Perceptual-Computing-Lab/openpose)
- [Release notes](https://github.com/CMU-Perceptual-Computing-Lab/openpose/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/cmu-perceptual-computing-lab-openpose
