Model or dataset
facebookresearch/sapiens avatar
facebookresearch/sapiens

Sapiens: Meta FAIR's Foundation Models for Human-Centric Vision

High-resolution models for human tasks.

5,425 stars323 forksPythonNOASSERTION

At a glance

What is it?
Sapiens is a family of vision models from Meta FAIR, pretrained on 300 million in-the-wild human images and natively trained at 1024x1024 resolution, that covers 2D pose estimation, body part segmentation, depth estimation, and surface normal estimation in a single framework. It was a Best Paper Candidate at ECCV 2024.
Who is it for?
Sapiens is the right choice for researchers and engineers who need high-resolution human-centric vision across multiple tasks and want a single pretrained backbone for all of them. The 1024x1024 native resolution and the 300M pretraining set make it well suited for tasks where fine detail matters, such as hand or face pose, detailed body segmentation, or surface normal estimation for avatar reconstruction.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 126 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Sapiens addresses in human-centric computer vision

Most computer vision models are trained on general image datasets and then finetuned for specific tasks. The Sapiens README describes a different approach: a model family pretrained specifically on 300 million in-the-wild human images, designed to extract high-resolution features from images containing people. The pretraining is done at a native 1024x1024 image resolution with a 16-pixel patch size, which the README states enables excellent generalisation to unconstrained, real-world conditions.

The result is a shared backbone that can be finetuned for several human-centric tasks without starting from a general-purpose ImageNet pretrained model. According to the README, the suite covers 2D pose estimation, body part segmentation, depth estimation, and surface normal estimation. A separate image encoder (pretrain) is also available for teams who want to use the backbone directly. The research was presented at ECCV 2024, where it was a Best Paper Candidate.

Model sizes and checkpoint structure

Sapiens provides multiple model sizes. The README's checkpoint directory layout lists four scales under the pretrain path: sapiens_0.3b, sapiens_0.6b, sapiens_1b, and sapiens_2b. Each size is represented by a corresponding checkpoint file (for example, sapiens_0.3b_epoch_1600_clean.pth). Task-specific checkpoints for pose, segmentation, depth, and surface normals are structured similarly, with per-size folders under the respective task directory.

Checkpoints are downloaded from Hugging Face at the facebook/sapiens collection. The README notes that users can be selective about which checkpoints they download, which matters because a 2B parameter model takes significant storage. After downloading, the SAPIENS_CHECKPOINT_ROOT environment variable should point to the sapiens_host folder that contains the organized checkpoint directories. The detector/ subdirectory holds the RTMPose-based person detector, which is used as a first stage before the Sapiens task models run.

Installing Sapiens: the Lite and Full paths

The README recommends the Lite installation for users running inference with existing checkpoints. The Lite setup requires only PyTorch, numpy, and cv2. The README states it offers 4x faster inference than the full installation. The first step for any installation is to clone the repository and set the root environment variable:

bash
git clone https://github.com/facebookresearch/sapiens.git
export SAPIENS_ROOT=/path/to/sapiens

The full installation replicates the training environment. Run the conda script from the _install directory:

bash
cd $SAPIENS_ROOT/_install
./conda.sh

This creates a conda environment named sapiens and installs all dependencies. The README recommends the full installation only for users who need to replicate the training setup. Finetuning for the four supported tasks (pose, segmentation, depth, surface normals) is documented in separate README files under docs/finetune/.

Finetuning support and the four human-centric tasks

Beyond running pretrained checkpoints, Sapiens supports finetuning for each of its four tasks. The README provides dedicated guides: POSE_README.md, SEG_README.md, DEPTH_README.md, and NORMAL_README.md under both the docs/ and the lite/docs/ paths. Each guide covers the finetuning procedure for that specific task. The lite/ directory mirrors the main structure and provides a Lite-specific version of each task's documentation, so the Lite installation path is not limited to the pretrain encoder.

The Lite inference path is the better-documented starting point for new users. The full training setup inherits from OpenMMLab, which the README acknowledges in its acknowledgements section. That dependency means the full environment replicates the OpenMMLab training infrastructure, which is a significant additional setup step not described in the README itself.

Limitations: licence uncertainty, Sapiens2, and hardware requirements

The main licence for Sapiens is not identified in the repository metadata (GitHub returns NOASSERTION for the licence field). The README states the project is licensed under the terms in the LICENSE file, and that portions derived from open-source projects are under Apache 2.0. Engineers who need clarity on commercial use must read the LICENSE file directly, as the README does not reproduce the licence terms.

The README notes that Sapiens2 is available at https://github.com/facebookresearch/sapiens2. The README does not describe what Sapiens2 adds or changes, so teams evaluating Sapiens for new projects should read the Sapiens2 repository before committing. Running 1024x1024 models at the 2B scale requires substantial GPU memory; the README does not specify hardware requirements, and users running the 0.3B Lite model and the 2B full model will have very different compute needs.

How Sapiens compares to MediaPipe for human analysis

MediaPipe, developed by Google, is the most widely deployed alternative for real-time human body analysis. It covers face detection, face mesh, hand tracking, pose estimation, and segmentation in a single framework designed for mobile and edge deployment. MediaPipe runs at high frame rates on CPU and mobile hardware, prioritising throughput and deployment breadth over output resolution and accuracy at high resolution.

Sapiens takes the opposite trade-off. It is designed for high-resolution, high-accuracy outputs on GPU hardware, not for real-time mobile deployment. The 1024x1024 native resolution and the large pretraining set produce finer output detail than MediaPipe's pose or segmentation, at the cost of requiring GPU inference and more substantial setup. For a researcher building an avatar pipeline, a motion capture system, or a high-quality dataset annotation tool where accuracy matters more than latency, Sapiens is the better fit. For a mobile application or a real-time web interaction that needs to run on a phone CPU, MediaPipe is the more practical choice.

Editorial conclusion

Sapiens is the right choice for researchers and engineers who need high-resolution human-centric vision across multiple tasks and want a single pretrained backbone for all of them. The 1024x1024 native resolution and the 300M pretraining set make it well suited for tasks where fine detail matters, such as hand or face pose, detailed body segmentation, or surface normal estimation for avatar reconstruction. The Lite installation path (PyTorch, numpy, and cv2 only) reduces the setup burden for inference. Check the LICENSE file before using in a commercial product, as the main licence type is not identified in the repository metadata; portions derived from OpenMMLab are under Apache 2.0. Sapiens2 has been released at https://github.com/facebookresearch/sapiens2. Teams starting new work should review that repository before committing to Sapiens. The last push here was on 2026-05-26.

Frequently asked questions

What tasks does the Sapiens model cover?

Sapiens covers 2D pose estimation, body part segmentation, depth estimation, and surface normal estimation, plus a pretrained image encoder that can be used directly. Each task has dedicated finetuned checkpoints available on Hugging Face.

How do I install Sapiens for running inference?

The recommended path is the Lite installation, which requires only PyTorch, numpy, and cv2 and is described as 4x faster than the full setup. Clone the repository, set SAPIENS_ROOT, and follow the Lite README at lite/README.md.

Is there a newer version of Sapiens available?

Yes. The README in this repository states that Sapiens2 is available at https://github.com/facebookresearch/sapiens2. Teams starting new work should review the Sapiens2 repository.

Official sources

  1. facebookresearch/sapiens on GitHub
  2. Issues
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/facebookresearch-sapiens.svg)](https://hysenlabs.com/projects/facebookresearch-sapiens)