Library / SDK
AmmarkoV/SAM3DBody-cpp avatar
AmmarkoV/SAM3DBody-cpp

SAM3DBody-cpp: C++ Inference Engine for Real-Time 3D Full-Body Reconstruction

Real-time 3D full-body reconstruction from a single camera, Multiperson BVH output, Pure C++ runtime, ONNX + ggml, 70-joint skeleton with hands.

677 stars99 forksCMIT

At a glance

What is it?
SAM3DBody-cpp is a standalone C++ library that takes a single BGR camera frame, runs YOLO person detection and a DINOv2 backbone through ONNX Runtime, and produces per-person 3D body meshes, 70-joint keypoints, and optionally BVH motion-capture files or MPEG ARF avatar containers, with no Python required at runtime.
Who is it for?
SAM3DBody-cpp is the right tool for engineers embedding real-time 3D body reconstruction in C++ applications or motion-capture pipelines that need output compatible with Blender or any Digital Content Creation tool that reads BVH. The hard constraint is the CUDA requirement: without a GPU, one backbone forward pass takes 5 to 15 seconds, making video processing impractical.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What SAM3DBody-cpp Provides and Who Uses It

SAM3DBody-cpp is a C++ inference engine for the SAM-3D-Body model that eliminates the Python dependency at runtime. Given a BGR image, it detects all persons in the frame, runs each crop through a DINOv2-ViT-H backbone and a transformer decoder, and produces 3D body model parameters, camera translation, and optionally the full 3D mesh and 70 keypoints.

The repository targets three types of users: C++ application developers who need to embed 3D body pose estimation without shipping a Python interpreter, motion-capture artists who want to drive Blender characters from a single ordinary camera using the included BVH export, and researchers who want to evaluate the SAM-3D-Body model outside a Python-centric ML stack.

Python support is available through frontends that call the compiled shared library via ctypes, and a CSV exporter for the 70 MHR keypoints is included. The distinction between runtime and frontends is deliberate: the heavy inference path (model loading, ONNX execution, LBS skinning) runs in C++, while Python can drive it through the ctypes interface when that is more convenient.

The Inference Pipeline from Image to Mesh

The pipeline is not 2D-to-3D pose lifting. The README is specific: the network regresses 3D body model parameters directly from image features, requiring no depth sensor, floor plane, or stereo camera.

Starting from a raw BGR image, the pipeline runs five steps. YOLO11m-pose detects person bounding boxes. Each crop feeds a DINOv2-ViT-H backbone that produces a spatial feature map of shape [1280, 32, 32]. A transformer decoder conditioned on the crop's ray directions and focal length compresses that feature map into a 1024-dimensional pose token. Two small feed-forward network heads running on CPU via ggml decode the token into 519 pose parameters (global orientation in 6D continuous rotation, per-joint Euler angles for 127 joints, SMPL-like shape betas for 45 coefficients, hand pose for 108 parameters, and face expression for 72 parameters) and three camera parameters. Those parameters drive linear blend skinning over 18,439 vertices to produce the full body mesh and 70 keypoints.

Focal length is estimated from the image diagonal and baked into the decoder's conditioning. An IoU tracker maintains stable identity across frames, allowing BVH files to carry consistent per-person identities across a video sequence.

Downloading Models and Running the First Video

Pretrained ONNX, GGUF, and LBS model files are hosted on HuggingFace. To fetch the CUDA model bundle into the onnx/ directory at the repository root:

bash
bash tools/fetch_model.sh shared cuda    # ~5.2 GB

The fetch script only downloads files that are absent, so an interrupted download can be resumed by running it again. For CPU-only machines use the cpu profile (~3.6 GB); for TensorRT use trt (~1.9 GB).

To process a video clip and export one BVH file per detected person:

bash
./scripts/video.sh --from clip.mp4 --bvh ./p.bvh --headless

This produces p_0.bvh, p_1.bvh, and so on. The --headless flag suppresses the display window, which is needed for server environments.

For CPU-only inference (if no CUDA GPU is available), the standard backbone.onnx uses BFloat16 weights that the ONNX Runtime CPU execution provider cannot load. The README provides separate CPU-compatible downloads: backbone_fp32.onnx (~3.2 GB data file) and decoder_fp16.onnx. Once those files are placed in onnx/, passing --cuda -1 picks them up automatically:

bash
./scripts/video.sh --from your_video.mp4 --cuda -1

BVH Motion-Capture Export for Blender

The --bvh flag is the main integration point for DCC workflows. When a BVH file path is passed, the system writes one BVH file per detected person (p_0.bvh, p_1.bvh, and so on). The README states that each file's joint OFFSETs are auto-resized to the actor's measured bone lengths, so the skeleton proportions match the person in the video rather than a fixed template.

A Blender plugin at blender/blender_bvh_plugin.py drives a MakeHuman-rigged character from the exported BVH. This is the documented path for animating Blender characters from the single-camera capture output. The BVH output is standard format and compatible with any DCC that reads BVH files, including MotionBuilder, Maya, and commercial BVH testing tools.

Identity stability across frames is maintained by the built-in 2D bounding box IoU tracker, which assigns a consistent person index to each detected person across frames. This is what allows a multi-person scene to produce stable, non-swapping BVH files rather than files where identities switch mid-sequence.

MPEG ARF Avatar Export

The --arf flag exports a different format: MPEG Avatar Representation Format (.arfz container) per detected person. Unlike BVH, which carries only the skeleton and joint rotations, ARF includes the mesh, skin weights, a personalized rest mesh, and optionally facial blendshapes when --dev-face is passed.

The ARF container is self-contained: a single .arfz file holds the base avatar and the per-frame Avatar Animation Unit stream. The README describes it as independent of --bvh, so both flags can be passed together to produce both formats from the same run:

bash
./scripts/offline_video.sh --from clip.mp4 --arf ./p.arfz

This produces p_0.arfz, p_1.arfz, and so on. The format is based on ISO/IEC 23090-39. The repository's knowledge/ARF.md documents the container layout, the JSON schema subset implemented, and how the implementation maps onto the specification.

GPU Requirement and CPU Inference Reality

The DINOv2-ViT-H backbone has approximately 630 million parameters. The standard backbone.onnx and decoder.onnx use BFloat16 weights, which require a CUDA GPU to execute because the ONNX Runtime CPU provider has no BF16 kernels.

The CPU-compatible alternatives (backbone_fp32.onnx at 3.2 GB and decoder_fp16.onnx) do run on CPU, but the README is explicit about performance: one backbone forward pass takes 5 to 15 seconds on a modern laptop CPU. Video processing at any meaningful frame rate is impractical without a CUDA GPU. Single images and low-frequency use cases (one frame per 30 seconds, for example) are described as feasible.

WSL2 users can achieve GPU acceleration if nvidia-smi works inside WSL: the README specifies installing only the CUDA toolkit for WSL2 (not the full Linux driver) and then using the standard backbone.onnx. The repository includes knowledge/WSL.md as a step-by-step guide for WSL2 setup.

A refined pose mode (--refined-pose) adds an iterative decoder that improves accuracy at the cost of an additional model download (~607 MB for pipeline_refined.gguf and iterative decoder graphs). This is fetched automatically on first use with the flag, or manually with bash tools/fetch_model.sh cuda refined.

License and Build Requirements

The repository is licensed under MIT, which permits commercial use, modification, and redistribution without restriction beyond attribution.

The last push was on 2026-09-21. The repository is not archived. Build requirements include CMake, the ONNX Runtime library, and a CUDA toolkit for GPU builds. CMake warns at configure time if the onnx/ directory is absent or the model zip is not found. The models themselves are not included in the repository and must be downloaded separately via tools/fetch_model.sh before the first run. The INSTALL file in the repository root covers build prerequisites in more detail.

A comparable alternative for single-camera 3D body estimation is MediaPipe Pose from Google, which also runs from a single camera without special hardware. MediaPipe Pose is cross-platform and has prebuilt packages for Python, JavaScript, and Android, but its output is a 33-landmark skeleton without mesh skinning or BVH export. SAM3DBody-cpp produces a full 18,439-vertex mesh with 70 keypoints and direct BVH and ARF export, making it more appropriate for DCC integration and motion-capture workflows at the cost of significantly higher hardware requirements.

Editorial conclusion

SAM3DBody-cpp is the right tool for engineers embedding real-time 3D body reconstruction in C++ applications or motion-capture pipelines that need output compatible with Blender or any Digital Content Creation tool that reads BVH. The hard constraint is the CUDA requirement: without a GPU, one backbone forward pass takes 5 to 15 seconds, making video processing impractical. Before building, run bash tools/fetch_model.sh shared cuda to download the model bundle, verify that onnx/ is populated, and confirm that your CMake configuration can find the CUDA toolkit and ONNX Runtime. WSL2 users should follow the dedicated WSL.md guide rather than a generic CUDA setup.

Frequently asked questions

Can SAM3DBody-cpp run without a CUDA GPU?

Yes, but with severe performance limitations. The standard model files require a CUDA GPU. CPU-compatible alternatives (backbone_fp32.onnx and decoder_fp16.onnx) are available as separate downloads, but the README states that a single backbone forward pass takes 5 to 15 seconds on a modern laptop CPU, making video processing impractical. Single-image use cases are described as feasible.

What is the BVH output format in SAM3DBody-cpp?

SAM3DBody-cpp writes a standard BVH motion-capture file per detected person when --bvh is passed. Each file's joint OFFSETs are auto-resized to the actor's measured bone lengths, identity is kept stable across frames by a 2D IoU tracker, and the output is compatible with Blender, MotionBuilder, and any DCC that reads BVH. A Blender plugin at blender/blender_bvh_plugin.py is included.

How do I download the ONNX model files for SAM3DBody-cpp?

Run bash tools/fetch_model.sh shared cuda to download the CUDA model bundle (approximately 5.2 GB) into the onnx/ directory. Use shared cpu for CPU-only machines (approximately 3.6 GB) or shared trt for TensorRT. The script resumes interrupted downloads and skips files that are already present.

Official sources

  1. AmmarkoV/SAM3DBody-cpp on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ammarkov-sam3dbody-cpp.svg)](https://hysenlabs.com/projects/ammarkov-sam3dbody-cpp)