Model or dataset
NVlabs/FoundationPose avatar
NVlabs/FoundationPose

FoundationPose: Unified 6D Pose Estimation and Tracking for Novel Objects

[CVPR 2024 Highlight] FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects

3,599 stars548 forksPythonNOASSERTION

At a glance

What is it?
FoundationPose is NVIDIA's unified model for 6D object pose estimation and tracking that works without fine-tuning on novel objects, given only a CAD model or a small set of reference images. It presented at CVPR 2024 as a Highlight and supports both model-based and model-free setups within the same framework.
Who is it for?
FoundationPose is the right choice for robotics manipulation and AR applications that need to track novel objects in real time without per-object fine-tuning. The Docker setup is the most reliable path to a working environment: the conda path requires matching PyTorch, CUDA toolkit, and PyTorch3D versions manually, and users with RTX 4090-class GPUs must switch to a community image that supports CUDA 12.1.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 154 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The 6D Pose Estimation Problem FoundationPose Addresses

6D object pose estimation means determining the three-dimensional position and three-dimensional orientation (six degrees of freedom combined) of a physical object in a camera image. The classic approach trains a specialized model per object, which requires annotated data for each new object and retraining before deployment.

FoundationPose eliminates the per-object training step. Given a CAD model of an object, it estimates the object's 6D pose in a new scene at test time without any fine-tuning on that object. This is the model-based setup. When no CAD model is available, a model-free setup accepts a small number of reference images instead, using a neural implicit representation to synthesize novel views from those references and keep the downstream pose estimation modules unchanged.

The README states this approach was evaluated on multiple public datasets involving challenging scenarios and that the unified approach outperforms methods specialized for each individual task. As of March 2024, the README listed it as ranked first on the BOP leaderboard for model-based novel object pose estimation.

Architecture: Refiner, Scorer, and Neural Implicit Representation

FoundationPose uses two neural network components whose weights must be downloaded before use. The refiner weight checkpoint is from 2023-10-28 and the scorer weight checkpoint is from 2024-01-11. Both are available from a Google Drive folder linked in the README.

For the model-free setup, the system generates novel views of the object from reference images using a neural implicit representation. This bridges the gap between having a CAD model and having only reference photos: the same downstream pose estimation and tracking modules handle both cases.

Pose estimation is run on the first frame, then the system automatically switches to tracking mode for the rest of the video. The tracked pose is refined on each subsequent frame. The first run is slower due to online CUDA kernel compilation; subsequent runs reuse compiled kernels.

For production deployments in ROS-based robotics pipelines, the README points to Isaac ROS Pose Estimation as the recommended alternative, which uses TensorRT for inference and is implemented in C++.

Setting Up the Environment: Docker or conda

The README recommends Docker as the primary setup path. Pull the prebuilt image and run the container:

bash
cd docker/
docker pull wenbowen123/foundationpose && docker tag wenbowen123/foundationpose foundationpose
bash docker/run_container.sh

On first launch, build the custom C++ extensions inside the container:

bash
bash build_all.sh

For newer GPUs such as the RTX 4090, the standard image does not work. The README directs users to a community image:

bash
docker pull shingarey/foundationpose_custom_cuda121:latest

For the conda path, create the environment from the provided file:

bash
conda env create -f environment.yml
conda activate foundationpose

After activating the environment, install PyTorch with the CUDA build matching the machine's driver, then install PyTorch3D and NVDiffRast from source with CUDA_HOME pointing at the toolkit:

bash
export CUDA_HOME=/usr/local/cuda
export PATH="$CUDA_HOME/bin:$PATH"
python -m pip install --no-build-isolation "git+https://github.com/facebookresearch/pytorch3d.git"
python -m pip install --no-build-isolation "git+https://github.com/NVlabs/nvdiffrast.git"

The --no-build-isolation flag is required so the source build can find the already-installed PyTorch. Finally, install remaining dependencies and build the mycpp extension:

bash
python -m pip install -r requirements.txt
bash build_all_conda.sh

Running the Demo and Evaluating on Public Datasets

After downloading the network weights and the demo data from the linked Google Drive folders, run the model-based demo:

bash
python run_demo.py

By default, paths are set in argparse. The demo runs the mustard bottle scene: pose estimation on the first frame, then tracking for the remaining frames. To run on other objects, change the paths in argparse. No retraining is required for different objects.

For evaluation on LINEMOD and YCB-Video datasets, separate scripts are provided:

bash
python run_linemod.py --linemod

Result visualizations are saved to the debug_dir specified via argparse. The debug folder contains per-frame visualization outputs from the tracking run.

Limitations: Installation Complexity and Hardware Requirements

The installation process has several manual steps that depend on exact CUDA version matching. PyTorch3D and NVDiffRast must be compiled from source, which requires nvcc. The CUDA_HOME variable must be set to the correct toolkit path. The PyTorch version and CUDA build must be consistent across the PyTorch install, PyTorch3D compile, and NVDiffRast compile steps.

The requirements.txt file lists warp-lang as a dependency, which is a separate NVIDIA package. The model-free setup adds Kaolin as an optional dependency with version requirements that must match the installed PyTorch and CUDA versions precisely, as described in the Kaolin documentation.

The pure Python implementation has no TRT acceleration. For applications requiring low-latency inference, the README explicitly recommends Isaac ROS Pose Estimation over this repository.

The network weight files are hosted on Google Drive rather than Hugging Face or a package registry. This means downloading them requires a browser or gdown, and they are not versioned through a standard package management system.

License and Maintenance Status

The repository's LICENSE file is listed as NOASSERTION, meaning the license is a custom NVIDIA license rather than a standard OSI-approved one. The code listings use a custom NVlabs license; the exact terms are in the LICENSE file.

The last push to the repository was on 2026-04-29. The paper was presented at CVPR 2024. The repository has no GitHub releases.

A related project, BundleSDF, provides 6-DoF tracking and 3D reconstruction of unknown objects and is cited in the README with a CVPR 2023 bibtex entry. FoundationPose builds on the neural implicit representation approach from BundleSDF for the model-free setup. Teams evaluating FoundationPose for a robotics use case should also check whether Isaac ROS Pose Estimation, which wraps FoundationPose with TRT inference, better fits their latency and platform requirements.

Comparison with Classic Pose Estimation Methods

Traditional 6D pose estimators such as PoseCNN and DeepIM require separate training on each target object using annotated data or 3D renders. DeepIM uses an iterative refinement loop similar in concept to FoundationPose's refiner network, but DeepIM was designed for known objects with precomputed templates rather than novel objects encountered at test time.

FoundationPose differs by generalizing across objects without retraining. The model-based setup requires a CAD model instead of training data. The model-free setup requires a small number of reference views. This test-time generalization is the primary architectural contribution compared to object-specific methods.

For use cases where the object set is fixed and small, a specialized model trained on that object set may still outperform FoundationPose, since a specialized model can overfit to the known geometry. FoundationPose trades specialization for generality: it covers novel objects at the cost of the installation complexity and the compute overhead of a foundation model at inference time.

Editorial conclusion

FoundationPose is the right choice for robotics manipulation and AR applications that need to track novel objects in real time without per-object fine-tuning. The Docker setup is the most reliable path to a working environment: the conda path requires matching PyTorch, CUDA toolkit, and PyTorch3D versions manually, and users with RTX 4090-class GPUs must switch to a community image that supports CUDA 12.1. Teams that need a ROS integration should use Isaac ROS Pose Estimation, which provides TRT inference and C++ acceleration rather than this pure Python implementation.

Frequently asked questions

What is 6D pose estimation?

6D pose estimation determines the full position and orientation of an object in three-dimensional space from a camera image: three coordinates for position and three for rotation, totaling six degrees of freedom. FoundationPose estimates this pose for novel objects without requiring per-object training.

What is FoundationPose?

FoundationPose is NVIDIA's unified model for 6D object pose estimation and tracking that works on novel objects at test time without fine-tuning. It accepts either a CAD model or a small number of reference images and handles both pose estimation on the first frame and tracking on subsequent frames within the same framework.

How do you use FoundationPose?

Set up the environment via Docker (recommended) or conda, download the network weights and demo data from the Google Drive links in the README, then run python run_demo.py. The demo uses argparse for paths; change the paths to evaluate on different objects without retraining.

What is the FoundationPose alternative for production ROS deployments?

The README points to Isaac ROS Pose Estimation as the production-grade alternative for ROS environments. It uses TensorRT acceleration and a C++ implementation, which provides lower inference latency than FoundationPose's pure Python code.

Official sources

  1. Issues
  2. NVlabs/FoundationPose on GitHub
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvlabs-foundationpose.svg)](https://hysenlabs.com/projects/nvlabs-foundationpose)