Open-source project
NVIDIA-ISAAC-ROS/isaac_ros_pose_estimation avatar
NVIDIA-ISAAC-ROS/isaac_ros_pose_estimation

Isaac ROS Pose Estimation: Three ROS 2 Nodes for 6-DoF Object Pose on Jetson

Deep learned, NVIDIA-accelerated 3D object pose estimation

502 stars57 forksC++Apache-2.0

At a glance

What is it?
The repository ships three ROS 2 packages for deep-learned 3D object pose estimation on NVIDIA hardware. FoundationPose handles unseen objects, DOPE is the time-tested option, and CenterPose regresses cuboid dimensions for known classes.
Who is it for?
Adopt this if you already run ROS 2 on Jetson or another NVIDIA GPU and need 6-DoF object pose as a topic in a perception graph. Do not adopt it if you need CPU-only inference, if you cannot supply a DOPE or CenterPose model trained for your object class, or if you expect pose estimation at camera frame rate.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 9 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What isaac_ros_pose_estimation actually solves

A robot arm that needs to pick up a mug has to know where that mug is in three dimensions, not just which pixels contain it. A 2D bounding box gives you a region of the image. Pose estimation gives you the object's position and orientation in 6 degrees of freedom, which is what a manipulation or navigation stack consumes.

This repository packages three separate deep learning approaches to that problem as ROS 2 nodes, all running DNN inference on the GPU. The README's comparison table draws the lines clearly. isaac_ros_foundationpose handles novel objects without retraining and the table marks it as the newest and highest quality but not the fastest. isaac_ros_dope is the fastest and the most mature, but it cannot handle a novel object and has no TAO support. isaac_ros_centerpose sits in between: faster than FoundationPose, better quality than DOPE, established, and it does support TAO.

The audience is narrow by design. You need ROS 2, an NVIDIA GPU, and a reason to fuse pose output with depth to get distance for navigation or manipulation. If you are prototyping on a laptop CPU, nothing here applies to you.

How the three nodes differ in mechanism

FoundationPose splits into two models. The refine model takes initial pose hypotheses and iteratively improves them, then hands the refined set to the score model, which picks the final estimate. The same refine model doubles as a tracker: given a new image and the previous frame's pose, it updates the estimate. The README states that this tracking path is more efficient than full pose estimation, with speeds exceeding 120 FPS on the Jetson Thor platform. That number is the project's own claim, not something verified here.

DOPE takes a different route. It needs a pre-trained model for a known object, plus that object's 3D bounding cuboid dimensions. Input images may need cropping and resizing to preserve aspect ratio and match the model's input resolution. Inference produces belief maps, and a decoder combines those maps with the specified object type to emit poses. NVIDIA provides a DOPE model trained on the HOPE dataset, which the README describes as research-oriented, built from toy grocery objects and 3D textured meshes for synthetic training. For your own objects, you retrain.

CenterPose does more in one pass: detection, 2D keypoints, 6-DoF pose up to scale, and regression of relative 3D cuboid dimensions. The distinction the README emphasizes is that CenterPose works on a known object class without knowing the instance. A chair model detects a chair it was never trained on specifically.

Installing and running a first pose estimation graph

The README does not reproduce install commands in the text provided here; it points to the Isaac ROS documentation site and the release branches. The repository is structured around release-4.6 artifacts, and the packages depend on Isaac ROS DNN Inference for Triton or TensorRT acceleration and on the DNN Image Encoder for preprocessing. Those are separate repositories you need alongside this one.

The top-level layout tells you what you are installing: separate directories for isaac_ros_centerpose, isaac_ros_dope, isaac_ros_foundationpose, a models install directory for FoundationPose, and isaac_ros_pose_proc. The FoundationPose model files are not in the repository itself; isaac_ros_foundationpose_models_install exists precisely because the models are fetched separately, and the README links the FoundationPose model to the NGC catalog. Expect that step to be a distinct download rather than part of the build.

Once a graph is running, the output is a pose topic you feed into your own fusion step with depth. The README is explicit that pose estimation is compute-intensive and not performed at the input camera's frame rate: one estimate is produced for a single frame and handed to navigation, with additional estimates computed at a lower frequency to refine navigation in progress.

The retraining requirement is the real constraint

The most common way to be disappointed by this repository is to pick DOPE expecting it to work on an arbitrary object. It will not. DOPE requires a pre-trained model, and the only one NVIDIA ships is trained on HOPE, a research dataset of toy grocery objects. The README is direct about this: to use DOPE for objects relevant to your application, the model needs to be trained with another dataset targeting those objects, and it links a tutorial for training your own DOPE model.

That training step is not incidental. It means data collection, annotation, and a training pipeline before you get a single pose out of the node. The README's example is telling: DOPE has been trained to detect dollies for a mobile robot that navigates under, lifts, and moves that type of dolly. That is a bespoke model for a bespoke task.

FoundationPose removes the retraining requirement for novel objects, which is why the table marks it as the new default for that case. But it is not the fastest option, and the models live outside the repository. CenterPose avoids retraining per instance but still needs a known object class. None of the three is a general-purpose detector you point at anything.

The second limitation is throughput. Because pose estimation runs below camera frame rate by design, a fast-moving object or a control loop that assumes per-frame poses will not get them. The README frames this as a resource trade-off, and it is, but it is also a hard ceiling on what the node can drive.

FoundationPose versus DOPE versus CenterPose: choosing between the three

These three packages are the alternatives to each other, and the repository's own table is the honest comparison. The axis that matters most is whether you can retrain.

If you cannot retrain and your object is genuinely novel, FoundationPose is the only option in this repository. It is built on the FoundationPose model from NVLabs, which the README describes as capable of both pose estimation and tracking on unseen objects without fine-tuning. You supply 3D bounding cuboid dimensions rather than a trained network. The cost is speed and model download complexity.

If you can retrain and you want the fastest inference, DOPE is the choice. It is described as time-tested, which in a robotics context usually means the failure modes are known. You trade flexibility for speed and predictability.

If you want something in between, CenterPose adds TAO support, which matters if you are already in the TAO training workflow. Its per-class, per-instance-agnostic behavior is a genuine middle ground: less work than training DOPE on every object, more structure than FoundationPose's cuboid-dimension input.

The wrong tool for all three is a CPU-only deployment, or a use case where you need pose at full camera rate for a fast-moving target.

Maintenance, licensing, and upgrade cost

The repository is not archived, and the last push was on 2026-08-19, which is recent enough that the project is being updated. Releases v4.6-0 (2026-08-19), v4.5-0 (2026-07-07), and v4.4-0 (2026-05-01) show a roughly two-month cadence across the last three versions. That cadence is worth noting if you pin versions: you will be upgrading more than once a year.

The release numbering is tied to the Isaac ROS release train, and the README references release-4.6 paths for its images and documentation links. That means an upgrade is not a single package bump. You move the whole Isaac ROS stack, including DNN Inference and the DNN Image Encoder, in step. A partial upgrade is likely to break the graph.

Licensing is Apache-2.0, which is permissive for commercial use, but the licence covers this repository's code. The FoundationPose model is distributed separately through the NGC catalog, and DOPE models may carry their own terms. Read the model licences before shipping. Nothing here is legal advice; check the actual licence files for the models you download.

Editorial conclusion

Adopt this if you already run ROS 2 on Jetson or another NVIDIA GPU and need 6-DoF object pose as a topic in a perception graph. Do not adopt it if you need CPU-only inference, if you cannot supply a DOPE or CenterPose model trained for your object class, or if you expect pose estimation at camera frame rate. Verify first that your GPU platform appears in the repository's performance table, that you have the FoundationPose or DOPE model files the packages expect, and that your target object falls into the category each node supports: novel objects for FoundationPose, trained classes for DOPE and CenterPose.

Frequently asked questions

What is pose estimation?

In this repository, pose estimation means predicting the 6-DoF pose of an object using GPU-accelerated DNN inference. The output prediction can be fused with corresponding depth to provide the 3D pose of an object and distance for navigation or manipulation.

What is NVIDIA Isaac ROS?

Isaac ROS is the broader NVIDIA robotics stack that this repository belongs to. This repository contains three ROS 2 packages for pose estimation, and they rely on Isaac ROS DNN Inference for Triton or TensorRT acceleration and on the Isaac ROS DNN Image Encoder for preprocessing.

What is pose estimation in robotics?

The README describes the output prediction as usable by perception functions when fused with corresponding depth to provide the 3D pose of an object and distance for navigation or manipulation. Pose estimation is a compute-intensive task and is therefore not performed at the frame rate of an input camera.

Official sources

  1. License: Apache-2.0
  2. NVIDIA-ISAAC-ROS/isaac_ros_pose_estimation on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvidia-isaac-ros-isaac-ros-pose-estimation.svg)](https://hysenlabs.com/projects/nvidia-isaac-ros-isaac-ros-pose-estimation)