# NVIDIA Isaac GR00T N1.7: an open vision-language-action model for humanoid robots

> GR00T N1.7 pairs a Cosmos-Reason2-2B vision-language backbone with a diffusion transformer action head, licensed Apache-2.0. Here is what the repository documents about installing it, fine-tuning it, and where it stops being the right tool.

**NVIDIA/Isaac-GR00T** — NVIDIA Isaac GR00T N1.7 -  A Foundation Model for Generalist Robots.

- Repository: https://github.com/NVIDIA/Isaac-GR00T
- Website: https://developer.nvidia.com/isaac/gr00t
- Stars: 8,146 · Forks: 1,480
- Language: Python
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-isaac-gr00t

## What GR00T N1.7 actually is, and the problem it addresses

Training a manipulation policy for one robot arm has been a solved-ish problem for years. Training one that transfers across a bimanual humanoid, a semi-humanoid and a tabletop arm is not, mostly because each robot reports joint states and end-effector poses in its own convention. GR00T N1.7 is NVIDIA's attempt at the second problem: an open vision-language-action model that takes language plus camera images as input and emits continuous action chunks, pretrained on a mixture that the README describes as bimanual, semi-humanoid and an expansive humanoid dataset, plus 20K hours of EgoScale human video. The intended user is not someone who wants a chatbot for a robot. It is a robotics engineer who has demonstrations on disk and wants a policy that generalizes beyond the exact trajectories recorded. The README frames the release as General Availability, with pre-trained weights, fine-tuning and inference code, and a deployment path through Gr00tPolicy accelerated with TensorRT. Cross-embodiment transfer is the whole selling point, and the mechanism it leans on is the relative end-effector action space introduced in N1.7, which is shared between robot and human data. That is the design decision worth understanding before you invest in data collection.

## Inside the architecture: VLM backbone plus diffusion transformer head

The architecture is two stacked components. A vision-language foundation model handles the multimodal input (language instructions and images), and a diffusion transformer head denoises continuous actions. The repository includes a schematic at media/model-architecture.png. N1.7 replaces the Eagle backbone used in N1.6 with Cosmos-Reason2-2B, which follows the Qwen3-VL architecture. The practical consequence of that swap is stated plainly in the release notes: the new backbone supports flexible resolution and encodes images in their native aspect ratio without padding. If your camera feeds are not square, that removes a preprocessing step that used to distort them. Actions in N1.7 are represented as deltas from the current end-effector pose rather than absolute target poses. The README calls this a key factor in cross-embodiment performance, and the reasoning is visible: a delta is a delta whether it came from a human hand in an EgoScale video or from a robot gripper, so the same action head can learn from both. The trade-off is that anything downstream of the policy, your controller, your safety limits, your coordinate frames, has to consume deltas correctly. Absolute-pose pipelines do not drop in unchanged.

## Installing GR00T N1.7 and running a first inference call

The packaging is unusual in one respect worth noting up front: the Python requirement is pinned to >=3.12,<3.13, so a 3.11 or 3.13 environment will not resolve. The repository ships a uv.lock and a pyproject.toml whose build backend is setuptools. The project name in pyproject.toml is gr00t. There are separate dependency markers for x86_64 and aarch64: deepspeed publishes wheels only for x86_64 Linux, and triton is listed explicitly for aarch64 (GB200) because it ships with torch on x86_64. The README points to the Installation section for the full procedure, and a docker/ directory exists in the repository alongside a .dockerignore.

The README's Installation section is the place to start, and the repository ships a uv.lock alongside pyproject.toml, so the dependency set is locked rather than resolved freely. That lock covers torch 2.9.0, torchvision 0.24.0 and transformers 4.57.3. On aarch64 Linux, note the comment in pyproject.toml: torchcodec has no aarch64 wheel on PyPI, so the repository expects a prebuilt wheel from scripts/deployment/dgpu/wheels/. If you are on Ubuntu 25.10 or 26.04, install an FFmpeg runtime below version 8 first, because torchcodec 0.8.0 supports FFmpeg 4 through 7 only.

The documented workflow has four steps: prepare data in the GR00T LeRobot format, run inference with the base model on a pretrain embodiment or with a fine-tuned checkpoint, fine-tune with launch_finetune.py, then evaluate open-loop before touching hardware. Fine-tuning examples live under examples/, including finetune.sh and per-benchmark directories such as examples/SO100/ and examples/LIBERO/. The README states that demo datasets are included for quick testing, which is the cheapest way to confirm the environment works before pointing it at your own recordings.

## The relative EEF action space is a migration, not a flag

The single largest source of friction for an existing N1.6 user is the action space change. N1.7 adopts relative end-effector actions shared across robot and human embodiments, and the README directs readers to getting_started/finetune_new_embodiment.md for guidance on configuring relative EEF for their own robot. That file is where the actual work lives. If your robot's state and action vectors were authored against absolute targets, you are not changing a config key; you are re-deriving the action representation and re-validating the controller that consumes it. The repository also notes that N1.7 moves to the gr00t_n1d7 model package, expands the state and action dimensions, and increases the model action horizon. All three of those are interface changes. Anyone with code that imports from gr00t_n1d6 will need to update it, and the README explicitly says to use the n1d6 branch when you need the N1.6 model package and runtime behaviour. That branch is the escape hatch, not a compatibility layer.

## Where GR00T N1.7 is the wrong choice

Three cases stand out. First, CPU-only deployment. Every dependency path in pyproject.toml assumes a GPU stack: torch 2.9.0, flash-attn wheels sourced through [tool.uv.sources], TensorRT export under scripts/deployment/. There is no documented CPU inference path. Second, non-Linux platforms. The dependency markers are conditioned on sys_platform == 'linux' and platform_machine values of x86_64 and aarch64. The comment above deepspeed says it publishes wheels only for x86_64 Linux. macOS and Windows are not addressed in what the repository documents. Third, and more subtly, teams that want a stable policy and no retraining. The README states that N1.7 delivers comparable performance to N1.6, with the improvement concentrated in generalization and language-following. If your N1.6 policy already works on your robot, the migration cost of the new model package and the relative EEF representation buys you generalization you may not need. Staying on n1d6 is a legitimate engineering decision, and the repository keeps that branch alive for exactly this reason. There is also a hardware truth that no README can soften: this is a model for robots. Without a robot, or a simulator like the ones referenced in the benchmark examples, you can fine-tune and evaluate open-loop, but you cannot verify the thing that matters.

## LeRobot integration and how it compares to writing your own policy

The obvious alternative is not another foundation model. It is a task-specific policy trained from scratch on your own demonstrations, which for a single robot and a narrow task is often competitive and far cheaper to debug. The difference in approach is data. A from-scratch policy learns only from your trajectories; GR00T N1.7 starts from pretraining on bimanual, semi-humanoid and humanoid data plus 20K hours of EgoScale human video, and fine-tuning adapts that prior to your embodiment. The bet is that the prior transfers, and the relative EEF action space is what makes the bet plausible, because human video and robot data land in the same action representation. Where the bet fails is when your task is far outside the pretraining distribution, or when your robot's kinematics cannot be expressed as end-effector deltas in a frame the model understands. On the data tooling side, the repository has a dedicated LeRobot Integration section and a GR00T LeRobot format, so the conversion path is documented rather than improvised. That matters more than it sounds: data format is where most fine-tuning attempts die, and having a named format with demo datasets attached shortens the loop considerably.

## Licence, maintenance and the cost of upgrading

GR00T N1.7 is licensed Apache-2.0, and the README states it is fully commercially licensable under that licence. That is a permissive licence, which removes the source-availability question for most commercial deployments, but it says nothing about the model weights themselves or about the third-party dependencies in pyproject.toml, several of which carry their own terms. Check ATTRIBUTIONS.md in the repository root before shipping. On maintenance: the last push to the default branch was on 2026-08-20, and the repository is not archived. The most recent release tag in the list is n1.6.1-release from 2026-04-23, with n1.7-release from 2026-04-18 and n1.6-release from 2026-04-15. Note the ordering: the N1.7 release tag predates the N1.6.1 tag, so if you are tracking tags to decide what to pin, read the dates rather than assuming the newest-looking number is the newest artefact. Upgrade cost between N1.6 and N1.7 is real and concentrated in three places: the model package rename, the expanded state and action dimensions, and the action horizon change. The README's own detailed-changes section is the checklist. Budget for a re-validation pass, not a version bump.

## Conclusion

Adopt GR00T N1.7 if you already have a robot arm or humanoid, demonstrations recorded in the GR00T LeRobot format, and an NVIDIA GPU to fine-tune on; the repository ships example configs for SO100, LIBERO, RoboCasa, SimplerEnv, DROID and GR1 tabletop tasks to start from. Do not adopt it if you need CPU-only inference, if your data is in a format you cannot convert, or if you want a model whose behaviour is unchanged from N1.6: the model package moved from gr00t_n1d6 to gr00t_n1d7 and the state and action dimensions expanded. Before committing, verify three things: that your Python is 3.12 (pyproject.toml pins >=3.12,<3.13), that your FFmpeg is below version 8 because torchcodec 0.8.0 does not support FFmpeg 8, and that your embodiment's action space is expressed as relative end-effector deltas, which is the representation N1.7 was trained around.

## FAQ

### What is NVIDIA Isaac GR00T N1.7?

It is an open vision-language-action model for generalized humanoid robot skills, released by NVIDIA as a General Availability version. It takes language and images as input and outputs continuous actions, combining a vision-language foundation model with a diffusion transformer head that denoises those actions.

### Is NVIDIA Isaac GR00T a VLA model?

Yes. The README describes GR00T N1.7 as an open vision-language-action (VLA) model for generalized humanoid robot skills that takes multimodal input, including language and images, to perform manipulation tasks.

### what is isaac gr00t

The project is NVIDIA's cross-embodiment robot foundation model, trained on bimanual, semi-humanoid and humanoid data and adaptable through post-training for specific embodiments, tasks and environments. N1.7 is the current GA release and is licensed Apache-2.0.

### what is nvidia isaac gr00t

It is NVIDIA's open vision-language-action model for generalized humanoid robot skills, distributed with pre-trained weights and reference code under Apache-2.0. The README states that fine-tuning and inference can be done with custom robot data or demonstrations.

## Sources

- [License: Apache-2.0](https://github.com/NVIDIA/Isaac-GR00T/blob/main/LICENSE)
- [NVIDIA/Isaac-GR00T on GitHub](https://github.com/NVIDIA/Isaac-GR00T)
- [Project website](https://developer.nvidia.com/isaac/gr00t)
- [README](https://github.com/NVIDIA/Isaac-GR00T/blob/main/README.md)
- [Releases](https://github.com/NVIDIA/Isaac-GR00T/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-isaac-gr00t
