Library / SDK
nv-tlabs/ardy avatar
nv-tlabs/ardy

ARDY pins three dependencies on purpose and compiles C++ on install

Official implementation of ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation (SIGGRAPH 2026).

952 stars117 forksPythonApache-2.0

At a glance

What is it?
ARDY is NVIDIA's reference implementation for interactive, text-driven human motion generation with kinematic constraints. Its dependency floor is written as a torch prerelease tag to stop pip from replacing a container build, transformers is held at one exact version for a vendored copy, and even the core install compiles a C++ extension.
Who is it for?
ARDY earns its setup only if you already have an NVIDIA GPU, a compiler toolchain and an approved Llama 3 access request, since all three are prerequisites rather than options. It earns its keep when you need streaming text-driven motion with root paths, waypoints, full-body keyframes or sparse joint constraints, rendered in a browser viewport at interactive rates.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 85 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The torch floor is a prerelease tag so pip leaves container builds alone

The dependency floor is written as `torch>=2.4.0a0`, not `torch>=2.4`. The comment beside it explains why: NGC containers ship torch as a prerelease, and `pytorch:24.07-py3` carries `2.4.0a0+...nv24.07`, a string that sorts below `2.4.0`. A plain floor would therefore make pip replace the container's torch with a PyPI wheel and break the matched torchvision and torch-tensorrt builds it ships with. The flip side is visible to everyone else: the `a0` suffix means a nightly or a release candidate satisfies the floor too, so nothing stops pip from choosing a pre-2.4.0 build. Version selection here is deliberate avoidance rather than a minimum, which is why the setup instructions ask you to install PyTorch for your own CUDA version before the package install rather than after it.

transformers is held at one exact version for a vendored copy

`transformers==5.8.1` is an equality pin, and the reason is local to the tree: `ardy/model/llm2vec` is a vendored copy that has only been tested with that version. That is a ceiling for every user of the package, not a suggestion, and it is the kind of pin that breaks on the first release after it. Two neighbouring entries are constrained for the same class of reason. `vector-quantize-pytorch>=1.25,<1.26` is capped because 1.26 and above declare `torch>=2.4`, which the NGC prerelease `2.4.0a0` does not satisfy, so pip would swap the container's torch; only FSQ is used from that library and its interface is unchanged across the cap. numpy is held below 2 with `numpy>=1.23,<2`. Three pins in one dependency list, each carrying a comment about what the resolver would otherwise do.

The core install compiles a C++ extension, so a compiler is required

`pip install -e .` is described as core model inference only, and it still builds something. The install compiles a bundled motion-correction extension from the `MotionCorrection/` directory at the repository root, which needs CMake 3.15 or newer and a C++17 compiler, installed on Ubuntu with `sudo apt install cmake build-essential`. pyproject therefore lists `cmake>=3.15` next to `setuptools>=61.0` in build-system requires, and setup.py carries a custom build command that runs `cmake --version` and raises `RuntimeError("CMake must be installed to build this package")` when the binary is missing. On Windows the same path reads `CMAKE_GENERATOR` from the environment and changes how output paths are passed when it detects mingw. A machine with Python and nothing else cannot install even the inference-only package.

bash
conda create -n ardy python=3.11 -y
conda activate ardy
# Install PyTorch for your CUDA version first — see https://pytorch.org/get-started/locally/
# For example:
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu126
pip install -e ".[all]"

The stated test environment is Ubuntu 22.04 with an RTX 4090, driver 575 and Python 3.11, while the metadata floor is Python 3.10 or newer.

The checkpoints do not carry the repository's Apache-2.0 license

The code is Apache-2.0, declared in pyproject with `license = {text = "Apache-2.0"}` and backed by a LICENSE file at the root. The published model weights do not use it. Every row of the checkpoint table links to the NVIDIA Open Model Agreement instead, and every row also names its training data, Bones Rigplay 1, which carries its own terms at the Bones site. An `ATTRIBUTIONS.MD` sits next to the LICENSE for the same reason. So a deployment that vendors the ARDY code and downloads a checkpoint is combining two licenses with a third data source, and no single sentence in the project says which of the three covers a given artifact. Anyone publishing a product on top of this has to read the model agreement separately from the repository license.

Two skeletons, two horizons, and frame rates that do not line up

Four checkpoints, two skeletons, two prediction horizons each. The Core models run at 20 FPS with horizons of 40 and 8 frames. The Unitree G1 models run at 25 FPS with horizons of 52 and 8 frames. Because the sample rates differ, a horizon of 40 and a horizon of 52 do not cover the same span of time, so a number read off one row cannot be carried to the other row without doing the arithmetic yourself, and the short-horizon variant is 8 in both cases. All four were released on July 10, 2026, which is also the date of the last push to the default branch. A version trained on the same data with the SOMA body model skeleton is announced as coming, with no date attached. The table gives skeleton, data, rate, horizon, date and license, and nothing about parameter count, memory use or frame time.

The demo needs gated Llama access before it needs a checkpoint

Checkpoints download themselves on first use, and the token requirement sits in front of that. The text encoder is built on the gated `meta-llama/Meta-Llama-3-8B-Instruct`, so a Hugging Face account must already have access granted on that model page, and a token must be supplied at runtime, either by running `hf auth login` or by pasting the token into `~/.cache/huggingface/token`. If the `hf` command is missing, the stated fix is `pip install --upgrade huggingface_hub`. Two model stores are therefore involved in one demo: the Llama weights you have to request and wait for, and the ARDY checkpoints you do not. `huggingface_hub>=1.0` and `safetensors>=0.4` are runtime dependencies for that path, and `gradio_client>=2.0` is there to talk to the optional standalone text-encoder service.

TensorRT compiles on first load and reaches outside PyPI at install time

The `[trt]` extra is the fast route and the demanding one. It needs an NVIDIA driver of 525 or newer, described as CUDA 12 capable with the CUDA runtime bundled through pip rather than installed separately, and it needs `pypi.nvidia.com` reachable while the install runs. Where those conditions are not met, the documented fallback is to install `.[demo]` and select a non-TensorRT acceleration mode inside the demo. The cost also lands at runtime: the first Load Model can take a few minutes when TensorRT compilation is enabled, which is a separate wait from the checkpoint download itself. The demo serves a fixed address, `http://localhost:2333`, with left-drag to rotate, right-drag to pan and scroll to zoom. Launching it is `python scripts/run_demo.py`; running `python scripts/run_text_encoder_server.py` in another terminal keeps the text encoder resident so it is not rebuilt on every launch.

Constraint data is provided for G1 only, and the quick start stops mid-word

The optional dataset is needed only for the kinematically constrained generation demos. It is Bones SEED, delivered as CSV for the G1 skeleton, with the matching text descriptions read from `seed_metadata_v004.csv` during sampling rather than stored alongside each clip. The expected layout is a `datasets/bones-seed/` directory at the repository root containing `g1/csv/` and `metadata/`, so the constraint data that ships with the project covers G1 and says nothing about the Core skeleton, even though Core is one of the two published skeletons. The quick start then walks the interface, opens the Model Directory dropdown in the Model tab, loads a checkpoint and notes that a default is loaded when none is chosen, and its final line reads "Start playback: P" and stops, leaving the rest of the key bindings unwritten.

Editorial conclusion

ARDY earns its setup only if you already have an NVIDIA GPU, a compiler toolchain and an approved Llama 3 access request, since all three are prerequisites rather than options. It earns its keep when you need streaming text-driven motion with root paths, waypoints, full-body keyframes or sparse joint constraints, rendered in a browser viewport at interactive rates. Before starting, check that the torch build you have is the CUDA-matched one you want, because the dependency floor is written specifically to stop pip replacing it, and note that the checkpoints ship under the NVIDIA Open Model Agreement rather than this repository's Apache-2.0. With no release tag published and the package still at version 0.2.0, pin a commit rather than a version, and expect the CMake step to be the first thing that fails on a machine without a C++17 toolchain.

Frequently asked questions

What does ARDY actually generate?

Human motion, autoregressively and online from text prompts, with kinematic constraints that hold over long horizons: root paths and waypoints, full-body keyframes, and sparse joint positions or rotations. The interactive demo renders it in a browser viewport with mouse and keyboard locomotion controls.

Do I need a Hugging Face token to run ARDY?

For the text encoder, yes, because it relies on the gated meta-llama/Meta-Llama-3-8B-Instruct model. Your account needs access granted on that model page, and a token supplied at runtime through `hf auth login` or `~/.cache/huggingface/token`. The ARDY checkpoints themselves download automatically when the demo first uses them.

Which license covers the ARDY checkpoints?

Not the repository's Apache-2.0. Every checkpoint links to the NVIDIA Open Model Agreement, and each was trained on the Bones Rigplay 1 dataset, which has its own terms. The repository code is Apache-2.0, and an ATTRIBUTIONS.MD file sits beside the LICENSE.

What does installing ARDY compile?

A bundled motion-correction C++ extension from the MotionCorrection directory, which needs CMake 3.15 or newer and a C++17 compiler. That applies to the core install as well, so `pip install -e .` is not a pure Python install and will fail without a build toolchain.

Which ARDY checkpoints are published?

Four, in two pairs. Core skeleton models at 20 FPS with horizons of 40 and 8 frames, and Unitree G1 models at 25 FPS with horizons of 52 and 8 frames. All four were trained on Bones Rigplay 1 and released on July 10, 2026, with a SOMA body model version announced as coming.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. nv-tlabs/ardy on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nv-tlabs-ardy.svg)](https://hysenlabs.com/projects/nv-tlabs-ardy)