# Kimodo: NVIDIA's Kinematic Motion Diffusion Model for Human and Robot Motions

> Kimodo is a kinematic motion diffusion model from NVIDIA trained on 700 hours of commercial mocap data, generating 3D human and humanoid robot motions from text prompts and kinematic constraints including full-body pose keyframes, end-effector positions, and 2D paths. It ships as a Python package with a CLI, an interactive authoring demo, and a benchmark suite.

**nv-tlabs/kimodo** — Official implementation of Kimodo, a kinematic motion diffusion model for high-quality human(oid) motion generation.

- Repository: https://github.com/nv-tlabs/kimodo
- Stars: 3,697 · Forks: 406
- Language: Python
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/nv-tlabs-kimodo

## What Kimodo Generates and Who It Is For

Kimodo is a diffusion model that generates 3D motion sequences for human and humanoid robot skeletons. It takes text prompts as primary input and additionally accepts a rich set of kinematic constraints: full-body pose keyframes, end-effector positions and rotations, 2D path curves, and 2D waypoints. The model was trained on 700 hours of commercially licensed optical mocap data from the Bones Rigplay 1 dataset.

The project is for researchers working on motion synthesis, robotics motion planning, game animation pipelines, and anyone building systems that need controllable, constraint-following 3D character motion. It is not a real-time inference system. A separate related project called ARDY, released by the same NVIDIA lab in July 2026, addresses real-time motion generation.

Kimodo provides a CLI for scripted generation, an interactive demo with a timeline interface for authoring motions by combining text prompts and constraint inputs, and a benchmark suite with test cases and evaluation code.

## Available Model Variants and Their Skeletons

Seven model variants are available, organized by skeleton and training dataset. The default recommendation from the README is Kimodo-SOMA-RP-v1.1, which uses the SOMA 77-joint skeleton and was trained on the full 700-hour Bones Rigplay 1 dataset. A parallel v1.1 variant trained on BONES-SEED (288 hours of publicly available mocap) exists for comparing against other models trained on that dataset.

Robot-focused variants include Kimodo-G1-RP-v1 and Kimodo-G1-SEED-v1, both targeting the Unitree G1 robot skeleton. SMPL-X variants are available for compatibility with the SMPL-X body model.

All models are hosted on Hugging Face and download automatically when first invoked from the CLI or the interactive demo. Manual downloading is not required.

The v1.1 models were released on 2026-04-10 primarily for compatibility with the new Motion Generation Benchmark and include minor quality improvements over v1. The initial v1 release was on 2026-03-16. A breaking change was introduced on 2026-03-19: model inputs and outputs now use the SOMA 77-joint skeleton (somaskel77), and code written for earlier snapshots must be updated.

## Installing Kimodo and Running the CLI

Kimodo supports installation via Docker (recommended for reproducibility) or local pip install. The pyproject.toml defines three install extras: demo (which adds the interactive authoring frontend via a pinned Viser fork), soma (which adds SOMA-X skeleton support), and all (which combines both).

The Docker image is defined in the Dockerfile at the root of the repository. It is based on nvcr.io/nvidia/pytorch:24.10-py3. The build step installs all dependencies from docker_requirements.txt:

```bash
SKIP_MOTION_CORRECTION_IN_SETUP=1 python -m pip install -r docker_requirements.txt
```

The docker-compose.yaml defines two services. The text-encoder service runs at port 9550 and exposes a Gradio endpoint:

```bash
python -m kimodo.scripts.run_text_encoder_server
```

The demo service runs the interactive authoring interface:

```bash
python -m kimodo.demo
```

Both services mount `${HOME}/.cache/huggingface` from the host into the container so Hugging Face model files are not re-downloaded between restarts. The pyproject.toml registers four CLI entry points: kimodo_gen for batch generation, kimodo_demo for the interactive demo, kimodo_textencoder for running the text encoder as a standalone service, and kimodo_convert for motion file conversion. Models download automatically on first use; no manual download step is required.

## Constraint Controls: What You Can Feed the Model

Beyond a text prompt, Kimodo accepts multiple constraint types that let authors pin specific aspects of the output motion. End-effector positions and rotations let you specify exactly where a hand, foot, or robot joint should be at a particular time step. Full-body pose keyframes let you author a specific pose for any frame and have the model fill in the motion between keyframes.

The 2D path input accepts a curve in the horizontal plane, which the model follows while generating upper-body motion from the text prompt. Waypoints constrain the character to pass through specific ground-plane locations at specified times.

This combination of constraints separates Kimodo from text-only motion synthesis models. An application that needs a character to pick up a specific object at a specific location can specify the end-effector position for that moment, rather than relying entirely on the text description to produce spatially accurate results.

The interactive demo provides a timeline interface for authoring these constraints graphically before running inference.

## The Motion Generation Benchmark

The repository includes a benchmark suite released in April 2026 alongside the v1.1 models. Test cases are hosted on Hugging Face at the nvidia/Kimodo-Motion-Gen-Benchmark dataset. The evaluation code measures text-following ability and constraint-following ability separately.

The benchmark is built on top of the BONES-SEED dataset, which is publicly available. This means other motion generation models can be evaluated against the same test cases and compared to Kimodo's published scores. A bug fix was released on 2026-05-03 that corrected an error in the averaged metrics calculation for constraint test cases.

The benchmark is an external contribution point for the research community: teams developing competing motion models can run it against their own checkpoints using the provided evaluation code.

## Limitations: GPU Requirements and License Constraints

Kimodo requires an NVIDIA GPU for inference. The Docker image uses the NVIDIA Container Toolkit via the deploy section in docker-compose.yaml. The text encoder service can be offloaded to CPU using the TEXT_ENCODER_DEVICE=cpu environment variable, which reduces VRAM requirements for the main generation step.

Installation has several dependencies that must be compiled from source: PyTorch3D and NVDiffRast both require nvcc and the CUDA toolkit. The CUDA_HOME environment variable must point to the correct toolkit installation. The requirements.txt comments note that CUDA_HOME must be set for the build to find nvcc.

The license situation varies by model variant. The SOMA-based and G1-based models are released under the NVIDIA Open Model License. The SMPL-X variant is under the NVIDIA R&D Model License (NVIDIA Internal Scientific Research and Development Model License), which has different commercial use terms.

The project has no support for CPU-only inference: motion generation on a machine without an NVIDIA GPU is not supported.

## Maintenance and Repository Activity

The last push to the repository was on 2026-09-22. The project is maintained by NVIDIA's SIL (Simulation and Interactive Learning) research lab. A CHANGELOG.md documents all changes. Contributing guidelines are in CONTRIBUTING.MD.

The repository uses a MANIFEST.in and a setup.py alongside pyproject.toml, which reflects a custom CMake extension for the MotionCorrection subpackage. The Dockerfile uses a build cache mount for pip to speed up iterative builds. The compose file healthchecks the text encoder service by polling the Gradio HTTP endpoint.

The ARDY project, released 2026-07-10, extends Kimodo's controllability model to real-time inference and is linked from the changelog. Users whose requirements include latency constraints should evaluate ARDY rather than Kimodo.

The repository includes a CONTRIBUTING.MD and a CODE_OF_CONDUCT.md. An ATTRIBUTIONS.MD documents third-party assets and dependencies used by the project. The pyproject.toml sets requires-python to 3.8 or higher. Core dependencies include hydra-core, omegaconf, transformers at version 5.1.0 (pinned, not a range), scipy, peft, einops, gradio at 6.8.0 or higher, trimesh, scenepic, and av for video I/O. The pinned transformers version means upgrades to that library require explicit intervention rather than automatic updates, which is a maintenance cost to track when new transformer versions fix security issues. The SOMA skeleton variant also requires py-soma-x, installed from the SOMA-X GitHub repository rather than from PyPI, which means network access to GitHub is required during install for that variant.

## Conclusion

Kimodo is a strong choice for researchers and motion graphics engineers who need text-controllable 3D motion on the SOMA 77-joint or Unitree G1 robot skeleton and have a machine that meets the NVIDIA GPU requirements. Users who need real-time motion generation should look at the separately released ARDY project from the same NVIDIA lab, which is described in Kimodo's changelog as a real-time model with the same controllability. The SMPL-X model variant is under an NVIDIA R&D license rather than the NVIDIA Open Model license, so commercial use of that specific variant requires verifying the terms separately.

## FAQ

### What is Kimodo?

Kimodo is NVIDIA's kinematic motion diffusion model that generates 3D human and robot motion sequences from text prompts and kinematic constraints such as end-effector positions, pose keyframes, and 2D paths. It was trained on 700 hours of commercially licensed optical mocap data.

### Is Kimodo free to use?

The code is released under Apache-2.0. Most model weights use the NVIDIA Open Model License. The SMPL-X model variant uses the NVIDIA R&D Model License, which has separate terms for commercial use. The full license text for each variant is linked from the Hugging Face model pages.

### Is NVIDIA Kimodo open source?

The code repository is open source under Apache-2.0. The model weights are released under NVIDIA-specific licenses (NVIDIA Open Model License or NVIDIA R&D Model License depending on the variant), which are not standard OSI-approved open-source licenses and may restrict certain commercial applications.

### How do you install NVIDIA Kimodo?

Install with pip install -e .[all] for a full local install including demo and SOMA support. For Docker, use docker-compose up to start the text encoder and demo services. PyTorch3D and NVDiffRast must be compiled from source and require the CUDA toolkit with CUDA_HOME set.

## Sources

- [Issues](https://github.com/nv-tlabs/kimodo/issues)
- [License: Apache-2.0](https://github.com/nv-tlabs/kimodo/blob/main/LICENSE)
- [nv-tlabs/kimodo on GitHub](https://github.com/nv-tlabs/kimodo)
- [README](https://github.com/nv-tlabs/kimodo/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nv-tlabs-kimodo
