# HY-Motion 1.0: Tencent's billion-parameter text-to-motion model, from install to first batch run

> HY-Motion 1.0 turns English action descriptions into skeleton-based 3D character animation using a Diffusion Transformer with flow matching. It ships as Python inference code plus Hugging Face weights, and it wants a GPU with more than 20GB of VRAM before it will do anything useful.

**Tencent-Hunyuan/HY-Motion-1.0** — HY-Motion model for 3D human motion or 3D character animation generation. 

- Repository: https://github.com/Tencent-Hunyuan/HY-Motion-1.0
- Website: https://hunyuan.tencent.com/motion
- Stars: 2,577 · Forks: 219
- Language: Python
- License: NOASSERTION
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/tencent-hunyuan-hy-motion-1-0

## What HY-Motion 1.0 generates, and who the output is for

The README describes HY-Motion 1.0 as a series of text-to-3D human motion generation models built on a Diffusion Transformer and flow matching. The output is skeleton-based 3D character animation, not rendered video and not a rigged mesh. That distinction decides the audience. If you are a game or film pipeline that already has a character rig, a retargeting step, and a place to put a motion clip, this model produces the clip. If you want a finished video file to post, the repository does not claim to produce one.

The stated goal is instruction following: the project says it is the first to scale DiT-based text-to-motion models to the billion-parameter level, and that this improves how well the generated motion matches the prompt. Two checkpoints are published. HY-Motion-1.0 is the standard model at 1.0B parameters, and HY-Motion-1.0-Lite is a 0.46B lightweight variant. Both were released on 2025-12-30, and both list a minimum VRAM requirement in the model zoo table: 26GB for the standard model and 24GB for Lite. Those numbers are worth reading twice, because the gap between them is small relative to the size difference between the two models. If you were hoping Lite would run on a 12GB card, the table says otherwise.

## Diffusion Transformer plus flow matching, and the three training stages

The architecture is a DiT backbone trained with flow matching rather than the DDPM-style objective older motion models used. The README does not spell out the sampling procedure, so the practical consequence for an integrator is limited to what the inference script exposes: a model path, a prompt directory, an output directory, and switches for the prompt-engineering stage.

Training is described in three stages. Large-scale pre-training covers over 3,000 hours of motion data to build a general motion prior. High-quality fine-tuning then uses 400 hours of curated 3D motion data to sharpen detail and smoothness. The third stage applies reinforcement learning from human feedback and reward models to improve instruction following and naturalness. The repository also contains an `ssae` directory, released on 2026-01-29, holding evaluation prompts and code for Structured Semantic Alignment Evaluation, a VLM-based metric for semantic alignment. That is a meaningful inclusion: the maintainers shipped the scoring method alongside the model, so you can check alignment claims with their own tooling rather than trusting the comparison figure in `assets/sotacomp.jpg`.

The repository layout is conventional for a research release. `hymotion/` holds the Python package, `ckpts/` holds weights or weight instructions, `examples/example_prompts/` holds sample inputs, and `local_infer.py` and `gradio_app.py` are the two entry points. There is no service layer, no queue, no REST API in what the README documents.

## Installing HY-Motion 1.0 and running your first batch of prompts

The README says the project supports macOS, Windows, and Linux, and directs you to install PyTorch from the official site first. The dependency list pins torch==2.5.1 and torchvision==0.20.1, so match your PyTorch install to that pair rather than taking whatever pip offers by default. The clone step uses git-lfs, which matters because the repository carries assets and checkpoint material through LFS.

```bash
git clone https://github.com/Tencent-Hunyuan/HY-Motion-1.0.git
cd HY-Motion-1.0/
# Make sure git-lfs is installed
git lfs pull
pip install -r requirements.txt
```

After the dependencies, weights are a separate step: the README points at `ckpts/README.md` for instructions on downloading the necessary model weights. Nothing in the top-level README gives a direct download command, so treat that file as the source of truth. Once weights are in place, batch inference runs through `local_infer.py` with a model path.

```bash
# HY-Motion-1.0
python3 local_infer.py --model_path ckpts/tencent/HY-Motion-1.0

# HY-Motion-1.0-Lite
python3 local_infer.py --model_path ckpts/tencent/HY-Motion-1.0-Lite
```

Prompts come from a directory of `.txt` or `.json` files passed via `--input_text_dir`, and results land in `--output_dir`, which defaults to `output/local_infer`. There is a trap here. The duration estimation and prompt rewriting stages run against a separate module, Text2MotionPrompter, and if you do not supply `--prompt_engineering_host` or `--prompt_engineering_model_path`, the README states you must also pass `--disable_duration_est` and `--disable_rewrite`, or the script raises an error because the host is unavailable. For a first run, passing both disable flags is the shortest path to output.

For interactive use, `python3 gradio_app.py` starts a web interface; the README says to open `http://localhost:7860` afterwards. If Gradio fails with a VRAM error even though the model itself fits, the README gives an escape hatch: run with prompt engineering disabled via `DISABLE_PROMPT_ENGINEERING=True python3 gradio_app.py`.

## The VRAM floor, the language requirement, and what the model will not do

The most concrete limitation is memory. The model zoo table lists 26GB minimum for the 1.0B checkpoint and 24GB for the 0.46B Lite checkpoint. The README offers three settings to reduce VRAM: `--num_seeds=1`, a text prompt under 30 words, and motion length under 5 seconds. Note that these are reductions applied on top of the table's floor, not a way to run on consumer hardware the table excludes. The table also explicitly excludes the VRAM needed for the LLM-based prompt engineering feature, which is why the Gradio app can fail on a machine where the motion model itself loads.

Language is the second constraint. The prompting guide says to use English, and recommends keeping prompts under 60 words. For other languages, the guide directs you to the Text2MotionPrompter to rewrite the prompt, which means a second model download and a second set of VRAM or hosting requirements. A team whose prompts are written in Chinese or Japanese is not running this model alone; it is running this model plus a rewriter.

The README's limitations list is truncated in the published text at a bullet reading "NOT Supported", so the full list of unsupported behaviours cannot be confirmed from what is available. What can be confirmed is the framing: the guide asks for action descriptions and detailed movements of limbs and torso, which points at locomotion, gestures and body mechanics rather than facial performance, lip sync or object interaction. If your requirement is a talking character with hand contact on a prop, this is the wrong tool, and the missing limitations list means you should test your specific case rather than assume it is covered.

## How it compares with the other open text-to-motion route

The obvious alternative is the earlier generation of open text-to-motion models, which are typically transformer or diffusion models in the tens to low hundreds of millions of parameters, trained on motion capture corpora such as HumanML3D, and released with a text encoder plus a motion decoder that you run on a single mid-range GPU. The difference in approach is scale and objective. Those models fit comfortably on a 12GB card and produce motion that follows short, literal prompts; the trade-off is weaker handling of longer or more compositional instructions.

HY-Motion 1.0 takes the opposite bet: a billion-parameter DiT with flow matching, three training stages ending in reinforcement learning from human feedback, and a published VLM-based alignment metric. The README claims superior instruction understanding against comparable open-source models, and the SSAE release gives you a way to check that claim on your own prompts. The cost of the bet is the 24GB to 26GB floor and the added prompt-engineering module for non-English input. There is also a middle option inside the project itself: HY-Motion-1.0-Lite at 0.46B, which the model zoo places only 2GB below the full model. If you are choosing between the two, the table suggests the Lite variant is not a way to escape the memory requirement, only a way to trade some quality for a small reduction.

## Maintenance, licensing, and what to verify before you commit

The repository is not archived, and the last push was on 2026-07-18. The release history in the README is short and recent: inference code and pretrained models on 2025-12-30, and the SSAE evaluation release on 2026-01-29. There are no retrieved releases beyond those entries, so there is no versioned changelog to track and no upgrade path documented. Upgrading in practice means pulling the repository again and re-reading `ckpts/README.md`, since the dependency pins in `requirements.txt` are exact for most packages and a re-pull can move them.

Licensing needs care. The repository carries a `License.txt` file, and the metadata reports the licence as NOASSERTION, which means the classifier could not map the file to a known identifier. The README does not restate the licence terms in the text available here. Model weights are hosted separately on Hugging Face, and the terms attached to those weights are a separate question from the terms attached to the code. Read `License.txt` and the model card on Hugging Face before you ship anything commercial; nothing here substitutes for that reading.

The dependency list also deserves a look before adoption. It pins `bitsandbytes==0.49.0`, `diffusers==0.26.3`, `transformers==4.53.3` and `fbxsdkpy==2020.1.post2`, and it adds an extra index URL pointing at a GitLab package registry. Exact pins make reproducibility easier and conflict resolution harder if your environment already carries different versions of diffusers or transformers.

## Conclusion

Adopt HY-Motion 1.0 if you already have a 24GB-or-larger GPU, English prompts, and a pipeline that consumes skeleton motion rather than rendered video. Do not adopt it if you need non-English prompts without a separate prompter model, or if your only machine is a laptop. Before committing, verify the VRAM figure against your own card by running the Lite checkpoint with --num_seeds=1 and a prompt under 30 words, and read ckpts/README.md to confirm which weights you actually need to download.

## FAQ

### What GPU do I need to run HY-Motion 1.0?

The model zoo table lists a minimum of 26GB VRAM for HY-Motion-1.0 and 24GB for HY-Motion-1.0-Lite. The table excludes the VRAM needed for the LLM-based prompt engineering feature, and the README suggests `--num_seeds=1`, prompts under 30 words and motion under 5 seconds to reduce requirements.

### Can I use HY-Motion 1.0 with non-English prompts?

The prompting guide says to use English and recommends prompts under 60 words. For other languages it directs you to the Text2MotionPrompter to rewrite the prompt, which is a separate module you have to download or host.

### Why does local_infer.py fail with a host unavailable error?

The duration estimation and prompt rewriting stages need the Text2MotionPrompter module. If you do not set `--prompt_engineering_host` or `--prompt_engineering_model_path`, the README states you must also pass `--disable_duration_est` and `--disable_rewrite`, otherwise the script raises an error.

### Does HY-Motion 1.0 output a rendered video or a 3D skeleton animation?

The README describes it as generating skeleton-based 3D character animations from text prompts, intended for integration into 3D animation pipelines. It does not describe a video rendering stage.

## Sources

- [Issues](https://github.com/Tencent-Hunyuan/HY-Motion-1.0/issues)
- [Project website](https://hunyuan.tencent.com/motion)
- [README](https://github.com/Tencent-Hunyuan/HY-Motion-1.0/blob/master/README.md)
- [Tencent-Hunyuan/HY-Motion-1.0 on GitHub](https://github.com/Tencent-Hunyuan/HY-Motion-1.0)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tencent-hunyuan-hy-motion-1-0
