# HunyuanVideo: Running Tencent's Text-to-Video Model Locally

> HunyuanVideo is Tencent's open text-to-video framework, shipping PyTorch model definitions, weights and sampling code. It is built for people with serious GPU memory who want the model on their own hardware, not a hosted playground.

**Tencent-Hunyuan/HunyuanVideo** — HunyuanVideo: A Systematic Framework For Large Video Generation Model

- Repository: https://github.com/Tencent-Hunyuan/HunyuanVideo
- Website: https://aivideo.hunyuan.tencent.com
- Stars: 12,576 · Forks: 1,339
- Language: Python
- License: NOASSERTION
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/tencent-hunyuan-hunyuanvideo

## What HunyuanVideo actually is, and who it is for

This repository is the reference implementation behind the paper "HunyuanVideo: A Systematic Framework For Large Video Generation Model". It contains PyTorch model definitions, pre-trained weights and inference and sampling code. That framing matters: it is a research release from Tencent, not a packaged desktop application. The README lists the project page, a web playground, the arXiv tech report and a HuggingFace model card, and the code itself lives in the hyvideo/ package with sample_video.py as the entry point.

The audience is narrow and specific. You need a machine with enough GPU memory to hold a 720p text-to-video diffusion transformer, and you need to be willing to download checkpoints from HuggingFace before anything runs. The repository also publishes a FP8 quantized weight file, released in December 2024, explicitly to save GPU memory, which tells you the full-precision path is demanding. If you want an API call or a browser tab, this is the wrong repository; the README points to a hosted playground for that.

## The pipeline: diffusion transformer, text encoder, VAE

The architecture visible in the repository follows the standard three-part latent video generation stack. A text encoder turns the prompt into conditioning, a diffusion transformer denoises a latent representation over a series of steps, and a VAE decodes the latent back into pixel frames. The topics on the repository (diffusion-models, diffusion-transformer, video-generation) match that reading.

The distinctive engineering choice is parallelism. The README lists "Multi-gpus Sequence Parallel inference (Faster inference speed on more gpus)" as a completed item, and the December 2024 release note says the parallel inference code is "powered by xDiT". That means the model is sharded across GPUs along the sequence dimension rather than replicated, which is how a single video generation job gets faster as you add devices. It also means a multi-GPU run is a coordinated job, not several independent processes.

A second mechanism worth noting is prompt rewriting. The repository links a separate HunyuanVideo-PromptRewrite model on HuggingFace, so the pipeline can expand a short user prompt into something the text encoder handles better. The README does not document when that rewriting is applied inside sample_video.py, so treat it as an optional component you may need to wire up yourself.

## Installing HunyuanVideo locally

The repository ships a requirements.txt at the top level, and the versions are pinned tightly. Start by installing those dependencies into a fresh environment; the pins include torch==2.6.0, diffusers==0.31.0 and transformers==4.46.3, so mixing in a newer diffusers is likely to break the sampling script.

```bash
pip install -r requirements.txt
```

After that, the weights are not in the repository. The README's December 2024 entry links a download guide at ckpts/README.md, and the HuggingFace organization page hosts the model under tencent/HunyuanVideo. Follow that guide rather than guessing at a directory layout, because the sampling script expects a specific checkpoint structure.

Once the checkpoints are in place, the top-level script is the entry point. The README does not reproduce a full invocation, so read the argument parser in sample_video.py before running it, and expect to supply at least a prompt and a checkpoint path. For an interactive session instead of a one-shot script, the repository also contains gradio_server.py, which corresponds to the "Web Demo (Gradio)" item in the open-source plan.

```bash
python gradio_server.py
```

Running that launches a local Gradio interface. The README does not state the port it binds to, so check the output of the command rather than assuming a default.

## Memory, speed and the FP8 escape hatch

The most concrete limitation is GPU memory. The repository shipped FP8 quantized weights in December 2024 with the stated purpose of saving GPU memory, and the community contribution list includes an entry called HunyuanVideoGP described as a "GPU Poor version". Two separate memory-reduction efforts in the same README is a strong signal about the default requirement.

Sequence parallelism helps throughput but not the per-device footprint in the way quantization does; it spreads the work, so more GPUs means faster wall-clock time rather than a smaller card. If you have one consumer GPU, the realistic path is the FP8 weights or one of the community quantizations, and the README does not promise identical output quality for those.

The dependency pins are a second constraint. torch==2.6.0 with numpy==1.24.4 and pandas==2.0.3 is a snapshot from a particular point in time. Upgrading any of them is an untested combination as far as the repository is concerned. There is no requirements lock beyond this file and no documented upgrade path.

## Community forks versus the reference implementation

The README devotes a full section to community contributions, and that section is more useful than most. ComfyUI users have two options: ComfyUI-HunyuanVideoWrapper by Kijai, described as FP8 inference with V2V and IP2V generation, and native support in ComfyUI itself. If your goal is to use HunyuanVideo inside an existing node graph, those are the practical routes, and they differ from the reference script in that they expose the model as composable nodes rather than a single sampling call.

For speed, the README lists TeaCache (cache-based acceleration), FastVideo with a consistency distilled model and sliding tile attention, and NaviCache. These take a different approach from the base repository: instead of running the full denoising schedule, they skip or approximate computation, trading some fidelity for latency. That is a real architectural difference, not a wrapper.

Quantization is covered by HunyuanVideo-gguf, and length extrapolation by RIFLEx. The takeaway is that the reference repository is the baseline, and most production-oriented concerns (VRAM, speed, integration) have been pushed into forks that the maintainers acknowledge but do not maintain.

## The 1.5 release and what it means for this repository

In November 2025 Tencent released HunyuanVideo-1.5, described in the news list as "a highly efficient and powerful new foundation model", as a separate repository. The same pattern applies to HunyuanVideo-I2V for image-to-video, HunyuanCustom for multimodal-driven customization, and HunyuanVideo-Avatar for audio-driven human animation. Each is its own repository building on this one.

That structure has consequences. This repository is the original text-to-video foundation, and the newer capabilities are not merged into it. Anyone starting fresh should decide early whether they need plain text-to-video or one of the derivative models, because the installation and checkpoint instructions differ per repository. The README here does not document a migration path from this codebase to 1.5.

The last push to this repository was on 2026-06-29, roughly three months before this writing, and the repository is not archived. The README's own news list, however, stops at the 1.5 announcement, so the visible development activity has moved elsewhere.

## Licence and upgrade cost

The repository's licence is reported as NOASSERTION, meaning GitHub could not match LICENSE.txt to a recognized template. Both LICENSE.txt and a separate Notice file exist at the top level, so the terms are spelled out in prose rather than a standard identifier. Read both before any commercial use, and note that model weights distributed via HuggingFace can carry terms distinct from the code. This is a description of what the repository contains, not legal advice.

The upgrade cost is the pinned dependency set. Because requirements.txt fixes torch, diffusers and transformers at specific versions, moving to a newer diffusers release means testing the sampling path yourself. There are no retrieved releases on this repository, so there is no changelog to read between versions. If you need stability, pin the same versions in your own environment and treat any bump as a code change to be validated.

## Conclusion

Adopt HunyuanVideo if you have the GPU memory for the full checkpoint, you are comfortable with pinned PyTorch dependencies, and you want the reference implementation rather than a wrapper. Do not adopt it if you need a one-command CPU install or a documented support policy, because the README offers neither. Before committing, check the ckpts/README.md download instructions, confirm your CUDA build matches torch==2.6.0, and read LICENSE.txt and Notice, since the repository declares NOASSERTION and the text-to-video weights carry their own terms.

## FAQ

### What is HunyuanVideo?

It is Tencent's open framework for large video generation models, published with a tech report. The repository contains PyTorch model definitions, pre-trained weights and inference and sampling code.

### How do I install HunyuanVideo?

Clone the repository, install the pinned dependencies with pip install -r requirements.txt, then download the checkpoints following the guide linked in the README at ckpts/README.md. The sampling entry point is sample_video.py, and gradio_server.py launches a local web demo.

### Is HunyuanVideo free?

The code and weights are published openly on GitHub and HuggingFace, so there is no per-generation fee for running it yourself. The cost is hardware: you need a GPU with enough memory for the model, which is why FP8 quantized weights and community memory-reduction forks exist.

### How do I use HunyuanVideo in ComfyUI?

The README lists two routes: ComfyUI-HunyuanVideoWrapper by Kijai, which supports FP8 inference plus V2V and IP2V generation, and native support in ComfyUI itself. Both are community or third-party integrations rather than code in this repository.

### Can I run HunyuanVideo locally?

Yes, that is what the repository is for, but the README's news list points to FP8 quantized weights released specifically to save GPU memory and to a community "GPU Poor version", which indicates the full-precision path needs substantial VRAM. Multi-GPU sequence parallel inference is supported for faster runs.

## Sources

- [Issues](https://github.com/Tencent-Hunyuan/HunyuanVideo/issues)
- [Project website](https://aivideo.hunyuan.tencent.com)
- [README](https://github.com/Tencent-Hunyuan/HunyuanVideo/blob/main/README.md)
- [Tencent-Hunyuan/HunyuanVideo on GitHub](https://github.com/Tencent-Hunyuan/HunyuanVideo)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tencent-hunyuan-hunyuanvideo
