Model or dataset
Tencent-Hunyuan/HunyuanImage-3.0 avatar
Tencent-Hunyuan/HunyuanImage-3.0

HunyuanImage-3.0: Tencent's Autoregressive MoE Model for Image Generation

HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation

3,279 stars191 forksPythonNOASSERTION

At a glance

What is it?
HunyuanImage-3.0 is an open-source 80-billion-parameter Mixture of Experts image generation model from Tencent Hunyuan that uses a unified autoregressive architecture to generate and edit images from text prompts or reference images, with a reasoning-capable variant and a distilled checkpoint for faster inference.
Who is it for?
HunyuanImage-3.0 is worth trying if you have access to CUDA 12.8 hardware and want a large open-source image generation model with integrated reasoning and image-to-image editing. The 80-billion-parameter base model requires substantial GPU memory; the distilled checkpoint is the practical starting point for most deployments, as it is designed for 8-step sampling and should run on hardware that cannot sustain a full inference pass of the base model.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 99 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What HunyuanImage-3.0 Does and Who It Targets

HunyuanImage-3.0 is a text-to-image and image-to-image generation model. The README describes it as targeting researchers, developers, and engineers who want to run large-scale image generation locally or integrate it into a pipeline. It is released as inference code with model weights on HuggingFace, not as a hosted API.

The project is from Tencent Hunyuan's research team. The repository holds inference code, scripts for running the Gradio demo, and documentation pointing to HuggingFace for the model weights. Three model checkpoints are available: the base text-to-image model, the Instruct variant which adds reasoning and image-to-image capabilities, and the Distil variant which is a distilled checkpoint designed for faster sampling.

The README's open-source plan lists multi-turn interaction as not yet released, so conversational image editing that maintains context across multiple turns is not available in the current public release.

The Unified Autoregressive Architecture and MoE Design

The README states that HunyuanImage-3.0 moves away from the Diffusion Transformer (DiT) architecture that is common in other large image generation models. Instead, it uses a unified autoregressive framework that processes text and image tokens together in the same model.

The model is described as the largest open-source image generation Mixture of Experts model, with 64 experts, 80 billion total parameters, and 13 billion parameters activated per token. The MoE design means each forward pass activates only a fraction of the total model capacity, which the README says enhances capacity and performance.

The Instruct variant adds reasoning capabilities: it interprets input images, applies world knowledge to understand user intent, and automatically elaborates on sparse prompts with contextually appropriate details before generating the output. The README describes this as the model leveraging its understanding to produce superior outputs from minimal input instructions.

The README describes the training approach as combining rigorous dataset curation with advanced reinforcement learning post-training. The stated outcome is an optimal balance between semantic accuracy and visual excellence, with the model demonstrating what the README calls exceptional prompt adherence and fine-grained detail rendering. The README does not quantify these claims with specific benchmark numbers in the visible portion of the text, but links to a technical report on arxiv (arxiv.org/pdf/2509.23951), published on September 28, 2025, for formal evaluation methodology and results. The README also includes a Showcase section with sample outputs from both the Instruct and base text-to-image variants.

Environment Setup and Installing Dependencies

The model requires Python 3.12 or higher and CUDA 12.8. The PyTorch installation must match the CUDA version exactly. The requirements.txt specifies:

bash
pip install torch==2.8.0 torchvision==0.23.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128

After installing PyTorch, install the remaining dependencies from requirements.txt. The core dependencies include:

code
einops==0.8.1
numpy==2.2.0
pillow==12.0.0
diffusers==0.35.2
transformers[accelerate,tiktoken]==4.57.1
huggingface_hub[cli]
loguru>=0.7.3

The README also lists an optional performance optimization: flashinfer-python==0.5.0, which the requirements.txt comments as providing up to 3x faster inference. This must be installed separately.

For the interactive demo, Gradio 4.21.0 or higher is required. The top-level run scripts in the repository are run_demo_instruct.sh for the Instruct variant, run_demo_instruct_distil.sh for the distilled model, and run_image_gen.py for the base text-to-image model.

Model Variants and Choosing a Checkpoint

There are three publicly available checkpoints. The base HunyuanImage-3.0 checkpoint handles text-to-image generation. The Instruct checkpoint adds the reasoning layer and image-to-image editing, accepting both text prompts and reference images as input. The Instruct-Distil checkpoint is a distilled version of the Instruct model designed for 8-step sampling.

The README released the Instruct and Instruct-Distil checkpoints on January 26, 2026, and the base text-to-image checkpoint and vLLM acceleration on October 30 and September 28, 2025 respectively. All weights are hosted on HuggingFace at the tencent/HunyuanImage-3.0, tencent/HunyuanImage-3.0-Instruct, and tencent/HunyuanImage-3.0-Instruct-Distil repositories.

For most deployment scenarios, the Instruct-Distil model is the practical starting point. Eight sampling steps significantly reduces the inference time and memory peak compared to the full Instruct model, while the README positions it as the checkpoint for efficient deployment. The full 80-billion-parameter Instruct model requires hardware that can hold 13 billion active parameters in GPU memory during each forward pass.

vLLM Acceleration and Gradio Interface

vLLM support was added on October 30, 2025 as a separate inference path documented in the vllm_infer/ directory. The README describes vLLM as providing significantly faster inference for deployments that need higher throughput. Standard transformers-based inference is documented in the main README; the vllm_infer/README.md covers the vLLM path.

The Gradio interface is accessible by running run_app.sh after setting up the environment. The README documents the app/ directory as containing the interactive demo code and lists the web interface as accessible after the launch script completes. The README does not document the exact URL the Gradio server starts on, but Gradio defaults to http://localhost:7860.

The pyproject.toml defines a console script named hunyuan-image that maps to run_image_gen:main, which is the command-line entry point for batch text-to-image generation.

Comparison with FLUX and Licensing

FLUX is an image generation model from Black Forest Labs, available in several variants including FLUX.1-dev and FLUX.1-schnell, and uses a Diffusion Transformer architecture. The README positions HunyuanImage-3.0 as achieving performance comparable to or surpassing leading closed-source models, but does not make a direct quantitative comparison to FLUX specifically. Both models are available for local inference on high-end GPU hardware.

The architectural difference is significant: HunyuanImage-3.0 uses a unified autoregressive MoE framework that processes text and image tokens in a single model, while FLUX uses a flow-based DiT design where the language understanding and image synthesis are handled through separate mechanisms. The Instruct variant of HunyuanImage-3.0 adds integrated reasoning over input images and supports multi-image fusion as an editing input, which is not a standard capability in FLUX. The Instruct-Distil checkpoint reduces HunyuanImage-3.0's inference cost toward the range where FLUX.1-schnell runs, though both still require substantial GPU hardware.

The LICENSE file is present in the repository, but the metadata identifies the license as NOASSERTION rather than a standard identifier. The pyproject.toml classifies it as 'License :: Other/Proprietary License'. The last push to the repository was on June 23, 2026.

Editorial conclusion

HunyuanImage-3.0 is worth trying if you have access to CUDA 12.8 hardware and want a large open-source image generation model with integrated reasoning and image-to-image editing. The 80-billion-parameter base model requires substantial GPU memory; the distilled checkpoint is the practical starting point for most deployments, as it is designed for 8-step sampling and should run on hardware that cannot sustain a full inference pass of the base model. Before running the model, check the LICENSE file in the repository directly, since the license metadata shows NOASSERTION and the terms for commercial use are not immediately clear from the README alone.

Frequently asked questions

Is HunyuanImage-3.0 open source?

The inference code and model weights are publicly available on GitHub and HuggingFace. However, the license metadata shows NOASSERTION and the pyproject.toml classifies it as Other/Proprietary License, so the terms for commercial use are not a standard open-source license. Read the LICENSE file in the repository before any commercial use.

How do you run HunyuanImage-3.0?

Install Python 3.12 or higher and CUDA 12.8, then install torch==2.8.0 with the CUDA 12.8 index URL. Download the model weights from HuggingFace using huggingface_hub, then run run_image_gen.py for text-to-image or run_demo_instruct.sh for the Instruct variant with reasoning and image-to-image support.

Is HunyuanImage-3.0 censored?

The README does not document content filtering settings or describe what categories of prompts the model declines. The README does not document this aspect of the model's behavior.

Is HunyuanImage-3.0 free to use?

The inference code and model weights are available at no cost for download and local use. The license is not a standard open-source identifier; the pyproject.toml lists it as Other/Proprietary License. The terms for commercial deployment should be verified in the LICENSE file.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Tencent-Hunyuan/HunyuanImage-3.0 on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tencent-hunyuan-hunyuanimage-3-0.svg)](https://hysenlabs.com/projects/tencent-hunyuan-hunyuanimage-3-0)