NVlabs/Sana: a linear diffusion transformer for 1024px images and 720p video
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
At a glance
- What is it?
- Sana is NVIDIA Labs' efficiency-oriented codebase for high-resolution image and video generation, with training and inference pipelines for SANA, SANA-1.5, SANA-Sprint, SANA-Video, SANA-WM and Sol-RL. It is a research repository first and a packaged library second, and the install path reflects that.
- Who is it for?
- Sana suits engineers who want to read the model code, retrain a checkpoint or run the Gradio app on a full NVIDIA GPU node, and who accept a pinned dependency tree in exchange for access to the same pipeline the papers use. It is the wrong choice if you need a small, stable library API you can vendor into a product, or if you have no CUDA-capable hardware, since the install targets torch 2.9.1 with cu128 wheels and the Dockerfile starts from nvcr.io/nvidia/pytorch:24.06-py3.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Sana is for, and who it is aimed at
Sana is an efficiency-oriented codebase for high-resolution image and video generation, and the README frames it as a complete training and inference pipeline rather than a single model. The repository hosts several generations of work under one roof: SANA, SANA-1.5, SANA-Sprint, SANA-Video, SANA-Video 2.0, SANA-WM, SANA-Streaming and Sol-RL. Each has its own documentation page under nvlabs.github.io/Sana.
The intended reader is someone who needs the model, the configs and the training scripts in the same tree. The pyproject.toml exposes two console entry points, sana-run and sana-upload, which points at a workflow where you run inference locally and push results or weights to Hugging Face. Training scripts, inference scripts for video, and a diffusion package all sit at the top level, so the expectation is that you will edit configs rather than call a frozen API.
The audience is therefore narrower than the demo pages suggest. If you only want to generate an image from a prompt, the hosted demos and the ComfyUI integration are the cheaper route. The repository is for people who want the weights, the architecture and the training loop in a form they can modify.
Linear attention and the architecture choices behind SANA-Video 2.0
The project name in the repository description is explicit about the mechanism: a linear diffusion transformer. The 2026/08 release note for SANA-Video 2.0 describes hybrid linear/softmax attention and Attention Residuals, which is the concrete architectural claim. Instead of running softmax attention at every layer of the diffusion transformer, the model mixes linear attention layers with softmax layers, and the video variant adds a residual scheme around attention.
The practical consequence is that cost grows differently with resolution and frame count. The 5B 720p checkpoint supports text-to-video and text-image-to-video, with 8-second outputs, and a separate 4-step DMD preview generates 720p in four denoising steps for 5-second and 8-second clips. Step count is the other lever: SANA-Sprint and the 4-step preview both attack sampling cost rather than per-step cost.
Sol-RL is a different kind of component. The release notes describe Sol Engine as an inference acceleration branch, with a 33B omni-modal audio+video DiT reported at 3.95x faster on GB200 and 4.52x on a GeForce RTX 5090, achieved without distillation, LoRA or a calibration pass. That branch lives at github.com/NVlabs/Sana/tree/sol-engine, not on main, so it is a separate checkout rather than a directory in the default tree. Treat those multipliers as the project's own reported numbers from its blog posts, not as independent measurements.
Installing Sana from the repository and running the Gradio app
The repository ships an environment_setup.sh at the top level and a Dockerfile, and the README points at the documentation site for details. The pyproject.toml requires Python 3.11 or newer and pins a deep stack: torch 2.9.1, torchvision 0.24.1, torchaudio 2.9.1, transformers 4.57.3, diffusers 0.37.0 or newer, triton 3.5.1, xformers 0.0.33.post2, timm 0.6.13, mmcv 1.7.2 and flash-linear-attention 0.4.2 or newer. There is also a pip extra-index-url for CUDA 12.8 wheels.
Install from a clone of the repository:
pip install -e .That installs the sana package and registers the sana-run and sana-upload console scripts. Expect a long resolve, and expect the pinned versions to fight any existing environment you have; a fresh virtualenv or the container is the sane starting point.
The Dockerfile is the more reproducible path. It builds from nvcr.io/nvidia/pytorch:24.06-py3, installs libgl1-mesa-glx and libglib2.0-0, copies pyproject.toml, diffusion, configs, sana, app and tools, runs environment_setup.sh, and then starts a Gradio app with a share flag and a 1024px config:
CMD ["python", "-u", "-W", "ignore", "app/app_sana.py", "--share", "--config=configs/sana_config/1024ms/Sana_1600M_img1024.yaml", "--model_path=hf://Efficient-Large-Model/Sana_1600M_1024px/checkpoints/Sana_1600M_1024px.pth"]The model_path uses the hf:// scheme, so the checkpoint is pulled from Hugging Face on first run rather than baked into the image. What you should see is the Gradio interface on the container's port once the weights finish downloading. The --share flag makes Gradio publish a public link, which is convenient for a demo box and a bad idea on a machine holding anything you care about.
For a non-container run, the same app entry point works from a shell:
python -u -W ignore app/app_sana.py --config=configs/sana_config/1024ms/Sana_1600M_img1024.yaml --model_path=hf://Efficient-Large-Model/Sana_1600M_1024px/checkpoints/Sana_1600M_1024px.pthDrop --share if you only need local access. The config path and the model path are the two arguments that matter; both are copied verbatim from the Dockerfile's CMD.
Where Sana gets in your way
The dependency list is the first obstacle. Pinning torch 2.9.1 alongside xformers 0.0.33.post2, triton 3.5.1, peft 0.18.0 and mmcv 1.7.2 means this package does not coexist politely with other ML projects in the same environment. mmcv in particular has a history of hard build requirements, and the Dockerfile installs libgl1-mesa-glx and libglib2.0-0 for a reason.
Hardware is the second. The README's demo badges advertise 6x3090 for the main SANA demo, 1x3090 for the 4-bit and ControlNet demos, and an H100-backed Replicate API. There is no CPU path documented, and the pinned CUDA 12.8 wheel index reinforces that. If you do not have an NVIDIA GPU, the honest answer is to use one of the hosted demos or the Replicate endpoint.
Scope is the third. This is a research repository with multiple model families in one tree, and the README does not document a stable public API for the Python package, nor does it describe rollback or version pinning for the checkpoints. The version field in pyproject.toml is 0.2.0 while the release tags run to v2.0.0, so the package version and the model version are not the same numbering scheme. Do not assume the library API is the interface the maintainers optimize for; the configs and scripts appear to be.
Sana against Stable Diffusion pipelines in diffusers
The most direct comparison is with the diffusers library, which Sana depends on at version 0.37.0 or newer. Diffusers gives you a stable, documented pipeline abstraction: you load a checkpoint, call the pipeline, and the library handles scheduling, offloading and dtype. Sana gives you the training and inference code for its own architectures, with configs in YAML and scripts at the top level.
The architectural difference is the attention mechanism. Standard diffusion transformers in the diffusers ecosystem generally use softmax attention throughout. Sana's stated design is a linear diffusion transformer, and SANA-Video 2.0 uses hybrid linear/softmax attention. That is the reason the project exists: to trade some modeling flexibility for cost that scales better at 1024px and 720p.
A second alternative is ComfyUI, which the README links through the ComfyUI_ExtraModels node pack. ComfyUI is a graph editor; you get visual pipelines and a large node ecosystem, and you give up the ability to change the model code. For someone who wants to run SANA-1.5 or SANA-Sprint without touching Python, that is a better fit than this repository. The repository is for the case where the node pack is not enough.
Maintenance, licensing and the cost of tracking this repo
The repository is not archived and the last push was on 2026-09-19, two days before this writing, so the tree is moving. The release cadence backs that up: v1.0.0 and v1.5.0 both landed on 2025-03-25, and v2.0.0 (SANA-Video and SANA-WM) landed on 2026-06-09, with further work in July, August and September 2026. Upgrading is not a matter of bumping a version number. Each release brings new model families with their own configs, and the pinned dependency set in pyproject.toml moves with them, so a jump from one generation to the next is closer to a reinstall than an upgrade.
The licence is Apache-2.0, declared in LICENSE and in the pyproject.toml classifier. Apache-2.0 permits commercial use and modification and includes a patent grant, which matters for a model repository. Two things it does not settle: the terms attached to individual checkpoints on Hugging Face, which are separate artifacts from the code, and the licences of dependencies such as CLIP, which pyproject.toml pulls from a git URL rather than a released package. Check the model card for whichever checkpoint you deploy; that is a separate question from the repository licence, and this is not legal advice.
Editorial conclusion
Sana suits engineers who want to read the model code, retrain a checkpoint or run the Gradio app on a full NVIDIA GPU node, and who accept a pinned dependency tree in exchange for access to the same pipeline the papers use. It is the wrong choice if you need a small, stable library API you can vendor into a product, or if you have no CUDA-capable hardware, since the install targets torch 2.9.1 with cu128 wheels and the Dockerfile starts from nvcr.io/nvidia/pytorch:24.06-py3. Before committing, check the model zoo page for the checkpoint that matches your resolution and step budget, and confirm which of the several projects in this repository (SANA-1.5, SANA-Sprint, SANA-Video 2.0, SANA-WM, SANA-Streaming, Sol-RL) your use case actually needs, because they share a repo but not a single entry point.
Frequently asked questions
What exactly is Sana?
Sana is an efficiency-oriented codebase from NVIDIA Labs for high-resolution image and video generation, providing complete training and inference pipelines. The repository covers SANA, SANA-1.5, SANA-Sprint, SANA-Video, SANA-Video 2.0, SANA-WM, SANA-Streaming and Sol-RL, with documentation at nvlabs.github.io/Sana/docs/.
What is SANA-WM?
The README describes SANA-WM as a 2.6B controllable world model supporting 720p, 1-minute video generation with 6-DoF camera control, positioned as a baseline for world modeling and embodied AI. Its Stage-1 training, covering bidirectional, chunk-causal and distillation training, was released in 2026/07.
How do I install NVlabs/Sana?
Clone the repository and run pip install -e . in an environment with Python 3.11 or newer, or build the provided Dockerfile, which starts from nvcr.io/nvidia/pytorch:24.06-py3 and runs environment_setup.sh. The pyproject.toml pins torch 2.9.1 with a CUDA 12.8 wheel index, so a fresh environment or the container is the practical route.
Can I run Sana without an NVIDIA GPU?
The repository does not document a CPU path. The README's demo badges list 6x3090 for the main demo, 1x3090 for the 4-bit and ControlNet demos, and an H100-backed Replicate API, and the dependency set targets CUDA 12.8 wheels. Hosted demos and the Replicate endpoint are the documented alternatives.
Does Sana work with ComfyUI?
The README links to the ComfyUI_ExtraModels node pack at github.com/lawrence-cj/ComfyUI_ExtraModels, which is the integration path for running Sana models from ComfyUI rather than from this repository's scripts. The repository itself does not document a built-in ComfyUI node.
What licence does Sana use?
The repository declares Apache-2.0 in its LICENSE file and in the pyproject.toml classifier. Individual model checkpoints on Hugging Face are separate artifacts and may carry their own terms, so check the model card for the checkpoint you intend to deploy.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvlabs-sana)