Model or dataset
Lightricks/LTX-2 avatar
Lightricks/LTX-2

LTX-2: running the Lightricks audio-video model locally

Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.

9,550 stars1,509 forksPythonNOASSERTION

At a glance

What is it?
The official Python package for LTX-2 inference and LoRA training is a workspace of three packages with a Hugging Face weight download behind it. Here is how it installs, what it costs in VRAM, and where it stops being the right tool.
Who is it for?
Adopt LTX-2 if you have a Linux machine with a CUDA GPU and enough disk for the roughly 66 GiB checkpoint set, and you want the reference implementation rather than a wrapper. Do not adopt it if you are on macOS or Windows and need the natten path, if you have no CUDA GPU, or if you want a graphical timeline instead of a CLI.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What LTX-2 is and who the package is for

LTX-2 is the code side of an audio-video generative model from Lightricks. The README describes it as a DiT-based audio-video foundation model that keeps synchronized audio and video, several performance modes and API access in one model rather than splitting them across separate tools. The repository is the official Python inference and LoRA trainer package for that model, so the audience is narrower than the model's marketing suggests: it is for people who want to run generation from a shell, script it, or fine-tune a LoRA, not for people who want a timeline editor.

The weights are not in the repository. They live in a separate Hugging Face repository, and the README points at both the LTX-2.5 model repository and a hosted playground for anyone who would rather not run it at all. That split matters when you plan a deployment: the Python package is small, the model is not, and the two have independent access rules. If you only want to try the model, the playground link is the cheaper path. If you want to fine-tune or batch render, you need the local package.

The README calls LTX-2.5 the recommended model. The repository also ships a MODELS-LTX-2.3.md file at the top level, so earlier model generations are documented in the same tree. Treat the version number in a filename as authoritative: the pipelines check it, and mixing a 2.3 checkpoint into a 2.5 pipeline is not something the documentation presents as supported.

How the pipeline is put together

The repository is a uv workspace. The pyproject.toml lists ltx-core and ltx-pipelines as workspace members and ltx-kernels as a path dependency that is deliberately excluded from the workspace. The comment in the file explains why: ltx-kernels compiles CUDA extensions, and keeping it out of the workspace means a plain uv sync never forces a CUDA toolchain onto the machine. You opt into the compiled kernels through a separate dependency group.

The generation path is a set of pipelines under packages/ltx-pipelines, each exposed as a module you run with python -m. They share a small set of inputs: a transformer, a text encoder, a video VAE, an audio VAE, and for some pipelines a spatial upsampler or a LoRA. The README is explicit that the text encoder is required by every pipeline and that it is bundled with the model, so there is no separate Gemma download step. It also warns that Google's stock Gemma 4 release is not a substitute, because loading checks the encoder's version against the one the checkpoint was trained with.

Two transformers exist and they are not interchangeable. The dev transformer is the full model and is what the guided two-stage pipelines (TI2Vid, Keyframe, A2Vid) use. The distilled transformer runs in far fewer steps and is what DistilledPipeline, DFRPipeline, ICLoraPipeline and DubItPipeline expect. The README states plainly that you should not pass the full dev transformer to DFR. Getting this wrong is the most likely first failure, because both files sit in the same diffusion_models directory with similar names.

There is also a decoder choice inside the video VAE. The diffusion decoder, NADiffusionDecoder, is described as better quality at the cost of longer decode time and more VRAM, and it is fastest with the natten extra. So quality, speed and memory are three dials on the same component, not independent settings.

Installing LTX-2 and rendering a first clip

The README's Quick Start assumes a Linux machine with CUDA. Clone the repository, then sync dependencies with the natten extra. The README notes that natten is Linux and CUDA only, and that on Windows and macOS it is skipped automatically with decoding falling back to a Triton or eager implementation, so the same command is safe to run everywhere even though you only get the fast backend on Linux.

bash
uv sync --extra natten

Next, authenticate with Hugging Face and download the five components the distilled path needs. The README gives this exact command, and notes the total is roughly 66 GiB. The CLI preserves the repository's folder layout under --local-dir, which is why every later path includes a subdirectory such as diffusion_models/ or vae/.

bash
hf auth login
hf download Lightricks/LTX-2.5 \
    diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \
    text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
    vae/ltx-2.5-video-vae-bf16.safetensors \
    vae/ltx-2.5-audio-vae-bf16.safetensors \
    latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
    --local-dir models/ltx-2.5

If the download returns 401 or 403, the README says to accept the model terms on Hugging Face and log in with a Read token, and that fine-grained tokens need the read gated repos scope enabled. That is an access problem, not a network problem, and retrying will not fix it.

With the weights in place, run the distilled pipeline. The README uses 121 frames and seed 4 in its example; the flags map to the files you just downloaded.

bash
uv run python -m ltx_pipelines.distilled \
    --transformer-path       models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \
    --text-encoder-path      models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
    --video-vae-path         models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors \
    --audio-vae-path         models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors \
    --spatial-upsampler-path models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
    --num-frames 121 \
    --seed 4

If the run dies on memory, the README points at two flags: --quantization fp8-cast and --offload with cpu or disk. Those trade speed for headroom, and the offload modes write to CPU memory or to disk respectively.

For production quality the README directs you to DFR, which reuses the same distilled transformer and adds a detailing IC-LoRA from a separate repository. Download it into models/ltx-2.5/loras, then run the DFR module with --detailing-lora added to the same arguments. Defaults for that path are 1024x1536 at 24 fps. The README notes UHD 4K is --width 3840 --height 2176 and calls out that it is not 2160, which is easy to mistype.

Where LTX-2 will not work for you

The clearest boundary is the platform. The natten extra is Linux and CUDA only. On Windows and macOS the sync command still succeeds, but you lose the fastest decoder backend and fall back to Triton or eager, which the README presents as a fallback rather than an equivalent. If your goal is the fastest available decode, a Mac is the wrong machine for this package.

The dependency pinning in pyproject.toml is unusually specific and worth reading before you file a bug. The file pins nvidia-cudnn-cu13 to 9.24.0.43 on Linux with a comment explaining that torch 2.13+cu132 pulls a cuDNN wheel missing libcudnn_engines_tensor_ir, that a system cuDNN then resolves the soname from /usr/lib and triggers CUDNN_STATUS_SUBLIBRARY_VERSION_MISMATCH, and that a wheel older than the one torch was built against aborts the process with a cublasLtGetVersion symbol error the first time torch selects its cuDNN SDPA backend. It also notes 9.25.0.15 is yanked and that the override is Linux-only because the wheel is CUDA and has no darwin build. If you maintain your own environment outside uv, you inherit this problem directly.

Storage is the second constraint. The README states the quick-start download is roughly 66 GiB, and that is before the DFR detailing LoRA, before the temporal upscaler that --temporal-upscalings 1 or 2 requires, and before any outputs. The disk offload option writes to disk as well. A workstation with a fast GPU and a small SSD will spend more time managing space than generating.

Finally, this is a CLI. There is no GUI in the repository, and the README does not document a rollback or version-pinning workflow for the model files themselves. If your team needs a visual editor, a hosted API, or a supported upgrade path between model generations, the local package is the wrong layer, and the playground or API access the README mentions is the better fit.

How it compares with ComfyUI-based video workflows

The obvious alternative for people who want to run open video models locally is a node-based front end such as ComfyUI, which is what a large share of the related searches point at. The difference is not quality, it is where the control lives. ComfyUI gives you a graph you can rewire, a queue, and a visual record of what you ran. LTX-2 gives you a Python module with named flags and a workspace you can import into your own code.

That matters in both directions. If you want to iterate on prompts and look at results, a graph is faster to work in. If you want to render a batch overnight from a script, or train a LoRA against the model, the Python package is the more direct route, and it is the thing the LoRA trainer ships in. The README also points at a hosted playground, which is the third option and the only one that needs no GPU at all.

A practical consequence: the distilled and dev transformer split, the IC-LoRA download, and the version check on the text encoder are all handled by the pipelines here. In a third-party front end, those choices become your responsibility, and a mismatched checkpoint may fail in ways the README of this repository does not describe.

Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-08-26, which is recent enough to treat the project as being worked on. Two releases are listed in quick succession, v1.2.0 on 2026-08-11 and v1.3.0 on 2026-08-26, so the cadence over that window is fast. The CHANGELOG.md at the top level is the place to look for what changed between them.

The licence field is NOASSERTION, meaning GitHub could not classify it automatically. The repository does carry a LICENSE file plus LICENSE-2 and LICENSE-2_x, and the model weights live in a separate Hugging Face repository with their own terms that you accept when you log in. Those two licences may not be the same, and the README does not reconcile them. Read both before you ship anything commercial; this article cannot tell you what they permit.

The upgrade cost is dominated by weights, not code. A new model generation means a new multi-tens-of-gigabytes download, and the text encoder's version check means you cannot mix a new encoder with an old transformer. Budget disk and time for that, not just a git pull.

Editorial conclusion

Adopt LTX-2 if you have a Linux machine with a CUDA GPU and enough disk for the roughly 66 GiB checkpoint set, and you want the reference implementation rather than a wrapper. Do not adopt it if you are on macOS or Windows and need the natten path, if you have no CUDA GPU, or if you want a graphical timeline instead of a CLI. Before committing, verify three things: that your GPU can hold the distilled transformer at your target resolution, that you have accepted the model terms on Hugging Face and hold a Read token, and that your disk can take the weight download plus whatever the disk offload mode writes.

Frequently asked questions

Is LTX-2 free?

The repository is publicly available and the README links to open weights on Hugging Face, but the licence field is NOASSERTION and the repository carries multiple licence files. The model weights have their own terms on Hugging Face that you accept at download time. Check both before assuming commercial use is covered.

Is LTX-2 open source?

The code is published in a public Lightricks repository and the README describes open access, with weights on Hugging Face. GitHub reports the licence as NOASSERTION rather than a named licence, so the exact terms come from the LICENSE files in the repository.

How do I install LTX-2 locally?

Clone the repository, run uv sync --extra natten, then download the model components with the Hugging Face CLI into models/ltx-2.5. The README notes the download is roughly 66 GiB and that natten is Linux and CUDA only, with Windows and macOS falling back to Triton or eager decoding.

How do I use LTX-2 locally?

Run one of the pipeline modules with python -m, passing paths to the transformer, text encoder, video VAE, audio VAE and any upsampler. The README's example is uv run python -m ltx_pipelines.distilled with --num-frames 121 and --seed 4, and it points at --quantization fp8-cast and --offload for low-memory GPUs.

How do I use LTX-2 in ComfyUI?

The repository documents no ComfyUI integration: the README describes running pipeline modules with python -m and points at a hosted playground as the alternative to local execution. Any ComfyUI workflow would be a third-party wrapper, and the transformer, text encoder and VAE version checks described here would become your responsibility.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lightricks-ltx-2.svg)](https://hysenlabs.com/projects/lightricks-ltx-2)
Community notes

Community notes