NVlabs/LongLive 2.0: NVFP4 parallel infrastructure for long video generation
Long Video Gen Infrastructure
At a glance
- What is it?
- LongLive 2.0 is NVIDIA's research infrastructure for training and serving long, autoregressive video generators, with NVFP4 weights, activations and KV cache plus sequence parallelism. It is a GPU-heavy research stack, not a drop-in video API.
- Who is it for?
- Adopt LongLive 2.0 if you are training or distilling an autoregressive long video model and already run multi-GPU CUDA nodes, because the NVFP4 and sequence-parallel paths are the reason the repository exists. Do not adopt it if you need a hosted endpoint, a CPU fallback, or a stable tagged release, since the repository publishes no releases and the README points to a 5B checkpoint path you must supply yourself.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What LongLive 2.0 actually solves, and for whom
Autoregressive video generation has a memory problem before it has a quality problem. A long clip is decoded frame by frame, and the KV cache grows with every step, so a model that looks fine at a few seconds stalls or runs out of memory at a few minutes. LongLive 2.0 is NVIDIA's answer to that specific bottleneck. The README describes it as an "NVFP4 Parallel Infrastructure for Long Video Generation", and the two nouns carry the design: 4-bit floating point for weights, activations and the KV cache, and parallelism across the sequence dimension for both training and inference.
The audience is narrow and worth stating plainly. This is for research teams who already own the model and the cluster. LongLive 1.0 shipped as a real-time interactive generator with published weights (LongLive-1.3B) and a VBench leaderboard placement; 2.0 is framed as infrastructure, aimed at people doing autoregressive teacher-forcing training, few-step DMD distillation, and quantized serving. If you want to call a hosted endpoint and get an MP4 back, this repository is not that, and the install path will not get you there.
NVFP4, sequence parallel and the async decode path
The mechanism is a stack of choices layered on a causal diffusion pipeline. Inference runs through CausalDiffusionInferencePipeline, which the quick-start example constructs from a normalized OmegaConf config and a merged checkpoint. The generator is what gets quantized: NVFP4 inference is described as W4A4, meaning both weights and activations are 4-bit, and the KV cache is quantized in the same format. Alongside that sits a BF16 path, and since the 2026.07.08 news entry, a TorchAO FP8 post-training quantization path (W8A8) derived from the BF16 checkpoint.
Parallelism is the second axis. The README lists balanced sequence parallel for T2V and I2V autoregressive training, sequence parallel inference, and async decoding. There is also a multi-shot attention sink, which matters because attention sinks are how LongLive 1.0 kept a streaming generator from drifting over long horizons; 2.0 extends that to multi-shot video. The 2026.05.25 entry enumerates the optimizations that produced its stated 18.6% throughput improvement: fused Triton RoPE and adaLN kernels, reduced KV-cache synchronization overhead, in-place quantized KV-cache updates, faster FP4 KV dequantization, pinned VAE transfers, and a safer LoRA-before-quantization ordering. That list is the most useful thing in the README for an engineer, because it tells you where the cost actually was.
One design detail deserves attention rather than applause. The changelog notes a "safer LoRA-before-quantization setup", which implies the ordering of adapter application and quantization was previously a source of trouble. Anyone loading LoRA weights onto an NVFP4 checkpoint should treat that ordering as a correctness constraint, not a style preference.
Installing LongLive 2.0 and running a first BF16 inference
The README does not inline install steps. It links to a full documentation site, with anchors for Installation, NVFP4 Setup, Training Modes, Inference and Data Organization, so the authoritative commands live at nvlabs.github.io/LongLive/LongLive2/docs. What the repository does give you is requirements.txt, and one packaging detail in it is worth reading before you start: torchao is pinned to 0.13.0 and the comment says it is "paired with the documented PyTorch 2.8.0 environment". Treat those two as a unit. The same file notes that TensorRT support is optional and requires NVIDIA's package index first, because nvidia-tensorrt is a placeholder on PyPI.
The clone itself has a trap. The default clone pulls every branch, including a demopage branch with large assets, so the README asks for a single-branch shallow clone:
git clone --single-branch --branch main --depth 1 https://github.com/NVlabs/LongLive.gitDependencies install from the pinned requirements file:
pip install -r requirements.txtThe quick-start example is BF16, not NVFP4, which makes it the right first target. It loads configs/inference.yaml, normalizes it, builds the pipeline on CUDA, loads a merged checkpoint, and moves the pipeline to bfloat16:
import torch
from omegaconf import OmegaConf
from pipeline import CausalDiffusionInferencePipeline
from utils.config import normalize_config
from utils.inference_utils import (
load_generator_checkpoint,
place_vae_for_streaming,
prepare_single_prompt_inputs,
save_video,
)
config = normalize_config(OmegaConf.load("configs/inference.yaml"))
device = torch.device("cuda")The example then disables gradients, constructs the pipeline, calls load_generator_checkpoint against a merged checkpoint path written in the README as LongLive-2.0-5B/model_bf16.pt, casts to bfloat16, and calls place_vae_for_streaming, which the comment says honors streaming_vae and vae_device when set. Inputs come from prepare_single_prompt_inputs with a text prompt, inference returns a video tensor, and save_video writes it out. The README truncates before showing the output filename, so check the docs for the exact save path. Note also that the checkpoint path is a local convention, not a download URL: the README does not say where that 5B checkpoint comes from.
Where LongLive 2.0 breaks down
The most obvious limitation is hardware. NVFP4 is a 4-bit floating point format tied to recent NVIDIA GPU generations, and every headline capability of 2.0 (W4A4 inference, NVFP4 KV cache, fused Triton kernels) assumes that class of hardware. There is no documented CPU path, no Apple Silicon path, and no fallback described for older cards. If your fleet is not in that generation, the FP8 PTQ path via TorchAO is the closest thing to a bridge, and it still starts from a BF16 checkpoint on a CUDA device.
Second, the release surface is thin. The repository shows no retrieved releases, so there are no version tags to pin against; you track main. For a research artifact that is normal, and for a production dependency it is a real cost, because a commit can change the quantization ordering or a kernel without a version number to anchor your build.
Third, the README is a launch page, not a manual. It states capabilities as checkboxes and links out for everything operational. It does not document rollback, does not describe failure modes for the quantized KV cache, and does not give memory figures per resolution or clip length. The 45.7 FPS figure in the 2026.05.13 release note carries no stated resolution, batch size, GPU model or clip length, so it is not a number you can plan capacity against. Anyone quoting it as a benchmark should read it as a project claim under unstated conditions.
Finally, the scope is autoregressive causal video generation. If your task is a single short clip from a text prompt with no temporal continuity requirement, a standard diffusion pipeline is simpler and this infrastructure's parallelism and streaming VAE machinery buys you nothing.
How it differs from a general diffusion pipeline
The natural comparison is a general-purpose diffusion video toolkit, Hugging Face diffusers being the obvious one. The difference is architectural, not cosmetic. A standard diffusers video pipeline denoises a fixed-length latent in one pass; memory scales with the clip you asked for, and you get the whole clip at the end. LongLive is causal: it generates autoregressively, keeps a KV cache across steps, and supports async decoding so frames can stream out while later frames are still being computed. That is what makes sequential prompts and user-guided long video possible at all.
The cost of that choice is everything in the requirements file. diffusers appears there pinned at 0.31.0, so LongLive is not an alternative to diffusers; it builds on it and adds a causal pipeline, quantization, and sequence parallelism on top. The trade is complexity for horizon. You accept a config system, a checkpoint-merging convention, quantized KV cache state, and multi-GPU sequence parallelism in exchange for generating past the point where a single-pass pipeline runs out of memory. If you never need that horizon, the added machinery is pure overhead, and a plain diffusers pipeline will be easier to debug.
LongLive 1.0 remains a separate branch, and the README also points to LongLive-RAG, a retrieval-augmented framework for long video generation released under its own repository. Those are different projects with different entry points, not configuration flags on 2.0.
Maintenance, licence and what an upgrade costs
The repository is not archived, and the last push was on 2026-09-07, which is recent enough that the codebase is moving. Movement cuts both ways here. The news entries show a steady cadence through 2026: FP8 inference in July, LongLive-RAG in June, I2V teacher-forcing and DMD distillation for Wan2.2-TI2V-5B in late May, an 18.6% throughput pass in late May, and the 2.0 release in mid May. That is an actively changing inference and training path, and the FP8 addition in particular means the quantization surface grew after 2.0 shipped.
Upgrade cost therefore concentrates in three places: the torchao and PyTorch pairing from requirements.txt, the quantization ordering relative to LoRA loading, and any local patches you made to the Triton kernels. Because there are no releases to pin, an upgrade means diffing main against your fork and re-validating numerics on your own data. Budget for that.
The licence is Apache-2.0, which is permissive and includes an explicit patent grant, and it is the same licence family used across much of NVIDIA's research output. That covers the repository code. It does not automatically cover model weights, which ship under their own terms, and the README does not state a licence for the LongLive-2.0-5B checkpoint it references. Check the weight distribution separately before you ship anything built on it. This is a description of what the repository states, not legal advice.
Editorial conclusion
Adopt LongLive 2.0 if you are training or distilling an autoregressive long video model and already run multi-GPU CUDA nodes, because the NVFP4 and sequence-parallel paths are the reason the repository exists. Do not adopt it if you need a hosted endpoint, a CPU fallback, or a stable tagged release, since the repository publishes no releases and the README points to a 5B checkpoint path you must supply yourself. Before committing, verify the documented PyTorch 2.8.0 pairing against torchao==0.13.0, confirm your GPU generation supports NVFP4, and check that the NVFP4 Triton kernels compile on your driver stack.
Frequently asked questions
What is LongLive 2.0 from NVIDIA?
It is an NVFP4 parallel infrastructure for long video generation, released by NVlabs. The README describes support for autoregressive T2V and I2V training, few-step DMD distillation, and NVFP4 or BF16 inference.
How do I install LongLive 2.0?
The README does not inline the install steps; it links to a documentation site with an Installation section and a separate NVFP4 Setup section. The repository ships requirements.txt, which pins torchao to 0.13.0 and notes it is paired with the documented PyTorch 2.8.0 environment.
Does LongLive 2.0 run on CPU or older GPUs?
The README documents no CPU path. Its headline capabilities are NVFP4 W4A4 inference, an NVFP4 KV cache, and fused Triton kernels, all of which assume CUDA-class hardware, and the quick-start example builds the pipeline on a CUDA device.
What is the difference between LongLive 1.0 and LongLive 2.0?
LongLive 1.0 is real-time interactive long video generation and now lives in the v1.0 branch, with published weights and a VBench leaderboard placement. LongLive 2.0 is framed as infrastructure, adding NVFP4 quantization, sequence parallelism, multi-shot attention sink and async decoding for training, distillation and inference.
What licence does LongLive use?
The repository is licensed under Apache-2.0. The README does not state a separate licence for the model weights it references, so the checkpoint terms need to be checked independently.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvlabs-longlive)