PiD: the decoder that denoises in pixel space, and the checkpoint matrix behind it
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
At a glance
- What is it?
- An NVIDIA research decoder that replaces a VAE or RAE decoder and produces super-resolved pixels in one diffusion pass. Three checkpoint variants and twelve backbone combinations decide what you actually get, and the matrix is not rectangular.
- Who is it for?
- PiD fits a researcher with a latent diffusion backbone who wants to swap the final decode step without retraining the generator, and who can afford an exact-version dependency environment. It does not fit someone who wants a plug-in without picking a variant, because the sharpest decoder and the highest resolution decoder are different files and are not available for every backbone.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 73 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
PiD replaces the decoder, not the generator
The scope claim is narrow and that is what makes it useful. PiD is a plug-and-play diffusion decoder that replaces a VAE or RAE decoder, taking latent representations and producing super-resolved pixels in a single pass. It reformulates latent-to-pixel decoding as a conditional pixel-space diffusion model, so decoding and upsampling become one generative module rather than two stages.
Two entry points reflect the two ways you arrive at a latent. from_ldm.py runs text or class through a latent diffusion model and then decodes, and from_clean.py takes an image, encodes it with a VAE, and decodes with PiD. Both take a --backbone argument.
The from_ldm path has a detail worth knowing: it captures the intermediate x_t at denoising steps you specify, which is early termination of the latent model, as well as the final clean x_0, then decodes each captured latent twice, once with the native VAE or RAE decoder as a baseline and once with PiD. So a single run produces its own comparison, which is the cheapest possible way to see whether the decoder is doing anything for your model.
What is not replaced is the generator. A backbone still has to be run first, which is why every configuration is expressed as a pair of backbone and checkpoint variant.
Three variants, and a matrix that is not rectangular
Every entry point takes --pid_ckpt_type with three values and 2k as the default.
The 2k variant is the original 2048px-trained decoder, trained at 2K resolution only, with aspect ratios typically 2048 by 2048, 2304 by 1728 for 4:3, 1728 by 2304 for 3:4, 2688 by 1536 for 16:9 and 1536 by 2688 for 9:16. The 2kto4k variant is the v1 decoder trained across a varying resolution range from 2K to 4K, and it is less sharp than 2k at 2048 pixels. The 2kto4k_v1pt5 variant is the v1.5 decoder for FLUX, FLUX.2 and Qwen-Image with a WAN2.1 VAE, also trained from 2K to 4K, with better colour accuracy, no grid artefacts in the corners, and more anime and small-face data in training. It beats 2kto4k overall and is still less sharp than 2k at 2048.
Availability is where it gets uneven. flux, flux2, flux2-klein-4b, flux2-klein-9b, zimage and zimage-turbo offer 2k and 2kto4k_v1pt5. sd3 offers 2k and 2kto4k with no v1.5. qwenimage and qwenimage-2512 offer only 2kto4k_v1pt5, so there is no plain 2k path for them. sdxl offers only 2kto4k. dinov2 as an RAE backbone and siglip as Scale-RAE offer only 2k.
So the sharpest decoder is unavailable for four backbones and the highest resolution decoder is unavailable for five. The exact path behind each pair lives in a checkpoint registry inside the package, with docs/checkpoints.md alongside it.
Two installs, and one of them skips the pinned base
There is a fast path for an environment that already has PyTorch with CUDA, transformers 4.57.x and diffusers 0.37. It installs only the utility dependencies the inference code imports eagerly:
pip install hydra-core omegaconf pyyaml \
attrs einops loguru termcolor fvcore iopath wandb \
imageio opencv-python-headless pandas \
safetensors sentencepiece boto3 botocoreThe point of that list is what it omits. torch, torchvision, numpy, diffusers and transformers are not in it, because you are expected to supply them, and the README says so as a condition rather than as a warning.
The clean path uses uv and pins Python 3.12:
uv python install 3.12
uv sync --frozen
source .venv/bin/activate
PYTHONPATH=. python verify_env.pyEither way, commands run from the repository root with PYTHONPATH set to the current directory, and verify_env.py is the check. It reports a pass line only after the required imports and the CUDA checks both succeed, which makes it the one command worth running before debugging anything else.
Checkpoints come from HuggingFace rather than a package index:
hf download nvidia/PiD --local-dir . --include "checkpoints/*"The --frozen flag on uv sync is what keeps a clean environment matching the committed lock rather than resolving fresh.
Exact pins on torch, numpy and diffusers, and two environment files
The dependency declaration is stricter than most research code. Several packages are pinned to an exact version: torch at 2.10.0, torchvision at 0.25.0, numpy at 1.26.4, opencv-python-headless at 4.11.0.86, diffusers at 0.37.1, transformers at 4.57.1, hydra-core at 1.3.2, omegaconf at 2.3.0, packaging at 24.2, and fvcore and iopath at their own exact builds. Others carry ranges, such as Pillow above 11.1 and below 12, einops above 0.8.1 and below 0.9, accelerate above 1.1 and below 2, and pandas above 2.2.3 and below 2.3.
The build backend is pinned too, setuptools at 83.0.0 and wheel at 0.47.0, which is unusual and means a build cannot pick up a patched release automatically.
Two dependencies are there for the infrastructure rather than the model. nvidia-ml-py is imported eagerly by the shared inference and distributed utilities, which is how GPU telemetry gets into a training run, and wandb carries the logging. boto3 and botocore are listed as optional S3 output support imported by the inference I/O registry, so writes to object storage are a supported path rather than a patch.
The Python requirement is narrow at above 3.12 and below 3.13. And the root holds both uv.lock and environment.yml, so there are two environment definitions for the same project, one for uv and one for conda, with no note about which is authoritative.
The lint recipe runs pre-commit twice when it fails
Task running is done with a justfile, and its shape is worth reading because one recipe behaves in a way that is easy to misread.
Setup installs pre-commit from a version floor of 4.3.0 with uv pip install and then registers the hook. Every other recipe depends on setup, including pre-commit itself, so the hook is ensured before anything runs. The recipe body is pre-commit run -a followed by a fallback that runs the identical command again, with an or between them.
That retry is undocumented in the file. It means a formatting or lint failure triggers two full passes over the repository before giving up, and because the recipe line carries no error-ignoring prefix, a second failure aborts the run. Whether the retry exists to catch a first-pass autofix, which rewrites files and then needs a second pass, or to work around a flaky hook is not stated anywhere.
Lint depends on that recipe, and format is an alias of lint rather than a separate target, so there is no separate formatting entry point. A third recipe prefixed with an underscore handles automatic fixes by running manual-stage ruff-fix, then ruff-noqa, then ruff-format, then ruff-noqa again. The duplicated ruff-noqa pass at the end suggests the noqa bookkeeping needs a second look after formatting, which is a reasonable thing to encode, though nothing in the file says so.
The default recipe simply lists the available recipes, which is a friendly touch and the reason the file opens cleanly on a fresh clone.
Version 0.1.0, Alpha, and no releases
The package version is 0.1.0 and its classifier is Development Status 3 - Alpha, which matches a research artifact whose public surface is still moving. Python support is declared for 3.12 only. There are no GitHub releases, so there is nothing to pin from a tag and nothing to diff against.
Licensing has two signals that do not agree. pyproject declares Apache-2.0 and a LICENSE file sits at the root of the tree, while the repository's own license metadata carries no recognised identifier. Nothing in the repository reconciles the two, so treat the file and the declared field as the evidence and confirm which governs before redistributing.
The repository is compact for a project of this size: pid/ holds the package, scripts/ the helper entry points, docs/ the guides including boogu_image.md and checkpoints.md, docker/ and scripts/ for containers, assets/ and figures/ for media, plus environment.yml, uv.lock, verify_env.py, the justfile and a pre-commit configuration.
The last commit on main is dated 2026-07-22, and the most recent README entry is from July 14. So the code has moved past its own news section by about a week, which is the normal state for a repository whose documentation doubles as a release log.
Training code and distilled checkpoints arrived ten weeks after the weights
The news section runs from May 25 to July 14, 2026, and the order of it explains how the project is meant to be used.
On May 25 came the paper, the code and the model weights, with PiD options for FLUX, FLUX.2, Z-Image, Z-Image-Turbo, SD3, DINOv2 and SigLIP. Two days later it landed in ComfyUI through an upstream pull request, which is the route for anyone who wants to try it without writing Python. On June 2 came checkpoints for SDXL, Qwen-Image and Qwen-Image-2512, alongside a codebase cleanup that removed unused code and a Torch.compile mode.
The training side arrived on July 9, in two parts: the training code itself, and PixelDiT with PiD v1.5 checkpoints in both distilled and undistilled form for the 2kto4k variant. A separate July 9 entry released the v1.5 checkpoints for FLUX with Z-Image and Z-Image-Turbo, FLUX.2 and Qwen-Image, with a release page for side-by-side comparison.
So the sequence is roughly ten weeks from public weights to public training code, and the distilled checkpoints arriving alongside the trainer is what makes the v1.5 variant reproducible rather than a black box.
The last feature, on July 14, adds optional Boogu-Image text-to-image support, covering both native Boogu generation and PiD decoding of Boogu's Flux-style VAE latents, which is the first sign of the decoder being used on latents it was not trained for.
Editorial conclusion
PiD fits a researcher with a latent diffusion backbone who wants to swap the final decode step without retraining the generator, and who can afford an exact-version dependency environment. It does not fit someone who wants a plug-in without picking a variant, because the sharpest decoder and the highest resolution decoder are different files and are not available for every backbone. Verify four things first. Read the checkpoint matrix rather than assuming availability, since sdxl and sd3 have no v1.5 variant, qwenimage and qwenimage-2512 have no plain 2k variant, and the RAE backbones dinov2 and siglip only have 2k. Decide which trade you want at 2048 pixels, because 2k is sharper there while 2kto4k and 2kto4k_v1pt5 reach 4K and the v1.5 variant adds better colour accuracy and fewer corner artefacts at the cost of sharpness. Choose your install path deliberately, since the fast path installs only what the inference code imports eagerly and skips the pinned base entirely. And settle the licensing question, because pyproject declares Apache-2.0 with a LICENSE file at the root while the repository metadata carries no recognised identifier. Version 0.1.0 is classified as Alpha, there are no GitHub releases, and the last commit is dated 2026-07-22.
Frequently asked questions
What does NVIDIA PiD do?
It is a plug-and-play diffusion decoder that replaces a VAE or RAE decoder, turning latent representations directly into super-resolved pixels in one pass. It reformulates latent-to-pixel decoding as a conditional pixel-space diffusion model, which unifies decoding and upsampling into a single generative module, while leaving the generator itself untouched.
How do I install NVIDIA PiD and check the environment?
If your environment already has PyTorch with CUDA, transformers 4.57.x and diffusers 0.37, install only the utility dependencies the inference code imports eagerly. Otherwise use uv with Python 3.12 and run uv sync --frozen. Either way, run from the repository root with PYTHONPATH=. and validate with verify_env.py, which prints a pass line once imports and CUDA checks succeed.
Which backbones and checkpoint variants does PiD support?
flux, flux2, flux2-klein-4b, flux2-klein-9b, zimage and zimage-turbo offer 2k and 2kto4k_v1pt5. sd3 offers 2k and 2kto4k, sdxl offers only 2kto4k, qwenimage and qwenimage-2512 offer only 2kto4k_v1pt5, and the RAE backbones dinov2 and siglip offer only 2k.
What is the difference between the PiD 2k and 2kto4k decoders?
The 2k variant was trained at 2048 pixels only and is the sharpest there, supporting typical aspect ratios from 1:1 to 9:16. The 2kto4k variant was trained across 2K to 4K but is less sharp at 2048. The 2kto4k_v1pt5 variant also reaches 4K for FLUX, FLUX.2 and Qwen-Image VAEs, with better colour accuracy and no corner grid artefacts, at the cost of sharpness against 2k.
What license is NVIDIA PiD released under?
pyproject declares Apache-2.0 and a LICENSE file exists at the repository root, while the repository license metadata carries no recognised identifier, and the two signals are not reconciled in the repository. The package version is 0.1.0, its classifier is Development Status 3 - Alpha, and there are no GitHub releases.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nv-tlabs-pid)