# Diffuman4D: sparse-view human view synthesis with spatio-temporal diffusion

> Diffuman4D is the official ICCV 2025 code release for turning four-view video of a person into a 44-camera image grid. It is research code with a real install path, a Hugging Face example scene, and one notable gap where a dependency was never open sourced.

**zju3dv/Diffuman4D** — [ICCV 2025] Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models

- Repository: https://github.com/zju3dv/Diffuman4D
- Website: https://diffuman4d.github.io
- Stars: 640 · Forks: 37
- Language: Python
- License: NOASSERTION
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/zju3dv-diffuman4d

## The gap Diffuman4D is trying to close

Dense multi-camera capture of a moving person is expensive. The DNA-Rendering dataset that Diffuman4D builds on uses 48 cameras per scene, which is not a rig most labs or studios can assemble for every shoot. The project asks a narrower question: given video from only four viewpoints, can a diffusion model synthesise the other 40 camera views well enough that a Gaussian splatting reconstruction stays consistent across both space and time? The README frames the output as "high-fidelity free-viewpoint rendering of human performances from sparse-view videos", and the repository name makes the 4D claim explicit. The intended audience is researchers and engineers who already have calibrated multi-view footage of a person and want to expand the camera count in software rather than in hardware. It is not a tool for casual video editing, and it assumes you can supply camera parameters, foreground masks and skeleton maps in the layout the preprocessing scripts expect.

## How the spatio-temporal diffusion model is wired

The pipeline has two halves. The first is data preparation: the repository expects a scene directory laid out as {scene_label}/{data_type}/{camera_label}/{frame_label}{file_ext}, with sibling folders for images, fmasks, skeletons and cameras, plus a sparse_pcd.ply and a transforms.json in nerfstudio format. That transforms.json is what makes the second half possible, because the generated views need camera poses to be reconstructed against. The second half is inference through inference.py, which is driven by Hydra. The README shows three experiment configurations: demo_3d produces a 3D image grid from 4 input cameras to 44 cameras, demo_4d_tiny produces a 4D grid from 4 cameras times 15 frames to 44 cameras times 15 frames, and demo_4d runs the full 4 cameras times 150 frames to 44 cameras times 150 frames. The scene label and data directory are passed as Hydra overrides, data.scene_label and data.data_dir, which is the standard hydra-core pattern and means you can swap scenes without editing Python. Results land in ./output/results/dna_rendering/0023_06 for the example scene. The generated image grid then becomes training data for a downstream splatting model rather than being the final artefact itself.

## Installing Diffuman4D and running the tiny 4D demo

The README gives a conda-first install. Python 3.12 is the stated version, and the inference dependencies come from requirements.txt while the reconstruction and data-processing helpers come from EasyVolcap, installed from git with --no-deps so it does not drag in a conflicting dependency set. Note that requirements.txt pins torch==2.7.1 and diffusers[torch]==0.33.1, so this is not a project you can drop into an arbitrary existing environment.

```bash
conda create -n diffuman4d python=3.12
conda activate diffuman4d
pip install -r requirements.txt
pip install git+https://github.com/zju3dv/EasyVolcap.git --no-deps
```

Next, pull the example scene. The download script takes a --repo_id and a --types argument, and the README passes the types as a JSON-style list string. The four types requested are images, fmasks, skeletons and cameras, which matches the directory layout described above.

```bash
python scripts/download/download_dataset.py --repo_id "krahets/diffuman4d_example" --types='["images", "fmasks", "skeletons", "cameras"]'
```

The README states that inference will try to fetch the pretrained model from Hugging Face automatically. If that fails behind a firewall, the documented fallback writes the weights into ./models/ with a specific local directory name.

```bash
hf download krahets/Diffuman4D --local-dir ./models/models--krahets--Diffuman4D
```

Finally, run the smallest configuration. The README recommends demo_3d or demo_4d_tiny for a single-GPU server, which is a useful admission that the full 4D run is heavy. The command below is the tiny 4D grid: 4 input cameras and 15 frames expanding to 44 cameras and 15 frames.

```bash
python inference.py exp=demo_4d_tiny data.scene_label=0023_06 data.data_dir=./data/datasets--krahets--diffuman4d_example
```

Sampling results are saved under ./output/results/dna_rendering/0023_06. To turn that grid into a splat, the README points at nerfstudio and the splatfacto trainer, invoked against the generated transforms.json.

```bash
ns-train splatfacto --data "./output/results/dna_rendering_tiny/0023_06/transforms.json"
```

One detail worth flagging: the inference output path in the README is dna_rendering/0023_06 while the ns-train example points at dna_rendering_tiny/0023_06. Those two paths do not match, so you should check the actual directory that inference.py creates before running the trainer.

## Where Diffuman4D stops being the right tool

The clearest limitation is documented by the authors themselves. Step 6 of the quick start says that LongVolcap has not been open sourced, and that the team will "attempt to provide" alternative 4D-Gaussian-Splatting reconstruction scripts. That means the 4D reconstruction half of the pipeline is not reproducible from this repository today. You can generate image grids, and you can train a static 3DGS model through nerfstudio, but the path from a 4D image grid to a 4D splat is a promise rather than a shipped component. If your goal is a working 4DGS pipeline, this repository gives you the view synthesis stage and leaves the rest to you. A second constraint is the data contract. The preprocessing expects EasyVolcap-style camera files, foreground masks and skeleton maps, and the re-annotated DNA-Rendering labels cover exactly those four categories. If your footage is not multi-camera and calibrated, or if you cannot produce foreground masks and skeletons, the inference step has nothing to condition on. The README also does not document rollback, checkpoint resumption, or what happens when the Hugging Face download is interrupted partway through, so treat long downloads as something you manage yourself.

## Diffuman4D against a plain 3DGS pipeline

The obvious alternative is to skip the diffusion stage entirely and train a Gaussian splatting model directly on the four input views with nerfstudio's splatfacto. The difference in approach is what each one assumes about the missing data. Direct splatting treats the four cameras as the whole observation set and fits geometry to them, which works when the views are close together and degrades into stretched, floaty geometry when they are not. Diffuman4D instead synthesises the intermediate views first, so the splatting model sees a denser camera ring that it did not have to invent. That is the entire value proposition: the diffusion model is doing novel view synthesis, and splatting is doing reconstruction on a richer input. The trade is that you inherit the diffusion model's errors. If the synthesised views are inconsistent with each other across frames, the splatting stage will faithfully reconstruct that inconsistency as flicker. A direct splat of four views is blurrier but honest about its uncertainty; Diffuman4D is sharper but its accuracy depends on a generative step you cannot fully audit. For a static single-frame capture, the diffusion stage buys you less, and direct splatting is the simpler choice.

## Maintenance, licence and the cost of staying current

The repository is not archived, and the last push was on 2026-09-15, two days before this writing, so the codebase is moving. There are no retrieved releases, which means there is no tagged version to pin against and no changelog to read before upgrading. That matters more than usual here because the dependency set is tightly pinned: torch==2.7.1, torchvision==0.22.1, transformers==4.49.0, diffusers[torch]==0.33.1, hydra-core==1.3.2. Upgrading any one of those is a coordinated change rather than a routine bump, and the Hydra configuration surface means experiment definitions live in configs/ and can shift between commits. The licence is reported as NOASSERTION, which means the repository's LICENSE file does not map to a standard SPDX identifier that GitHub recognises. I am not in a position to interpret it, and this is not legal advice, but for anyone planning to use the weights or the re-annotated DNA-Rendering labels commercially, reading LICENSE and the dataset's own terms on Hugging Face before building on either is the sensible order of operations. The pretrained model is distributed separately at krahets/Diffuman4D, so its terms may differ from the code's.

## Conclusion

Adopt Diffuman4D if you already work with multi-camera human performance data and want to test whether a spatio-temporal diffusion model can fill in the missing viewpoints for a 3DGS or 4DGS pipeline. Skip it if you need a packaged application, a training recipe, or a supported 4DGS path, because the README states that LongVolcap is not open sourced and that the alternative 4D-Gaussian-Splatting scripts are still only planned. Before committing, verify three things on your own hardware: that the demo_4d_tiny configuration completes on a single GPU, that the automatic Hugging Face model download succeeds or that the manual hf download command places weights where inference.py expects them, and that the DNA-Rendering raw images you hold can be extracted with scripts/download/extract_dnar_images.py, since the re-annotated labels cover only masks, skeletons and cameras.

## FAQ

### Does Diffuman4D work with a single GPU?

The README recommends running exp=demo_3d or exp=demo_4d_tiny if you are using a single-GPU server, which implies the full exp=demo_4d run expects more. It does not state a minimum VRAM figure, so you have to measure that yourself.

### What data layout does Diffuman4D expect for inference?

The README describes a structure of {scene_label}/{data_type}/{camera_label}/{frame_label}{file_ext}, with folders for images, fmasks, skeletons and cameras, alongside a sparse_pcd.ply and a transforms.json in nerfstudio format.

### Can Diffuman4D reconstruct a 4DGS model from its output?

Not from this repository alone. The README states that LongVolcap has not been open sourced and that the authors will attempt to provide alternative 4D-Gaussian-Splatting reconstruction scripts, so only the 3DGS path through nerfstudio splatfacto is documented end to end.

## Sources

- [Issues](https://github.com/zju3dv/Diffuman4D/issues)
- [Project website](https://diffuman4d.github.io)
- [README](https://github.com/zju3dv/Diffuman4D/blob/main/README.md)
- [zju3dv/Diffuman4D on GitHub](https://github.com/zju3dv/Diffuman4D)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zju3dv-diffuman4d
