Model or dataset
nv-tlabs/XCube avatar
nv-tlabs/XCube

nv-tlabs/XCube: sparse 3D voxel generation whose dependency installs from an unmerged pull request

[CVPR 2024 Highlight] XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies

549 stars42 forksPythonNOASSERTION

At a glance

What is it?
XCube is NVIDIA's CVPR 2024 highlight model for high-resolution sparse 3D voxel grids, generating up to 1024 cubed feed-forward. The research claim is clear; the install is the fragile part, since fVDB comes from OpenVDB pull request 1808 with a setup.py copied in from the repository's own assets folder.
Who is it for?
XCube fits a researcher who needs large sparse 3D samples or scene-scale reconstructions and is prepared to build a GPU environment from source on Linux, accepting that some advertised modes are placeholders. It is a poor fit if you need macOS or Windows, a tagged release, or the conditional sampling paths, since the single-scan condition is marked coming soon.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 108 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The fVDB install fetches an unmerged pull request

OpenVDB is not installed from a release. The setup clones the upstream repository and then fetches a pull request as a local branch:

bash
git clone https://github.com/AcademySoftwareFoundation/openvdb.git
cd openvdb
git fetch origin pull/1808/head:feature/fvdb
git checkout feature/fvdb
rm fvdb/setup.py && cp ../assets/setup.py fvdb/
cd fvdb && pip install .

Two of those lines are doing unusual work. The fetch pulls a pull request head into a branch named feature/fvdb, so the dependency is an unmerged contribution rather than a tagged release, and a closed or rebased pull request leaves you with a branch this repository does not expect. The install then deletes the fVDB setup.py that arrived with the clone and copies a replacement out of this repository's own assets/ directory, so the dependency's packaging is governed by files here rather than by upstream.

The comment on that block is a hardware floor rather than a version note. fVDB is described as a 3D learning framework that requires a GPU later than Ampere. Mesh extraction is a separate build in ext/nksr-cuda, installed with python setup.py develop.

Linux only, and the Docker route is conda over someone else's image

The platform statement is one sentence: only Linux is currently supported, with support for other platforms welcomed. Everything below that line assumes a Linux toolchain, from the conda environment file to the CUDA extension build.

There is a Docker route, but it is a suggestion rather than a recipe. The advice is to take a base image from a fork of OpenVDB maintained by one of the authors, on its fw/fvdb branch, and apply the conda setup over it. So the image supplies the dependency and the repository supplies the environment, which is the same division as the manual install, arranged so you do not have to compile fVDB yourself.

The conda part has an optional accelerator. Three commands update conda in the base environment, install conda-libmamba-solver, and set the solver to libmamba, offered as a quality of life improvement rather than a requirement. The environment itself comes from environment.yml at the top of the tree and is activated as xcube, and the training entry points, train.py and test.py, sit beside it at the repository root.

Every sampling command starts with a bare none

The inference scripts have an argument shape worth reading before you change anything. Every documented invocation opens with a bare word:

bash
python inference/sample_shapenet.py none --category chair --total_len 20 --batch_len 4 --ema --use_ddim --ddim_step 100 --extract_mesh

That first positional is none, and it is where a conditioning mode would go. The ShapeNet script takes --category chair, car or plane, and the Waymo and Objaverse scripts pass the same none under their own names. Nothing in the released examples fills that position with anything else.

The remaining flags are the readable part. total_len 20 and batch_len 4 set how many steps run and how many run at a time, --ema selects the exponential moving average weights, --use_ddim with --ddim_step 100 selects the sampler and its step count, and --extract_mesh asks for meshes rather than voxels alone. Visualisation is a separate script: visualize_object.py for ShapeNet and Objaverse, visualize_scene.py for Waymo, each taking a results path and an id.

Mesh extraction moved out of the VAE and the refinement network is gone

The release notes state two deliberate differences between this code and the paper, and both are about what was left out.

The refinement network is omitted for cleaner code. The stated consequence is that results may vary slightly, and that the differences are not significant. The second change is structural: mesh extraction has been moved from the VAE to post-processing, so the autoencoder no longer emits the mesh and a later stage does.

That matters if you intend to compare numbers. A figure produced by a system that fuses refinement and extraction inside the VAE is not the same computation as this repository's, and the mesh you receive is now the output of a separate stage you can inspect on its own.

Beyond the scripts themselves, data preparation and what the text calls useful tricks live in MISC.md rather than in the main file. The repository also carries a datagen/ directory and an assets/ directory, and the latter is what the fVDB install copies its setup.py from, so it does double duty.

Two training stages, dense 16 cubed then sparse 128 cubed

Training is staged, and the stage names describe the voxel grid rather than the schedule. Stage 1, labelled coarse, trains an autoencoder per category on a 16x16x16 dense grid and then a latent diffusion model at the same resolution:

bash
python train.py ./configs/shapenet/chair/train_vae_16x16x16_dense.yaml --wname 16x16x16-kld-0.03_dim-16 --max_epochs 100 --cut_ratio 16 --gpus 8 --batch_size 32

Stage 2, labelled fine, moves to a 128x128x128 sparse grid, where the run name becomes 512_to_128-kld-1.0, the batch size drops from 32 to 8, gradient accumulation of 2 makes up the difference, and eight GPUs are still used. The diffusion stage in between uses batch size 8 with accumulation of 4 and evaluates every 5 epochs.

The naming is doing documentation work. A wname of 16x16x16-kld-0.03_dim-16 records the resolution, the KL weight and the latent dimension in one token, so a run is readable from its output name without opening the config. Every documented command goes through train.py at the repository root with the config path as its first argument, and the configs are grouped by dataset and category underneath configs/.

Waymo data and conditional sampling are both placeholders

Three capabilities in this release are unfinished, and they are the ones named in the abstract rather than the ones you run on the first day.

The helper that downloads the pretrained checkpoints is labelled temporarily unavailable, which leaves the documented route as fetching the checkpoints from Hugging Face and placing them under checkpoints yourself. The Waymo training data is listed as coming soon, so ShapeNet is the only dataset with a download link, on Hugging Face or Google Drive. And the Waymo single-scan condition is annotated coming soon inside the inference section, even though scene completion from a single scan is one of the tasks named in the abstract.

What does work is unconditional sampling on all three datasets, ShapeNet per category, Waymo and Objaverse, each with the same flag set described above.

The data layout has one sharp edge. ShapeNet is expected as an extracted folder at ../data/shapenet, which is outside the repository, or you change _shapenet_path in configs/shapenet/data.yaml. The default is therefore a relative path that requires moving the dataset rather than the checkout.

100 metres at 10 centimetres, with no test-time optimization

The scale claims are specific, and they are what the architecture exists for. The model generates millions of voxels with a finest effective resolution of up to 1024 cubed, in a feed-forward fashion, without time-consuming test-time optimization. That last clause is the contrast: the usual alternative in this area optimizes something per instance after sampling, and XCube does not.

The mechanism is a hierarchical voxel latent diffusion model that generates progressively higher resolution grids coarse to fine, using a custom framework built on the VDB data structure. VDB is the sparse voxel database, which is also the dependency you install from a pull request, so a data structure choice reaches all the way into the build instructions.

The scene-scale result is stated separately: large outdoor scenes at 100m by 100m with a voxel size as small as 10cm, which is what a 10cm grid over that footprint implies in voxel count. Grid attributes are arbitrary rather than fixed to occupancy, and that is what lets one framework carry the listed tasks of user-guided editing, single-scan scene completion and text-to-3D.

No releases, and the newest work points at two successors

There are no GitHub releases, so nothing is pinned and the version history is the commit log alone. The license field reads NOASSERTION while the tree carries a LICENSE.txt, and anyone planning commercial use is pointed at the NVIDIA Research Licensing form rather than at a self-serve grant, which is the arrangement to check first. Questions about the model are directed to two of the six authors instead of to the issue tracker.

The dates tell their own story. Code and model were released on 18 June 2024, the most recent news entry is dated 11 December 2024, and the last push to the default branch main is dated 18 June 2026. Between those two, the project twice points at work that is not itself: InfiniCube, which extends XCube to unbounded 3D generation, and SCube, a NeurIPS 2024 work extending XCube on large-scale scene reconstruction.

The repository is not archived and carries 21 open issues, so the code is present and the questions are still being read.

Editorial conclusion

XCube fits a researcher who needs large sparse 3D samples or scene-scale reconstructions and is prepared to build a GPU environment from source on Linux, accepting that some advertised modes are placeholders. It is a poor fit if you need macOS or Windows, a tagged release, or the conditional sampling paths, since the single-scan condition is marked coming soon. Before you start, check three things: whether you have a GPU later than Ampere, which is the stated floor for fVDB, whether the pull request the build fetches is still open, since your checkout depends on its head branch, and whether the results you need are unconditional sampling, given that the released code also omits the refinement network and moves mesh extraction to post-processing.

Frequently asked questions

What does XCube generate?

High-resolution sparse 3D voxel grids with arbitrary attributes, producing millions of voxels at a finest effective resolution of up to 1024 cubed in a feed-forward fashion without test-time optimization. It is also demonstrated on large outdoor scenes at 100m by 100m with voxels as small as 10cm.

Which platforms does XCube support?

Linux only at the moment, and the README states that support for other platforms is welcome. Docker users are pointed at a base image from an OpenVDB fork on its fw/fvdb branch and then told to apply the conda setup over it.

How do I install the fVDB dependency XCube needs?

Clone AcademySoftwareFoundation/openvdb, fetch pull request 1808 into a branch named feature/fvdb, check it out, delete the bundled fvdb/setup.py and copy the replacement from this repository's assets/ folder, then pip install from that directory. fVDB requires a GPU later than Ampere.

Why does the released XCube code differ from the paper?

Two differences are stated: the refinement network is omitted for cleaner code, which may cause slight variations in results, and mesh extraction has been moved from the VAE into post-processing.

How do I get the pretrained XCube checkpoints?

Download them from Hugging Face and place them under checkpoints. The convenience script inference/download_pretrain.py is labelled temporarily unavailable in the release.

Official sources

  1. Issues
  2. nv-tlabs/XCube on GitHub
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nv-tlabs-xcube.svg)](https://hysenlabs.com/projects/nv-tlabs-xcube)