Open-source project
wgsxm/PartCrafter avatar
wgsxm/PartCrafter

PartCrafter: two inference scripts, one flag that calls out, and a requirements line that stops mid sentence

[NeurIPS 2025] PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers

2,486 stars165 forksPythonMIT

At a glance

What is it?
The NeurIPS 2025 mesh generation model ships two inference entry points, two optional flags that send the input photo to an external model, a pinned CUDA 12.4 wheel set tested on one GPU, and a system requirements line that never finishes.
Who is it for?
PartCrafter is a research release rather than a product: the code is MIT licensed, the checkpoints live on Hugging Face, and the documentation is written around the machine its authors tested.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 174 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The object path pulls two weight sets, the scene path pulls one

Object generation runs through `scripts/inference_partcrafter.py`, and the first run fetches two sets of weights on its own. The PartCrafter model comes from the `wgsxm/PartCrafter` repository on Hugging Face and lands in `pretrained_weights/PartCrafter`. The second download comes from a different account entirely, `briaai/RMBG-1.4`, and lands in `pretrained_weights/RMBG-1.4`, which is the background removal step rather than part of the generative model.

Scene generation is a separate script, `scripts/inference_partcrafter_scene.py`, and it pulls exactly one set, PartCrafter-Scene from `wgsxm/PartCrafter-Scene`. So the object path depends on a second publisher's upload and the scene path does not, which matters if you are counting what has to stay reachable.

The repository itself has no GitHub releases at all. The paper sits on arXiv as 2506.05573, the weights sit on Hugging Face, the demo is a Hugging Face Space credited to someone other than the authors, and the project page is a separate site. Nothing inside the tree carries a version number to pin.

The scene asset note describes a prefix its own example does not use

For the object examples the docs explain the filename convention: names begin with the recommended part count, `np3` meaning three parts, and the sample images come from Objaverse and ABO in `./assets/images`. The scene section repeats that sentence word for word, same `np3` example, while saying its images come from 3D-Front and live in `./assets/images_scene`.

The one scene command on display does not match the rule stated beside it:

code
python scripts/inference_partcrafter_scene.py \
  --image_path assets/images_scene/np6_0192a842-531c-419a-923e-28db4add8656_DiningRoom-31158.png \
  --num_parts 6 --tag dining_room --render

The asset is prefixed `np6` and the flag asks for six parts, one part more than the sentence two lines above describes. The object path has no such mismatch, since `assets/images/np3_2f6ab901c5a84ed6bbdf85a67b22a2ee.png` is passed with `--num_parts 3`. Output directories follow `--tag` in both cases, `./results/robot` and `./results/dining_room`.

Two optional flags put the input photo into an API request

The base object command needs no credentials:

code
python scripts/inference_partcrafter.py \
  --image_path assets/images/np3_2f6ab901c5a84ed6bbdf85a67b22a2ee.png \
  --num_parts 3 --tag robot --render

The other two features are prefixed with a key. `--part_suggest` defaults to the model `gemini-3-flash-preview`, sends the image to a VLM that proposes a part count, and can be repointed with `--part_provider gemini --part_model gemini-3-flash-preview`. `--style_transfer` defaults to `gemini-3.1-flash-image-preview` and exists because the model was trained on rendered Objaverse images, so a real photograph has to be converted to that look before it is fed in. Both can be combined on one photo with `--part_suggest --style_transfer --rmbg --render`.

Two details are worth planning around. The stylized intermediate is written to disk as `styled_input.png` in the output directory, so the converted image stays behind as an artifact. And the provider layer is an open extension point: a new file in `src/utils/providers/` implementing `suggest_num_parts()` and/or `stylize_for_objaverse()` is the whole documented cost of adding another backend. Note that `--rmbg` itself is local, using the downloaded RMBG weights.

The wheel set is pinned to one CUDA build and one tested machine

The stack is stated as `torch-2.5.1+cu124` and `python-3.11`, with a note that other versions should also work. The commands pin those exact versions:

code
conda create -n partcrafter python=3.11.13
conda activate partcrafter
pip install torch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 --index-url https://download.pytorch.org/whl/cu124

Then the repository is cloned and one shell script does the rest, `bash settings/setup.sh`. Graphics libraries for a conda environment without root access are handled separately with `conda install -c conda-forge libegl libglu pyopengl`.

The platform story is narrow and honest about it. The installation is described as tested on Debian 12 with NVIDIA H20 GPUs, and Windows is not covered by the main branch at all: users are pointed at pull request 24 and at a contributor's fork on a `windows-main` branch, with thanks given for that contribution. The three CUDA 12.4 wheels, the Python patch level and the single GPU model are the entire platform contract on offer.

The system requirements line opens and does not finish

The last section in the file begins with the words `A CUDA-enabled GPU with at` and ends there. The card it was about to name, the memory figure it was about to give, and whatever followed the minimum specification are all absent from the visible text.

That gap sits directly under the installation instructions, which is the wrong place for it, because the rest of the document gives a precise hardware story in a different register: one distribution, one GPU model, one CUDA build. A reader who wants to know whether their card qualifies has three facts to work from, the pinned `cu124` wheel index, the mention of NVIDIA H20 as the tested device, and the fragment above. Nothing in the file converts those into a requirement.

The same shape appears elsewhere. The paper is dated to arXiv on 2025-06-09 and the project is described as fully open sourced on 2025-07-13, while the data question is still open in the checklist. Precise dates and version numbers for the code, silence on the hardware floor.

The checklist is four fifths done and the news stops in September 2025

Five tracked items, four checked. Released are the inference scripts, the training code with its data preprocessing scripts, the pretrained checkpoints at both object and scene level, and the hosted demo. Unchecked is the preprocessed dataset. Training code and the scripts that build training data are therefore public while the data they consume is not, which narrows what a third party can reproduce without assembling the source corpora themselves.

The news entries run on a tight clock: arXiv on 2025-06-09, full open sourcing on 2025-07-13, a Windows guide in a fork on 2025-07-20, the scene version trained on 3D-Front on 2025-07-23, a Hugging Face demo on 2025-08-15, and NeurIPS 2025 acceptance on 2025-09-18. The last entry is the acceptance. The most recent commit on the default branch is dated 2026-04-16, so roughly seven months of work sit past the end of the news list with nothing added to it, including anything that would close the dataset item.

Every documented choice is a flag, though the tree carries a configs directory

The top level holds `assets/`, `configs/`, `datasets/`, `scripts/`, `settings/` and `src/` beside the README and the license file. The quick start then names exactly one path in that layout, `settings/setup.sh`, and routes everything else through command line flags: `--image_path`, `--num_parts`, `--tag`, `--render`, `--rmbg`, `--part_suggest`, `--style_transfer`, `--part_provider`, `--part_model`, `--style_provider`, `--style_model`. No configuration file is described anywhere in the document.

So a `configs/` directory sits in the tree while every documented knob is a flag, and nothing explains what is in it. The same pattern appears with `datasets/`: the folder exists, the checklist still marks the preprocessed dataset as unreleased, and the file listing gives no breakdown of what the folder holds.

That combination, a tree laid out for configuration and data that the text does not describe, is what a reader has to resolve on their own before trusting the layout. It also means reproducing a run means remembering the flag combination, since nothing in the docs saves one.

Support is one address, and the author list uses two undefined markers

The contact line offers a single address, `[email protected]`, or an issue, and that is the whole support surface. No discussion forum, no group, no maintainer list appears in the file.

The author header lists seven names and decorates some of them with markers, an asterisk after the first two and a dagger after the third, without a key anywhere in the document saying what either symbol means. Equal contribution and corresponding author are the usual readings of that notation, but the file does not say so, and the affiliations themselves are only implied by the personal page links.

Elsewhere the credit is split cleanly. The Hugging Face demo is credited to a third party, the Windows fork to another, and the news entries thank both by name on the dates they landed. The project page, the arXiv entry, the two model repositories and the video are the remaining pointers, all external to the repository.

Editorial conclusion

PartCrafter is a research release rather than a product: the code is MIT licensed, the checkpoints live on Hugging Face, and the documentation is written around the machine its authors tested. Anyone planning a run should confirm three things first, the GPU model and memory the requirements line was going to name, whether the preprocessed dataset gap matters for their use, and how comfortable they are with `--part_suggest` and `--style_transfer` handing the input photo to an external API. It suits evaluation on a CUDA box, not unattended pipelines.

Frequently asked questions

how to use partcrafter

Two entry points exist. Objects run through scripts/inference_partcrafter.py with --image_path and --num_parts, and scenes run through scripts/inference_partcrafter_scene.py. Required weights download automatically on the first run, and output goes to ./results under the value passed to --tag.

is partcrafter free

The repository carries an MIT LICENSE file at its top level. The paper, the project page, the two model repositories on Hugging Face and the hosted demo are all linked from the file, and the repository itself publishes no GitHub releases.

Does PartCrafter need an API key to run?

Not for the base commands. Both the part suggestion example and the style transfer example are prefixed with GEMINI_API_KEY=your_key, so a key is only needed when --part_suggest or --style_transfer is passed.

Can PartCrafter be installed on Windows?

The main branch does not document a Windows install. The docs point Windows users at pull request 24 and at a contributor's fork that keeps a windows-main branch, credited to that contributor.

What GPU does PartCrafter need to run?

The system requirements section opens with A CUDA-enabled GPU with at and stops there, naming no card and no memory figure. The only hardware mentioned anywhere is an NVIDIA H20, on Debian 12, as the tested configuration.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. wgsxm/PartCrafter on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/wgsxm-partcrafter.svg)](https://hysenlabs.com/projects/wgsxm-partcrafter)