MeshFlow: artistic mesh generation with MeshVAE and a flow-matching DiT
Repository for the CVPR 2026 paper MeshFlow Efficient Artistic Mesh Generation via MeshVAE and Flow-based Diffusion Transformer by Weiyu Li, Antoine Toisoul, Tom Monnier, Roman Shapovalov, Rakesh Ranjan, Ping Tan and Andrea Vedaldi.
At a glance
- What is it?
- MeshFlow is Meta AI and HKUST's CVPR 2026 research release for turning an input mesh or point cloud, optionally with a reference image, into a new artistic mesh. It ships inference scripts, a Gradio app and a checkpoint bundle, but no training code and a licence that is not a standard SPDX identifier.
- Who is it for?
- Adopt MeshFlow if you have a CUDA GPU, a mesh or point cloud to condition on, and you want to reproduce or build on a CVPR 2026 method rather than ship a product. Skip it if you need training code, a permissive standard licence, or CPU-only inference, because the repository ships inference scripts only and its licence is NOASSERTION.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 28 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What MeshFlow generates, and from what input
MeshFlow is the code release for the CVPR 2026 paper "MeshFlow: Efficient Artistic Mesh Generation with MeshVAE and Flow-based Diffusion Transformer", listed in the README as a Highlight. The task it targets is narrow and specific: given an existing mesh or point cloud, and optionally a reference image, produce a new mesh that looks hand-modelled rather than like a smoothed scan. The README describes the result as "artist-like meshes" produced in about one second.
The input geometry is not decoration. It is what the README calls the "RoPE geometry condition": the pipeline samples the input surface and encodes its spatial structure into rotary position embeddings for the transformer. That means MeshFlow is not a text-to-3D generator. If you have no geometry to start from, the pipeline has nothing to condition on. The reference image is the optional half of the signal, and it only steers the result when you supply it.
The intended audience is researchers and technical artists working on generative geometry. The repository is a research release: it contains inference_vae.py, inference_dit.py, gradio_app.py and evaluate.py, plus the meshflow package. There is no training script in the top-level listing, so reproducing the training run is not something this repository supports.
MeshVAE, the DiT denoiser and the DINOv3 image path
The README splits the system into four named modules, and the split explains the data flow. MeshFlowVAE encodes mesh topology into continuous latents and decodes vertices, normals and adjacency back out. MeshFlowDiT performs flow matching on those latents, using voxel RoPE and optional image cross-attention. DINOv3Encoder produces the visual tokens for image conditioning. MeshFlowPipeline wires the three together: surface sampling, then flow matching, then VAE decode.
The practical consequence of a VAE latent stage is that generation happens in a compressed space, and the decoder is what determines whether the output mesh has usable connectivity. The README lists adjacency among the decoded outputs, which is why the pipeline can return a mesh rather than a point cloud.
Image conditioning is where the dependency graph gets awkward. The README states that model.pth does not bundle DINOv3 weights, and that the visual encoder backbone is loaded from a local DINOv3 hub checkout on first reference-image use. So a mesh-only or point-cloud-only run does not touch DINOv3 at all, while any run with --ref_image or a Gradio image upload does. DINOv3 weights are gated: you clone facebookresearch/dinov3 into ~/.cache/torch/hub/facebookresearch_dinov3_main, then request access through Meta's download page and fetch the checkpoint with wget rather than a browser. That approval step is a real gate on the image-conditioned path, not a formality.
Installing MeshFlow and running a first generation
The README's Quick Start is three commands: clone the repository, enter it, install requirements.txt. The requirements file pins torch==2.8.0 and torchvision==0.22.1 against the PyTorch cu128 wheel index, so it assumes a CUDA 12.8 build unless you change the index URL as the file's own comment suggests.
git clone https://github.com/facebookresearch/meshflow.git
cd meshflow
pip install -r requirements.txtInstalling dependencies is not enough. The README says to download the MeshFlow checkpoint bundle and place it under ckpt/meshflow/, which must contain config.yaml and model.pth. The repository does not ship those files, so the directory has to be created and populated before any script will load a model.
mkdir -p ckpt/meshflow
# download config.yaml and model.pth into ckpt/meshflow/The first real use is the Python pipeline. The README's example loads the bundle from ckpt/meshflow, passes an input .ply as the RoPE geometry condition, leaves image as None, and exports a .glb. Note the README's own annotation: guidance_scale is only effective when image is provided, so in this mesh-only call the value 2.5 does nothing.
from meshflow.pipelines import MeshFlowPipeline
pipeline = MeshFlowPipeline.from_pretrained(
"ckpt/meshflow",
device="cuda",
dtype="fp16",
)
mesh = pipeline.run(
mesh="path/to/input.ply",
image=None,
steps=28,
guidance_scale=2.5,
seed=42,
)
mesh.to_trimesh().export("output.glb")If you prefer a browser, the README also documents a local Gradio app. Omit --model_path and it uses a local ckpt/meshflow/ when present, otherwise it downloads config.yaml and model.pth from the Hugging Face repository into ~/.cache/meshflow/. torch.compile is on by default on CUDA, and --no-compile turns it off.
python gradio_app.py --gpu 0 --dtype fp16
python gradio_app.py --model_path ckpt/meshflow --num_verts 4096The --num_verts slider only has an effect when the model config sets denoiser_model.use_proj_cond_on_temb to true, per the README, and even then it only roughly controls the generated vertex count. There is also a hosted demo at facebook/meshflow on Hugging Face Spaces if you want to see the output before installing anything.
Where MeshFlow is the wrong tool
The first limitation is structural: this repository is inference only. The top-level entries are evaluation and inference scripts plus the meshflow package, and the README documents no training entry point. If your goal is to train MeshVAE or the DiT on your own mesh corpus, MeshFlow gives you the architecture to read and the checkpoints to run, not the pipeline to reproduce. That is normal for a paper release, and it should be priced in before you plan around it.
The second is the licence. The repository metadata reports NOASSERTION, which means the licence could not be mapped to a standard SPDX identifier. The LICENSE file exists at the repository root, but nothing in the README summarises its terms. For a research group that is a minor annoyance; for anyone embedding the model in a product it is the first thing to resolve, and the README will not resolve it for you.
The third is hardware and dependency weight. The example calls device="cuda" and dtype="fp16", the requirements pin CUDA 12.8 wheels, and torch.compile is enabled by default in the Gradio app. There is no documented CPU path. Add the gated DINOv3 download on top, and the image-conditioned workflow has an external approval step that the mesh-only workflow does not.
Finally, MeshFlow conditions on geometry. If your input is a text prompt or a single photograph with no accompanying mesh or point cloud, this is not the model to reach for; the RoPE condition is the core of how it works, not an optional extra.
How MeshFlow differs from TRELLIS-style image-to-3D pipelines
The obvious comparison is with image-to-3D generators in the TRELLIS family, which take one or more images and produce a mesh with no geometry input required. That is a different contract. Those systems solve the harder problem of inferring shape from appearance alone, and they pay for it in fidelity control: you cannot tell them what silhouette to keep.
MeshFlow inverts the setup. The input mesh or point cloud carries the geometry through surface sampling and voxel RoPE, and the reference image, when present, is cross-attended as a visual condition. The README's guidance_scale flag is described as CFG on visual conditioning and is only effective when --ref_image is set, which confirms the image is a conditioning signal layered on top of geometry rather than the primary driver.
That makes MeshFlow closer in spirit to a retopology or stylisation tool than to a generator from nothing. If you already have a base mesh and want a variant with different surface character, the geometry-conditioned design is an advantage. If you have only a photograph, TRELLIS-style pipelines start where MeshFlow cannot. The two are complements, not substitutes, and the choice follows from what you already hold.
Maintenance, dependencies and what a licence of NOASSERTION means for you
The repository is not archived, and the last push was on 2026-09-02, which is recent. There are no retrieved releases, so there is no tagged version to pin against; you are tracking the main branch. For a paper release that is typical, and it means an upgrade is a git pull plus a re-read of requirements.txt rather than a version bump.
Upgrade cost concentrates in the pinned dependencies. torch==2.8.0 and torchvision==0.22.1 are tied to the cu128 wheel index, and the file's own comment tells you to change the index URL for your CUDA or PyTorch build. diffusers==0.38.0 and transformers==5.0.0rc3 are pinned tightly enough that a major bump in either could break the pipeline. The transformers pin is a release candidate, which is worth noting if your environment refuses pre-release versions. point-cloud-utils==0.34.0 is needed only by evaluate.py, and gradio==6.6.0 plus plotly==6.5.2 only by gradio_app.py, so a lean inference install can skip those.
On licensing: NOASSERTION is a signal to read LICENSE yourself, not a summary. I am not giving legal advice, and the README does not state terms. What can be said concretely is that the checkpoint bundle is hosted separately from the code, so code and weights may carry different terms, and the DINOv3 weights come from a separate gated download with its own agreement. Three artefacts, potentially three sets of conditions.
Editorial conclusion
Adopt MeshFlow if you have a CUDA GPU, a mesh or point cloud to condition on, and you want to reproduce or build on a CVPR 2026 method rather than ship a product. Skip it if you need training code, a permissive standard licence, or CPU-only inference, because the repository ships inference scripts only and its licence is NOASSERTION. Before you commit, read LICENSE in the repository root, confirm the checkpoint bundle downloads into ckpt/meshflow/, and check that your PyTorch build matches the cu128 wheels pinned in requirements.txt.
Frequently asked questions
What is MeshFlow and what does it generate?
MeshFlow is the code release for the CVPR 2026 paper on efficient artistic mesh generation, built from MeshVAE and a flow-matching diffusion transformer. It takes an input mesh or point cloud, optionally with a reference image, and generates a new artist-like mesh, which the README describes as taking about one second.
How do I install MeshFlow?
Clone the repository, install requirements.txt, then download the checkpoint bundle so that ckpt/meshflow/ contains config.yaml and model.pth. The README's Quick Start uses pip install -r requirements.txt, and the pinned torch wheels target CUDA 12.8 unless you change the index URL.
Does MeshFlow need DINOv3 for image conditioning?
Yes, but only when you use a reference image. The README states that mesh or point-cloud-only inference does not load DINOv3, and that model.pth does not bundle DINOv3 weights; the visual encoder is loaded from a local DINOv3 hub checkout on first reference-image use, and access to those weights is granted through Meta's download page.
Can I run MeshFlow without a GPU?
The README does not document a CPU path. Its examples pass device="cuda" and dtype="fp16", the requirements pin CUDA 12.8 wheels, and the Gradio app enables torch.compile by default on CUDA.
Does MeshFlow include training code?
The repository's top-level entries are inference scripts, a Gradio app, evaluate.py and the meshflow package, and the README documents no training entry point. It is an inference release with pretrained checkpoints.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookresearch-meshflow)