Open-source project
microsoft/TRELLIS.2 avatar
microsoft/TRELLIS.2

microsoft/TRELLIS.2: image-to-3D generation with O-Voxel latents

Native and Compact Structured Latents for 3D Generation

11,324 stars1,363 forksPythonMIT

At a glance

What is it?
TRELLIS.2 is Microsoft's 4B-parameter image-to-3D model, released under MIT with a Linux-only setup script and a 24GB GPU requirement. It is a research release, and the README is honest about where it has and has not been run.
Who is it for?
Adopt TRELLIS.2 if you have a Linux workstation with an NVIDIA GPU of at least 24GB and your output is a textured mesh you can export, because the pipeline is a Python library and the O-Voxel conversion path is the part that is hard to replace elsewhere. Do not adopt it if you need Windows, CPU-only inference, or a service with an uptime commitment; the repository ships no server and the README tests Linux only.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 82 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What TRELLIS.2 generates, and who the 4B checkpoint is for

TRELLIS.2 is a 4B-parameter generative model for image-to-3D. You give it a picture, it produces a 3D asset with geometry and full PBR materials. The repository describes the output as high-resolution and fully textured, and it lists Base Color, Roughness, Metallic and Opacity as the surface attributes the model handles, which is a wider attribute set than color-only reconstruction. That last point matters for anyone whose downstream renderer expects a roughness map rather than a flat albedo.

The intended user is not someone who wants a web app. The repository is a Python package with a setup shell script, a training script, example scripts, and a Hugging Face Space for people who just want to try it. The README states the code is tested only on Linux, with an NVIDIA GPU of at least 24GB, verified on A100 and H100. That is a research-lab or studio workstation profile, not a laptop profile. If your only machine has an 8GB card, the model does not fit and no amount of configuration changes that.

The design target is topology. The README claims the O-Voxel representation handles open surfaces such as clothing and leaves, non-manifold geometry, and internal enclosed structures. Those three cases are exactly where iso-surface field extraction tends to fail, so the claim is specific rather than generic. Whether it holds on your input is something you have to check yourself, because the README gives no failure gallery.

How O-Voxel and the sparse VAE replace iso-surface fields

The mechanism described in the README is a sparse voxel structure called O-Voxel, which the project calls field-free. That phrase is the architectural claim: instead of learning a continuous occupancy or signed-distance field and extracting a surface at some threshold, the model works with voxel structure directly. The README says this avoids lossy conversion, which is the usual complaint about marching-cubes style extraction from a field.

On top of that sits a Sparse 3D VAE with 16x spatial downsampling, which compresses assets into a compact latent space. The generative backbone is described as vanilla DiTs, meaning the project did not need a bespoke transformer variant to reach its numbers. Two separate stages are implied by the timing table: shape and material are broken out separately at each resolution.

The processing side is where the architecture shows its practical value. The README states that converting a textured mesh to O-Voxel takes under 10 seconds on a single CPU, and converting O-Voxel back to a textured mesh takes under 100ms on CUDA. It describes both directions as rendering-free and optimization-free. That is a different proposition from fitting a field to a mesh, which is an optimization loop and can take minutes. The asymmetry is worth noting: encoding is CPU-bound and slow-ish, decoding is GPU-bound and fast. If your workflow repeatedly round-trips assets, the decode path is the one you will feel.

Installing TRELLIS.2 on Linux with setup.sh

Installation is a recursive clone followed by a single setup script. The recursive flag is not optional: the repository lists o-voxel/ as a top-level entry and carries a .gitmodules file, so the O-Voxel component arrives as a submodule.

bash
git clone -b main https://github.com/microsoft/TRELLIS.2.git --recursive
cd TRELLIS.2

The setup script takes flags rather than a single install mode. The README's recommended invocation creates a conda environment named trellis2 and pulls in the basic dependencies plus flash-attn, nvdiffrast, nvdiffrec, cumesh, o-voxel and flexgemm. Run it with a leading dot so it executes in your current shell, because it needs to activate the environment.

bash
. ./setup.sh --new-env --basic --flash-attn --nvdiffrast --nvdiffrec --cumesh --o-voxel --flexgemm

The script's own help output lists each flag individually, so you can drop the ones you do not need or install them one at a time if a build fails:

bash
. ./setup.sh --help

Three constraints from the README are easy to miss. The trellis2 environment defaults to PyTorch 2.6.0 with CUDA 12.4, and the recommended CUDA Toolkit is 12.4. If you have more than one toolkit installed, export CUDA_HOME before running setup, for example export CUDA_HOME=/usr/local/cuda-12.4. And if your GPU does not support flash-attn, such as a V100, install xformers manually and set ATTN_BACKEND to xformers. The README warns that installation takes a while because of the dependency count.

A first image-to-3D run from example.py

The repository ships example.py, which the README calls the minimal example. It is short and it shows the three moving parts: an environment map, the pipeline, and the input image. The first lines set two environment variables before importing anything else, and the ordering is deliberate. OPENCV_IO_ENABLE_OPENEXR must be set before cv2 is imported or the EXR read will fail, and PYTORCH_CUDA_ALLOC_CONF is set to expandable_segments:True, which the comment in the file says can save GPU memory.

python
import os
os.environ['OPENCV_IO_ENABLE_OPENEXR'] = '1'
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"  # Can save GPU memory
import cv2
import imageio
from PIL import Image
import torch
from trellis2.pipelines import Trellis2ImageTo3DPipeline
from trellis2.utils import render_utils
from trellis2.renderers import EnvMap
import o_voxel

The environment map is built from an EXR file under assets/hdri/. The example uses assets/hdri/forest.exr, converted from BGR to RGB and moved to CUDA as a float32 tensor:

python
envmap = EnvMap(torch.tensor(
    cv2.cvtColor(cv2.imread('assets/hdri/forest.exr', cv2.IMREAD_UNCHANGED), cv2.COLOR_BGR2RGB),
    dtype=torch.float32, device='cuda'
))

The pipeline is then loaded by name from Hugging Face and moved to the GPU. The checkpoint identifier is microsoft/TRELLIS.2-4B, and the README points to the model card on Hugging Face for details. The published example is truncated in the README after loading the input image, so the exact call that produces the mesh is not shown there; read example.py in the repository for the rest. The README's timing table gives an expectation for what you should see: roughly 3 seconds total at 512 cubed (2s shape plus 1s material), about 17 seconds at 1024 cubed, and about 60 seconds at 1536 cubed, measured on an H100.

python
pipeline = Trellis2ImageTo3DPipeline.from_pretrained("microsoft/TRELLIS.2-4B")
pipeline.cuda()

There is also a shape-conditioned texturing path, with example_texturing.py and app_texturing.py in the repository, and training code in train.py. The README roadmap marks image-to-3D inference, the 4B checkpoints, the Space demo, the texturing inference code and the training code as released.

Where TRELLIS.2 will not fit: Windows, small GPUs, and the demo path

The hardest limitation is stated plainly: the code is currently tested only on Linux. Not "recommended on Linux", tested only on Linux. If you are on Windows, the README gives you nothing, and the setup script is a bash script with conda assumptions. WSL is not mentioned in the README either way, so treat it as unverified rather than supported.

The second limitation is memory. At least 24GB of GPU memory is necessary, full stop. The example's expandable_segments setting is a memory-saving measure, not a way to fit a 4B model on a smaller card. Verified hardware is A100 and H100, which are data-center parts. A consumer card with enough VRAM may work, but the README does not claim it does.

The third is the attention backend. flash-attn is the default, and GPUs that do not support it need xformers installed manually plus ATTN_BACKEND set to xformers. This is a real fork in the install path, and getting it wrong produces a failure at runtime rather than at setup.

Fourth, there is no server in the repository. The top-level entries include app.py and app_texturing.py, which are Gradio-style demo applications, but nothing that looks like a deployment target. If you need an HTTP API with concurrency limits and a health check, you are building that yourself on top of Trellis2ImageTo3DPipeline. And the README does not document rollback or version pinning for the checkpoint, so reproducibility across model updates is on you.

TRELLIS.2 versus field-based reconstruction pipelines

The obvious comparison is with the earlier TRELLIS line and with field-based pipelines generally. The README frames the difference as field-free versus iso-surface extraction. In a field-based pipeline you train or fit a continuous function, then extract geometry at a threshold, and that extraction is where thin structures, open surfaces and internal cavities get lost or merged. TRELLIS.2's claim is that working in O-Voxel space avoids the lossy conversion step entirely, and that the same representation carries material attributes rather than only geometry.

The practical difference shows up in the conversion timings. Under 10 seconds on a single CPU for mesh to O-Voxel, under 100ms on CUDA for O-Voxel back to mesh, both rendering-free and optimization-free. A fitting-based approach is an optimization loop, so the cost scales with how hard the asset is to fit rather than being roughly constant. If your pipeline ingests many assets, that difference compounds.

The trade-off is ecosystem. A field representation is a familiar interchange format with many consumers; O-Voxel is specific to this project, and the o-voxel submodule is the only implementation referenced here. Adopting TRELLIS.2 means adopting its representation at the boundaries of your pipeline, not just its generator. If you only need the mesh at the end, that cost is small. If you wanted to hand the intermediate representation to other tools, check whether o-voxel exposes what you need before you commit.

Licence, maintenance and what an upgrade costs

The repository is MIT licensed, with the LICENSE file at the top level and the badge in the README. MIT is permissive: it allows commercial use and modification with attribution and no warranty. That covers the code in the repository. It does not automatically cover the pretrained weights, which live on Hugging Face under microsoft/TRELLIS.2-4B with their own model card. The README directs you to that model card for details, so read it separately before shipping anything. This is a description of where the terms live, not legal advice.

The repository is not archived, and the last push was on 2026-07-10. The roadmap shows every listed item checked off: paper, image-to-3D inference code, 4B checkpoints, the Space demo, texturing inference code, and training code. There are no retrieved releases, so there is no tagged version to pin against. That is the upgrade cost in one sentence: without release tags, tracking upstream means tracking the main branch and the Hugging Face checkpoint identifier, and the README does not document rollback. If you need a frozen artifact, clone at a commit and record it.

Setup cost is front-loaded and large. The README warns the install may take a while due to the dependency count and suggests installing flags one at a time if something breaks. Budget an afternoon for a first install on a fresh machine, and treat the CUDA_HOME and ATTN_BACKEND settings as part of your environment documentation rather than one-off shell history.

Editorial conclusion

Adopt TRELLIS.2 if you have a Linux workstation with an NVIDIA GPU of at least 24GB and your output is a textured mesh you can export, because the pipeline is a Python library and the O-Voxel conversion path is the part that is hard to replace elsewhere. Do not adopt it if you need Windows, CPU-only inference, or a service with an uptime commitment; the repository ships no server and the README tests Linux only. Before you commit, verify three things in this order: that CUDA_HOME resolves to the toolkit version you intend, that the attention backend matches your GPU generation (flash-attn by default, xformers via ATTN_BACKEND otherwise), and that the Hugging Face checkpoint microsoft/TRELLIS.2-4B fits your card. If the third check fails, nothing else in the setup matters.

Frequently asked questions

Is TRELLIS.2 from Microsoft?

Yes. The repository is microsoft/TRELLIS.2 and the pretrained checkpoint is published as microsoft/TRELLIS.2-4B on Hugging Face.

Can you run TRELLIS.2 locally?

Yes, but the README states the code is tested only on Linux and requires an NVIDIA GPU with at least 24GB of memory, verified on A100 and H100. Installation is a recursive clone followed by the setup.sh script.

How do you install TRELLIS.2?

Clone the repository with the recursive flag, then run the setup script with flags for the components you need. The README's example creates a conda environment named trellis2 with flash-attn, nvdiffrast, nvdiffrec, cumesh, o-voxel and flexgemm, defaulting to PyTorch 2.6.0 with CUDA 12.4.

How much does TRELLIS.2 cost?

The repository is MIT licensed and the README links the 4B checkpoint on Hugging Face, so there is no listed price for the code or weights. Your real cost is hardware: an NVIDIA GPU with at least 24GB of memory, plus the time the README says the dependency install takes.

How do you use TRELLIS.2 for image-to-3D generation?

The repository's example.py is the minimal path: set the OpenEXR and CUDA allocator environment variables before importing, build an EnvMap from an EXR under assets/hdri/, load Trellis2ImageTo3DPipeline.from_pretrained with microsoft/TRELLIS.2-4B, call pipeline.cuda(), then pass your image. The README's example is truncated after loading the image, so read the full script for the generation call.

Official sources

  1. Issues
  2. License: MIT
  3. microsoft/TRELLIS.2 on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-trellis-2.svg)](https://hysenlabs.com/projects/microsoft-trellis-2)