# TRELLIS: Microsoft's 3D Asset Generation Model from Text or Image Prompts

> TRELLIS is a large 3D asset generation model from Microsoft Research, presented at CVPR 2025 as a Spotlight paper, that converts text or image prompts into 3D assets in multiple output formats including Radiance Fields, 3D Gaussians, and meshes. It uses a unified Structured LATent (SLAT) representation and Rectified Flow Transformers as the generation backbone.

**microsoft/TRELLIS** — Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).

- Repository: https://github.com/microsoft/TRELLIS
- Website: https://trellis3d.github.io
- Stars: 13,739 · Forks: 1,349
- Language: Python
- License: MIT
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-trellis

## What TRELLIS generates and what makes it distinct

Most 3D generation tools produce a single output format: a mesh, a point cloud, or a neural radiance field. TRELLIS generates all of them from one generation pass. The README states that TRELLIS takes in text or image prompts and generates high-quality 3D assets in various formats including Radiance Fields, 3D Gaussians, and meshes.

The key architectural feature is the SLAT representation: Structured LATent. SLAT is a unified 3D representation that can be decoded into multiple output formats without re-running the generation. This means a single generation pass produces a SLAT, and the renderer can produce a mesh, Gaussian splatting representation, or Radiance Field from that single latent without starting over.

TRELLIS also supports local 3D editing: generating variants of the same object or modifying specific regions of a generated 3D asset. The README describes this as flexible output format selection and local 3D editing capabilities which were not offered by previous models at this scale.

The model was presented at CVPR 2025 as a Spotlight paper. The paper is available at arxiv.org/abs/2412.01506 and an interactive demo is hosted on Hugging Face Spaces at Microsoft/TRELLIS.

## The SLAT representation and generation architecture

The SLAT (Structured LATent) representation is the core technical contribution that enables TRELLIS's multi-format output. Traditional 3D generation models encode geometry in format-specific ways, which means changing the output format requires a different model or re-generation. SLAT encodes geometry in a format-agnostic structured space that different decoders can read.

The generation backbone is Rectified Flow Transformers. Rectified Flow is a continuous normalizing flow training method that trains the model to generate samples by following straight trajectories in latent space, which tends to produce more stable training and high-quality samples compared to diffusion models with curved trajectories. The README describes the Rectified Flow Transformers as tailored for SLAT as the powerful backbone of the generation system.

The pre-trained models are available on Hugging Face. The README lists the primary model as TRELLIS-image-large with 1.2 billion parameters. The project trained on a dataset of 500K diverse objects called TRELLIS-500K, released with the repository. The README notes that models up to 2 billion parameters are provided.

The text-conditioned model (TRELLIS-text) was released in March 2025. The README notes explicitly that it is always recommended to generate 3D assets by first generating images using text-to-image models and then using the TRELLIS-image model. Text-conditioned models are described as less creative and detailed due to data limitations.

## Installing TRELLIS: hardware requirements and the setup script

The hardware requirements are specific: the README requires an NVIDIA GPU with at least 16GB of memory and states that the code has been verified on NVIDIA A100 and A6000 GPUs. On GPUs without flash-attn support (like the NVIDIA V100), the ATTN_BACKEND environment variable must be set to xformers.

The software prerequisites are Linux (Windows is not fully tested), CUDA Toolkit 11.8 or 12.2, Conda for dependency management, and Python 3.8 or higher.

Clone the repository with submodules:

```sh
git clone --recurse-submodules https://github.com/microsoft/TRELLIS.git
cd TRELLIS
```

Install dependencies using the setup script:

```sh
. ./setup.sh --new-env --basic --xformers --flash-attn --diffoctreerast --spconv --mipgaussian --kaolin --nvdiffrast
```

The --new-env flag creates a new conda environment named `trellis`. Without it, the script installs into the current environment. The --flash-attn flag installs flash attention support; remove it for GPUs that do not support flash-attn and set ATTN_BACKEND=xformers instead. The README notes the installation may take a while due to the large number of dependencies.

If CUDA Toolkit is installed in multiple versions, set PATH to the correct version before running setup.sh. For example, `export PATH=/usr/local/cuda-11.8/bin:$PATH` for CUDA 11.8.

## Using TRELLIS: the Gradio apps and the example scripts

The repository provides both Gradio web apps and Python example scripts for running inference. The main image-to-3D Gradio app is app.py. The text-to-3D app is app_text.py.

For scripted use, example.py shows the image-to-3D workflow. The README notes that Gaussian export was added to app.py in December 2024. Multi-image conditioning (using multiple reference images to constrain the generation) is available in example_multi_image.py; the README describes it as a tuning-free algorithm that may not give the best results for all input images.

Variant generation (producing multiple versions of the same object with controlled variation) is demonstrated in example_variant.py. This was released along with the TRELLIS-text model in March 2025.

The TRELLIS-500K dataset and toolkits for data preparation are in the dataset_toolkits/ directory, with the DATASET.md file covering the data format. Training code is available in train.py for researchers who want to fine-tune or re-train the model.

## Output formats and their use cases

TRELLIS decodes SLAT into three output formats, each suited for different downstream applications.

Radiance Fields (Neural Radiance Fields or NeRF-style outputs) are suitable for photorealistic rendering and novel view synthesis, but they are not directly editable in standard 3D tools. They represent the scene volumetrically rather than as a surface mesh.

3D Gaussians (the 3D Gaussian Splatting format) are an explicit scene representation optimized for real-time rendering from novel viewpoints. They render fast and produce high-quality images, but they are also not a traditional polygon mesh and require specialized viewers or renderers.

Meshes are the most universally compatible format: they work in 3D modeling software (Blender, Maya, 3ds Max), game engines, and 3D printing workflows. The mesh output from TRELLIS may require post-processing (decimation, unwrapping, or cleanup) before use in a production pipeline.

Because TRELLIS generates all three from the same SLAT, a practitioner can generate once and export in the format most appropriate for their downstream tool without re-running the model.

## TRELLIS versus Shap-E: the difference in approach

Shap-E, released by OpenAI, is an earlier text-and-image-conditioned 3D generation model. It generates 3D objects represented as implicit functions (NeRF) or mesh outputs. Like TRELLIS, it accepts both text and image conditioning.

The key difference is architectural: Shap-E generates directly in output-specific latent spaces, while TRELLIS uses the SLAT intermediate representation to enable flexible multi-format decoding. TRELLIS also operates at a larger scale (up to 2 billion parameters) compared to Shap-E. The CVPR 2025 paper describes TRELLIS as significantly surpassing existing methods at similar scales, based on the paper's own evaluation.

Practically, TRELLIS's multi-format output (mesh, Gaussian, Radiance Field from one generation) is a concrete advantage over Shap-E's single-format generation for teams that need the same 3D asset in multiple representations. Both models run on GPU and require non-trivial hardware.

## License, maintenance, and scope of the repository

TRELLIS is MIT licensed. The last push to the repository was on 2026-06-26. The repository is not archived, though there are no GitHub releases.

The training code release in March 2025 and the TRELLIS-500K dataset are significant additions for the research community. A team that wants to train a domain-specific variant (for example, a model trained on architectural objects or biological structures) has the full training pipeline available in the repository.

Windows support is not guaranteed. The README references a community issue tracking Windows setup steps but notes they are not fully tested. Linux is the tested platform. ARM-based systems and non-NVIDIA GPUs are not mentioned as supported.

## Conclusion

TRELLIS is appropriate for researchers and practitioners who need a high-quality 3D asset generation model from image prompts, have access to an NVIDIA GPU with at least 16GB VRAM, and can run Linux. The README recommends using image-to-3D rather than text-to-3D for better results: generate an image first with a text-to-image model, then run TRELLIS on that image. Text-conditioned models are described in the README as less creative and detailed due to data limitations. Windows is not fully tested. The last push to the repository was on 2026-06-26.

## FAQ

### What is TRELLIS's SLAT representation?

SLAT stands for Structured LATent, a unified 3D intermediate representation that TRELLIS uses during generation. SLAT can be decoded into multiple output formats (Radiance Fields, 3D Gaussians, and meshes) from a single generation pass, without re-running the model for each format.

### Can TRELLIS run on Windows?

The README states the code is currently tested only on Linux. A community issue in the repository tracks partial Windows setup steps, but the README notes those steps are not fully tested. Linux with a CUDA-capable NVIDIA GPU is the supported configuration.

### How large are the TRELLIS pretrained models?

TRELLIS-image-large has 1.2 billion parameters and is the primary image-to-3D model. The README states that pretrained models up to 2 billion parameters are available on Hugging Face. The models were trained on the TRELLIS-500K dataset of 500,000 diverse 3D objects.

## Sources

- [Issues](https://github.com/microsoft/TRELLIS/issues)
- [License: MIT](https://github.com/microsoft/TRELLIS/blob/main/LICENSE)
- [microsoft/TRELLIS on GitHub](https://github.com/microsoft/TRELLIS)
- [Project website](https://trellis3d.github.io)
- [README](https://github.com/microsoft/TRELLIS/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-trellis
