# image-to-3d-pipeline: run four open-source image-to-3D models on one input and score the results

> A JavaScript pipeline that reconstructs meshes from AI-generated renders with TRELLIS, TripoSR, stable-fast-3d and Hunyuan3D-2, then ranks them through a fixed Blender inspection protocol. The ranking is only as good as the protocol behind it.

**dreamers-laboratory/image-to-3d-pipeline** — Reconstruct 3D meshes from images with several open-source models and score which did it best.

- Repository: https://github.com/dreamers-laboratory/image-to-3d-pipeline
- Stars: 278 · Forks: 22
- Language: JavaScript
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/dreamers-laboratory-image-to-3d-pipeline

## The choice problem this pipeline is built around

Single-image-to-3D reconstruction has a dozen open-source models and no obvious winner. That sentence from the README is the whole premise. If you have one concept image and four candidate reconstruction tools, the usual workflow is to try one, look at the result, and either accept it or try the next one. Nothing about that process is comparable, because each run uses a different input, a different camera, and a different person's judgement of what looks right.

The pipeline targets a narrow audience: engineers and technical artists who are already comfortable running Python inference scripts and Blender headless, and who need to justify a model choice to someone else. It is not aimed at someone who wants a hosted upload-and-download service. The repository is explicit that all source imagery is AI-generated synthetic renders of a fictional submersible, so the worked example is a controlled object with clean edges rather than a scanned real-world subject.

## Multiview reasoning happens before any geometry

The mechanism that makes the comparison fair is upstream of the models. According to the README, the pipeline generates eight viewpoints from one concept image first: port profile, three-quarter, bow-on, stern-on and the rest. Every reconstruction model in the repository starts from that same sheet, which is why their outputs can be placed side by side at all.

From there the stages are numbered and mostly separate scripts. tools/preprocess_object_masks.py produces transparent object-only masks from the source renders. tools/run_trellis_multiview.py runs multi-image TRELLIS, and the runners for the other candidates live in 07-experiments/. Evaluation is where the design gets opinionated: tools/run_reconstruction_eval.sh renders each GLB through the same fixed Blender orbit and builds contact sheets for a 0-2 visual rubric documented in 07-experiments/HUNYUAN_2MV_EXPERIMENT_PROTOCOL.md. The final stage loads the chosen GLB into 05-web-explorer/, a Three.js scene with BVH collision.

The README names the six moves as mask, infer, bake, render, reject, export. Reject is a stage, not an afterthought, which is the part most reconstruction demos leave out.

## Installing it and running the first reconstruction

The repository does not ship model weights or environments. The README states that large artifacts are excluded: Python environments, model checkpoints, intermediate PLY and GLB outputs, and the explorer's full runtime asset set. So the first step is fetching the third-party tools at the exact commits used in the evaluation.

```bash
./setup.sh
```

That script clones four reconstruction models into tools/vendor/ at pinned commits. They are cloned rather than vendored because each carries its own license, which the README says you accept when you fetch it. Expect four separate dependency trees to satisfy afterward; the README does not describe a single unified environment for all four.

If you want to see output before installing anything, the examples/ directory contains three sample source renders and one reconstructed GLB, submersible-v2-stochastic.glb, about 1.5 MB. The README recommends inspecting it without running the pipeline. For the explorer specifically, 05-web-explorer/README.md describes how to supply the runtime assets that are not committed, and examples/ provides a mesh you can use for that.

```bash
cd 05-web-explorer
```

The README points to that directory for local run instructions rather than reproducing them at the top level, so the exact dev-server command lives there, not here. What you should see once assets are in place is the chosen mesh inside a navigable underwater scene with BVH collision, plus the six-stage build strip described in the README's screenshot caption.

## What the 0-2 rubric can and cannot tell you

The evaluation scores meshes on a 0-2 visual rubric after rendering every candidate through the same fixed Blender orbit. Fixing the orbit is the right call: it removes camera angle as a variable, and it means a mesh cannot look better simply because it was photographed from a flattering side. The contact sheets in 06-evaluation/ exist so a human can compare candidates under identical conditions.

A 0-2 scale is coarse. Three levels cannot express the difference between a mesh with a slightly noisy surface and one with a collapsed thin feature, unless the rubric document defines those cases precisely. Whether it does is something you should read 07-experiments/HUNYUAN_2MV_EXPERIMENT_PROTOCOL.md to determine. The README does not summarize the rubric's level definitions.

The deeper limit is stated by the project itself, in the caption under the mesh viewer screenshot: a depth map estimates distance for visible pixels and cannot reveal what the camera never saw. That is not a bug in any of the four models. It is the boundary of the method, and it means the back of your object is inferred, not measured. A ranking produced here tells you which model hallucinates more plausibly, not which one is geometrically correct.

## When this is the wrong tool

If you need metric accuracy, stop. Nothing in the pipeline as described calibrates scale against a known reference, and the evaluation is a visual rubric rather than a measurement against ground truth geometry. A mesh that scores well here can still be the wrong size and the wrong curvature.

If your inputs are real photographs, the pipeline's own example does not cover you. The source imagery is AI-generated synthetic renders with clean backgrounds, which is exactly the condition that makes object masking in tools/preprocess_object_masks.py tractable. Real photos bring shadows, reflections, clutter and inconsistent exposure, and the README does not document how the preprocessing handles any of that.

There is also a licensing trap that has nothing to do with mesh quality. Two of the four tools are not permissively licensed. stable-fast-3d is under the Stability AI Community License, and Hunyuan3D-2 is under the Tencent Hunyuan community license, which the README describes as territorially restricted. Reviewing the ranking and then shipping the winning mesh may mean shipping under a license you did not intend to accept. The README tells you to review each license before use, particularly those two, and that is the correct instruction rather than a formality.

## Against a single-model workflow

The obvious alternative is to pick one reconstruction tool and build around it. TripoSR, for instance, is a single-image reconstruction model under the MIT license, and its own repository documents how to run it directly. If you choose that route, you get one dependency tree, one license, one inference script, and no evaluation harness.

The difference in approach is not the models, it is the protocol. A single-model workflow answers the question can I get a mesh. This pipeline answers the question which of these four meshes is better, and it answers it by fixing the input sheet, fixing the render orbit, and scoring against a written rubric. That costs you four sets of weights, four environments, and Blender in the loop.

The trade is only worth making if the choice itself matters to you. If you have already decided on TRELLIS and you just need meshes, the evaluation stages are overhead. If you are writing a recommendation that someone will challenge, the contact sheets and the fixed orbit are the evidence.

## Maintenance, licensing and what the repository excludes

The repository is not archived, and the last push was on 2026-09-02. That is recent enough that the pinned commits in setup.sh should still resolve, but the pinned-commit approach means the pipeline will not silently pick up upstream changes to any of the four models. You upgrade by editing setup.sh, not by running a package manager. Whether that is a feature or a burden depends on whether you want reproducibility or upstream fixes.

On licensing, the split is clean and worth restating. The code in this repository is Apache-2.0, and the third-party tools fetched by setup.sh are excluded from that grant and remain under their own licenses. Apache-2.0 on the pipeline does not extend to the meshes you produce with the non-permissive tools, and it does not extend to the tools themselves. That is a fact about the file layout, not legal advice; the specific terms of the Stability AI and Tencent licenses are outside this repository and outside this review.

The practical maintenance cost sits in the excluded artifacts. Python environments, checkpoints and intermediate outputs are all yours to provision, and 05-web-explorer/README.md is the only place the runtime asset requirement is documented. A clone of this repository does not run end to end without that work.

## Conclusion

Adopt it if you already have a small set of clean, object-only render sheets and you want a defensible ranking rather than a single model's output, because the fixed Blender orbit and the 0-2 visual rubric are the part most single-model demos skip. Do not adopt it if you need metric geometry, if your inputs are real photographs with uncontrolled lighting, or if you plan to ship stable-fast-3d or Hunyuan3D-2 commercially, since those two carry their own licenses with commercial and territorial restrictions. Verify three things first: that setup.sh clones all four tools at the pinned commits on your platform, that Blender is available to run tools/run_reconstruction_eval.sh, and that you can supply the explorer's runtime assets per 05-web-explorer/README.md, which the repository deliberately excludes.

## FAQ

### What is a 3D pipeline?

In this repository it is a fixed sequence of stages that starts with source renders and ends with a mesh in a browser: mask, infer, bake, render, reject, export. Each stage is a separate script or directory, so the output of one feeds the next rather than being fused into a single tool.

### How to use image to 3D model in image-to-3d-pipeline?

Run ./setup.sh to clone the four reconstruction tools at pinned commits, preprocess your renders with tools/preprocess_object_masks.py, then run tools/run_trellis_multiview.py or the runners in 07-experiments/ and score the results with tools/run_reconstruction_eval.sh. The README notes that model weights and environments are not included and must be provisioned separately.

### Can ChatGPT convert image to 3D model like image-to-3d-pipeline does?

The repository does not discuss ChatGPT or any hosted conversational service. Its reconstruction step runs local open-source models, TRELLIS, TripoSR, stable-fast-3d and Hunyuan3D-2, cloned by setup.sh at pinned commits.

### What does a 3D image mean in the context of image-to-3d-pipeline?

Here the input is not a single 3D image but a sheet of eight generated viewpoints derived from one concept image: port profile, three-quarter, bow-on, stern-on and others. Every reconstruction model starts from that same sheet so their outputs are comparable.

## Sources

- [dreamers-laboratory/image-to-3d-pipeline on GitHub](https://github.com/dreamers-laboratory/image-to-3d-pipeline)
- [Issues](https://github.com/dreamers-laboratory/image-to-3d-pipeline/issues)
- [License: Apache-2.0](https://github.com/dreamers-laboratory/image-to-3d-pipeline/blob/main/LICENSE)
- [README](https://github.com/dreamers-laboratory/image-to-3d-pipeline/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dreamers-laboratory-image-to-3d-pipeline
