# SimFoundry: turning a tabletop video into an OmniGibson scene

> NVIDIA's SimFoundry is a three-pipeline Python system that reconstructs a physics-ready simulation scene from a short real-world video, then augments it with object variants and task proposals. It is powerful and heavy: Linux, an NVIDIA GPU, gated model access and roughly 250 GB of disk.

**NVlabs/SimFoundry** — Modular and Automated Scene Generation for Policy Learning and Evaluation

- Repository: https://github.com/NVlabs/SimFoundry
- Stars: 401 · Forks: 29
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvlabs-simfoundry

## The gap SimFoundry fills between a phone video and a simulator

Building a manipulation scene by hand is slow. Someone models each object, assigns mass and friction, places it on a table, and writes a task description. Do that for fifty variations of a kitchen counter and the work outgrows the experiment. SimFoundry's answer is to start from footage. The README states that it "turns a short real-world video into a physics-ready simulation scene in under an hour, with no manual annotation required", and that pointing it at a tabletop yields segmentation, geometry reconstruction, textured meshes, and an OmniGibson scene with physical parameters, digital cousin variations and task proposals.

The intended user is a robotics researcher evaluating policies in simulation, not a game artist. The output is an OmniGibson scene, so the tool is only useful to people already inside that ecosystem or willing to enter it. The README also notes that example inputs live in docs/assets/example_videos/ and that capture tips are in scripts/pipeline/README.md, which tells you the input quality matters more than the tooling around it.

One honest framing: the under-an-hour figure is a claim in the README, not a benchmark. It also excludes the install, which the README describes as taking a while, and excludes the gated-model approval wait, which it warns can take time.

## Three pipelines, thirteen stages, and where the foundation models sit

The architecture is a staged pipeline rather than a service. Pipeline A, Reconstruction, runs 13 stages: video processing, depth estimation, ground segmentation, object decomposition, mesh generation, pose estimation, physics compilation, and USD/OmniGibson export. Pipeline B, Augmentation, generates digital cousin variations across geometry, topology and visual appearance, and proposes manipulation tasks. Pipeline C, Application, loads the result into OmniGibson for policy evaluation, teleoperation data collection, and smoke testing.

The modularity claim is specific rather than decorative: each stage is an independently swappable component, so replacing a segmentation or mesh model does not require rewriting the pipeline. The README repeats this as the design rationale, that as foundation models improve the pipeline improves with them. That is a real property of the layout, and it is also a maintenance burden, because the interfaces between stages become the thing you must keep stable.

Stage 7, mesh shape generation, is the memory bottleneck. The README states the default needs about 29 GiB, and that on a 24 GiB card you should pass s7_mesh.low_vram=true. Other streamed stages budget VRAM as a fraction of the card, 90% by default, which is why the same command is documented as working unchanged on 24 GiB and 96 GiB GPUs. The --max-vram-gb flag exists to pin an absolute cap instead.

## Installing SimFoundry and reconstructing your first scene

The requirements are not negotiable: Linux with an NVIDIA GPU and CUDA, Mamba or Conda with mamba, ffmpeg, roughly 250 GB of free disk space for a full install, a Hugging Face account, and either a Google Cloud project or a Gemini API key. The build step is a shell script, and the README notes it takes a while, or that you can hand it to an agent using docs/AGENT_INSTALL.md.

```bash
bash scripts/installation/install_everything.sh
```

Before that script is useful you need access to gated Hugging Face models: facebook/sam3, facebook/dinov3-vitl16-pretrain-lvd1689m and briaai/RMBG-2.0, with black-forest-labs/FLUX.1-Kontext-dev as optional. Approval can take time, so request it first. The VLM stages run on Google Cloud Vertex AI, which is why authentication is a separate step.

```bash
export GCLOUD_PROJECT=<your-gcp-project>
gcloud auth application-default login
hf auth login
```

If you have no GCP project, the README gives an alternative: generate a Gemini API key at AI Studio and run export GEMINI_API_KEY=<your-key> instead. There is also an interactive helper that covers all services at once, bash scripts/installation/login_services.sh. Checkpoints come next:

```bash
bash scripts/installation/download_checkpoints.sh --default
```

If you are already logged in to Hugging Face, the README says you can fold this into the first step with bash scripts/installation/install_everything.sh --checkpoints. Articulation support is a separate optional install, bash scripts/installation/install_articulate.sh.

The first real run reconstructs a scene from a video:

```bash
bash scripts/pipeline/A_reconstruction/run.sh \
  --scene-name my_scene \
  --video-fpath /path/to/video.mov
```

Expect a scene named my_scene in the output tree. On a 24 GiB card, add -- s7_mesh.low_vram=true. To include articulation decomposition, add --detect-articulation, which needs the optional articulate environments. Augmentation and smoke testing are separate scripts: B_augmentation/run.sh generates cousins and task proposals, and C_application/run.sh with --mode smoke-random loads the scene in OmniGibson. A unified dispatcher, scripts/pipeline/run.sh, takes A_reconstruction, B_augmentation or C_application as its first argument and forwards --help.

## The setup.py that installs nothing

Read setup.py before you plan a deployment. It declares name="simfoundry", version 0.1.0, python_requires='>=3.10', and an install_requires list that is empty. The package data it ships is narrow: configs/*.yaml and configs/modality/*.json, with a comment explaining that configs holds data rather than code. Everything else, torch, open3d, hydra-core, diffusers, transformers and the rest, lives in requirements.txt and the per-feature requirement files, requirements_dev.txt, requirements_hunyuan.txt and requirements_teleop.txt.

That split is a deliberate choice and it has consequences. pip install . from the repository gives you the Python package and its bundled configs, not a working pipeline. The conda environments built by the installation scripts are the supported path, and the README points to docs/INSTALL.md for the full details. If your team standardises on pip and a lockfile, you are working against the project's grain.

There is also a patches/ directory and a PATCH_PROVENANCE.md at the repository root, which suggests third-party code is patched in place rather than forked. The README does not document rollback for those patches, and neither does the material describe how they interact with upstream updates.

## Where SimFoundry is the wrong tool

The constraints are structural, not cosmetic. First, it is Linux and CUDA only. There is no CPU path described, and no macOS or Windows instructions in the README. Second, the disk figure is roughly 250 GB for a full install, which rules out most laptops and many shared workstations. Third, three of the core models are gated on Hugging Face, so a first run depends on approvals you do not control. Fourth, the VLM stages depend on Google Cloud Vertex AI or a Gemini API key, which means an external service sits inside your reconstruction loop; if that quota or key lapses, stages fail.

The output format is the sharpest boundary. SimFoundry produces OmniGibson scenes. If your evaluation harness is MuJoCo, PyBullet or a custom renderer, the reconstruction result is not directly usable, and the README does not describe exporters to other scene formats. Pipeline C exists precisely because the scene is meant to be consumed by OmniGibson.

The README's own news table is another limitation worth reading literally. V0 covers rigid-body and articulation generation. Automated background generation, and robotics data generation, training and evaluation, are listed as coming soon. So the augmentation and application surface is narrower today than the project description implies.

## How SimFoundry differs from hand-authored asset pipelines

The obvious alternative is the conventional route: author or buy 3D assets, assemble scenes in a simulator editor, and annotate physics by hand. That approach gives you exact control, reproducible assets under your own version control, and no external model dependencies. It also scales linearly with the number of scenes, and it cannot capture the specific objects on your own table.

A second alternative is single-model reconstruction tooling, where one mesh generator is pointed at an image and a mesh comes out. That is simpler to install and has no pipeline to maintain, but the result is a visual mesh, not a physics-ready scene. SimFoundry's contribution is the compilation step: physics parameters, sanity checking in a simulator, and the augmentation stage that produces digital cousins. The README frames the difference as modularity, that each stage is independently swappable, which is the opposite trade-off from a monolithic reconstruction tool: more moving parts, but each one replaceable.

The practical comparison is about where your time goes. With hand authoring it goes into asset creation. With a single-model tool it goes into fixing meshes and adding physics yourself. With SimFoundry it goes into installation, gated access, and keeping a thirteen-stage pipeline running.

## Maintenance, licensing and what to check before you depend on it

The last push to the repository was on 2026-08-27, and the repository is not archived. The news table shows an initial open-source release on 2026-08-14 and example scenes and assets on 2026-08-26, so the project is early and moving. There are no retrieved releases, which means no tagged version to pin against; setup.py carries version 0.1.0. If you adopt it, plan to track the main branch and to re-run the installation scripts when the environment definitions change, because the conda environments are part of the artifact, not just the Python code.

Licensing is Apache-2.0 for the project itself, per the LICENSE file and the SPDX header in setup.py. That is permissive, but it does not cover everything the pipeline downloads. The repository carries THIRD_PARTY_LICENSES.md, THIRD_PARTY_NOTICES.md and a third_party_notices/ directory, and the pipeline pulls gated models from Hugging Face under their own terms. The README does not state the licence of generated meshes or of assets derived from the digital cousin stage. Treat that as an open question to resolve with whoever handles licensing at your organisation rather than as settled by the Apache-2.0 badge.

## Conclusion

Adopt SimFoundry if you already run OmniGibson or Isaac-based simulation and you need scene variety that manual asset authoring cannot supply at the rate your policy evaluation demands, and if you can spare a Linux machine with an NVIDIA GPU, roughly 250 GB of disk, and the patience to clear three gated Hugging Face approvals before the first run. Do not adopt it if you need a scene in minutes, if you have no CUDA GPU, or if you want a single pip-installable library: setup.py declares no install_requires, so the Python package alone will not run the pipeline, and install.sh plus scripts/installation/install_everything.sh are the real entry points. Before committing, verify three things: that your GPU can hold the stage 7 mesh shape generation step, which the README says needs about 29 GiB unless you pass s7_mesh.low_vram=true; that your Google Cloud project or Gemini API key is accepted by the VLM stages; and that the license files in third_party_notices/ are compatible with how you intend to ship anything derived from the generated assets.

## FAQ

### Does SIM stand for simulation in SimFoundry?

The README does not expand the name or state what SIM stands for. It describes the project as turning a short real-world video into a physics-ready simulation scene, so the word simulation appears in the description, but no acronym is defined.

### How do I install SimFoundry?

The README's quick start runs bash scripts/installation/install_everything.sh, then authenticates with gcloud auth application-default login and hf auth login, then downloads checkpoints with bash scripts/installation/download_checkpoints.sh --default. It requires Linux with an NVIDIA GPU and CUDA, Mamba or Conda with mamba, ffmpeg, and roughly 250 GB of free disk space.

### What are the hardware requirements for SimFoundry?

The README lists Linux with an NVIDIA GPU and CUDA, Mamba or Conda with mamba, ffmpeg, and roughly 250 GB of free disk space for a full install. Streamed stages budget VRAM as a fraction of the card, 90% by default, so the documented commands work on a 24 GiB or a 96 GiB GPU, though stage 7 mesh shape generation needs about 29 GiB unless you pass s7_mesh.low_vram=true.

### Which models does SimFoundry need access to?

The README lists facebook/sam3, facebook/dinov3-vitl16-pretrain-lvd1689m and briaai/RMBG-2.0 as gated Hugging Face models you must request access to, with black-forest-labs/FLUX.1-Kontext-dev as optional. VLM stages run on Google Cloud Vertex AI, or you can use a Gemini API key instead.

## Sources

- [Issues](https://github.com/NVlabs/SimFoundry/issues)
- [License: Apache-2.0](https://github.com/NVlabs/SimFoundry/blob/main/LICENSE)
- [NVlabs/SimFoundry on GitHub](https://github.com/NVlabs/SimFoundry)
- [README](https://github.com/NVlabs/SimFoundry/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvlabs-simfoundry
