# LiteReality-Agent needs an iPhone scan, two logged-in agent CLIs, and gates that fail open loudly

> LiteReality-Agent turns a RoomPlan capture from an iOS scanner into a MuJoCo-ready room, staged as a deterministic scene init followed by an agentic authoring loop. The dependency file carries its own post-mortem about two quality gates that used to pass when they had no chance to.

**LiteReality/LiteReality-Agent** — LiteReality-Agent: Turn the real world into simulation-ready environments.

- Repository: https://github.com/LiteReality/LiteReality-Agent
- Website: https://litereality.github.io/agent/
- Stars: 546 · Forks: 52
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/litereality-litereality-agent

## The input is an iPhone RoomPlan capture, not a video file

The pipeline does not accept footage you already have. It starts from the LiteReality Scanner app on the App Store, where one walkthrough captures the RGB frames, the depth and the RoomPlan `room.usdz`, and you upload that capture to the machine you will run on. Each scan directory is a specific shape: `frame_*.jpg` plus `frame_*.json` plus `room.usdz`. Whatever a general camera rig or a phone video would have produced does not fit, and no other input format appears anywhere in the project.

For people without a scan, the project points at a separate repository of example rooms and gives you the exact two commands:

```bash
git clone https://github.com/LiteReality/example-scans.git
uv run litereality run example-scans/<scan>
```

That clone is also the fastest way to see the tool work, since it removes both the scanner and the capture step. Everything downstream, from detection to rendering, assumes the capture is already inside `LR_SCANS_DIR`.

## BLENDER_PATH takes the install folder, not the binary

The one configuration value that is called out three times is `BLENDER_PATH`, and each time the note says it points at the directory containing the `blender` binary, not the binary itself. The example value is `/Applications/Blender.app/Contents/MacOS`, which is exactly the kind of path where the binary sits one level down, so a value that looks complete fails in a way that is hard to read. The requirement is Blender 5.x, tested on 5.1.

The platform requirement is narrower than the tooling suggests. The tested configurations are macOS on Apple Silicon and Linux with a GPU of at least 24 GB. Beyond that, the stated prerequisites are `uv`, an image-generation API key, and the agent CLIs described below. On Linux the GPU number belongs to the local model path; on the hosted path described next, the machine running the pipeline never loads a GPU library at all.

## Two agent CLIs split the work, and both must already be signed in

Logged-in agent CLIs on `PATH` are a requirement, not a suggestion. `codex` with GPT-6 is the default for room authoring, materials, object refinement and independent visual review. `claude` remains the default for reconstruction reasoning and procedural asset initialization. Two vendors on two halves of one reconstruction means two subscriptions, two logins kept alive, and a specific division of labour that the project can override per role through `models.env`.

The fourth prerequisite is billed separately: an image-generation API key for reference images, `OPENAI_API_KEY` by default or `GEMINI_API_KEY` with `LR_IMAGE_PROVIDER=gemini`, at a stated cost typically under $1 per scene. So a reconstruction carries a model subscription on each side plus a small per-scene image bill. The image provider choice has to agree with `LR_IMAGE_PROVIDER` in `models.env`, which the env example repeats, and the runtime dependencies carry `anthropic`, `openai`, `google-genai` and `claude-agent-sdk` to match.

## The heavy models run on Modal, and setup deploys a workspace

TRELLIS and GroundingDINO need a GPU, so by default they run hosted on Modal. The install is two commands:

```bash
uv sync --frozen --extra modal --group dev
cp .env.example .env
```

The env file then takes a Modal token pair, and `uv run litereality setup` deploys the apps, once per workspace. `SANITY_DEEP=1 uv run python sanity.py` is the readiness check that follows. The env example states the payoff plainly: filling in the tokens selects hosted TRELLIS for image to 3D plus hosted GroundingDINO and DINOv2, so the pipeline process never loads torch and needs no NVIDIA GPU. No Hugging Face account or token is required, because the weights are not fetched locally.

The naming is overridable, which matters if a workspace already uses those app names: `MODAL_TRELLIS_APP`, `MODAL_TRELLIS_FUNCTION`, `MODAL_DINO_APP`, `MODAL_DINO_FUNCTION` and `MODAL_ENVIRONMENT` all have defaults. A named profile in `~/.modal.toml` from `modal setup` works in place of the token pair, and `TRELLIS_PYTHON` points at a local GPU TRELLIS interpreter on Linux with NVIDIA hardware, used only when no Modal credentials are set. The detection runtime has its own `auto` and `local` choice.

## Two physics dependencies are mandatory because their gates fail open

The dependency list explains itself in comments that read like incident reports. `python-fcl` is marked NOT optional because the publish path runs `room_qc.correct` on every default run and only records a warning when it exits non-zero, so a missing FCL turned the deterministic clash gate into a silent no-op. `coacd` is marked NOT optional for the same reason with a worse failure mode: without it every concave link falls back to its own convex hull and the gate still reports a clean pass.

Both comments end with the same argument for accepting the weight. Prebuilt wheels exist for every interpreter and platform the project supports, so the cost is a few megabytes and no compiler, and the alternative is shipping a gate that quietly does nothing. A collision gate that reports success because its collision library is absent is worse than a missing gate, and a convex-hull fallback that passes the solver gate hides exactly the concave geometry the articulated assets were generated to have.

The truncated tail of the second comment is worth noting as a project fact: the sentence explaining the `coacd` failure does not finish on the published page.

## Scene init is deterministic, and authoring is the half that loops

A reconstruction has two stages. Scene init turns the capture into a seed room and is deterministic. Authoring takes that seed, compares it against the capture, and edits until it matches, and that half is agentic. Either can run alone, and the installed CLI is named as the only supported pipeline entry point:

```bash
uv run litereality run  scans/<scan>
```

Scene init alone is `--through seed`. Authoring alone is `uv run litereality stage author run/<scan> --force --polish --live`. `--polish` layers object refinement and materials on top of plain authoring. The model-driven quality pass is a separate flag, `--quality-pass`, for a specific reason: it is the longest agent pass on a run and nothing downstream reads its output. `--live` shows the build in real time next to the agent's trace, and the viewer starts before the room exists, waits for it, and prints its url again once the first build lands.

## The MuJoCo export writes one scene file, and --shake measures what moved

The finished room becomes a MuJoCo scene made of bodies rather than one baked mesh, each body carrying the mass, inertia, friction, colliders and joints its generated asset was compiled and solver-gated with. Two commands produce and check it:

```bash
uv run litereality stage simulate run/<scan>            # -> realism_authoring/mujoco/scene.xml
uv run litereality stage simulate run/<scan> --shake     # ...and measure what actually moves
```

`--shake` is the honest half: it perturbs the scene and measures what actually moves, which is the only way to catch an object whose physics were never compiled. The write-up of where each number comes from, and of what happens to an object with no compiled physics, lives in `doc/Sim-Ready-intergration/Mujoco.md`, whose path carries that spelling.

The general-purpose outputs sit alongside it: `room_preview/Room.glb` with materials baked and clips intact for Blender, Unity, Unreal or the web, `room_preview/Room.blend` as a Blender scene, and `room/` for the editable room itself. `uv run litereality view run/<scan>` opens the viewer.

## Everything simulation-related landed in three weeks, at version 0.0.1

The change log is four entries. On 2026-08-01 came LiteReality-Agent 0.0, turning room scans into interactive 3D. On 2026-09-07 layout agents settled the noisy layout of a raw scan into a collision-free layout before anything was built from it. On 2026-09-09 articulated objects began carrying their own mass, inertia, colliders and joints following Articraft, exporting to URDF and MJCF. On 2026-09-12 every reconstructed room began loading as a MuJoCo scene with objects as rigid bodies and the whole room exporting to URDF and MJCF. So the simulation path is three weeks old.

The project file says version 0.0.1 and classifies itself as Development Status 3, Alpha, and the repository publishes no releases, so there is no tag to pin and the only stable reference is a commit. The last push was on 2026-09-29. Python support is `>=3.10,<3.13`, so 3.13 and 3.14 are refused at install time while the classifiers stop at 3.12.

## Conclusion

Judge it as research code at version 0.0.1 with no published releases, so pin a commit rather than a tag. It suits anyone who already has scans from the LiteReality Scanner app, agent CLIs signed in on PATH, a Blender 5.x install and a per-scene image API budget. Before trusting a reconstruction, check that python-fcl and coacd actually installed, since both quality gates report success when they are missing.

## FAQ

### What does LiteReality-Agent need as input?

A capture made with the LiteReality Scanner app, where one walkthrough records the RGB frames, the depth and the RoomPlan room.usdz. Each scan directory holds frame_*.jpg, frame_*.json and room.usdz, and LR_SCANS_DIR points at the folder holding them. Example rooms can be cloned from a separate example-scans repository.

### Does LiteReality-Agent need a local GPU?

Not on the default path. TRELLIS and GroundingDINO run hosted on Modal, and the env example states that filling in the token pair means the pipeline process never loads torch and needs no NVIDIA GPU, so an Apple Silicon Mac is enough. The local alternative needs Linux with a GPU of at least 24 GB.

### Which agent CLIs does LiteReality-Agent drive?

codex with GPT-6 is the default for room authoring, materials, object refinement and independent visual review, while claude stays the default for reconstruction reasoning and procedural asset initialization. Per-role overrides live in models.env, and both CLIs have to be logged in on your PATH.

### What does the --shake flag do in LiteReality-Agent?

It runs the simulate stage and measures what actually moves, complementing the default behaviour of writing realism_authoring/mujoco/scene.xml. The write-up of where each physics number comes from is in doc/Sim-Ready-intergration/Mujoco.md.

### Is there a released version of LiteReality-Agent?

No release is published. The project file says version 0.0.1 with a Development Status of Alpha, the news entries label the first release LiteReality-Agent 0.0, and Python support is capped at >=3.10,<3.13.

### Why are python-fcl and coacd required by LiteReality-Agent?

Both are marked not optional because the quality gates around them fail open. A missing python-fcl turns the deterministic clash gate into a silent no-op, and without coacd every concave link falls back to its own convex hull while the gate still reports a clean pass. Prebuilt wheels keep the cost to a few megabytes.

## Sources

- [Issues](https://github.com/LiteReality/LiteReality-Agent/issues)
- [License: Apache-2.0](https://github.com/LiteReality/LiteReality-Agent/blob/main/LICENSE)
- [LiteReality/LiteReality-Agent on GitHub](https://github.com/LiteReality/LiteReality-Agent)
- [Project website](https://litereality.github.io/agent/)
- [README](https://github.com/LiteReality/LiteReality-Agent/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/litereality-litereality-agent
