Model or dataset
manycore-research/SpatialLM avatar
manycore-research/SpatialLM

SpatialLM: structured indoor modeling from point clouds with a 3D language model

[NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling

4,744 stars401 forksPythonNOASSERTION

At a glance

What is it?
SpatialLM turns point clouds into walls, doors and oriented object boxes using a small language model. It is a research release with a heavy CUDA build, and the licence is not a standard open source one.
Who is it for?
Adopt SpatialLM if you already have point clouds in a z-up, axis-aligned frame and you want structured layout output without writing a detection pipeline yourself. Skip it if you need a permissively licensed component, a CPU-only path, or a supported product with an upgrade contract; the licence field is Llama3.2 and the repository is a research release.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 98 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap SpatialLM fills between raw point clouds and usable scene layout

A point cloud is a pile of coordinates. Almost nothing downstream of it wants coordinates. A robot planner wants walls and doors. A simulator wants oriented boxes with category labels. A dataset pipeline wants a text file it can diff. The usual route between those two ends is a hand-built stack of plane fitting, clustering and manual category rules, retuned for every sensor and every room. SpatialLM replaces that stack with a language model that emits the layout as text.

The README frames the target as embodied robotics, autonomous navigation and complex 3D scene analysis. That is a research audience, and the repository reads that way: train.py, eval.py, visualize.py, a configs directory, and separate finetuning instructions in FINETUNE.md. If you need a hosted API that returns a floor plan, this is not it. If you are building the pipeline that such a product would sit on, the pieces are all here.

How the point cloud becomes text: encoder, LLM, structured output

SpatialLM is a multimodal architecture with two halves. A point cloud encoder turns geometry into tokens the language model can consume, and a decoder-only LLM produces the structured scene description. The README describes the output as architectural elements (walls, doors, windows) plus oriented object bounding boxes with semantic categories.

The 1.1 release changed the encoder. According to the release news, SpatialLM1.1 incorporates the Sonata point cloud encoder and doubles the point cloud resolution relative to 1.0. That is why the install instructions split into two dependency groups: torchsparse belongs to the 1.0 line, while Sonata pulls in flash-attn, torch-scatter and spconv. The 1.1 models also support detection with user-specified categories, which the README attributes to the flexibility of the LLM rather than to a fixed classification head.

One constraint is stated plainly and is easy to miss. Input point clouds must be axis-aligned with the z-axis as the up axis. The README calls this orientation crucial for consistency across datasets. Nothing in the inference command corrects a tilted cloud, so if your reconstruction comes out in a camera frame, you own that alignment step before the model ever sees the file.

The repository ships four checkpoints. SpatialLM1.1-Llama-1B and SpatialLM1.1-Qwen-0.5B are the current pair, with SpatialLM1.0-Llama-1B and SpatialLM1.0-Qwen-0.5B still listed.

Installing SpatialLM and running inference on a test scene

The README states the tested environment: Python 3.11, PyTorch 2.4.1 and CUDA 12.4. The pyproject file is slightly wider, allowing Python between 3.10 and 3.12 and pinning torch to ^2.4.1+cu124 from the PyTorch wheel index. Poetry is the dependency manager, and the poe tasks wrap the awkward native builds.

Start by cloning and creating the conda environment with the CUDA toolkit, then install the Python dependencies:

bash
git clone https://github.com/manycore-research/SpatialLM.git
cd SpatialLM

conda create -n spatiallm python=3.11
conda activate spatiallm
conda install -y -c nvidia/label/cuda-12.4.0 cuda-toolkit conda-forge::sparsehash

pip install poetry && poetry config virtualenvs.create false --local
poetry install

Then pick the dependency group that matches the model you intend to run. The README warns that both builds take a while, and the pyproject tasks confirm why: install-torchsparse installs from the mit-han-lab GitHub repository, and install-sonata chains flash-attn, torch-scatter and spconv.

bash
# SpatialLM1.0 dependency
poe install-torchsparse

# SpatialLM1.1 dependency
poe install-sonata

For a first real run, pull one preprocessed scene from the testset rather than reconstructing your own cloud. The README points to point clouds reconstructed from RGB videos with MASt3R-SLAM, stored in the SpatialLM-Testset dataset on Hugging Face.

bash
huggingface-cli download manycore-research/SpatialLM-Testset pcd/scene0000_00.ply --repo-type dataset --local-dir .

Inference is a single script call with a point cloud path, an output path and a model identifier. The model_path argument accepts the Hugging Face repository name directly, so the weights download on first use.

bash
python inference.py --point_cloud pcd/scene0000_00.ply --output scene0000_00.txt --model_path manycore-research/SpatialLM1.1-Qwen-0.5B

What you should see is a text file named scene0000_00.txt containing the structured layout. The repository also ships visualize.py and depends on rerun-sdk, so the layout can be inspected alongside the geometry rather than read as raw coordinates.

Where SpatialLM breaks down or is the wrong choice

The orientation requirement is the first failure mode and the least forgiving. A cloud that is not axis-aligned with z up will produce layout output that is confidently wrong rather than obviously broken, because the model has no way to report that the input frame is off. There is no validation step in the documented inference command.

The second is the build. The 1.1 path requires compiling flash-attn, torch-scatter and spconv against a specific CUDA and PyTorch combination. The pyproject pins torch to ^2.4.1+cu124 and transformers to a narrow range between 4.41.2 and 4.46.1, with tokenizers capped below 0.20.4. That is a research environment, not a dependency set that will absorb an unrelated upgrade in your project. If your stack runs a newer PyTorch, you are choosing between this model and that upgrade.

The third is scope. SpatialLM is an indoor modeling model. The README's examples are rooms and architectural elements. Nothing in the repository suggests outdoor scenes, terrain or large-scale mapping are in scope, and the category detection in 1.1 is described in terms of indoor objects. Treat it as a room-scale tool.

Finally, the licence. The pyproject declares license = "Llama3.2" while the repository metadata reports NOASSERTION. Those two signals disagree about how the project should be classified, and the practical consequence is that you cannot treat this as a permissive dependency without reading LICENSE.txt and the model terms on Hugging Face yourself.

SpatialLM compared with training your own detector on point clouds

The realistic alternative is not another named project. It is the conventional route: take the same point cloud, run plane segmentation for walls and floors, cluster the residual points, fit oriented boxes, and train a small classifier for categories. That pipeline is deterministic, debuggable and cheap to run on a CPU.

The difference in approach is where the knowledge lives. A classical pipeline encodes your assumptions in code: thresholds, cluster sizes, plane tolerances, a category list. SpatialLM encodes them in weights. That is the trade. You get category flexibility without retraining because the LLM conditions on the categories you specify, which is exactly what SpatialLM1.1 advertises. You also get a model that can produce a layout for a room type you never wrote rules for.

In exchange, you give up determinism and a CPU path. The classical pipeline fails in ways you can trace to a line number. SpatialLM fails by emitting plausible text, and the README does not document a confidence signal or a way to reject a bad prediction. For a fixed sensor in a fixed building type, the classical route is often the better engineering choice. SpatialLM earns its place when the input varies and the category set changes.

Maintenance, upgrades and what the version split costs you

The last push to the repository was on 2026-06-26. The most recent tagged releases are v0.1.0 and v0.1.1, both dated 2025-06-10. So the codebase has moved since the last tag without a corresponding release, which means pinning to a tag and pinning to main are different decisions.

The upgrade cost is concentrated in the 1.0 to 1.1 transition. Moving to 1.1 means the Sonata encoder, the flash-attn and spconv builds, and double the point cloud resolution. It is not a version bump you apply casually; the dependency group changes, and the pyproject shows install-torchsparse and install-sonata as separate tasks precisely because the 1.0 and 1.1 lines do not share a build. If you already have a working 1.0 environment, staying there is a defensible position, since the 1.0 checkpoints remain listed.

On licensing, the two signals in the repository disagree, and the decision affects redistribution, not just internal use. Read LICENSE.txt and the model card terms on Hugging Face before you ship anything built on these weights. This is a factual caution about what the repository states, not legal advice.

Editorial conclusion

Adopt SpatialLM if you already have point clouds in a z-up, axis-aligned frame and you want structured layout output without writing a detection pipeline yourself. Skip it if you need a permissively licensed component, a CPU-only path, or a supported product with an upgrade contract; the licence field is Llama3.2 and the repository is a research release. Before committing, download the testset scene and run inference.py with the 0.5B model to confirm your point cloud orientation matches what the model expects.

Frequently asked questions

What is SpatialLM and what does it output?

SpatialLM is a 3D large language model that processes point cloud data and generates structured 3D scene understanding outputs. Those outputs include architectural elements such as walls, doors and windows, plus oriented object bounding boxes with semantic categories.

How do I install SpatialLM?

The README documents a conda environment with Python 3.11 and CUDA 12.4, then poetry install, then either poe install-torchsparse for the 1.0 line or poe install-sonata for the 1.1 line. The README notes that building the torchsparse wheel and the flash-attn wheel each take a while.

Which SpatialLM models are available on Hugging Face?

Four checkpoints are listed: SpatialLM1.1-Llama-1B, SpatialLM1.1-Qwen-0.5B, SpatialLM1.0-Llama-1B and SpatialLM1.0-Qwen-0.5B. The 1.1 models double the point cloud resolution and use the Sonata encoder.

What is the SpatialLM dataset and testset?

The SpatialLM-Dataset became available on Hugging Face in September 2025, and SpatialLM-Testset holds preprocessed point clouds reconstructed from RGB videos using MASt3R-SLAM. The README uses a testset scene as the example input for inference.

What orientation must the input point cloud have?

The README states that input point clouds are considered axis-aligned with the z-axis as the up axis, and calls this orientation crucial for consistency across datasets and applications. No correction step is documented in the inference command.

Official sources

  1. Issues
  2. manycore-research/SpatialLM on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/manycore-research-spatiallm.svg)](https://hysenlabs.com/projects/manycore-research-spatiallm)