manhattan_sdf: Manhattan-world SDF reconstruction from ScanNet scenes
Code for "Neural 3D Scene Reconstruction with the Manhattan-world Assumption" CVPR 2022 Oral
At a glance
- What is it?
- manhattan_sdf is the CVPR 2022 reference implementation of a neural SDF that constrains indoor rooms to Manhattan-world planes. It ships a Conda environment, a YAML config per scene, and three commands for training, mesh extraction and evaluation.
- Who is it for?
- Adopt manhattan_sdf if you are reproducing the CVPR 2022 paper, comparing against COLMAP, ACMP, NeRF, UNISURF, NeuS or VolSDF on ScanNet, or studying how a Manhattan-world prior changes surface extraction in a room. Do not adopt it as a general-purpose reconstruction service: the README documents one dataset, one config per scene, and no packaging, versioned release or API.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What manhattan_sdf assumes about the room it reconstructs
The project implements the CVPR 2022 Oral paper "Neural 3D Scene Reconstruction with the Manhattan-world Assumption" by Guo, Peng, Lin, Wang, Zhang, Bao and Zhou. The assumption is in the title: indoor scenes are largely built from mutually perpendicular planes, so floors, walls and ceilings can be treated as axis-aligned structure rather than free-form geometry. A signed distance field fitted without that constraint tends to round off wall-floor junctions and blur flat surfaces, because the network has no reason to prefer a plane over a slightly curved surface. manhattan_sdf adds the prior so flat regions stay flat.
The audience is narrow and academic. The repository is a research codebase from the ZJU 3D Vision group, not a product. It expects you to have ScanNet scene data, a CUDA-capable GPU, and the patience to edit YAML. If you want a reconstruction tool with a user interface and a support channel, this is the wrong repository. If you are writing a paper that compares neural surface reconstruction methods on indoor scans, the authors have already done part of your work: the README points to docs/RESULTS.md, which holds their quantitative and qualitative results plus the trained ScanNet models, and a baseline table covering COLMAP, COLMAP*, ACMP, NeRF, UNISURF, NeuS and VolSDF. That baseline dump is arguably the most reusable artefact in the repository for anyone who is not reproducing the method itself.
How the code is laid out: configs, lib, and two entry points
The top level contains run.py and train_net.py, plus configs/, lib/, docs/, assets/ and environment.yml. Training and evaluation are separated: train_net.py drives optimisation, while run.py dispatches other task types selected with a --type flag, which is why mesh extraction and evaluation are invoked through run.py rather than a dedicated script.
Configuration is per scene. The README's examples all reference configs/scannet/0050.yaml, and the data preparation step tells you the extracted ScanNet data must sit under data/ with a path consistent with line 38 of that config file. That single detail explains most first-run failures: the config hard-codes where it expects the scene, so a directory named differently or extracted one level too deep will not be found. The exp_name argument in each command ties a run to an output directory, and the same exp_name is passed to training, mesh extraction and evaluation so that the later steps read the checkpoint written by the first. If you change exp_name between commands, extraction will look for a run that does not exist.
The lib/ directory holds the model and data code, and the README's acknowledgement section credits VolSDF, neurecon, COLMAP and a customised COLMAP implementation used as a submodule in NerfingMVS. That lineage is worth knowing before you read the source: the SDF formulation descends from VolSDF, and the camera and point-cloud side of the pipeline is COLMAP-derived rather than written from scratch here.
Installing manhattan_sdf and running one ScanNet scene
The README gives a two-command setup. The environment is defined in environment.yml at the repository root, so the Conda file is the only dependency manifest you need to read before starting.
conda env create -f environment.yml
conda activate manhattanAfter activation, the environment is named manhattan. The README does not list the Python or CUDA versions separately, so environment.yml is the source of truth for what will be installed.
Data comes from a Google Drive folder linked in the README, containing the processed ScanNet scene data evaluated in the paper. Extract it into data/ and make sure the resulting path matches line 38 of configs/scannet/0050.yaml. The README also points to docs/CUSTOM.md for running on your own captures, which is the document to read if your data is not one of the released ScanNet scenes.
With the scene in place, training is one command:
python train_net.py --cfg_file configs/scannet/0050.yaml gpus 0, exp_name scannet_0050The gpus argument takes a comma-separated device list, and exp_name names the run. Note the trailing comma after the device list in the README's example; it is part of the documented invocation, so copy the line rather than retyping it.
Once training finishes, extract a mesh:
python run.py --type mesh_extract --output_mesh result.obj --cfg_file configs/scannet/0050.yaml gpus 0, exp_name scannet_0050This writes an OBJ file to the path given by --output_mesh, reading the checkpoint from the matching exp_name. Evaluation uses the same pattern with a different task type:
python run.py --type evaluate --cfg_file configs/scannet/0050.yaml gpus 0, exp_name scannet_0050The README does not state expected runtime, memory use or output quality, so treat the first run as an experiment rather than a benchmark.
Where manhattan_sdf is the wrong tool
The Manhattan-world prior is a constraint, and constraints fail on data that violates them. A room with a slanted ceiling, a curved staircase, a spiral wall or heavy non-planar clutter is exactly the case the method was not designed for, and the README offers no fallback mode that disables the assumption. If your scenes are outdoor, or indoor but not box-like, the prior is working against you.
The second limitation is scope. There is no packaged release: the repository's recent releases are empty, so installation means cloning and building the Conda environment from environment.yml. There is no Python package on an index, no CLI installed to your PATH, and no version pin to depend on. Every command is run from the repository root with relative config paths, which makes the code awkward to embed in a larger pipeline.
Third, the configuration is per scene. Each ScanNet scene in the paper has its own YAML under configs/scannet/, and the data path is hard-coded inside it. Scaling to a new dataset means writing configs and, per docs/CUSTOM.md, preparing data in the expected format. The README does not document a batch mode that sweeps a directory of scenes, so orchestration is your problem.
Finally, the repository is a paper artefact. The README does not document rollback, checkpoint compatibility between code revisions, or what happens if you resume a run with a changed config. The last push to main was on 2026-09-15, so the code is not abandoned, but there are no tagged releases to anchor a reproducible build against.
manhattan_sdf against COLMAP and the neural baselines it ships results for
The obvious non-neural alternative is COLMAP, which the README credits and which appears in the baseline table in docs/RESULTS.md alongside COLMAP*, ACMP, NeRF, UNISURF, NeuS and VolSDF. The difference is not a matter of one being newer. COLMAP is a multi-view stereo pipeline: it matches features across images, estimates poses, and produces a dense point cloud or mesh from photometric consistency. It needs no training run and no GPU-heavy optimisation, and it makes no assumption about right angles. Its failure mode is textureless walls, where matching has nothing to latch onto, and it will happily reconstruct a curved surface because it never assumed flatness.
manhattan_sdf inverts both properties. It optimises a neural signed distance field, which means a training run per scene and a checkpoint to manage, and it bakes in the planarity prior, which helps on textureless indoor walls and hurts on anything that is not a Manhattan plane. The README's decision to publish baseline numbers for COLMAP, COLMAP*, ACMP, NeRF, UNISURF, NeuS and VolSDF suggests the authors expect exactly this comparison, and docs/RESULTS.md is where to look before deciding whether the neural route is worth the compute for your scenes. If your evaluation is on ScanNet interiors, the table may already answer the question without you running anything.
Licence status and what upgrading costs
The repository's licence is reported as NOASSERTION, meaning GitHub could not match the LICENSE file to a known template. The LICENSE file exists at the top level, so there is a licence text to read, but its terms are not summarised anywhere in the README and the project page does not restate them. Read LICENSE yourself before using the code in anything beyond academic reproduction, and note that the repository also depends on third-party work: the acknowledgements name VolSDF, neurecon, COLMAP and a customised COLMAP fork used inside NerfingMVS. Those components carry their own terms, and the README does not consolidate them.
Upgrade cost is low in frequency and high in friction. With no published releases, there is no changelog and no version to pin, so moving to a newer commit means diffing the repository yourself. Configs are tied to the code that reads them: the data path lives in the YAML, and the README gives no guarantee that a config written for one revision keeps working after a pull. Checkpoints are likewise unversioned. The practical approach is to record the commit hash you built against and keep the Conda environment frozen, because there is no release tag to fall back to.
Editorial conclusion
Adopt manhattan_sdf if you are reproducing the CVPR 2022 paper, comparing against COLMAP, ACMP, NeRF, UNISURF, NeuS or VolSDF on ScanNet, or studying how a Manhattan-world prior changes surface extraction in a room. Do not adopt it as a general-purpose reconstruction service: the README documents one dataset, one config per scene, and no packaging, versioned release or API. Before committing, read docs/CUSTOM.md to confirm your capture fits the expected layout, check that configs/scannet/0050.yaml line 38 resolves to a real scene directory, and confirm your GPU and CUDA version against environment.yml, since the repository publishes no release and therefore no compatibility matrix.
Frequently asked questions
How do I install manhattan_sdf?
The README gives two commands: conda env create -f environment.yml, then conda activate manhattan. There is no package on an index and no published release, so installation means cloning the repository and building from environment.yml.
What data does manhattan_sdf need to run?
The README links a Google Drive folder with the processed ScanNet scene data evaluated in the paper, which you extract into data/ so the path matches line 38 of configs/scannet/0050.yaml. For your own captures, the README points to docs/CUSTOM.md for the instructions.
How do I extract a mesh with manhattan_sdf after training?
Run python run.py --type mesh_extract --output_mesh result.obj --cfg_file configs/scannet/0050.yaml gpus 0, exp_name scannet_0050. The exp_name must match the training run, since extraction reads the checkpoint from that run.
Can manhattan_sdf reconstruct rooms that are not Manhattan-shaped?
The method is built around the Manhattan-world assumption, and the README documents no mode that disables it. Scenes with slanted ceilings, curved walls or otherwise non-planar structure are outside what the paper targets.
Does manhattan_sdf publish trained models and baseline results?
Yes. The README states that quantitative and qualitative results plus trained ScanNet models are provided in docs/RESULTS.md, along with baseline results for COLMAP, COLMAP*, ACMP, NeRF, UNISURF, NeuS and VolSDF.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zju3dv-manhattan-sdf)