Hypersim: 74,619 synthetic indoor images with per-pixel geometry and a 1.9TB download
Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
At a glance
- What is it?
- Hypersim is Apple AI/ML Research's photorealistic synthetic dataset for indoor scene understanding, with 461 scenes and 74,619 publicly released images carrying per-pixel depth, normals and semantic labels. The ground truth triangle meshes require a TurboSquid purchase, and the Python toolkit targets Python 3.7.9.
- Who is it for?
- Use Hypersim when your research needs per-pixel ground truth for depth, surface normals, semantic instance labels or intrinsic image decomposition, and when you have the storage and bandwidth for a 1.9TB dataset. Skip it if you need the 3D meshes without a TurboSquid purchase, or if your environment requires a Python version newer than 3.7.9.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 28 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
461 scenes and 74,619 public images, with 2,781 manually excluded frames
The full Hypersim rendering run produced 77,400 images from 461 indoor scenes, but the public release holds 74,619 of them. Before publishing, the authors manually excluded images containing people and prominent logos, and all excluded frames are listed in `ml-hypersim/evermotion_dataset/analysis/metadata_images.csv`. Researchers who need to account for the gap between 77,400 and 74,619 have a precise record to work from.
The 461 scenes come from a collection created by professional artists. Each scene follows the naming pattern `ai_VVV_NNN`, where VVV is the volume number and NNN is the scene number within the volume. A single scene can have more than one camera trajectory, named cam_00, cam_01 and so on, and each trajectory holds one or more frames numbered from frame.0000 upward. That naming convention runs through the download script and the data loading utilities; everything in the codebase that accesses a file locates it by scene name, trajectory and frame number.
The dataset targets tasks where per-pixel ground truth labels are difficult or impossible to obtain from real photographs. Depth estimation, surface normal prediction, and semantic instance segmentation all require annotations that are impractical to produce at scale from captured images. Hypersim generates those labels directly from the rendering process, which provides ground truth for every pixel without a human annotation step. That is the core advantage of a synthetic dataset over a manually annotated one: the labels are exact and complete across the full 461-scene collection.
Four sub-images per frame: color, reflectance, illumination and the non-diffuse residual
Every frame in Hypersim is stored in more than one form inside the `scene_cam_XX_final_hdf5` directory. Four HDF5 files cover the image decomposition for a single frame: `frame.IIII.color.hdf5` holds the color image before any tone mapping, `frame.IIII.diffuse_reflectance.hdf5` holds the diffuse reflectance component (which many researchers call albedo), `frame.IIII.diffuse_illumination.hdf5` holds the diffuse illumination, and `frame.IIII.residual.hdf5` holds the non-diffuse residual that captures view-dependent lighting effects.
This four-way decomposition is the property that separates Hypersim from datasets that store only the final rendered image. Researchers working on intrinsic image decomposition get ground truth for all three sub-images from a single frame, without any estimation step, because those sub-images are the components the renderer used to produce the color output. The residual term specifically captures the portion of the image that varies with viewing angle, which is absent from standard diffuse-only decompositions.
All four image types are in HDF5 format. Preview images for each scene and camera are stored separately under `scene_cam_XX_final_preview`, as lower-fidelity files intended for browsing rather than for training.
Geometry per frame: depth in meters, normals in two reference frames
Geometric ground truth lives in a parallel directory, `scene_cam_XX_geometry_hdf5`. Each frame carries depth, world-space position, and surface normals in four variants. Depth is stored as Euclidean distances from the optical center of the camera, in meters, rather than as z-buffer depth. The files the renderer produces per frame are:
├── scene_cam_XX_geometry_hdf5 # lossless HDR image data that does not require accurate shading
│ ├── frame.IIII.depth_meters.hdf5 # Euclidean distances (in meters) to the optical center of the camera
│ ├── frame.IIII.position.hdf5 # world-space positions (in asset coordinates)
│ ├── frame.IIII.normal_cam.hdf5 # surface normals in camera-space (ignores bump mapping)
│ ├── frame.IIII.normal_world.hdf5 # surface normals in world-space (ignores bump mapping)
│ ├── frame.IIII.normal_bump_cam.hdf5 # surface normals in camera-space (takes bump mapping into account)The pair of bump-aware and bump-agnostic normals covers two distinct needs: an algorithm that wants the raw geometric surface shape uses `frame.IIII.normal_cam.hdf5`, while one that needs the appearance-level detail including surface texture uses `frame.IIII.normal_bump_cam.hdf5`. Both are within the same directory, produced in one render pass.
Camera information lives in `camera_keyframe_orientations.hdf5` and `camera_keyframe_positions.hdf5` inside each scene's cam_XX directory under _detail. Before any metric computation, `metadata_scene.csv` in the _detail directory provides the scale factor to convert asset units into meters.
Downloading 1.9TB: the Python script, ZIP partitioning and the Windows caveat
The entire dataset is roughly 1.9TB, split across a few hundred ZIP files between 1GB and 20GB each. The download script fetches each ZIP by URL:
python code/python/tools/dataset_download_images.py --downloads_dir /Volumes/portable_hard_drive/downloads --decompress_dir /Volumes/portable_hard_drive/evermotion_dataset/scenesThe script contains the URL list for every ZIP file. On Windows, the README notes that the script requires modification because it depends on the `curl` and `unzip` command-line utilities that are not standard on that platform.
Thomas Germer contributed an alternative download script at `contrib/99991` that can download subsets of files from within each ZIP archive rather than whole ZIP files. For researchers who need only a portion of the 461 scenes, that script avoids downloading gigabytes of data for scenes they will not use. The full 1.9TB download is impractical for projects that target a small subset of the dataset.
The geometry data, camera metadata and semantic label files are all part of the normal download. Ground truth triangle meshes are not included.
Python 3.7.9 and an osx-64 conda spec that predates current Python releases
The `requirements.txt` file carries a conda spec for `platform: osx-64` and specifies these package versions:
python=3.7.9=h26836e1_0
h5py=2.10.0=py37h3134771_0
joblib=1.1.1=py37hecd8cb5_0
matplotlib=3.3.1=0
pandas=1.1.3=py37hb1e8313_0
scikit-learn=0.23.2=py37h959d312_0
scipy=1.5.2=py37h912ce22_0Python 3.7 is no longer a supported release. The `osx-64` conda platform does not cover Apple Silicon Macs or Linux systems, so a researcher working on either cannot use this conda spec without modification. The repository gives no guidance on adapting the requirements to other platforms.
Researchers who load the HDF5 data files using code they write themselves are not bound to Python 3.7.9. Any version of h5py that reads HDF5 format opens the files. The version restriction is a constraint on the provided toolkit, not on the file format itself.
This gap between the toolkit requirements and current Python environments is common in datasets published alongside conference papers. Hypersim was presented at ICCV 2021, and the Python 3.7 specification aligns with the development environment at that time.
Triangle meshes require a TurboSquid purchase separate from the image download
The download script retrieves image data, depth maps, normals and metadata. The ground truth triangle meshes for each scene are not part of any download. The README says to obtain them by purchasing the asset files from TurboSquid, via the Evermotion collection storefront, but states no price.
Those meshes are the physical geometry underlying each scene, and without them a researcher has the rendered images, depth maps and camera parameters but not the surface geometry that produced the rendering. Algorithms that require raw mesh access, such as 3D reconstruction evaluation or physics simulation against the original scene geometry, hit this boundary directly.
The dataset description says it relies exclusively on publicly available 3D assets. In that context, publicly available refers to the source of the assets, meaning they are not proprietary to Apple. It does not mean they are free. For any workflow that requires mesh data, TurboSquid is the only documented route, and the cost of the Evermotion packs is not stated in the repository.
For researchers who work exclusively with rendered images and do not need the raw geometry, the TurboSquid purchase is never required. The download script gives them everything they need.
CC BY-SA 3.0 covers the data; a custom license governs the toolkit code
Two separate licenses apply to Hypersim. The image dataset and its labels are released under the Creative Commons Attribution-ShareAlike 3.0 Unported License. That license requires attribution when you use or redistribute the data, and requires any derivative works to carry the same CC BY-SA 3.0 terms.
The Python toolkit code uses a different license. GitHub classifies it as 'Other' because the text in `LICENSE.txt` does not match a standard SPDX identifier. The README does not summarize the code license terms. Reading `LICENSE.txt` at the repository root is necessary before using or redistributing the toolkit in a project.
For researchers who load the HDF5 data files using their own code, only the Creative Commons license applies. For anyone who copies or builds on functions from the repository's Python scripts, the custom license in `LICENSE.txt` is the governing document.
The paper citation in the README asks attribution to the ICCV 2021 paper by Roberts, Ramapuram, Ranjan and colleagues, separately from the license terms. The BibTeX entry in the README cites the 2021 International Conference on Computer Vision, and the paper is available at arxiv.org under identifier 2011.02523.
Editorial conclusion
Use Hypersim when your research needs per-pixel ground truth for depth, surface normals, semantic instance labels or intrinsic image decomposition, and when you have the storage and bandwidth for a 1.9TB dataset. Skip it if you need the 3D meshes without a TurboSquid purchase, or if your environment requires a Python version newer than 3.7.9. Before adopting, read LICENSE.txt for the code license, because the data is CC BY-SA 3.0 but the toolkit license is custom and not summarized in the README. The subset download script at contrib/99991 is a practical first step if you need only a fraction of the 461 scenes.
Frequently asked questions
What is the Hypersim dataset?
Hypersim is a photorealistic synthetic dataset for indoor scene understanding with 74,619 publicly released images from 461 scenes, including per-pixel depth, surface normals, semantic instance labels and a four-way image decomposition into reflectance, illumination and a non-diffuse residual.
How do I download the Hypersim dataset?
Run the Python script at code/python/tools/dataset_download_images.py with the --downloads_dir and --decompress_dir flags. The full dataset is roughly 1.9TB and is split into hundreds of ZIP files between 1GB and 20GB each.
What Python version does the Hypersim toolkit require?
The requirements.txt file specifies Python 3.7.9 in a conda environment for the osx-64 platform. Reading the HDF5 data files with any h5py version does not strictly require Python 3.7, but the provided toolkit targets it.
Does Hypersim include 3D mesh data?
The download script provides image, metadata and geometry channel files only. Ground truth triangle meshes must be purchased separately from TurboSquid.
What license covers the Hypersim dataset?
The image dataset is under the Creative Commons Attribution-ShareAlike 3.0 Unported License. The Python toolkit code uses a custom license; the full text is in LICENSE.txt at the repository root.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apple-aiml-research-ml-hypersim)