Open-source project
sou350121/Spatial-Intelligence-Handbook avatar
sou350121/Spatial-Intelligence-Handbook

Spatial-Intelligence-Handbook puts SLAM and VIO problems side by side across six embodiments

空間智能的跨 embodiment 手冊 —— 把 SLAM / VIO / 3D 表徵 / 感測器堆疊 / 部署坑 在 機械臂 / 無人機 / 自駕 / 人形 / 水下 之間橫向比較,底層共用 3DGS / VGGT / depth foundation。照見三冊 perception 端(VLA-Handbook + Physics-Controllable-Generation-Handbook 姊妹倉)。

30 stars0 forksPythonNOASSERTION

At a glance

What is it?
A documentation repository organised around a rule most spatial AI surveys break: every cross-embodiment article has to clear an admission bar before it is written, and each paper that passes is filed as a five axis coordinate rather than a bullet. No code to run, which is the point.
Who is it for?
Use this if you are choosing a representation or a sensor stack for an embodied system and the survey literature for your field has never mentioned your field. Read the crossing section first, since that is where the content no single field survey contains, and check the five axis coordinates rather than trusting the entry counts, which grow daily with the ingestion pipeline.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Every crossing article has to clear a five condition admission bar

The crossing directory is presented as the part of the repository that does not exist elsewhere, and it comes with a written threshold rather than a guideline. An article admitted there must span at least three embodiments. Each embodiment it touches must carry a paper source. At least one engineering number has to appear, so a claim cannot be argued purely in prose. And the piece must end with a boundary section that says where the comparison stops holding. Five wedges currently fill it: migration between SLAM and VIO across platforms, a sensor stack matrix, a scale comparison running from one centimetre to a thousand kilometres, a representation migration comparison, and an atlas of failure modes. That last one is pitched explicitly as a source of paper ideas rather than as a summary of existing ones.

The Atlas is a coordinate stream, not a reading list

Papers that pass the ingestion pipeline are not collected as citations. Each one is filed under five ontology axes: the problem, the representation, the sensor, the paradigm, and time. The result is a machine readable file alongside the human readable overview, and the stated purpose of keeping coordinates rather than titles is to make drift visible. The paradigm axis is the demonstration: on the seed corpus the counts run 23 for geometric, 76 for learned, 55 for hybrid, 94 for VLA, and 31 for world model as policy, with the note that the newest breakthroughs land in that smallest bucket. Because a pipeline adds entries on working days, those numbers are a snapshot rather than a fixed figure, and the axis is meant to be watched as it moves rather than quoted once.

code
geometric              ██████·················· 23
learned                ███████████████████····· 76
hybrid                 ██████████████·········· 55
VLA                    ████████████████████████ 94
world-model-as-policy  ████████················ 31

A working day pipeline writes new dissections into the tree

The repository argues its own freshness against the usual fate of a survey, and the mechanism is specific. Every working day one arxiv pipeline pulls the newest submissions and assigns each a rating from a three symbol scale covering work, tooling and reading. Selected papers are then written up as full dissections and committed, and each new entry has its five axis coordinates folded into the atlas. The dissections are meant to carry representation selection reasoning, key hyperparameters, shape sanity checks, and deployment traps, with the gap between reading a paper and running its code called out explicitly. So the growth model is editorial rather than automated summarisation: a machine narrows the field, and a person decides what is worth a long page.

foundations/ is the shared toolbox: thirteen zones and 139 articles

The lower layer collects what every embodiment ends up needing, on the argument that academia slices the literature by method while this repository slices it by the problem that keeps recurring. The zone count and article count are stated on the page: thirteen zones holding 139 articles. Five are surfaced as entry points. Feed forward 3D dissects the model that won a best paper award at a 2025 computer vision conference and is read as a paradigm signal rather than a single result. The 3DGS family is framed as the representation that displaced NeRF, with a hundredfold improvement claimed over six months. Depth foundation models are covered through the case of Depth Anything v2, specifically for the trap between relative and metric depth. Sensor physics is marked as an exclusive axis covering the engineering accounting of size, weight and power that survey papers leave out, and spatial maths holds the shared skeleton of pose algebra, bundle adjustment and inertial preintegration.

embodiments/ has six lanes, and only one is the maintainer's deep anchor

Six lanes sit at the application layer: manipulation, humanoid legged, ground mobile, driving, aerial, and marine. The page is explicit about the imbalance. Aerial is marked as the maintainer's depth anchor and is described as running one and a half to two times deeper than the other lanes, which is a useful warning for anyone weighing the coverage of the remaining five. Marine is positioned differently again, as the contrasting case where vision degrades. Augmented and virtual reality are deliberately not a lane of their own; they appear only as comparison cases inside the crossing material. For readers arriving from the vision language action side, a bridge directory documents the contract on both ends of the gap between a feature cloud and an action head, which is where most of the integration work actually sits.

benchmarks/ dissects conditions and gaps instead of publishing a leaderboard

The benchmark directory covers four standard sets and is organised by what each one fails to tell you. Each entry is broken into the conditions under which a number was produced, the failure modes of the method class, and the gap between the benchmark and a machine in the real world. That last axis is the one that distinguishes this directory from a results table, and it is the reason the entry point for sensor and hardware teams is the sensor physics zone rather than the leaderboards. A companion deployment directory carries the practical half: hardware selection, multimodal time synchronisation, calibration, compute budget, and failure modes. The cheat sheet directory collects a timeline, a representation quick reference, the sensor budget matrix, and the current ontology version, which the page pins at v3.2.

It is one of three handbooks, and the border is drawn at perception

The repository positions itself as the perception third of a set of three that cross-reference each other, with the same paper reading differently depending on which volume you are holding. The action volume covers action policies built on diffusion, flow or reinforcement learning methods. The generation volume covers physics controllable generation, spanning video world models, differentiable simulation and neural surrogates. This one covers world representation: 3D Gaussian splatting, feed forward reconstruction, depth models, and sensor physics. The stated purpose of keeping them separate is that the interfaces between them are hard, which is why each boundary has a dedicated bridge directory rather than a chapter. Content is offered under a Creative Commons attribution licence, and the contribution page asks for paper readings, field experience, sensor selection measurements and cross embodiment comparisons.

Editorial conclusion

Use this if you are choosing a representation or a sensor stack for an embodied system and the survey literature for your field has never mentioned your field. Read the crossing section first, since that is where the content no single field survey contains, and check the five axis coordinates rather than trusting the entry counts, which grow daily with the ingestion pipeline. Do not expect to install anything: the deliverable is prose plus a machine readable atlas file, and the readable version lives on the project's documentation site. Before relying on any single figure, note that the pages were written by one maintainer, whose depth anchor is aerial robotics.

Frequently asked questions

What is spatial intelligence in simple words?

In this repository the term means machine perception rather than human ability. The handbook covers SLAM, visual inertial odometry, 3D reconstruction, sensor stacks and deployment traps, read side by side across manipulation, aerial, driving and marine platforms on one shared representation layer.

Who is a famous person with spatial intelligence?

The repository profiles no individuals. The named references are technical: 3D Gaussian splatting, NeRF, a feed forward reconstruction model, Depth Anything v2, VINS-Mono, the ScanNet++, EuRoC, nuScenes and AQUALOC datasets, and organizations including World Labs, Cosmos, Skydio, Wayve and Tesla.

What jobs are good for people with spatial intelligence?

Careers are not discussed. What the handbook does offer is role based entry points, including a crossing section for algorithm engineers changing embodiment, a failure mode atlas for researchers hunting a paper idea, a sensor physics zone for hardware teams, an aerial lane, a spatial maths zone for pose algebra and bundle adjustment, and a bridge directory for people arriving from the vision language action side.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sou350121-spatial-intelligence-handbook.svg)](https://hysenlabs.com/projects/sou350121-spatial-intelligence-handbook)