Open-source project
changh95/visual-slam-roadmap avatar
changh95/visual-slam-roadmap

The Visual-SLAM roadmap is eleven folders of Markdown, and its level 1 links into level 2

Roadmap to become a Visual-SLAM developer in 2026

1,799 stars177 forksAstroMIT

At a glance

What is it?
A year-stamped learning roadmap published as an Astro site, covering camera geometry first and classical monocular SLAM second, then working outward to deep learning, event cameras and world models. The programming and camera device pages for the beginner level live in the level 2 directory, which tells you the folder boundary and the curriculum boundary were drawn separately.
Who is it for?
The roadmap is worth using as a syllabus if you are starting from zero and want to see the whole field in one list, because the ordering argument is the thing it offers, from camera geometry through classical pipelines to learned systems. Three cautions.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 76 days ago.
What is it written in?
Mainly Astro, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Level 1 programming pages are stored in the level 2 directory

The beginner section lists three programming topics, and every one of them points into the second level's folder. The C++ entry, described as pointers and object orientation, resolves to `level-02-getting-familiar/cpp.md`. Python resolves to `level-02-getting-familiar/python.md`, and Bash over Linux, described as basic terminal usage, resolves to `level-02-getting-familiar/bash-linux.md`. The same thing happens with the camera device material further down that section. Lens, sensor and the exposure settings all resolve to a single page inside the level 2 folder, and the lens distortion reference inside the camera calibration entry points there as well. Meanwhile the mathematics and geometry pages of the same beginner level sit where the level numbering would predict, under the level 1 directory. So the split between what a beginner needs and what a beginner can postpone is a split inside one directory rather than between two, which is harmless for a reader following the links and awkward for anyone restructuring the tree, since moving a topic means choosing which level owns the file rather than following the existing convention.

Headings and folder names follow two different naming schemes

The section titles and the directories that hold them are not the same strings. The heading for the second level is Getting Familiar with SLAM, while the directory is named for getting familiar. The third level is headed Monocular SLAM and lives in a monocular folder, the fourth is RGB-D and its folder drops the hyphen, and the sixth is VIO over VINS in a directory written as two short tokens joined by a hyphen. The navigation table compounds this by using its own anchor names, so the RGB-D entry points at an anchor written as level 4 rgb d visual slam while the folder is level 4 rgbd slam, and the collaborative entry points at an anchor naming multi robot where the heading says only collaborative. None of this breaks the links, since each one is internally consistent, but it does mean the site is assembled by hand rather than generated from the file tree, and it is why a rename of any level means touching the heading, the anchor and the directory.

The last three levels are not really SLAM any more

The progression starts conventionally. Level 1 is mathematics, programming and image fundamentals, level 2 is the anatomy of a SLAM system, level 3 is classical monocular work across feature based, direct and semi direct methods plus structure from motion and dynamic scenes, level 4 is RGB-D with dense tracking and volumetric or surfel fusion, level 5 applies deep learning with learned frontends and differentiable backends. From there the list widens rather than deepens. Level 6 fuses cameras with inertial units and contrasts filtering against optimisation, level 7 covers stereo for metric scale, level 8 covers multiple robots and map merging, level 9 covers LiDAR and tight camera, LiDAR and inertial fusion, level 10 handles event cameras for high dynamic range and fast motion, and level 11 moves from SLAM maps to learned world representations. Read as a sequence, the argument is that a visual SLAM engineer in this cycle ends up working on spatial representations rather than on pose graphs.

Geometry comes before image processing, which is the reverse of the usual order

Level 1 opens with mathematics and then goes straight into projective geometry: the pinhole camera model and image projection, camera calibration split into intrinsic and extrinsic parameters, rigid body motion expressed through Euler angles, quaternions and rotation matrices along with projective space and vanishing points, homogeneous transformation, epipolar geometry leading to the essential and fundamental matrices, and triangulation. It then covers camera models beyond pinhole by name, including the Kannala-Brandt fisheye model, the double sphere model, omnidirectional cameras and rolling shutter awareness. Only after all of that does the image data section appear, and it is a short one: colour and grayscale images, thresholding, Gaussian blur, a Harris corner detector, Sobel and Canny edges. Most computer vision courses start with the image operators and defer the geometry, so a reader coming from that direction will notice the inversion and should budget extra time for the geometry block.

The engineering column is where the job market is addressed, and it is marked optional

Level 2 splits programming into core and optional, and the optional half is aimed squarely at employers. Core material covers C and C++ including modern standards, data structures, compilers and the build tools CMake, Makefile and Ninja, plus Git, OpenCV in both its C++ and Python forms, Python for deep learning and scripting, shell work including ssh and terminal multiplexers, the maths libraries with Eigen and the optimisation libraries Ceres-solver, GTSAM and g2o, binding layers through PyBind11 and nanobind, the robotics frameworks, and containers. The optional half adds instruction set level parallelism and OpenMP and CUDA, edge deployment covering TensorRT and ONNX export of learned frontends with Jetson benchmarking, mobile development for both major platforms, C# with Unity and HoloLens, continuous integration, and simulation with Gazebo and Isaac Sim. Marking these optional is a deliberate claim: none of them is needed to understand the field, and each one changes what kind of job you can be hired for.

A year-stamped site with no version, no changelog and no releases

What is published is a rendered site plus one Markdown file per topic. The repository records Astro as its primary language, keeps the site sources in a site directory, the page sources in another, and images in a third, and the eleven level directories sit alongside those at the top level. The README is titled as a developer roadmap for 2026, the description frames it as a roadmap to become a Visual-SLAM developer in 2026, and the live version is published on a personal domain rather than on a documentation host. The licence is MIT, which is unusual for a personal roadmap and generous for one, and there are no GitHub releases at all, so there is nothing to pin and no release notes to read. The last recorded commit on main is dated 2026-07-19, which means the 2026 stamp is a statement of intent rather than a maintenance guarantee.

The author argues the barrier is breadth, not mathematics

The framing matters more than the list. The opening says Visual-SLAM is often portrayed as a difficult topic, with many people believing that good C++ and a deep grasp of mathematics are required, and then notes that there are few beginner courses, especially outside English. The purpose section says the roadmaps are meant to give an overview and to guide anyone confused about where to start. The note to beginners makes the central claim explicitly: the entry barrier is high not because the mathematics is difficult but because it takes equipping yourself with several different kinds of skill, and you do not have to learn everything to get going. That is why the list is long and shallow rather than short and deep, and it is also why the optional column exists. A reader who wants a research sequence and a reader who wants a hiring plan are both being served by the same tree, which is why the folder boundary does not line up with the level boundary.

Editorial conclusion

The roadmap is worth using as a syllabus if you are starting from zero and want to see the whole field in one list, because the ordering argument is the thing it offers, from camera geometry through classical pipelines to learned systems. Three cautions. It is stamped for 2026 and carries no version, no changelog and no releases, so there is nothing to pin and no way to tell whether a given topic page has moved since you read it. The levels describe a field rather than a project, so nothing here runs and nothing measures. And the folder boundaries do not match the curriculum boundaries, which means a reader who forks it to reorganise the material will have to decide for themselves which level owns the programming pages and the camera device pages.

Frequently asked questions

What does the Visual-SLAM roadmap actually cover?

Eleven levels: beginner mathematics, programming and camera fundamentals, then the anatomy of a SLAM system, classical monocular methods, RGB-D with dense fusion, deep learning applied to SLAM, visual inertial odometry and its filter versus optimisation split, stereo, collaborative and multi robot mapping, LiDAR and visual LiDAR fusion, event cameras, and world models with spatial representations.

What does level 1 of the Visual-SLAM roadmap ask a beginner to know?

C++ pointers and object orientation, Python and shell basics, probability and statistics with Gaussian distributions and Bayes, linear algebra through to singular value decomposition and eigenvalues, logarithm and exponential, differentiation and Taylor expansion, then projective geometry including the pinhole model, calibration, rigid body motion, epipolar geometry and triangulation, camera models beyond pinhole, and basic image operations.

Where are the level 1 programming pages stored in the repository?

Under the level-02-getting-familiar directory, together with the camera device pages for lens, sensor and exposure settings, while the mathematics and geometry pages are under level-01-beginner. The folder boundary and the curriculum boundary were drawn separately.

Which maths and optimisation libraries does the roadmap name?

Eigen as the maths library, and Ceres-solver, GTSAM and g2o as the optimisation libraries, all at level 2. Level 2 also lists CMake, Makefile and Ninja, PyBind11 and nanobind for binding, OpenCV in C++ and Python, ROS and ROS2, and Docker.

Is the Visual-SLAM roadmap a software package I can install?

No. It is a static site built with Astro plus one Markdown file per topic across eleven level directories, published on a personal domain and released under the MIT licence. The repository publishes no releases, so there is no version to pin.

Official sources

  1. changh95/visual-slam-roadmap on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/changh95-visual-slam-roadmap.svg)](https://hysenlabs.com/projects/changh95-visual-slam-roadmap)