# pySLAM: A Python and C++ Visual SLAM Pipeline for Feature and Reconstruction Experiments

> pySLAM is a hybrid Python/C++ Visual SLAM framework supporting monocular, stereo and RGB-D cameras, with swappable local and global features, depth prediction and semantic segmentation. It is a research baseline, not a product, and the README says so.

**luigifreda/pyslam** — pySLAM is a hybrid Python/C++ Visual SLAM pipeline supporting monocular, stereo, and RGB-D cameras. It provides a broad set of modern local and global feature extractors, multiple loop-closure strategies, a volumetric reconstruction module, integrated depth-prediction models, and semantic segmentation capabilities for enhanced scene understanding.

- Repository: https://github.com/luigifreda/pyslam
- Stars: 3,425 · Forks: 546
- Language: Python
- License: GPL-3.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/luigifreda-pyslam

## What pySLAM Is For, and Who It Is Actually For

pySLAM is a Visual SLAM pipeline that supports monocular, stereo and RGB-D cameras. The problem it addresses is not that SLAM is unsolved, but that comparing two SLAM design choices usually means comparing two codebases. If you want to know whether SuperPoint with a VLAD aggregator beats ORB with a Bag of Words vocabulary on your corridor sequence, the honest answer requires running both in the same tracker, the same map, the same loop-closure logic. pySLAM is built around that: a single Python environment where local features, global descriptors, depth models and segmentation models are configuration choices rather than forks.

The README states the audience directly: pySLAM serves as a flexible baseline framework to experiment with VO/SLAM techniques, local features, descriptor aggregators, global descriptors, volumetric integration, depth prediction and semantic mapping. It also says pySLAM is a research framework and a work in progress. That second sentence matters more than the feature list. This is a project for people who read source code when something breaks, not for people who file a ticket and wait.

It is also explicitly a hybrid Python/C++ project. The sparse-SLAM core exists in both languages with custom pybind11 bindings, and the README claims the two implementations are interoperable: maps saved by one can be loaded by the other. That is an unusual design and it is the most interesting thing here, because it means you can prototype in Python and then move the same map into the C++ tracker for speed.

## How the Pipeline Is Put Together

The architecture follows the conventional SLAM split, with the front end responsible for features and matching and the back end responsible for optimization and map maintenance. On the front end, pySLAM collects local features through a common interface, which is what makes swapping detectors and descriptors cheap. On the loop-closure side, it supports descriptor aggregators including visual Bag of Words (BoW, iBow), Vector of Locally Aggregated Descriptors (VLAD), and image-wise global descriptors such as SAD, NetVLAD, HDC-Delf, CosPlace, EigenPlaces and Megaloc. Those two families solve the same problem differently: BoW-style aggregation builds a vocabulary from local descriptors, while global descriptors produce one vector per image and can skip the vocabulary entirely. The README has a section titled vocabulary-free loop closing, which is the practical consequence of that second family.

Optimization is delegated. The README lists built-in support for both g2o and GTSAM, plus custom Python bindings for features not available in the original libraries. That is a deliberate choice not to reimplement pose-graph optimization, and it means your results depend on which engine you select.

Around the core sit three optional subsystems. A volumetric reconstruction pipeline processes depth and color images through volumetric integration, supporting voxel grid models with semantic support, TSDF with voxel hashing, and incremental Gaussian Splatting. A depth prediction module integrates models including DepthPro, DepthAnythingV2, DepthAnythingV3, RAFT-Stereo and CREStereo, which matters because it lets a monocular sequence feed the dense reconstruction path that would otherwise need a depth sensor. A segmentation suite covers DeepLabv3, Segformer, CLIP, DETIC, EOV-SEG, ODISE, RFDETR and YOLO for semantic mapping. There is also a separate feed-forward pipeline for multi-image 3D inference supporting DUSt3R, Mast3r, MV-DUSt3R, VGGT, Robust VGGT, DepthFromAnythingV3 and Fast3R, which is a different mode of operation from incremental SLAM and should not be confused with it.

## Installing pySLAM and Running a First Sequence

The repository provides install scripts rather than a single pip command, and that is a signal about the dependency weight. The top level contains install_all.sh, pyenv-create.sh, pyenv-activate.sh, build_cpp_core.sh and cuda_config.sh, plus a pixi.toml and pixi.lock for environment resolution. The README documents install paths for Ubuntu, macOS and Docker.

The Python requirement is explicit in pyproject.toml: requires-python is >=3.11.9. Start by creating the environment with the project's own script.

```bash
bash pyenv-create.sh
bash pyenv-activate.sh
```

After activation, the C++ core needs to be compiled separately. This is the step that most often fails on a fresh machine, because it pulls in the compiler toolchain and the third-party graph libraries.

```bash
bash build_cpp_core.sh
```

The main entry points are top-level scripts. For visual odometry, which runs the tracker without loop closure or global optimization, the README names main_vo.py. For full SLAM with loop closing, it names main_slam.py.

```bash
python3 main_vo.py
python3 main_slam.py
```

Both take a dataset and configuration parameters. The repository ships config.yaml at the top level and a settings directory, and the README has a section on selecting a dataset and different configuration parameters. Built-in support covers more than ten dataset types. Before running a full SLAM session, check the loop-detection configuration: the README documents a specific failure mode where the vocabulary for the selected front-end descriptor type is missing, and it also provides a way to verify vocabulary compatibility. If you switch the front-end descriptor and do not switch the vocabulary, loop closure will not behave as expected. For a first run, use main_vo.py on a supported dataset, confirm the trajectory appears in the viewer, and only then enable loop closing.

## Where pySLAM Breaks Down

The dependency list in pyproject.toml is pinned aggressively, and that is the first real constraint. Entries include absl-py==1.4.0, astor==0.8.1, decorator==4.4.2, gast==0.3.3, google-pasta==2.0.0, Keras-Applications==1.0.8, Keras-Preprocessing==1.1.2, numpy>=1.24.3, scikit-image==0.21.0, matplotlib==3.7.5 and networkx==3.1. Several of those are old TensorFlow-era packages. In a shared environment, these pins will collide with anything else that wants a newer version of the same packages. The project's answer is the pyenv scripts and pixi, which means pySLAM wants its own environment and does not play well as a guest in yours.

Second, the loop-closure configuration is not self-validating. The README devotes a subsection to verifying loop detection and vocabulary compatibility, and another to the case of a missing vocabulary for the selected front-end descriptor. A configuration that runs without error can still fail to close loops, and you will see it as drift rather than as a crash. That is a worse failure mode than a stack trace.

Third, the README describes pySLAM as a research framework and a work in progress. There are no retrieved releases for this repository, so there is no versioned artifact to pin against beyond the version string in pyproject.toml, which reads 2.10.7. If you need a dependency with a support contract, a deprecation policy or a compatibility matrix, this is the wrong tool. It is also the wrong tool if your goal is production robot deployment: the feature set is oriented toward experimentation, and the optional subsystems pull in large model weights and GPU-oriented dependencies.

Finally, the Python/C++ interoperability claim deserves scrutiny rather than trust. The README says maps saved by one implementation can be loaded by the other, and points to pyslam/slam/cpp/README.md for details. Whether that holds across every map feature you use, including semantic labels and volumetric data, is not stated in the top-level README.

## pySLAM Against ORB-SLAM and PyCuVSLAM

The obvious comparison is ORB-SLAM, which is the reference point most people arrive with. The difference in approach is feature policy. ORB-SLAM is built around ORB, and its loop closure is built around a Bag of Words vocabulary derived from ORB descriptors. That tight coupling is what makes it fast and predictable, and it is also what makes it hard to ask what a learned feature would do instead. pySLAM inverts that: the front end is a pluggable interface, the aggregator can be BoW, VLAD or a global descriptor, and the depth source can be a sensor or a predicted depth map. You trade a tuned, coherent system for a testbed where the components are separable.

PyCuVSLAM sits at the other end. It is a CUDA-accelerated visual SLAM library, which means the acceleration is the product and the platform constraint is the price. pySLAM's C++ core is an optional speed path with pybind11 bindings, not a GPU-first design, and the README presents it as a choice between high-performance/speed and high-flexibility modes rather than as a performance claim against other libraries.

A fair way to frame it: if your question is how do I get a reliable trajectory today, ORB-SLAM or a CUDA-backed library is the shorter path. If your question is which descriptor, aggregator and depth model combination works on my data, pySLAM is built to answer that question, and the cost is environment fragility and a configuration surface that can silently misbehave.

## Maintenance, Licence and Upgrade Cost

The repository is not archived, and the last push was on 2026-08-23. That is recent enough that the codebase is moving, but the absence of retrieved releases means there is no tagged artifact to depend on. Upgrades therefore mean pulling the master branch, and the pinned dependencies in pyproject.toml mean a pull can invalidate an environment that took an afternoon to build. Budget for rebuilding the environment, not just for pulling.

The licence is GPL-3.0. That is a copyleft licence, and the practical implication is that distributing a product that links pySLAM's code carries source-disclosure obligations for the combined work. For academic use and internal experimentation this is usually unproblematic. For a commercial product, the licence choice is a design constraint that should be settled before writing code against the API, not after. This is not legal advice; if you are shipping something, have someone qualified read the LICENSE file and the licences of the bundled third-party components under thirdparty/.

One more cost worth naming: the optional subsystems are not free to enable. Depth prediction and segmentation pull in model weights and inference runtimes, and the volumetric path with Gaussian Splatting adds further dependencies. A pySLAM install that only runs monocular VO is a much smaller commitment than one that runs the full semantic and volumetric stack.

## Conclusion

Adopt pySLAM if you are a researcher or graduate student who needs one environment where local features, global descriptors, loop closure, depth prediction and semantic segmentation can be swapped and compared on the same sequence. Do not adopt it if you need a supported library with versioned releases and a compatibility guarantee, or if you want a turnkey mapping product for a robot fleet. Before committing, verify three things: that your Python version is at least 3.11.9 as pyproject.toml requires, that the C++ core builds on your machine with build_cpp_core.sh, and that a vocabulary exists for the front-end descriptor you intend to use, because the README documents a missing-vocabulary failure mode for loop detection.

## FAQ

### What is the difference between SLAM and vSLAM, and where does pySLAM fit?

vSLAM is visual SLAM, meaning localization and mapping driven by camera images rather than by other sensors. pySLAM is a visual SLAM pipeline supporting monocular, stereo and RGB-D cameras, so it sits in the vSLAM category.

### What is vSLAM used for?

It estimates a camera trajectory and builds a map of the environment from images. pySLAM covers that core plus optional volumetric reconstruction, depth prediction and semantic segmentation for scene understanding.

### Are SLAM and LiDAR the same thing?

No. LiDAR is a sensor; SLAM is the estimation problem. pySLAM works from cameras, and where a depth sensor is unavailable it can use integrated depth prediction models such as DepthPro or DepthAnythingV2 instead.

## Sources

- [Issues](https://github.com/luigifreda/pyslam/issues)
- [License: GPL-3.0](https://github.com/luigifreda/pyslam/blob/master/LICENSE)
- [luigifreda/pyslam on GitHub](https://github.com/luigifreda/pyslam)
- [README](https://github.com/luigifreda/pyslam/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/luigifreda-pyslam
