Library / SDK
Robbyant/lingbot-map avatar
Robbyant/lingbot-map

LingBot-Map: streaming 3D reconstruction from video with a feed-forward transformer

A feed-forward 3D foundation model for reconstructing scenes from streaming data.

17,131 stars1,911 forksPythonApache-2.0

At a glance

What is it?
LingBot-Map is a Python 3D foundation model that reconstructs scene geometry from a live frame stream rather than from a batch of images. The installation path is short, but the model weights and a CUDA GPU are prerequisites, and the batch renderer pins PyTorch 2.8.0 for NVIDIA Kaolin wheels.
Who is it for?
Adopt LingBot-Map if you have a CUDA 12.8 machine, a downloaded lingbot-map checkpoint, and a pipeline that already produces ordered video frames; the interactive viser viewer at port 8080 makes the first run cheap to evaluate. Do not adopt it if you need CPU-only inference, if you plan to build the batch rendering pipeline on a PyTorch version other than 2.8.0, or if you expect a hosted service, since the README lists no hosted endpoint and no Android or mobile client.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 22 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What LingBot-Map reconstructs, and for whom

The problem LingBot-Map targets is incremental: a camera keeps producing frames, and you want geometry and camera poses as the stream continues, without re-running a global optimisation over every frame seen so far. The README describes the project as "a feed-forward 3D foundation model for streaming 3D reconstruction", which places it in the family of learned reconstruction models rather than classical structure-from-motion pipelines that refine a bundle adjustment problem to convergence.

The intended user is someone with an ordered frame sequence and a CUDA machine. The repository ships example scenes under example/courthouse, example/loop and example/university, an interactive demo in demo.py, an offline rendering pipeline under demo_render/, and evaluation scripts under benchmark/ for KITTI and Oxford Spires. That combination points at research and evaluation work first, and at product integration second. If your input is a folder of unordered photographs, the streaming framing is a mismatch; if your input is a video or a live camera, it fits.

Anchor context, pose-reference window and trajectory memory

The architectural claim in the README is a Geometric Context Transformer that "architecturally unifies coordinate grounding, dense geometric cues, and long-range drift correction within a single streaming framework through anchor context, pose-reference window, and trajectory memory". Those three names describe the mechanism at a high level: a persistent anchor that keeps the reconstruction in a consistent coordinate frame, a bounded window of recent poses used as a reference while new frames arrive, and a memory of the trajectory that lets the model correct drift over long runs instead of accumulating it.

Inference is feed-forward, and the attention implementation is where the streaming behaviour lives. The project uses paged KV cache attention, with FlashInfer as the recommended backend and PyTorch SDPA as a fallback selected with --use_sdpa. Paging the KV cache is what allows a long sequence to be processed without holding every key and value in a contiguous buffer. The README states the model reaches roughly 20 FPS at 518x378 resolution on sequences exceeding 10,000 frames. That figure comes from the project's own description and is worth treating as a target to reproduce on your hardware, not a guarantee.

The release notes record two KV cache bugs that were fixed in 2026: an SDPA cache bug fixed on 2026-06-28, and a FlashInfer bug fixed on 2026-04-24 where --keyframe_interval > 1 silently cached non-keyframes. The second entry states that pose and reconstruction quality should be better when running with more than 320 frames. Anyone pinning an older commit is pinning those bugs.

Installing LingBot-Map and running the courthouse scene

Installation starts with a dedicated conda environment on Python 3.10, which is the minimum the pyproject.toml declares (requires-python = ">= 3.10").

bash
conda create -n lingbot-map python=3.10 -y
conda activate lingbot-map

The README recommends PyTorch 2.8.0 with CUDA 12.8 wheels. The reason given is concrete: NVIDIA Kaolin, which the batch rendering pipeline requires, has prebuilt wheels for torch-2.8.0_cu128. If you only intend to run demo.py, a newer PyTorch is acceptable, but the batch renderer then needs Kaolin built from source.

bash
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu128
pip install -e .

The editable install pulls the runtime dependencies listed in pyproject.toml: Pillow, huggingface_hub, einops, safetensors, opencv-python, tqdm and scipy. Visualization is an optional extra, installed as pip install -e ".[vis]", which adds viser, trimesh, matplotlib, onnxruntime and requests.

FlashInfer supplies the paged KV cache attention and is the recommended backend. The README notes it is a pure-Python package that JIT-compiles CUDA kernels on first use, so one wheel covers multiple CUDA and PyTorch versions.

bash
pip install --index-url https://pypi.org/simple flashinfer-python

If FlashInfer is absent, the model falls back to SDPA through --use_sdpa. Weights are not in the repository; they are downloaded from the Hugging Face or ModelScope repositories linked in the model table, where three checkpoints are listed: lingbot-map-long for long sequences and large scenes, lingbot-map as the balanced checkpoint used in the paper and benchmarks, and lingbot-map-stage1 as the stage-1 training checkpoint.

With a checkpoint on disk, the first run is one command. It opens an interactive viser viewer on http://localhost:8080.

bash
python demo.py --model_path /path/to/lingbot-map.pt \
    --image_folder example/courthouse --mask_sky

The --mask_sky flag applies sky masking, one of the demo options the README documents. For sequences longer than roughly 3000 frames the README points at windowed inference, and for very long videos at the offline pipeline under demo_render/, whose entry point is batch_demo.py.

Where LingBot-Map is the wrong tool

The hard dependency is a CUDA GPU. Nothing in the README describes a CPU path, so a laptop without an NVIDIA device is not a target for the demo. The FlashInfer route compiles CUDA kernels on first use, which means the first inference pays a JIT cost before steady-state throughput appears.

Memory is the second constraint. The README's own table of contents separates windowed inference for sequences over 3000 frames from the default streaming mode, and the performance section is where the project discusses memory. Long sequences are handled, but not for free, and the choice between the balanced and long checkpoints is a real trade-off rather than a formality: the long checkpoint is described as better suited to long sequences and large scenes, while the balanced one trades all-around performance across short and long sequences.

Version pinning is the third. If you need the batch rendering pipeline, PyTorch 2.8.0 is effectively fixed by the Kaolin wheel availability. Teams already standardised on a different PyTorch release will either build Kaolin from source or give up the offline renderer.

Finally, the project publishes no hosted service. The homepage field is empty, and the README's links point to an arXiv paper, a PDF, a project page, Hugging Face and ModelScope. Anyone searching for a LingBot-Map app or an Android client will not find one in this repository.

LingBot-Map against iterative optimisation pipelines

The contrast the README draws is with "existing streaming and iterative optimization-based approaches". The difference is structural. An iterative optimiser such as a bundle-adjustment-based SLAM or structure-from-motion system maintains an explicit map and refines it, which gives strong geometric consistency but makes cost grow with the size of the problem and pushes you toward keyframe selection and marginalisation heuristics.

LingBot-Map replaces that with a learned feed-forward pass plus a paged KV cache. There is no per-scene optimisation loop to converge, so latency is roughly constant per frame instead of growing with the map. The trade is that accuracy is bounded by the checkpoint rather than by the data you feed it: you cannot let it run longer to squeeze out a better result, and a scene far outside the training distribution has no fallback mechanism. The stage-1 checkpoint is also described as loadable into the VGGT model for bidirectional inference, which suggests the authors expect comparison against that line of work.

For evaluation, the repository includes benchmark scripts for KITTI and Oxford Spires, with preprocess/oxford.py to prepare Oxford Spires data. That is the honest way to compare: run the same sequences through both approaches rather than trusting either project's headline numbers.

Licence and the cost of keeping up

The repository is Apache-2.0, with the text in LICENSE.txt. Apache-2.0 permits commercial use and modification and includes an explicit patent grant, which matters if you ship a product built on the model. Two things to check yourself rather than assume: whether the downloaded checkpoints carry the same terms as the code, since weights are hosted on Hugging Face and ModelScope rather than in this repository, and whether the datasets referenced by the benchmark scripts have their own licences that constrain redistribution of derived results. This is not legal advice.

On maintenance, the GitHub metadata shows the repository is not archived, but the last push date was not available in the repository information reviewed, so no claim about how frequently it is updated can be made. What the README does show is a dense release cadence in 2026: bug fixes on 2026-04-24 and 2026-06-28, an accelerated path noted on 2026-04-27, benchmark scripts on 2026-05-25, and a long-video demo on 2026-04-29. Version 0.1.0 in pyproject.toml indicates the package has not reached a stable release, so expect the CLI flags and the checkpoint set to move. Budget for re-testing after upgrades, particularly around the attention backend: the two KV cache fixes both changed numerical behaviour, which means results are not comparable across those boundaries.

Editorial conclusion

Adopt LingBot-Map if you have a CUDA 12.8 machine, a downloaded lingbot-map checkpoint, and a pipeline that already produces ordered video frames; the interactive viser viewer at port 8080 makes the first run cheap to evaluate. Do not adopt it if you need CPU-only inference, if you plan to build the batch rendering pipeline on a PyTorch version other than 2.8.0, or if you expect a hosted service, since the README lists no hosted endpoint and no Android or mobile client. Before committing, verify that your GPU has enough memory for the sequence length you care about, confirm whether the FlashInfer or SDPA attention backend suits your hardware, and check the Apache-2.0 terms in LICENSE.txt against how you intend to ship the outputs.

Frequently asked questions

What is LingBot-Map?

It is a feed-forward 3D foundation model for reconstructing scenes from streaming data, built around a Geometric Context Transformer. The README describes it as unifying coordinate grounding, dense geometric cues and long-range drift correction in one streaming framework.

How do you use LingBot-Map?

Create a Python 3.10 conda environment, install PyTorch 2.8.0 with the CUDA 12.8 wheels, run pip install -e ., optionally install flashinfer-python, then run demo.py with --model_path and --image_folder. The demo opens a viser viewer at http://localhost:8080.

What is lingbot map?

It is the same project: a Python package named lingbot-map, version 0.1.0, distributed under Apache-2.0, with model checkpoints hosted on Hugging Face and ModelScope rather than in the repository.

Official sources

  1. Official README
  2. Project repository