LingBot-Vision: Self-Supervised ViT Backbones for Dense Spatial Perception
Self-supervised learning for spatial perception
At a glance
- What is it?
- LingBot-Vision is a family of self-supervised Vision Transformer backbones trained with a boundary-centric masked modeling objective, designed as drop-in visual encoders for depth estimation, semantic segmentation, and video object segmentation. It is aimed at computer vision researchers and robotics engineers who need strong frozen features for dense downstream tasks without task-specific pretraining.
- Who is it for?
- LingBot-Vision is a strong choice for dense prediction tasks such as depth estimation, segmentation, and video object segmentation where frozen backbone features are sufficient and task-specific pretraining is too costly. Researchers who need a well-understood baseline with published documentation should consider DINOv2 first, since LingBot-Vision's technical report and downstream evaluation results are still limited to the paper.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 84 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Boundary-Centric Self-Supervised Pretraining for Dense Tasks
LingBot-Vision is a family of self-supervised Vision Transformer (ViT) backbones built for dense spatial perception. Standard self-supervised ViT pretraining tends to produce semantically rich features that classify objects well but blur boundaries and spatial gradients. LingBot-Vision addresses this by using masked boundary modeling as its primary training objective: during pretraining, the teacher model discovers boundary tokens, and the student learns to reconstruct masked features while being guided by both the semantic and boundary structure of the teacher.
The result is a backbone that the README describes as capturing semantic grouping and geometric structure at the same time. Frozen patch tokens from LingBot-Vision can be fed directly into lightweight dense readout heads for tasks such as depth estimation, semantic segmentation, and video object segmentation, without fine-tuning the backbone itself.
The library targets computer vision researchers who want to evaluate boundary-aware pretraining on dense benchmarks, and robotics engineers who use LingBot-Depth 2.0 (which uses LingBot-Vision as its visual encoder at the ViT-L/16 and ViT-g/16 scales). The repository includes backbone weights for all four size variants and a PCA visualization demo.
The Masked Boundary Modeling Objective
The pretraining approach in LingBot-Vision is described in the technical report at arxiv.org/abs/2607.05247. The core idea is that boundary tokens, discovered by the teacher model, carry a signal about where object regions transition. By including this signal in the masking and reconstruction target, the model learns patch representations that are sensitive to edges and geometric transitions, not just semantic categories.
The training pipeline follows a teacher-student distillation setup at the ViT-g/16 scale (roughly 1.1B parameters). The teacher is trained on the full boundary-centric objective. From this giant teacher, smaller student backbones are distilled: ViT-L, ViT-B, and ViT-S. All released checkpoints are backbone-only .pt files, stored as model.pt in each Hugging Face model repository.
Config files for each variant are packaged under lingbot_vision/configs/ and selected automatically when calling load_pretrained_backbone. The model is downloaded from Hugging Face on first use if a local checkpoint is not provided.
Installing LingBot-Vision and Loading a Backbone
LingBot-Vision requires Python 3.10 or later and PyTorch 2.0 or later. A CUDA-capable GPU is recommended for large-model inference. Clone the repository and create a conda environment:
git clone https://github.com/robbyant/lingbot-vision.git
cd lingbot-visionconda create -n lingbot-vision python=3.10 -y
conda activate lingbot-visionInstall dependencies and the package itself:
python -m pip install -r requirements.txt
python -m pip install -e .The dependencies from requirements.txt are torch, torchvision, numpy, opencv-python-headless, pillow, omegaconf, and huggingface_hub.
To run the PCA visualization demo with the large backbone:
./scripts/run_pca_demo.sh \
--config-file lingbot_vision/configs/lbot_vision_vitl.yaml \
--ckpt /path/to/model.pt \
--input examples/example.png \
--out outputs/pca_demo \
--size 512 \
--mode square \
--dtype bf16The demo maps the top three PCA components of the patch tokens to RGB and writes both PCA-only and input/PCA panel visualizations to the output directory. For CPU-only inference, pass --dtype fp32 --device cpu. The README notes that images are resized to the specified size, aligned to the model's patch size, and normalized with ImageNet statistics.
Model Variants: ViT-S Through ViT-g
LingBot-Vision ships four backbone sizes. LingBot-Vision-Giant uses ViT-g/16 with SwiGLU, fp32 RoPE, and 4 register tokens, with an embedding dimension of 1536. It is the highest-quality option for dense features but the most expensive at inference time. LingBot-Vision-Large, marked as the recommended default in the README, uses ViT-L/16 distilled from Giant with embedding dimension 1024. It balances dense feature quality with practical inference cost. LingBot-Vision-Base uses ViT-B/16 with embedding dimension 768 for balanced inference cost. LingBot-Vision-Small uses ViT-S/16 with embedding dimension 384 and is described as suitable for lightweight demos and downstream use.
All weights are backbone-only .pt checkpoints available on Hugging Face under the robbyant collection and on ModelScope. The load_pretrained_backbone function accepts a variant parameter: giant, large, base, or small; if omitted, it defaults to large. You can also pass a local directory or an explicit Hugging Face model repo path to load_pretrained_backbone.
Patch tokens from any variant have shape [B, H * W, C], where H and W are the patch-grid dimensions and C is the embedding dimension for that variant.
Downstream Applications: Depth, Segmentation, and Video
The README documents four dense downstream categories where frozen LingBot-Vision features are evaluated: dense feature visualization, depth estimation, semantic segmentation, and video object segmentation.
For depth estimation, the README states that frozen patch tokens expose spatial structure to lightweight dense readouts. This matches the use case of LingBot-Depth 2.0, which replaces its encoder with LingBot-Vision at ViT-L/16 and ViT-g/16 scales and trains on a corpus the README describes as 150M RGB-D samples. The README notes substantial performance gains over the previous version on mirror and glass scenes, where raw sensor depth is typically missing.
For video object segmentation, the README describes a training-free approach: token matching and label propagation using frozen features, without adapting the backbone to the video task. This makes LingBot-Vision useful as a zero-shot spatial feature extractor for tracking applications.
None of the downstream evaluation numbers are included in the README. The technical details and benchmark results are in the paper at arxiv.org/abs/2607.05247.
LingBot-Vision Versus DINOv2 as a Dense Backbone
DINOv2, released by Meta AI Research, is a widely used self-supervised ViT model that also produces strong frozen features for dense tasks. DINOv2 uses a teacher-student distillation approach with a self-supervised objective combining several loss terms. It has been broadly evaluated on depth estimation and segmentation benchmarks and has published results across multiple datasets.
The core difference in approach is the training objective. DINOv2 does not include an explicit boundary-centric term; its patch features are semantically rich but not specifically trained to distinguish region transitions. LingBot-Vision's masked boundary modeling adds that signal during pretraining. Whether this difference translates into consistent gains on specific downstream benchmarks depends on the task and dataset; those comparisons are in the LingBot-Vision paper, not in the README.
For practitioners who need a backbone with extensive third-party evaluation, reproducible baselines, and broad community adoption, DINOv2 is the more tested choice. LingBot-Vision is a better fit for teams building on the LingBot ecosystem (LingBot-Depth, LingBot-VLA) or specifically interested in boundary-aware spatial features.
Apache-2.0 License and Project Maintenance
LingBot-Vision is licensed under Apache-2.0, which permits use in commercial and proprietary products without requiring the source of derivative works to be open. The LICENSE and LEGAL.md files are in the top-level repository. The pyproject.toml confirms the Apache-2.0 license and Python 3.10 requirement.
The last push to the repository was on July 8, 2026. The repository has no GitHub releases. Updates are taken directly from the main branch. The paper is at arxiv.org/abs/2607.05247 and the Hugging Face collection at huggingface.co/collections/robbyant/lingbot-vision. Weights are also mirrored on ModelScope under the Robbyant collection.
Editorial conclusion
LingBot-Vision is a strong choice for dense prediction tasks such as depth estimation, segmentation, and video object segmentation where frozen backbone features are sufficient and task-specific pretraining is too costly. Researchers who need a well-understood baseline with published documentation should consider DINOv2 first, since LingBot-Vision's technical report and downstream evaluation results are still limited to the paper. Anyone running inference on CPU should expect slow runtimes at large scales: a CUDA GPU is recommended for the large and giant variants.
Frequently asked questions
What dense vision tasks can LingBot-Vision backbones be used for?
According to the README, frozen LingBot-Vision features work for depth estimation, semantic segmentation, video object segmentation via token matching, and dense feature visualization. LingBot-Vision is also used as the visual encoder for LingBot-Depth 2.0 at ViT-L/16 and ViT-g/16 scales.
How large is the LingBot-Vision-Giant model?
LingBot-Vision-Giant uses a ViT-g/16 backbone with roughly 1.1 billion parameters, SwiGLU activations, fp32 RoPE, and 4 register tokens, with an embedding dimension of 1536. The README describes it as the highest-quality option for dense features and recommends ViT-L for most practical uses.
Do I need to fine-tune LingBot-Vision for downstream tasks?
No. The README describes LingBot-Vision as a drop-in visual encoder for dense downstream tasks using frozen patch tokens. The PCA visualization demo and the depth estimation use case both operate on frozen features without adapting the backbone. Task-specific heads are trained on top of the frozen backbone.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/robbyant-lingbot-vision)