# MambaVision: running NVIDIA's hybrid Mamba-Transformer vision backbone from PyPI

> MambaVision is an official PyTorch implementation of a hybrid Mamba-Transformer backbone for image classification, detection and segmentation. The Hugging Face path is a few lines of code; the repository also ships a Dockerfile and separate training trees for detection and segmentation.

**NVlabs/MambaVision** — [CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone

- Repository: https://github.com/NVlabs/MambaVision
- Website: https://arxiv.org/abs/2407.08083
- Stars: 2,235 · Forks: 150
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvlabs-mambavision

## What MambaVision is for, and who should reach for it

MambaVision is a vision backbone, not an application. It produces image features, and those features can be attached to a classification head or fed to a detector or segmenter. The repository describes it as a hybrid Mamba-Transformer backbone with a hierarchical architecture that employs both self-attention and mixer blocks. The intended user is someone training or fine-tuning a vision model who wants a different accuracy/throughput trade-off than a plain ViT or a plain convolutional network offers.

The paper's claim is a SOTA Pareto-front in Top-1 accuracy and throughput. That claim is about a specific measured setting, and the README does not restate the measurement conditions, so treat it as the authors' result rather than a portable guarantee. The practical reason to look at this repository is the combination of released checkpoints, a pip package, and separate code trees for object detection and semantic segmentation. If you only need a classifier on ImageNet-like data, the Hugging Face path is enough. If you need dense prediction, the extra trees matter more than the classification accuracy number.

## The hybrid block: why self-attention sits next to the mixer

The architectural point is the mixer block. The README describes it as creating a symmetric path without SSM to enhance the modeling of global context. State space models process a sequence with a recurrence-like scan, which is efficient but does not give every token direct access to every other token in the way attention does. MambaVision keeps both: self-attention layers and mixer blocks arranged in a hierarchy, so the network has stages rather than a single uniform stack.

That hierarchy is what makes the backbone reusable beyond classification. The Hugging Face feature-extraction path exposes the outputs of each of the four stages plus the final averaged-pool features that are flattened. The README states the multi-scale stage outputs are the ones used for downstream tasks such as classification and detection. So the design choice to keep four stages is not cosmetic; it is the reason the same checkpoint can be attached to a detection head. A flat transformer backbone would give you one resolution and leave you to build the pyramid yourself.

The trade-off is dependency weight. mamba-ssm is a pinned dependency, and it is the component that ties the project to CUDA hardware. That is the price of the SSM path, and it is not optional in the default install.

## Installing MambaVision from PyPI and classifying one image

The README's quickest route is the pip package. Installing it pulls the pinned dependencies listed in setup.py, including transformers, timm and mamba-ssm.

```bash
pip install mambavision
```

Once installed, the model loads through the transformers auto classes. The README uses trust_remote_code=True because the architecture is resolved from the checkpoint repository rather than from a class shipped inside transformers itself.

```python
from transformers import AutoModelForImageClassification

model = AutoModelForImageClassification.from_pretrained("nvidia/MambaVision-T-1K", trust_remote_code=True)
```

The README's end-to-end example builds a preprocessing transform from the model's own config values rather than hardcoding ImageNet statistics, which is worth copying because the config carries mean, std, crop_mode and crop_pct. It then moves the model and the input to CUDA and reads the logits.

```python
from PIL import Image
from timm.data.transforms_factory import create_transform
import requests

model.cuda().eval()
url = 'http://images.cocodataset.org/val2017/000000020247.jpg'
image = Image.open(requests.get(url, stream=True).raw)
input_resolution = (3, 224, 224)
transform = create_transform(input_size=input_resolution, is_training=False,
                             mean=model.config.mean, std=model.config.std,
                             crop_mode=model.config.crop_mode, crop_pct=model.config.crop_pct)
inputs = transform(image).unsqueeze(0).cuda()
outputs = model(inputs)
logits = outputs['logits']
predicted_class_idx = logits.argmax(-1).item()
print("Predicted class:", model.config.id2label[predicted_class_idx])
```

With the README's sample image, the printed label is brown bear, bruin, Ursus arctos. If you see a different label, the preprocessing is the first thing to check, not the weights.

The repository also carries a Dockerfile based on pytorch/pytorch:2.6.0-cuda12.6-cudnn9-devel, which installs tensorboardX, mamba-ssm, timm, einops, transformers, requests and Pillow and copies the code into /app. Note that the Dockerfile pins timm==1.0.9 while setup.py pins timm==1.0.15; that divergence is visible in the two files and is worth resolving deliberately rather than by accident.

## Feature extraction and the dense-prediction trees

For anything other than classification, use the AutoModel path. The README shows that this returns the four stage outputs plus the flattened pooled features, and states that the stage outputs are what downstream detection and classification consume. The input resolution is not fixed: the README says MambaVision supports any input resolutions, and the example sets input_resolution explicitly as a (3, H, W) tuple. Changing that tuple is the supported way to change resolution, provided the transform is rebuilt with it.

The repository layout confirms the ambition beyond classification: object_detection/ and semantic_segmentation/ are top-level directories, and the news entries announce that detection code and models were released on 2025-06-07 and semantic segmentation code and models on 2025-06-10. Those trees are separate from the pip package, so installing mambavision does not give you a training pipeline for either task. You get the backbone and the pretrained weights; the task heads and training loops live in those directories and are the part you should read before assuming a supported training recipe.

## Where MambaVision is the wrong choice

The licence is the first hard boundary. setup.py declares the licence as NVIDIA Source Code License-NC. That is a non-commercial licence, and it is the reason the README points business inquiries at an NVIDIA Research Licensing form. If your use is commercial, this repository is not a drop-in, regardless of how well the model performs. The repository's LICENSE file is the authoritative text and the metadata here marks it NOASSERTION, so read the file rather than trusting any summary, including this one.

The second boundary is hardware. mamba-ssm is pinned at 2.2.4 and the requirements file asks for torch>=2.6.0+cu124. The README examples call model.cuda(). There is no documented CPU or Apple-silicon path. If you need to run inference on a laptop without an NVIDIA GPU, this is the wrong tool, and the absence of such instructions in the README is itself the signal.

The third is version pinning. transformers is pinned at 4.50.0 and timm at 1.0.15 in setup.py. The README's examples depend on config attributes such as crop_mode and crop_pct being present on the loaded model, and trust_remote_code=True means the checkpoint repository supplies code that runs in your process. In a project with other pinned ML dependencies, expect to spend time on resolution rather than on modelling.

## How it differs from a plain ViT backbone

The obvious comparison is a plain vision transformer from timm, which is also the library MambaVision uses for its transforms. A ViT applies self-attention uniformly across the whole stack. MambaVision splits the work: mixer blocks carry part of the sequence modelling, and self-attention is retained alongside them, arranged in four hierarchical stages. The stated motivation is global context modelling without relying on attention everywhere.

That difference has practical consequences. A ViT backbone typically gives you a single-resolution feature map, so detection or segmentation work requires building a pyramid on top. MambaVision's four stage outputs are already multi-scale, which is why the same checkpoint serves detection and segmentation heads. On the other side, a timm ViT installs without mamba-ssm, so it runs wherever PyTorch runs, including CPU. The choice is between a heavier, CUDA-oriented dependency graph with built-in multi-scale features and a lighter, more portable one that leaves the pyramid to you.

If you want an SSM-free comparison inside the same family of ideas, a hierarchical convolutional or transformer backbone in timm is the fair test, because it isolates whether the Mamba mixer is buying you anything on your data.

## Maintenance, releases and what upgrading costs

The repository is not archived, and the last push was on 2026-03-11. The most recent release listed is v1.2.0 from 2025-07-22, following a pip release 1.1.0 from 2025-03-29. setup.py reports version 1.2.0, so the packaging metadata and the tagged release agree.

The upgrade cost is dominated by the pins. Because mamba-ssm, transformers and timm are all fixed versions, moving any one of them is a coordinated change, not a patch bump. The Dockerfile and setup.py disagree on timm (1.0.9 versus 1.0.15), which means the container and the pip package are not the same environment; pick one as your reference and align the other. The Python requirement is >=3.9, and the classifiers in setup.py list 3.7 through 3.13, which is broader than the python_requires floor and should not be read as tested coverage.

On licensing, the NVIDIA Source Code License-NC is a non-commercial licence and the README routes commercial inquiries to NVIDIA Research Licensing. That is a business decision to make before you build on the weights, and it is not something to settle by reading a summary.

## Conclusion

Adopt MambaVision if you need a pretrained backbone with hierarchical multi-scale features for classification, detection or segmentation, and you can accept the non-commercial NVIDIA Source Code License-NC plus the CUDA-bound mamba-ssm dependency. Skip it if you need a permissively licensed model or CPU-only inference. Verify the licence text, your CUDA and PyTorch combination, and whether the detection and segmentation trees match your framework before committing.

## FAQ

### What is MambaVision?

MambaVision is the official PyTorch implementation of a hybrid Mamba-Transformer vision backbone, accepted to CVPR 2025. It is a backbone that produces image features for classification, and the repository also ships detection and segmentation code and models.

### How do I install MambaVision?

The README's quick start installs the pip package with pip install mambavision, which pulls the pinned dependencies from setup.py. The repository also provides a Dockerfile based on a PyTorch CUDA image.

### Can I use MambaVision models through Hugging Face?

Yes. The README loads checkpoints such as nvidia/MambaVision-T-1K with AutoModelForImageClassification.from_pretrained and trust_remote_code=True, and AutoModel for feature extraction. The feature-extraction path returns the four stage outputs plus flattened pooled features.

### Does MambaVision require a GPU?

The README examples move the model to CUDA with model.cuda(), and mamba-ssm is a pinned dependency while requirements.txt asks for torch>=2.6.0+cu124. No CPU or Apple-silicon path is documented.

### What licence does MambaVision use?

setup.py declares the NVIDIA Source Code License-NC, a non-commercial licence, and the README directs business inquiries to an NVIDIA Research Licensing form. The repository's LICENSE file is the authoritative text.

## Sources

- [Issues](https://github.com/NVlabs/MambaVision/issues)
- [NVlabs/MambaVision on GitHub](https://github.com/NVlabs/MambaVision)
- [Project website](https://arxiv.org/abs/2407.08083)
- [README](https://github.com/NVlabs/MambaVision/blob/main/README.md)
- [Releases](https://github.com/NVlabs/MambaVision/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvlabs-mambavision
