VisionMamba: A PyTorch Implementation of Bidirectional Visual State Space Models
Implementation of Vision Mamba from the paper: "Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model" It's 2.8x faster than DeiT and saves 86.8% GPU memory when performing batch inference to extract features on high-res images
At a glance
- What is it?
- Vision Mamba is a Python package that reimplements the Vision Mamba paper's bidirectional state space model in PyTorch. It ships a single Vim class, installs from PyPI, and comes with no training script.
- Who is it for?
- VisionMamba suits researchers and engineers who want to instantiate the paper's bidirectional state space model in PyTorch without writing the SSM blocks themselves, and who are comfortable with a beta package that has no training script. It is the wrong choice if you need a trained checkpoint, a training loop, or a documented accuracy number, because the README provides none of those.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What VisionMamba Is For
VisionMamba is an implementation of the paper Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model. The README states the model is 2.8x faster than DeiT and saves 86.8% GPU memory when performing batch inference to extract features on high-res images. Those figures come from the paper's framing and are quoted in the README, not reproduced by any benchmark script in the repository.
The package targets engineers who want to instantiate a visual state space model in PyTorch and run a forward pass over image tensors, without reimplementing the bidirectional scan themselves. The pyproject.toml classifies it as Development Status 4 - Beta, so it is a research artifact rather than a production library. The README's own Todo list is honest about the gap: creating a training script for ImageNet is still an unchecked box, as is a variant for facial recognition.
How the Vim Model Is Structured
The public surface is one class, Vim, exported from the vision_mamba package. The constructor takes dim, heads, dt_rank, dim_inner, d_state, num_classes, image_size, patch_size, channels, dropout and depth. Input is a standard image tensor shaped (batch_size, channels, height, width), and the forward pass returns a tensor whose shape the README prints rather than describes.
Two constructor arguments deserve attention because they are not self-explanatory. dt_rank is documented in the README comment as the rank of the dynamic routing matrix, and d_state as the dimension of the state vector. These are the state space model's core knobs: dt_rank controls the low-rank parameterization of the input-dependent step size, and d_state sets how much latent state each scan carries. The heads argument is named after attention heads even though the architecture is a state space model, which is an artifact of the code borrowing a transformer-style signature. Nothing in the README explains how heads interacts with the SSM blocks, so treat it as an open question to resolve by reading vision_mamba/ in the repository.
The dependency list in pyproject.toml is short: python ^3.10, zetascale, einops and torch, with a poetry-core build backend. requirements.txt additionally lists swarms, which is not in the poetry dependency block. That mismatch is worth knowing before you pin versions from one file or the other.
Installing VisionMamba and Running a First Forward Pass
The README gives one installation command. It installs the published package from PyPI:
pip install vision-mambaAfter that, the usage example constructs the model and runs a forward pass on a random tensor. This is the README's own snippet, and it is the fastest way to confirm the install works. Note the Python version constraint: pyproject.toml declares python = "^3.10", so run this under Python 3.10 or newer.
import torch
from vision_mamba import Vim
x = torch.randn(1, 3, 224, 224)
model = Vim(
dim=256,
heads=8,
dt_rank=32,
dim_inner=256,
d_state=256,
num_classes=1000,
image_size=224,
patch_size=16,
channels=3,
dropout=0.1,
depth=12,
)
out = model(x)
print(out.shape)The README prints both out.shape and the tensor itself. The shape is the thing to check: with num_classes=1000 and a single image in the batch, you should get a classification-style output over 1000 classes, not a feature map. If you are extracting features rather than classifying, num_classes is the argument to change, and the README does not document which intermediate tensor to tap instead. The repository also contains an example.py at the top level, which is the second place to look for a runnable entry point.
Where VisionMamba Falls Short
There is no training script. The README lists Create training script for imagenet as an unfinished todo item, which means the package gives you a model definition and nothing else. You supply the data pipeline, the loss, the optimizer schedule and the evaluation loop. For a project whose headline claim is about inference efficiency on high-resolution images, the absence of a training path means you cannot reproduce the paper's reported numbers from this repository alone.
There are no released checkpoints. The recent releases list is empty, and the README never mentions pretrained weights or a model hub. If your goal is to classify images tomorrow, this package does not get you there; you would need to train from scratch or load weights from elsewhere.
The dependency situation is also loose. torch is unpinned in both pyproject.toml and requirements.txt, and the two files disagree on whether swarms is required. There is no lock file in the repository listing. A fresh install therefore resolves whatever the current torch release is, which is a reproducibility risk for a paper implementation.
Finally, the package metadata is thin in ways that matter for evaluation. The pyproject keywords are artificial intelligence, deep learning, optimizers and Prompt Engineering, which describe a different kind of project. The documentation URL points back at the GitHub repository. There is no separate docs site to fall back on when the README runs out.
VisionMamba Versus VMamba and Other Visual State Space Models
The obvious alternative is VMamba, which appears in the related searches as VMamba: visual state space model. Both descend from the Mamba state space model line, but they differ in how they handle the two-dimensional structure of an image. Vision Mamba, the paper this package implements, uses bidirectional state space sequences and is framed in the README around inference speed and GPU memory against DeiT. VMamba is a separate project with its own codebase; it is not a module inside this package, and the README here does not reference it. If you are deciding between them, the practical difference is that this repository gives you a compact Vim class and an MIT licence, while VMamba is a distinct implementation you would evaluate on its own terms.
The other comparison worth naming is the vision transformer. The related searches include vision mamba vs vision transformer, and the README's efficiency claims are stated relative to DeiT, which is a vision transformer. The architectural difference is the point: a transformer attends over all patch pairs, while a state space model scans the patch sequence with a recurrent state. That is why the memory claim exists at high resolution, where attention cost grows with the number of patches. The trade-off is that the scan imposes an ordering on patches, which is why the model is bidirectional in the first place.
Maintenance Status and Licence
The repository is not archived. The last push was on 2026-08-29, which is recent enough that the codebase is still moving. There are no retrieved releases, so version 0.1.0 in pyproject.toml is the only version identifier available, and there is no changelog in the repository listing to tell you what changed between pushes.
Upgrade cost is low but not zero. The package has three or four direct dependencies and no lock file, so a pip install vision-mamba will pull current torch. If you pin torch yourself, the unpinned declaration in pyproject.toml will not fight you. The lint tooling in the repository (ruff at line-length 70, black at line-length 70 with target py38) is inconsistent with the declared python = "^3.10" floor, which suggests the formatting config was copied from another project rather than tuned here.
The licence is MIT, declared in both LICENSE and the pyproject.toml license field. MIT permits commercial use and modification with attribution. That is a permissive licence, but it says nothing about the paper's own terms or about any weights you might obtain elsewhere. If you plan to ship a product built on this code, the licence question is settled at the code level and open at the model level.
Editorial conclusion
VisionMamba suits researchers and engineers who want to instantiate the paper's bidirectional state space model in PyTorch without writing the SSM blocks themselves, and who are comfortable with a beta package that has no training script. It is the wrong choice if you need a trained checkpoint, a training loop, or a documented accuracy number, because the README provides none of those. Before adopting it, check that pip install vision-mamba resolves on your Python version and that the Vim constructor's dim, depth and patch_size values produce the output shape your downstream code expects.
Frequently asked questions
What is VisionMamba?
It is a PyTorch implementation of the Vision Mamba paper, which describes a bidirectional state space model for visual representation learning. The package exposes a single Vim class that takes image tensors and returns model output.
How does VisionMamba compare with a vision transformer?
The README states the model is 2.8x faster than DeiT and saves 86.8% GPU memory when performing batch inference to extract features on high-res images. DeiT is a vision transformer, so the comparison is stated in terms of inference speed and memory rather than accuracy.
How do I install VisionMamba?
The README gives one command, pip install vision-mamba, and the package requires Python 3.10 or newer according to pyproject.toml. After installing, import Vim from vision_mamba and construct it with the arguments shown in the README.
Does VisionMamba ship pretrained weights or a training script?
No. The README lists Create training script for imagenet as an unfinished todo item, and there are no retrieved releases or any mention of checkpoints. You get the model definition only.
What licence does VisionMamba use?
MIT, declared in both the LICENSE file and the pyproject.toml license field. That permits commercial use and modification with attribution, though it does not address the paper or any external weights.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kyegomez-visionmamba)