Open-source project
ChaoningZhang/MobileSAM avatar
ChaoningZhang/MobileSAM

MobileSAM: swapping SAM's 611M parameter encoder for a 5M Tiny-ViT

This is the official code for MobileSAM project that makes SAM lightweight for mobile applications and beyond!

5,884 stars588 forksJupyter NotebookApache-2.0

At a glance

What is it?
The official repository for Faster Segment Anything, where the image encoder is replaced and the mask decoder is left untouched, so existing SAM projects can switch with minimal change.
Who is it for?
MobileSAM is a one-substitution project, and that is both its strength and its limit. Because the mask decoder, the pre-processing, the post-processing and every interface are inherited unchanged from SAM, a project already built on SAM can point at a different encoder and keep everything else, which is why adapters appeared in projects as varied as Grounding-SAM, Inpaint-Anything, AnyLabeling and the Stable Diffusion WebUI extension within days of the release.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 155 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One substitution, stated precisely

The README's central claim is narrow and checkable: MobileSAM keeps exactly the same pipeline as the original SAM and changes only the image encoder. The heavyweight ViT-H encoder, 611M parameters at 452ms per image, is replaced by a Tiny-ViT of 5M parameters at 8ms.

The mask decoder is unchanged, and the README gives the numbers for both sides rather than asking you to take it on faith. Original SAM and MobileSAM both use a prompt-guided mask decoder of 3.876M parameters running at 4ms. Add those together and the whole pipeline comparison is 615M parameters at 456ms for SAM against 9.66M parameters at 12ms for MobileSAM, split 8ms encoder and 4ms decoder.

Those numbers are what makes the substitution viable rather than cosmetic. A mask decoder is small enough to keep, and an image encoder is large enough that shrinking it by two orders of magnitude is where the time goes.

The README also states the quality position carefully, saying MobileSAM performs on par with the original SAM at least visually, and backs the visual comparison with example masks produced from a point prompt and from a box prompt.

Why adapting from SAM is close to a drop-in change

The FAQ entry on adapting is the reason downstream projects were able to support MobileSAM so quickly. Since MobileSAM keeps exactly the same pipeline as the original SAM, pre-processing, post-processing and all other interfaces are inherited from the original. The assumption is that everything is the same except a smaller image encoder, and the README concludes that projects using SAM can adapt to MobileSAM with almost zero effort.

The ecosystem timeline in the README is the evidence. Within about a week of the release, dates from 2023-06-27 through 2023-07-03, the list records Grounding-SAM adding support with a Grounded-MobileSAM variant, AnyLabeling using it for auto-labelling, the Stable Diffusion WebUI segment-anything extension supporting it, Inpaint-Anything adopting it for faster inpainting, Personalize-SAM adding one-shot personalisation on top, joliGEN using it for lightweight mask refinement with diffusion and GAN, and MobileSAM-in-the-Browser running it in a browser on a local PC or phone. SegmentAnythingin3D used it for 3D segmentation.

Two of those entries are worth a second look because they answer questions people actually ask. MobileSAM-in-the-Browser is the existence proof that this runs client-side, which is what the ONNX export support in the README is aimed at. And Grounded-MobileSAM pairs the small encoder with a text grounding stage, which is the shape most current promptable segmentation work takes.

MobileSAMv2 and the different problem it solves

The repository name covers two projects, and the README pins them to separate papers. MobileSAM, on arXiv as Faster Segment Anything Towards Lightweight SAM for Mobile Applications, replaces the heavyweight image encoder for faster segment anything, or SegAny. MobileSAMv2, Faster Segment Anything to Everything, replaces the grid-search prompt sampling in SAM with object-aware prompt sampling for faster segment everything, or SegEvery.

That is a different axis of improvement. The first project makes one prompt cheap. The second makes many prompts cheap, which matters when you want masks for an entire image rather than the object you clicked on. Grid search over prompt positions means evaluating the decoder repeatedly across a grid; object-aware sampling means choosing fewer, better candidate points.

The v2 work lives in its own `MobileSAMv2/` directory at the top of the repository, separate from the `mobile_sam/` package, which is consistent with the two being separately importable.

Neither paper's numbers are reproduced in the README beyond the framing above, so if you are choosing between v1 and v2 the decision is about which problem you have: a fast interactive segmenter on constrained hardware, or a fast batch segmenter over whole images.

The FastSAM comparison and what mIoU is measuring

The README positions MobileSAM against FastSAM, a concurrent lightweight alternative, on two axes. Size and speed: MobileSAM is around 7 times smaller and around 5 times faster, with whole pipeline figures of 68M parameters at 64ms for FastSAM against 9.66M at 12ms for MobileSAM.

The second axis is more interesting methodologically. The README argues MobileSAM aligns better with the original SAM than FastSAM does, and explains why the comparison is not symmetric: FastSAM is suggested to work with multiple points, so the authors compare mIoU using two prompt points at different pixel distances.

| Pixel distance | 100 | 200 | 300 | 400 | 500 | | FastSAM mIoU | 0.27 | 0.33 | 0.37 | 0.41 | 0.41 | | MobileSAM mIoU | 0.73 | 0.71 | 0.74 | 0.73 | 0.73 |

Higher mIoU indicates higher alignment with SAM's own output. The shape of the numbers matters as much as the values: MobileSAM's alignment stays flat across pixel distances while FastSAM's climbs steadily, which is consistent with FastSAM losing agreement as prompts move apart from each other. Flat is the desirable property here, since it means the lightweight encoder does not degrade as prompts spread out.

This is a self-reported comparison in a project README, so treat it as a claim about the design rather than as an independent benchmark.

Packaging, dependencies and the demo

`setup.py` is short and tells you how the package is meant to be consumed. It registers as `mobile_sam` version 1.0, has no required install dependencies, and excludes `notebooks` from `find_packages`. Optional extras carry the weight: `all` pulls matplotlib, pycocotools, opencv-python, onnx and onnxruntime, while `dev` pulls flake8, isort, black and mypy.

python
setup(
    name="mobile_sam",
    version="1.0",
    install_requires=[],
    packages=find_packages(exclude="notebooks"),
)

An empty `install_requires` means PyTorch is not pulled in for you, which is the right call given the CUDA build matrix, but it also means `pip install mobile_sam` alone will not run anything. The README states the real requirements separately: `python>=3.8`, `pytorch>=1.7` and `torchvision>=0.8`, with a link to the PyTorch install page and a strong recommendation to install both with CUDA support.

The onnx and onnxruntime entries in the `all` extra are the ones that line up with the ONNX export support the README highlights, and with the mobile deployment story the project name implies.

For trying it, there are two paths. A CPU demo runs on Hugging Face, where the README says it takes around 3s on a Mac i5 CPU and is slower on the hosted interface. A local demo lives in the `app/` directory of the repository. There is also a `notebooks/` directory, which is excluded from the package for the reason the setup file makes obvious.

Repository state and where the model weights live

The tree is small and legible. Alongside `mobile_sam/`, `MobileSAMv2/`, `app/`, `notebooks/`, `scripts/` and `assets/`, there is a `weights/` directory, which is where a model repository keeps its checkpoints, and `linter.sh` for the style tooling mirrored by the dev extras in `setup.py`.

The governance files are the usual research-repository set: `LICENSE`, `CONTRIBUTING.md`, `CODE_OF_CONDUCT.md` and a `Member.txt`. The licence is Apache-2.0, and `setup.py` carries a Meta Platforms copyright header on the source, which is a reminder that this code descends directly from SAM's own codebase.

On maintenance, the picture is mixed and worth stating plainly. The repository is not archived, it has 5,876 stars and 587 forks, and the last push was on 2026-05-05. There are no tagged releases, so the only version marker is the `1.0` in `setup.py`. The README's training question also carries a note that the training code will be available soon, which tells you this is a repository organised around inference and integration rather than around reproducing the model.

With 119 open issues, the backlog is the largest of any project in this batch, which is what you would expect from a research artefact that many downstream tools depend on. Projects that adapted it are listed in the README, and that list, rather than the repository's own issue count, is the better guide to how actively the idea is being used.

Editorial conclusion

MobileSAM is a one-substitution project, and that is both its strength and its limit. Because the mask decoder, the pre-processing, the post-processing and every interface are inherited unchanged from SAM, a project already built on SAM can point at a different encoder and keep everything else, which is why adapters appeared in projects as varied as Grounding-SAM, Inpaint-Anything, AnyLabeling and the Stable Diffusion WebUI extension within days of the release. The numbers behind that claim are in the README: the whole pipeline drops from 615M parameters and 456ms to 9.66M and 12ms, with 8ms of that in the encoder and 4ms in the decoder. MobileSAMv2 extends the same idea to batch prompt sampling for segment everything, replacing grid search with object-aware sampling. The open questions are elsewhere: the repository has 119 open issues and no tagged releases, and `pushedAt` is 2026-05-05, so start with the `mobile_sam/` package and the `app/` demo rather than expecting an actively versioned library.

Frequently asked questions

How does the SAM model work?

SAM splits the work into an image encoder and a prompt-guided mask decoder. The encoder turns the image into features, and the decoder turns a prompt such as a point or a box into a mask. MobileSAM keeps the decoder identical to SAM at 3.876M parameters and replaces only the encoder, going from a 611M ViT-H taking 452ms to a 5M Tiny-ViT taking 8ms.

What license is MobileSAM released under?

Apache-2.0, as recorded for the repository and consistent with `setup.py` in the tree. Because the code descends from Meta's SAM, the `setup.py` source header also carries a Meta Platforms copyright notice.

What is the difference between MobileSAM and MobileSAMv2?

MobileSAM replaces SAM's heavyweight image encoder with a lightweight Tiny-ViT to speed up segment anything, or SegAny. MobileSAMv2 instead replaces the grid-search prompt sampling with object-aware prompt sampling to speed up segment everything, or SegEvery. They live in separate directories, `mobile_sam/` and `MobileSAMv2/`.

Official sources

  1. ChaoningZhang/MobileSAM on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/chaoningzhang-mobilesam.svg)](https://hysenlabs.com/projects/chaoningzhang-mobilesam)