# EfficientSAM3: distilled SAM3 encoders that ship as checkpoints, not as a package

> A research distillation pipeline that shrinks SAM3's vision and text encoders into student models published as .pt files. Its install notes ask for Python 3.10 and PyTorch 2.0 while the build metadata requires Python 3.12 and torch 2.7, and its declared version is 0.1.0 while the releases are at v0.4.0.

**SimonZeng7108/efficientsam3** — EfficientSAM3 compresses SAM3 into lightweight, edge-friendly models via progressive knowledge distillation for fast promptable concept segmentation and tracking.

- Repository: https://github.com/SimonZeng7108/efficientsam3
- Website: https://simonzeng7108.github.io/efficientsam3/
- Stars: 688 · Forks: 57
- Language: Jupyter Notebook
- License: not declared
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/simonzeng7108-efficientsam3

## The install notes ask for Python 3.10 and torch 2.0, the build metadata requires 3.12 and 2.7

Two different environment contracts live in this repository. The prerequisites under the installation heading state Python 3.10+, PyTorch 2.0+, and CUDA 11.8+ for GPU support. The build metadata in pyproject.toml declares requires-python as >=3.12, and the torch pin appears inside the stage1 extra as torch>=2.7.0 with torchvision>=0.18.0 beside it. So a reader who trusts the prose installs on 3.10 and hits the metadata wall at build time, and a reader who trusts the metadata is told nothing at all about CUDA in the place it matters. The rest of the metadata is unusually specific in a different way: the base dependencies are research utilities with floors, namely timm>=1.0.17, numpy>=1.26.4, tqdm, ftfy==6.1.1, regex, iopath>=0.1.10, typing_extensions, huggingface_hub and psutil. Note that torch is not among them. A base install gives you no tensor library at all.

## The project declares version 0.1.0 while its releases run to v0.4.0

The metadata names the distribution efficientsam3 at version 0.1.0 and describes it as EfficientSAM3 research utilities and Stage-1 distillation pipeline. The published history is three tags: v0.4.0-efficientsam3ft-20260611 for the stage 3 fine-tuned models and the full PCS release, v0.3.0-efficientsam3.1-20260413 for the EfficientSAM3.1 and SAM3.1-LiteText image models, and v0.2.0-sam3litetext-20260218 for SAM3-LiteText. The tag names carry their own dates, so the sequence is readable, but nothing in the repository reconciles 0.1.0 with v0.4.0, and the description in the metadata stops at stage 1 while the newest tag is a stage 3 release. The homepage tells the same story twice over: the metadata records the GitHub repository as its Homepage URL, while the repository record points at a GitHub Pages site.

## Only stage1* and sam3* packages are installed, so the rest of the tree stays out of site-packages

One line decides what an install actually contains: include is set to stage1* and sam3*. The root holds data/, docs/, eval/, images/, sam3/, sam3_checkpoints/, stage1/, stage1_geometry_finetune/ and stage3/, plus five separate readme files named README.md, README_dataset.md, README_stage1.md, README_stage1_finetune.md and README_stage3.md. Everything outside the two include patterns stays where it is, so the stage 3 pipeline and the stage 1 geometry fine-tuning code are present in the clone but not importable from an installed copy. That is a sensible split for a research tree and an easy trap for anyone who pip installs and then looks for the evaluation scripts. A checkpoint directory sitting at the root, sam3_checkpoints/, points the same way: this repository is arranged for reading and reproducing, not for shipping. The same split shows up in the dependency extras. The stage1 extra is the heavy one, carrying decord, mmengine, pycocotools, yacs, opencv-python, scikit-image, tensorboard, einops, hydra-core, submitit, fvcore, fairscale, mmcv and pyyaml in addition to torch, and a separate sam1 extra exists for one dependency, segment-anything, needed only by the SAM1-compatibility wrapper under sam3/backbones/efficientvit. The pip name and the import name differ there, which the metadata comments point out.

## The quick start example ends in the middle of a statement

The worked example builds a model and then stops short. Reproduced as given:

```python
from sam3.model_builder import build_efficientsam3_image_model
from sam3.model.sam3_image_processor import Sam3Processor
from PIL import Image

# Load EfficientSAM3 TV-M model (uses TinyViT vision encoder + MobileCLIP-S0 text encoder)
model = build_efficientsam3_image_model(
    checkpoint_path="efficientsam3_tinyvit.pt",
    backbone_type="tinyvit",
    model_name="11m",
    text_encoder_type="MobileCLIP-S0",
    text_encoder_context_length=16,
    load_from_HF=False,
)

# Process image
processor = Sam3Processor(model)
image = Image.open("your_image.jpg").convert("RGB")
state = proc
```

The last line assigns from a name that is never defined, so the segment that would actually produce output from the opened image is absent. The second example, for SAM3-LiteText, stops inside the string value of text_encoder_type, partway through MobileC. Two details in the complete example are worth copying deliberately: model_name is set to 11m, an internal configuration label rather than a model size, and load_from_HF is False even though every download in the model zoo is a HuggingFace URL, so the checkpoint has to sit on disk first.

## The decoder and Other columns shrink too, although only the encoders are described as compressed

The prose frames the work as compressing the vision encoder and the text encoder. The tables say more happened. For the full models, the reference ImageSAM3 is given as vision 463M, text 354M, transformer 30.3M and other 14.2M, totalling 861.5M. The students come out at vision 22.2M for EV-M, 25.6M for RV-M and 28.3M for TV-M, but the decoder also falls from 30.3M to 21.0M and other from 14.2M to 3.5M, giving totals of 89.2M, 92.7M and 95.3M, described as 90 percent, 89 percent and 89 percent smaller. So a third of the reduction arrives from the mask decoder and the segmentation head plus scoring. The tables also disagree with themselves about naming: the column header reads Decoder while the explanatory note calls the same column Transformer.

## Context length 16 and 32 variants have identical parameter counts

The SAM3-LiteText table pairs every MobileCLIP variant with two context lengths, and the counts do not move between them. LiteText-S0-16 and LiteText-S0-32 both list vision 463.0M, text 42.5M, decoder 30.3M, other 14.2M and 550.0M parameters, 36 percent smaller than ImageSAM3. LiteText-S1-16 and LiteText-S1-32 both sit at 571.0M with text 63.5M, and LiteText-L-16 and LiteText-L-32 both sit at 631.3M with text 123.8M. The file names confirm the pairing, with ctx16 and ctx32 appended to each checkpoint, and the quick start sets text_encoder_context_length to 16. Since the parameter total is unchanged, the context length costs nothing in weights and only affects how much text the encoder sees, which makes 32 the better default for caption-heavy prompts and 16 the cheaper one to run.

## SAM3-LiteText is announced released in February and live on HuggingFace in April

The update log carries two entries for the same model. The 2026/02/18 line says SAM3-LiteText released, reducing the text encoder by 88 percent with similar performance. The 2026/04/19 line says SAM3-LiteText is live on HuggingFace and accepted by ICMR2026, with thanks to two named contributors. So the model existed for roughly two months before it was reachable on the hub, and anyone who followed the February announcement found no download there. The release titles add a naming wrinkle: the v0.3.0 tag is titled EfficientSAM3.1 and SAM3.1-LiteText, while the model zoo, the update log and the checkpoint paths all call the same weights SAM3-LiteText. The newest tag is also titled with a full PCS release, and PCS is not defined anywhere in the visible text.

## Stage 1 weights arrived in three separate drops between December and January

The older updates, folded into a details block, show how the stages were released piecemeal. On 2025/12/02 the stage 1 image encoder weights appeared, named RepViT, TinyViT and EfficientViT. On 2025/12/08 the stage 1 text encoder weights appeared, named MobileCLIP S0, S1 and MobileCLIP2 L. On 2026/01/11 stage 1 geometry-prompt fine-tuned weights followed. Six days later in the same month the sequence continued with the SAM3-LiteText release, and the newest drop, dated 2026/06/11, is the stage 3 fine-tuned models EV-M, RV-M and TV-M trained on 5 percent of SA1B data with SACap labels. The last push to the repository is dated 2026-08-11, two months after that newest tag, so the branch has moved past everything published.

## Conclusion

Treat this as a research repository with usable checkpoints rather than as a library. The distilled students are real and the numbers are laid out per model, which makes it easy to pick one: EV-M at 89.2M, RV-M at 92.7M, TV-M at 95.3M, or a LiteText variant that trades away the 463M vision encoder only if you do not need it. Before you build anything on it, resolve three inconsistencies yourself. Follow the stricter of the two environment statements, Python 3.12 and torch 2.7, not the 3.10 and 2.0 in the prose. Treat the Apache-2.0 line in the build metadata as the only statement of terms, since the repository records no license and the root carries no LICENSE file. And copy the quick start by hand rather than trusting it, because the worked example stops in the middle of a statement.

## FAQ

### What does EfficientSAM3 actually compress?

The vision encoder and the text encoder, into EfficientViT, RepViT or TinyViT students paired with MobileCLIP variants. The full models are EV-M at 89.2M parameters, RV-M at 92.7M and TV-M at 95.3M, compared with 861.5M for ImageSAM3.

### How do I install EfficientSAM3 from source?

Three commands: clone the repository, change into the efficientsam3 directory, then install the stage1 extra in editable mode. The listed prerequisites are Python 3.10+, PyTorch 2.0+ and CUDA 11.8+ for GPU support, though the build metadata requires Python 3.12 or newer.

### Which license applies to EfficientSAM3?

The build metadata declares Apache-2.0 in its license field. The repository record shows no license, and no LICENSE file appears among the top level entries, so that declaration is the only statement of terms available.

### Does SAM3-LiteText shrink the vision encoder as well as the text encoder?

No. It keeps the SAM3 vision encoder at 463.0M and replaces only the text encoder, which drops from 354M to 42.5M, 63.5M or 123.8M depending on the MobileCLIP variant. That text encoder reduction is where the 88 percent figure comes from.

### Where are the EfficientSAM3 checkpoints published?

On HuggingFace, one .pt file per row of the model zoo, with full models under an efficientsam3_ft path and the LiteText variants under sam3_litetext. The quick start sets load_from_HF to False, so the file is expected on disk at the path given to checkpoint_path.

## Sources

- [Issues](https://github.com/SimonZeng7108/efficientsam3/issues)
- [Project website](https://simonzeng7108.github.io/efficientsam3/)
- [README](https://github.com/SimonZeng7108/efficientsam3/blob/main/README.md)
- [Releases](https://github.com/SimonZeng7108/efficientsam3/releases)
- [SimonZeng7108/efficientsam3 on GitHub](https://github.com/SimonZeng7108/efficientsam3)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/simonzeng7108-efficientsam3
