# SAM 2: Meta's Promptable Segmentation Model for Images and Video

> SAM 2 is a transformer-based foundation model from Meta AI FAIR that segments objects in still images and video sequences using point, bounding box, or mask prompts. A streaming memory buffer tracks objects across frames so only a single annotated frame is required per target, but the model demands a CUDA-capable GPU and PyTorch 2.5.1 or higher.

**facebookresearch/sam2** — The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.

- Repository: https://github.com/facebookresearch/sam2
- Stars: 19,938 · Forks: 2,557
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/facebookresearch-sam2

## Promptable Segmentation in Still Images and Video

SAM 2 addresses one specific problem: given a visual prompt on a video frame or a static image, produce a binary segmentation mask for the target object across the entire clip without requiring per-frame annotation. The model treats static images as single-frame videos, so the same inference code path handles both cases. The target users are computer vision researchers building annotation pipelines, video editing tool developers who need stable object masks across frames, and teams building semi-automatic labeling tools for domains such as medical imaging or robotics. For applications that only ever process single photos and need no frame-to-frame consistency, SAM 2 carries more weight than a lighter segmentation library would require.

## Streaming Memory and the Four Checkpoint Sizes

The README describes SAM 2 as a simple transformer architecture with streaming memory for real-time video processing. Rather than reprocessing all previous frames when a new one arrives, the model stores encoded representations of already-seen frames in a memory bank. When it processes the current frame, it attends to both the current-frame features and the memory bank, which lets it recover a tracked object that was momentarily occluded or left the frame.

The model was trained on the SA-V dataset, which the README describes as the largest video segmentation dataset collected to date, built through a model-in-the-loop data engine that iterated human corrections back into training. Four SAM 2.1 checkpoint sizes are available: sam2.1_hiera_tiny.pt, sam2.1_hiera_small.pt, sam2.1_hiera_base_plus.pt, and sam2.1_hiera_large.pt. The SAM 2.1 variants are distinct from the original SAM 2.0 checkpoints and require the latest model code from the repository.

## Installing SAM 2 and Running First Inference

SAM 2 requires Python 3.10 or higher. PyTorch 2.5.1 or higher and TorchVision 0.20.1 or higher must be installed before the package itself. The README recommends a fresh Anaconda environment and installing PyTorch via pip from pytorch.org. The install sequence is:

```bash
git clone https://github.com/facebookresearch/sam2.git && cd sam2
pip install -e .
```

To run the Jupyter example notebooks, which also need matplotlib, OpenCV, and eva-decord, pass the additional extra:

```bash
pip install -e ".[notebooks]"
```

The install step compiles a custom CUDA kernel. If `nvcc` is absent, the build prints `Failed to build the SAM 2 CUDA extension`; the package still installs and basic inference still works, but some post-processing features are limited. The README directs Windows users to WSL with Ubuntu rather than native Windows.

After installation, download a checkpoint. The convenience script fetches all four SAM 2.1 sizes at once:

```bash
cd checkpoints && \
./download_ckpts.sh && \
cd ..
```

With a checkpoint on disk, the image predictor takes a few lines of Python:

```python
import torch
from sam2.build_sam import build_sam2
from sam2.sam2_image_predictor import SAM2ImagePredictor

checkpoint = "./checkpoints/sam2.1_hiera_large.pt"
model_cfg = "configs/sam2.1/sam2.1_hiera_l.yaml"
predictor = SAM2ImagePredictor(build_sam2(model_cfg, checkpoint))

with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
    predictor.set_image(<your_image>)
    masks, _, _ = predictor.predict(<input_prompts>)
```

The `SAM2ImagePredictor` API is designed to resemble the original SAM interface. For video, `SAM2VideoPredictor` handles multi-object tracking and allows adding new objects after tracking has started, a capability added in the December 2024 update.

## VOS Compilation and Multi-Object Tracking

The December 2024 update added support for compiling the entire model graph with `torch.compile` in video object segmentation mode. Activating it requires passing `vos_optimized=True` to `build_sam2_video_predictor`. The README states this produces a major speedup for VOS inference. The option is off by default because compilation adds a one-time startup cost. Teams running repeated inference on long video clips will see the most benefit; teams processing a small number of short clips are unlikely to recoup the overhead.

The same December update changed `SAM2VideoPredictor` to support independent per-object inference, removing the earlier constraint that all tracked objects had to be prompted at the start. New objects can now be added mid-sequence. The training and fine-tuning code lives under `training/` with its own `training/README.md`. The web demo front end and Flask back end are under `demo/` and can be launched with Docker Compose, mapping the frontend to port 7262 and the backend to port 7263.

## Where SAM 2 Falls Short

SAM 2 has no documented CPU-only inference path. The model loads onto CUDA by design, and the example code targets CUDA with bfloat16. Teams working on hardware without a compatible NVIDIA GPU cannot run the model as documented. The CUDA extension build failure is silent enough that users may not notice reduced post-processing functionality until they hit a code path that depends on it.

For purely static images without any tracking requirement, SAM 2 carries a larger footprint than necessary. The README does not document model quantization or ONNX export, so deploying the model in a memory-constrained or latency-sensitive environment requires additional engineering not covered by the repository. The docker-compose web demo requires a GPU with NVIDIA container runtime; the compose file requests `driver: nvidia, count: 1` and will not start correctly without it.

The repository has no GitHub releases. The last code push was on 2026-05-30. Version management relies on the SAM 2.1 checkpoint naming scheme rather than formal release tags.

## SAM 2 Compared to the Original Segment Anything Model

The original SAM (facebookresearch/segment-anything) handles static images only. SAM 2 adds the video tracking path through the streaming memory architecture and the SA-V training data. For purely static-image use cases, SAM 2's image predictor exposes an API that the README describes as closely resembling SAM. The practical difference is that SAM was trained on the SA-1B image dataset, while SAM 2 was trained on a combination of image and video data. Teams already integrated with the SAM Python API can migrate to SAM 2 without rewriting their inference code, but they must upgrade PyTorch to 2.5.1 and replace the original checkpoints with the SAM 2.1 variants. If you previously installed the SAM 2.0 version, the README notes you must first run `pip uninstall SAM-2`, pull the latest code, and reinstall.

## License and Docker Deployment

SAM 2 is released under the Apache-2.0 license. The custom CUDA kernel included in the repository carries a separate license stored in LICENSE_cctorch, so deployments that bundle the compiled extension should review both files. The docker-compose.yaml at the repository root provides a two-service demo deployment. The frontend builds from `demo/frontend` and serves on port 7262. The backend builds from the repository root using `backend.Dockerfile` and serves on port 7263. Key environment defaults include `SERVER_ENVIRONMENT=DEV`, `GUNICORN_WORKERS=1`, `GUNICORN_THREADS=2`, `VIDEO_ENCODE_CODEC=libx264`, `VIDEO_ENCODE_CRF=23`, and `VIDEO_ENCODE_FPS=24`. Anyone deploying the demo beyond a local test environment should review the `SERVER_ENVIRONMENT` default before exposing the service.

## Conclusion

SAM 2 is the right choice for computer vision teams that need semi-automatic video annotation or object tracking where a single prompted frame drives the full clip, provided they have a CUDA GPU and PyTorch 2.5.1 or higher. It is the wrong tool for CPU-only inference environments, for Windows machines without WSL, and for projects where a lighter segmentation library covers the static-image case without the weight of a full video foundation model. Before committing, verify that the SAM 2 CUDA extension builds cleanly, since a silent failure at install time leaves some post-processing features disabled with no runtime warning.

## FAQ

### What is SAM 2 used for?

SAM 2 is used for promptable visual segmentation: given a point, bounding box, or mask prompt on an image or video frame, it returns a binary segmentation mask for the target object. For videos, a streaming memory buffer tracks the object across frames so only a single annotated frame is needed per target.

### What is the difference between SAM 2 and SAM3?

The README does not document a model called SAM3. SAM 2 is described as the successor to the original Segment Anything Model, adding video object segmentation through a streaming memory architecture and training on the SA-V video dataset.

### What is facebookresearch/sam2?

facebookresearch/sam2 is the GitHub repository for SAM 2, Meta AI FAIR's foundation model for promptable visual segmentation in images and videos. It provides model code, four pre-trained SAM 2.1 checkpoints, Jupyter notebooks, training code, and a deployable web demo with Docker Compose.

## Sources

- [facebookresearch/sam2 on GitHub](https://github.com/facebookresearch/sam2)
- [Issues](https://github.com/facebookresearch/sam2/issues)
- [License: Apache-2.0](https://github.com/facebookresearch/sam2/blob/main/LICENSE)
- [README](https://github.com/facebookresearch/sam2/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/facebookresearch-sam2
