# SAM-Track: Segment and Track Anything with SAM and DeAOT

> SAM-Track pairs Segment Anything key-frame masks with DeAOT propagation to track multiple objects through video, either interactively or automatically. The repository is a research codebase with a Gradio WebUI, a Dockerfile, and a 2026-07-03 last push.

**z-x-yang/Segment-and-Track-Anything** — An open-source project dedicated to tracking and segmenting any objects in videos, either automatically or interactively. The primary algorithms utilized include the Segment Anything Model (SAM) for key-frame segmentation and Associating Objects with Transformers (AOT) for efficient tracking and propagation purposes.

- Repository: https://github.com/z-x-yang/Segment-and-Track-Anything
- Stars: 3,139 · Forks: 355
- Language: Jupyter Notebook
- License: AGPL-3.0
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/z-x-yang-segment-and-track-anything

## What SAM-Track solves and who it is for

Video object segmentation usually forces a choice. You either annotate every frame by hand, or you accept a detector that misses objects when they leave the frame and come back. SAM-Track takes a third route: you define the objects once on a key frame, and the tracker carries those masks forward. The README describes the split plainly. SAM handles automatic and interactive key-frame segmentation, while DeAOT (Decoupling features in Associating Objects with Transformers, NeurIPS 2022) handles multi-object tracking and propagation. The pipeline also lets SAM detect and segment new objects as the video runs.

The intended user is someone with a video and a question about what moves through it. The README lists street views, AR, cells, animations and aerial shots among the demo scenarios. The interactive WebUI accepts clicks and strokes for one object, and later versions add text prompts and multi-object selection. That matters for cell microscopy, where a threshold-based segmenter cannot separate touching cells, and for aerial footage, where object appearance changes with altitude.

The repository is not a product. It is a research codebase with a Colab notebook, a technical report on arXiv (2305.06558) and a set of tutorial markdown files. The README states that version 1.0 of the WebUI is a developer version and invites bug reports. Treat that as the honest description of maturity.

## How SAM and DeAOT split the work

The architecture is a two-stage loop. On a key frame, SAM produces masks, either from a click, a stroke, a text prompt, or from Grounding-DINO detection in demo_instseg.ipynb. Those masks become the initial object set. DeAOT then propagates them through subsequent frames, matching the current frame against reference frames held in long-term memory.

The memory design is the part worth understanding before you tune anything. The README documents two arguments for AOT-L: long_term_memory_gap and max_len_long_term. The gap controls how often the model adds a new reference frame to long-term memory. A smaller gap means more reference frames and more matching work; the README says a proper value helps performance, which is a polite way of saying you have to tune it. max_len_long_term caps how many memory frames are kept. When the cap is reached, the oldest frame is discarded and a new one is added, which the README frames as a guard against memory explosion on long videos.

That cap is a real trade-off, not a free win. Discarding the oldest reference frame means an object that leaves the scene early and returns much later may be matched against a memory that no longer contains its earlier appearance. The repository exposes the knob rather than solving the problem. If your footage has long occlusions, that is the parameter to watch.

The top-level layout reflects the same split. SegTracker.py and aot_tracker.py sit beside the sam/ and aot/ directories, with app.py as the WebUI entry point and seg_track_anything.py as the script-level interface.

## Installing SAM-Track and running a first track

The repository ships a Dockerfile, which is the shortest path because it pins the CUDA and PyTorch versions together. The base image is pytorch/pytorch:2.0.1-cuda11.8-cudnn8-devel, so you need an NVIDIA GPU and a working container runtime before anything else.

```bash
docker build -t sam-track .
docker run --gpus all -p 7860:7860 sam-track
```

The Dockerfile sets CMD to python app.py, so the container starts the Gradio WebUI directly. The README does not state the port that app.py binds to, so check the Gradio launch output in the container logs rather than assuming a mapping.

If you prefer a local install, the Dockerfile lists the exact dependency set, including transformers==4.30.2, timm==0.4.5, gradio==3.39.0 and opencv-python==4.10.0.84. The SAM package is installed from the vendored directory, and GroundingDINO is cloned from IDEA-Research and installed in editable mode.

```bash
pip install -e sam
pip install -e GroundingDINO --no-build-isolation
```

The README points to prepare.py and a set of checkpoints, and the tutorials cover the WebUI versions. The repository does not document a single download command for the model weights, so read prepare.py before running it and confirm which checkpoints it expects. For a first real use, the Colab notebook linked at the top of the README is the lowest-friction way to see the pipeline work before you commit GPU time locally.

## Where SAM-Track breaks down

The clearest limitation is compute. SAM and DeAOT both run on GPU, and the Dockerfile assumes CUDA 11.8. There is no documented CPU path. If your only machine is a laptop without an NVIDIA GPU, this project is the wrong tool, and the Colab notebook is the only realistic way to try it.

The second limitation is the long-term memory cap. As described above, max_len_long_term discards old reference frames. On a long video with repeated entrances and exits, tracking identity can drift. The README presents the cap as a memory safeguard, which it is, but it is also an accuracy ceiling that the user inherits.

The third is operational. The README does not document rollback, migration between versions, or a supported upgrade path between the 1.0, 1.5 and 1.6 WebUI versions. The tutorial files are versioned separately, which suggests the interfaces changed enough to need separate instructions. If you build a workflow on the 1.6 audio-grounding feature, you are on your own when the next version lands.

Finally, the licence. The project is AGPL-3.0. The repository includes a licenses.md file, which suggests third-party components carry their own terms. AGPL-3.0 has network-use implications that matter if you plan to expose a modified version as a service. That is a question for your own counsel, not something the README resolves.

## SAM-Track compared with tracking-by-detection pipelines

The obvious alternative is a tracking-by-detection stack: run a detector such as YOLO on every frame, then associate detections across frames with a tracker. That approach is well understood, runs faster in many configurations, and does not need a key-frame annotation step. Its weakness is the one SAM-Track targets. A detector only sees what its training labels cover, and it produces boxes, not pixel masks. If you need per-pixel masks for a cell or a drone target, you are back to a separate segmentation model.

SAM-Track inverts the order. Segmentation comes first, on a key frame, and tracking propagates those masks. The advantage is that the object set is defined by you, not by a label taxonomy, so it can include a specific cell or a specific vehicle. The cost is that the first frame needs human input unless you use the Grounding-DINO path in demo_instseg.ipynb, and that path depends on the detector's vocabulary anyway.

A second comparison point is the broader Segment Anything ecosystem. The upstream SAM repository is the segmentation model alone; it does not track across frames. SAM-Track's contribution is the composition with DeAOT and the interactive layer on top. If your problem is single-image segmentation, the upstream model is the simpler dependency and you do not need this repository at all.

## Maintenance, versions and what the licence means for you

The last push to the default branch was on 2026-07-03, so the repository is not abandoned. The most recent tagged release is v1.6 from 2024-04-25, which added the audio-grounding feature that tracks the sound-making object in a video's soundtrack. The gap between the last release tag and the last push is worth noting: work continues on main without a corresponding tag, so pinning to a release gives you a stable point but not the newest code.

Upgrade cost is dominated by the pinned dependency set. transformers==4.30.2, timm==0.4.5 and gradio==3.39.0 are all older than current releases. Moving any of them forward means re-testing the SAM and DeAOT integration, and the repository does not document a compatibility matrix. The Dockerfile is the safest way to hold a known-good combination.

On the licence, AGPL-3.0 is a strong copyleft licence with a network clause. The practical implication for engineers is that distributing a modified version, or offering it as a network service, triggers source-availability obligations. The repository also carries licenses.md, which implies bundled components have their own terms. Read both files before you build anything commercial on top, and take advice from someone qualified rather than from a README.

## Conclusion

SAM-Track fits research groups and engineers who need multi-object video masks and can supply a CUDA GPU, since the Dockerfile pins pytorch 2.0.1 with cuda11.8. It does not fit teams that need a supported product with documented rollback, or CPU-only deployments. Verify the checkpoint download step in prepare.py and the AGPL-3.0 licence implications first.

## FAQ

### What does SAM-Track use SAM for?

SAM handles automatic and interactive key-frame segmentation, producing the initial masks for the objects you want to follow. DeAOT then propagates those masks through the rest of the video.

### Can I run SAM-Track without a GPU?

The Dockerfile is built on pytorch/pytorch:2.0.1-cuda11.8-cudnn8-devel and the README documents no CPU path. The Colab notebook is the practical way to try it without local GPU hardware.

### How do I install SAM-Track?

The repository includes a Dockerfile that installs the pinned dependencies, builds the vendored sam package with pip install -e sam, and clones GroundingDINO. Building that image is the shortest documented path.

### What are long_term_memory_gap and max_len_long_term?

They are AOT-L arguments. The gap sets how often a new reference frame enters long-term memory, and max_len_long_term caps how many frames are stored, discarding the oldest when the cap is reached to avoid memory explosion on long videos.

### What is SAM-Track licensed under?

The repository is AGPL-3.0 and also ships a licenses.md file, which indicates bundled third-party components carry their own terms. Check both before commercial use.

## Sources

- [Issues](https://github.com/z-x-yang/Segment-and-Track-Anything/issues)
- [License: AGPL-3.0](https://github.com/z-x-yang/Segment-and-Track-Anything/blob/main/LICENSE)
- [README](https://github.com/z-x-yang/Segment-and-Track-Anything/blob/main/README.md)
- [Releases](https://github.com/z-x-yang/Segment-and-Track-Anything/releases)
- [z-x-yang/Segment-and-Track-Anything on GitHub](https://github.com/z-x-yang/Segment-and-Track-Anything)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/z-x-yang-segment-and-track-anything
