Model or dataset
bytedance/Sa2VA avatar
bytedance/Sa2VA

Sa2VA: pixel-level grounded understanding by pairing SAM-2 with a multimodal LLM

Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)

1,677 stars133 forksPythonApache-2.0

At a glance

What is it?
Sa2VA is a research codebase from ByteDance that combines SAM-2 segmentation with MLLM backbones such as InternVL and Qwen-VL. It is aimed at teams reproducing or extending dense grounded understanding, not at drop-in production inference.
Who is it for?
Adopt Sa2VA if you are reproducing dense grounded segmentation research or building on the Pixel-LLM family, and you are willing to manage a uv-locked environment and download gated weights yourself. Do not adopt it if you need a packaged inference server or a ComfyUI node, since the README documents neither.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 22 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Sa2VA fills between segmentation models and chat models

A conventional segmentation network returns masks. A multimodal LLM returns sentences. Neither alone answers a prompt like "segment the object the user is describing" while keeping the conversation going across frames. Sa2VA is built for exactly that overlap: the README describes it as "a unified model that marries SAM-2 with MLLMs for dense grounded understanding of images and videos," covering referring segmentation, grounded conversation, visual prompting, and image/video chat.

The audience is narrow and identifiable. This is a research repository, published as IEEE TPAMI 2026, with sibling projects (VRT, SAMTok, SaSaSa2VA, Pixel-SAIL) layered on top of the same environment. If you are evaluating it as an application backend, note that the README points to Hugging Face model collections and a project page rather than to a serving stack. If you are reproducing a paper or training a variant, the repository is organized around that workflow.

How the Pixel-LLM repository is laid out

The top level separates shared infrastructure from per-paper code. `projects/` holds `sa2va`, `vrt_sa2va`, `samtok`, `sasasa2va` and `pixel_sail`; `setup_env.sh` sits at the repository root, and so do `sa2va_eval/`, `tools/`, `vlm/` and `third_parts/`. The environment is defined once, under `projects/sa2va`, and the README states it is shared across the projects. That is the main architectural decision visible from the outside: one dependency graph for five research efforts.

The Sa2VA model itself is a composition. A SAM-2 (and, per `.env.example`, a SAM3 grounding encoder) supplies mask generation; an MLLM backbone supplies language understanding. The README lists InternVL2.5/3 and Qwen2.5-VL/Qwen3-VL as supported backbones, and `.env.example` names a concrete default path, `pretrained/qwen3vl/Qwen3-VL-4B-Instruct`. The `third_parts/sam3` directory is vendored, and the comment in `.env.example` explains why: using the vendored SAM3 tracker keeps `transformers` pinned at 4.57.1. That pin is the kind of detail that determines whether your own fine-tuning code coexists with this repository.

Installing Sa2VA with uv and running your first grounded prompt

Dependencies are managed with `uv`. The README gives the install command for the tool itself, then a helper script that places the virtualenv in `/tmp` and symlinks it back into the project.

bash
curl -LsSf https://astral.sh/uv/install.sh | sh
bash setup_env.sh                 # projects/sa2va, --extra=latest
# bash setup_env.sh sa2va legacy  # InternVL2.5 or earlier

The second argument selects the extra. Use `legacy` when you are working with InternVL2.5 or earlier; otherwise the default `latest` extra is what the script picks. If you prefer to do it by hand, the README shows the equivalent two commands from inside the project directory:

bash
cd projects/sa2va
uv sync --extra=latest            # or --extra=legacy

Before running anything that downloads gated weights, copy the environment template. The README states that `setup_env.sh` loads it automatically.

bash
cp .env.example .env   # then edit .env

The keys you will most likely fill in are `HF_TOKEN` (for gated models such as Qwen-VL and for datasets) and `OPENROUTER_API_KEY`, which the comment describes as optional unless you run the OpenRouter-based evaluators used by the VRT and VER eval scripts. Training paths are also environment-driven: `SA2VA_QWEN3VL_PATH` defaults to `pretrained/qwen3vl/Qwen3-VL-4B-Instruct` and `SA2VA_DATA_ROOT` defaults to `./data/`. Activate the environment with `source projects/sa2va/.venv/bin/activate` and run training or evaluation from the repository root, as the README instructs. What you should see after `uv sync` is a resolved environment matching `uv.lock`; the README does not document the console output of a first training run, so treat that step as project-specific and follow the per-project README under `projects/sa2va`.

Ref-SAV, gated weights and the friction you should expect

The clearest limitation is data and access. `.env.example` sets `SA2VA_INCLUDE_REFSAV=0` and comments that it should be set to 1 "only after SA-V (sam_v_full) is downloaded from Meta to enable Ref-SAV." In other words, a documented capability is off by default and depends on an external download that the repository does not host. The same pattern applies to weights: `HF_TOKEN` exists because some models are gated.

Hardware assumptions are baked into the defaults. `SA2VA_ACCUM` defaults to 8, described as the "16-gpu recipe"; the comment states that for 32 GPUs you set it to 4 to keep the effective batch at 128. That is a real constraint on anyone planning to fine-tune on a single workstation. The dependency pin on `transformers` 4.57.1, imposed by the vendored SAM3 tracker, is another one: if your surrounding stack needs a different version, you are choosing between them.

Where is this the wrong tool? If you need a segmentation API today, a standalone SAM-family checkpoint plus a thin wrapper will get you further with less setup. Sa2VA's value is the joint language-and-mask behaviour, and you pay for it with a research-grade environment.

How Sa2VA differs from OMG-LLaVA and from SAM-2 alone

OMG-LLaVA is the natural comparison point, and the difference is architectural. OMG-LLaVA attaches a general perception module to an LLM and asks the LLM to emit pixel-level queries that the perception model executes. Sa2VA's README frames the design as marrying SAM-2 with MLLMs, and the repository exposes the segmentation side as a configurable encoder: `.env.example` documents `SA2VA_SAM3_PATH` and `SA2VA_SAM3_IMG_SIZE` (native 1008) for the SAM3 grounding encoder used by `sa2va_qwen3_4b_sam3.py`. The backbone is likewise swappable across InternVL2.5/3 and Qwen2.5-VL/Qwen3-VL, which is a different bet from a single fixed perception stack: more combinations to test, more freedom to match an existing model you already run.

Against SAM-2 on its own, the difference is scope rather than mechanism. SAM-2 segments what you point at; it does not hold a conversation about what it segmented. Sa2VA adds that layer, and the sibling projects show where it leads: SAMTok reduces a mask to a two-token interface any MLLM can emit, and VRT builds object-level grounded reasoning on top, shipping the VRT-Bench evaluation set and VRT-80k training data.

Maintenance, licensing and what an upgrade actually costs

The repository is not archived, and the last push was on 2026-09-08, which is recent. Releases are tagged `v1` (2025-10-16) and `v2` (2025-10-17), with the README describing Sa2VA v2 and naming TPAMI 2026, CVPR 2026 and arXiv 2025 venues across the family. The practical upgrade cost is the lockfile. Because dependencies live in `pyproject.toml` with every transitive package pinned in `uv.lock`, moving between versions is a `uv sync` against a different lock rather than a manual `pip install` reconciliation. The README makes the reproducibility argument explicitly: the environment is "recreated exactly with one `uv sync`." The trade-off is that you inherit the pins, including `transformers` 4.57.1 whenever the vendored SAM3 tracker is in play.

On licensing, the repository is Apache-2.0. That covers the code in this repository. It does not automatically cover model weights, datasets, or the vendored `third_parts/sam3` code, each of which may carry its own terms; the README does not spell those out, so check the model and dataset cards on Hugging Face before you ship anything. This is a description of what the repository states, not legal advice.

Editorial conclusion

Adopt Sa2VA if you are reproducing dense grounded segmentation research or building on the Pixel-LLM family, and you are willing to manage a uv-locked environment and download gated weights yourself. Do not adopt it if you need a packaged inference server or a ComfyUI node, since the README documents neither. Before committing, verify that the uv extra you need (latest or legacy) matches your backbone, that your HuggingFace token can reach the gated Qwen-VL weights, and whether SA2VA_INCLUDE_REFSAV must be flipped to 1 for Ref-SAV.

Frequently asked questions

What is Sa2VA?

Sa2VA is a unified model from ByteDance that combines SAM-2 with multimodal LLMs for dense grounded understanding of images and videos, covering referring segmentation, grounded conversation, visual prompting and image/video chat. The README describes it as the core of the wider Pixel-LLM family of projects.

Which MLLM backbones does Sa2VA support?

The README lists InternVL2.5/3 and Qwen2.5-VL/Qwen3-VL as supported backbones. The environment template defaults to a Qwen3-VL path, pretrained/qwen3vl/Qwen3-VL-4B-Instruct.

How do I install Sa2VA?

Install uv, then run bash setup_env.sh from the repository root, or cd into projects/sa2va and run uv sync --extra=latest. Copy .env.example to .env and fill in HF_TOKEN before downloading gated models, since setup_env.sh loads that file automatically.

Is Sa2VA licensed for commercial use?

The repository is Apache-2.0, which covers the code. The README does not state the terms for model weights, datasets or the vendored third_parts/sam3 code, so those need to be checked separately.

Official sources

  1. bytedance/Sa2VA on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bytedance-sa2va.svg)](https://hysenlabs.com/projects/bytedance-sa2va)