SAM-Audio: Meta's Promptable Audio Separator, and What It Costs to Run
The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
At a glance
- What is it?
- SAM-Audio isolates a chosen sound from a mix using text, video masks or time spans. It installs with one pip command, but the large checkpoint is gated and the dependencies come from git.
- Who is it for?
- Adopt SAM-Audio if you need promptable separation where the prompt can be a noun phrase, a mask on a video frame, or a time span, and you can accept a CUDA GPU plus gated checkpoint access. Do not adopt it if you only need stem splitting on music, where Demucs-style fixed stem models are the simpler fit, or if you cannot authenticate to Hugging Face.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 127 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SAM-Audio separates, and for whom
Most audio separation tools answer one question: give me the vocals, or give me the drums. SAM-Audio answers a different one. You name the sound you want, in whatever form is easiest to express, and the model returns that sound as one file and everything else as another. The README describes it as "a foundation model for isolating any sound in audio using text, visual, or temporal prompts."
The audience follows from that. If you have a field recording with a car horn buried in traffic and you can type "car honking", you get the horn. If you have video and the target is visible, you can point at it with a mask instead of describing it. If you already know the timestamp, you can hand over a span. That flexibility is the product. A dubbing pipeline, a dataset-cleaning pass, or a research evaluation of separation quality are the natural fits. Someone who just wants four fixed stems out of a song is not the target user, and the model table in the README, which scores general SFX, speech, speaker, music, and instrument categories separately, is a hint that the design goal is breadth of target rather than one stem type.
How the three prompt paths reach the same separation call
The architecture visible in the repository is a thin Python package around a pretrained checkpoint. The sam_audio package exposes two classes, SAMAudio and SAMAudioProcessor, and the flow is: build a batch from your inputs, move it to CUDA, call model.separate, and read result.target and result.residual. The processor is what normalizes the different prompt types into that batch.
Text prompting passes descriptions. The README is explicit about phrasing: "To match training, please use lowercase noun-phrase/verb-phrase (NP/VP) format for text", with "thunder" given as the preferred form over a full sentence. Visual prompting fills descriptions with an empty string and instead supplies masked_videos built by processor.mask_videos from frames and a mask. Span prompting passes anchors such as [["+", 6.3, 7.0]], where the sign marks whether the span is included or excluded.
Two arguments in separate() change the cost profile. predict_spans lets the model find the time regions itself from the text description, which the README says is "especially helpful for separating non-ambience sound events." reranking_candidates generates k candidate outputs and picks one using ranking models: CLAP for text similarity, Judge for precision, recall and faithfulness, and ImageBind for visual prompting. The README states plainly that these two settings "have a large impact on performance" and that better results come "at the cost of latency and memory." The defaults in the basic example are the cheap path: predict_spans=False, reranking_candidates=1. Underneath, the model and the Judge both depend on Perception-Encoder Audio-Visual (PE-AV), a separate checkpoint family hosted on Hugging Face.
Installing SAM-Audio and running a first text prompt
The README lists two requirements: Python >= 3.11 and a CUDA-compatible GPU, the latter recommended rather than strictly required by the text. pyproject.toml pins requires-python = ">=3.11" and pulls in torch, torchaudio, torchvision, torchcodec, torchdiffeq, transformers>=4.54.0, einops, numpy, pydub and audiobox_aesthetics. Four dependencies are installed straight from git rather than PyPI: dacvae, imagebind, laion-clap and perception-models. That detail matters for anyone building in a sandboxed CI environment where git access to github.com is restricted.
Install from the repository root:
pip install .Before any inference, the checkpoint has to be unlocked. The README warns that you must request access on the SAM Audio Hugging Face repo and then authenticate. It points to the huggingface_hub quick-start and gives hf auth login as the example command after generating an access token:
hf auth loginThe basic example then loads the large checkpoint, builds a batch, and separates:
from sam_audio import SAMAudio, SAMAudioProcessor
import torchaudio
import torch
model = SAMAudio.from_pretrained("facebook/sam-audio-large")
processor = SAMAudioProcessor.from_pretrained("facebook/sam-audio-large")
model = model.eval().cuda()
batch = processor(audios=["<audio file>"], descriptions=["<description>"]).to("cuda")
with torch.inference_mode():
result = model.separate(batch, predict_spans=False, reranking_candidates=1)
torchaudio.save("target.wav", result.target.cpu(), processor.audio_sampling_rate)
torchaudio.save("residual.wav", result.residual.cpu(), processor.audio_sampling_rate)After this runs you should have two files on disk: target.wav holding the isolated sound and residual.wav holding the rest. Replace the description with a lowercase noun phrase such as "man speaking", following the README's NP/VP guidance. The examples directory holds three notebooks, text_prompting.ipynb, span_prompting.ipynb and visual_prompting.ipynb, which are the fastest way to see the visual and span paths in full.
The gated checkpoint and the reranking latency bill
Two constraints will decide whether SAM-Audio fits a project, and neither is hidden.
The first is access. The checkpoints are not open downloads. The README instructs you to request access on the Hugging Face repo and authenticate before use. That is a manual approval step sitting in the middle of any automated build. If your pipeline needs to spin up a fresh container and pull weights without human intervention, that gate is the thing to solve first, not the model code.
The second is the quality and cost dial. The README repeats the warning that predict_spans and reranking_candidates have a large impact. Turning on span prediction and raising reranking_candidates to 8 makes the model generate eight candidate separations and score them with CLAP, Judge or ImageBind before returning one. That is more forward passes, more memory, and more wall-clock time per clip. Nothing in the README quantifies the increase, so the only honest way to size it is to run both configurations on your own audio and time them. The default example deliberately sits at the cheap end.
There is also a phrasing constraint that is easy to miss. The README asks for lowercase NP/VP text to match training. A natural sentence like "the thunder in the background" is not the recommended input. Teams that hand user-written captions to the model should expect to rewrite them into short noun phrases first.
SAM-Audio versus Demucs-style stem separation
The comparison people reach for is Demucs, and the difference is in the interface, not just the weights. Demucs-style separators expose a fixed set of stems, typically vocals, drums, bass and other. You do not describe what you want; the model already decided which categories exist. SAM-Audio inverts that. The target is defined at call time by a text description, a mask over video frames, or a time anchor, and the output is always a pair: target and residual.
That inversion has consequences. A fixed-stem model can be evaluated once against a known track list, and its behaviour is stable across inputs. SAM-Audio's output depends on how you phrase the prompt, which is why the README spends a sentence on NP/VP formatting and why the reranking path exists at all: with many candidate interpretations, the model needs a scorer to pick between them. If your task is "split this song into its standard parts", the fixed-stem approach removes a variable you would otherwise have to manage. If your task is "pull this one specific sound out of this recording", the fixed-stem model has no way to express it. The README's own model table, which reports separate scores for speech, speaker, music and instrument categories, shows the project is aiming at the second task.
Maintenance, licence and the cost of upgrading
The repository is not archived, and the last push was on 2026-05-26, which is within six months of the current date. There are no retrieved releases, and pyproject.toml still declares version 0.1.0, so there is no tagged version history to pin against. That matters more than usual here because four dependencies are fetched from git branches rather than released packages: dacvae, imagebind, laion-clap, and perception-models, the last of which is pinned to the unpin-deps branch. A git dependency has no version number to freeze, so a rebuild months later can resolve to different code than the one you validated. Anyone deploying this should record the commit hashes their environment resolved to at install time.
The licence file is named LICENSE and the README states the project is "licensed under the SAM License." The repository metadata reports the licence as NOASSERTION, meaning the classifier could not map it to a standard identifier. The practical point is that this is not a recognized OSI licence such as Apache-2.0 or MIT, and the SAM License text is the only authoritative source for what it permits. Whether the terms fit a commercial product, a hosted service, or redistribution of derived weights is a question for the LICENSE file and, if the stakes are high, for counsel. The README does not summarize the terms.
Editorial conclusion
Adopt SAM-Audio if you need promptable separation where the prompt can be a noun phrase, a mask on a video frame, or a time span, and you can accept a CUDA GPU plus gated checkpoint access. Do not adopt it if you only need stem splitting on music, where Demucs-style fixed stem models are the simpler fit, or if you cannot authenticate to Hugging Face. Before committing, request access to facebook/sam-audio-large, run the text prompting example with predict_spans=False and reranking_candidates=1 to measure latency, then compare against predict_spans=True with reranking_candidates=8 on your own audio.
Frequently asked questions
What is SAM-Audio?
SAM-Audio is a foundation model from Meta research for isolating any sound in audio using text, visual, or temporal prompts, as described in the README. The repository provides inference code, checkpoint links, and example notebooks.
How do I install SAM-Audio?
The README requires Python >= 3.11 and a CUDA-compatible GPU, and gives pip install . as the install command from the repository root. You must also request access to the checkpoints on the SAM Audio Hugging Face repo and authenticate, for example with hf auth login.
How do I use Meta SAM-Audio?
Load SAMAudio and SAMAudioProcessor from the facebook/sam-audio-large checkpoint, build a batch with the processor from an audio file and a description, then call model.separate inside torch.inference_mode. The result exposes target for the isolated sound and residual for everything else, which you save with torchaudio.save.
Is SAM-Audio open source?
The README says the project is licensed under the SAM License, which is not one of the standard OSI identifiers, and the repository metadata reports the licence as NOASSERTION. The LICENSE file is the authoritative source for what the terms allow.
Is SAM-Audio free?
The repository does not state a price for the code or the checkpoints. Access to the checkpoints is gated: the README requires you to request access on the SAM Audio Hugging Face repo and authenticate before downloading them.
How does SAM-Audio compare with Demucs?
Demucs-style separators expose a fixed set of stems such as vocals, drums and bass, while SAM-Audio takes the target as a prompt: a lowercase noun phrase, a mask over video frames, or a time span. SAM-Audio always returns a target and a residual rather than named stems.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookresearch-sam-audio)