seamless_communication: Meta's SeamlessM4T Models for Speech and Text Translation
Foundational Models for State-of-the-Art Speech and Text Translation
At a glance
- What is it?
- seamless_communication is a Python library from Meta's FAIR team that packages the Seamless family of multilingual translation models. It covers around 100 languages and supports five translation task types, but its core dependency restricts installation to Linux x86-64 and Apple Silicon machines.
- Who is it for?
- Research engineers and ML practitioners who need multilingual speech-to-speech or streaming translation at a research scale, and who are running Linux x86-64 or Apple Silicon hardware, have a solid starting point here. Developers who need only speech-to-text transcription on general hardware should look at OpenAI Whisper instead: it covers more platforms without the fairseq2 constraint.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 22 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
SeamlessM4T as the Foundation of the Seamless Family
seamless_communication is the Python repository for a family of translation models developed by Fundamental AI Research (FAIR) at Meta. The centerpiece is SeamlessM4T, described in the README as a foundational all-in-one massively multilingual and multimodal machine translation model. It supports around 100 languages across both speech and text modalities.
The library ships three distinct components built on top of SeamlessM4T. SeamlessExpressive is a speech-to-speech translation model that the README says preserves elements of prosody such as speech rate and pauses, while also preserving voice style across languages. SeamlessStreaming supports simultaneous translation and streaming automatic speech recognition for around 100 languages. A fourth model simply called Seamless combines expressive and streaming capabilities into a unified system for expressive, real-time speech-to-speech translation.
SeamlessM4T v2 introduced the UnitY2 architecture, which the README describes as improving over v1 in both translation quality and inference latency for speech generation tasks. The W2v-BERT 2.0 speech encoder, described in section 3.2.1 of the accompanying paper, was open-sourced alongside the v2 release and is central to the Seamless models' speech processing pipeline. All four models are also available as HuggingFace Spaces demos.
Five Translation Tasks Covered by SeamlessM4T
SeamlessM4T handles five distinct translation tasks, and knowing which task a workflow needs determines which CLI command or model configuration to use.
Speech-to-speech translation (S2ST) takes audio as input and returns translated audio. Speech-to-text translation (S2TT) takes audio and returns translated text. Text-to-speech translation (T2ST) takes text and returns translated audio in another language. Text-to-text translation (T2TT) translates text to text across languages. Automatic speech recognition (ASR) transcribes speech within a single language without translation.
SeamlessStreaming adds simultaneous capabilities to S2ST, S2TT, and ASR, meaning it can begin producing output before the full input has been received. SeamlessExpressive only covers S2ST, since its core goal is preserving prosody and voice character through the audio-to-audio pipeline.
The README also notes that SeamlessM4T is available through the HuggingFace Transformers library, which provides a different integration path for users already working within that ecosystem. The setup.py entry points register four CLI commands: m4t_predict, m4t_evaluate, m4t_finetune, and m4t_prepare_dataset, as well as expressivity_predict and expressivity_evaluate for SeamlessExpressive.
Installing seamless_communication and Running First Inference
The README describes fairseq2 as a prerequisite with a direct caveat: pre-built packages are available only for Linux x86-64 and Apple Silicon Mac computers. fairseq2 also depends on libsndfile, which may not be installed by default and must be added separately through the system package manager before the Python install.
With those prerequisites in place, the library installs from the repository root:
pip install .The README notes that Whisper is automatically installed as a dependency for computing inference metrics. Whisper in turn requires the command-line tool ffmpeg, which must be installed separately from most system package managers.
Once installed, running inference uses the m4t_predict command. For speech-to-speech translation:
m4t_predict <path_to_input_audio> --task s2st --tgt_lang <tgt_lang> --output_path <path_to_save_audio>For text-to-text translation:
m4t_predict <input_text> --task t2tt --tgt_lang <tgt_lang> --src_lang <src_lang>The README directs readers to the inference README under src/seamless_communication/cli/m4t/predict for the full list of supported languages per task and modality. The repository also includes a Seamless_Tutorial.ipynb notebook from the NeurIPS 2023 Seamless EXPO, covering the entire model suite.
Platform and Dependency Constraints
The fairseq2 dependency is the sharpest practical constraint in this repository. It has pre-built wheels only for Linux x86-64 and Apple Silicon Mac, which rules out Windows, Linux ARM, and Intel Mac without a source build of fairseq2. The README explicitly flags this and points to the fairseq2 repository for instructions when installation fails.
The setup.py lists a specific version constraint for fairseq2: fairseq2==0.2.*. This is a pinned minor version, which means upgrading to a newer fairseq2 minor release would require an update to seamless_communication itself. The same file pins datasets==2.18.0 and simuleval~=1.1.3, and requires sonar-space==0.2.*. These version pins reduce the chance of API breakage but also mean the library may not install alongside other packages that require different versions of those dependencies.
The Python minimum is 3.8, as set in pyproject.toml's mypy section and confirmed by setup.py's python_requires field. The dependency on PyTorch comes through torchaudio, which also means GPU drivers and a compatible CUDA version are relevant for hardware acceleration, though the README does not document GPU configuration directly.
OpenAI Whisper as a Narrower Alternative
OpenAI Whisper is the most direct comparison for the ASR and speech-to-text use cases. Whisper is an open-source speech recognition model that transcribes and translates speech using a standard pip install, without requiring fairseq2 or a specific CPU architecture. It runs on Windows, Linux, and macOS without the hardware restrictions that affect seamless_communication.
The difference in scope is significant. Whisper handles ASR and speech-to-text translation but does not offer speech-to-speech output, prosody preservation, or simultaneous streaming translation. SeamlessExpressive's core goal of preserving how something is said across languages has no equivalent in Whisper's design. SeamlessStreaming's simultaneous output also addresses a different requirement than Whisper's batch transcription model.
For a team that needs only to transcribe audio or produce text translations from speech, Whisper is a simpler path that runs on more platforms. For teams that need speech-to-speech output, voice style preservation, or simultaneous translation output, seamless_communication is the more relevant tool, provided the platform constraint is acceptable.
Licence Terms and Repository Status
The repository contains two distinct licence files: SEAMLESS_LICENSE and MIT_LICENSE. The setup.py describes the licence as Creative Commons. The NOASSERTION value in the repository metadata reflects that the licence is custom and not a standard SPDX identifier. Developers integrating these models into commercial products should read the SEAMLESS_LICENSE and ACCEPTABLE_USE_POLICY files directly, as the acceptable use terms are separate from the code licence.
The last push to this repository was on 2026-09-08. The repository is not archived. It has no GitHub releases: the release channel is the HuggingFace model hub rather than GitHub tags.
The repository includes demo directories under demo/ for SeamlessM4T v2, SeamlessExpressive, and a DINO-based component, suggesting the codebase supports both research use and demonstration deployments. The tests/ directory and pyproject.toml with pytest configuration indicate test coverage exists, though the scope of that coverage is not documented in the README.
Editorial conclusion
Research engineers and ML practitioners who need multilingual speech-to-speech or streaming translation at a research scale, and who are running Linux x86-64 or Apple Silicon hardware, have a solid starting point here. Developers who need only speech-to-text transcription on general hardware should look at OpenAI Whisper instead: it covers more platforms without the fairseq2 constraint. Before committing to seamless_communication, verify that fairseq2 0.2.x builds successfully in your environment and that libsndfile is installed, since those two steps account for most reported installation failures.
Frequently asked questions
What translation tasks does SeamlessM4T support?
SeamlessM4T supports five tasks: speech-to-speech translation (S2ST), speech-to-text translation (S2TT), text-to-speech translation (T2ST), text-to-text translation (T2TT), and automatic speech recognition (ASR). SeamlessStreaming adds simultaneous streaming capabilities for S2ST, S2TT, and ASR.
Which platforms can install seamless_communication?
Installation requires Linux x86-64 or Apple Silicon Mac, because the core fairseq2 dependency only ships pre-built wheels for those two platforms. The README notes that source builds are possible for other setups but directs readers to the fairseq2 repository for those instructions.
What is the difference between SeamlessExpressive and SeamlessStreaming?
SeamlessExpressive is a speech-to-speech model focused on preserving prosody, speech rate, pauses, and voice style across languages. SeamlessStreaming supports simultaneous translation, starting output before the full input is received, and covers S2ST, S2TT, and ASR tasks.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookresearch-seamless-communication)