Model or dataset
bytedance/SALMONN avatar
bytedance/SALMONN

bytedance/SALMONN: a multi-branch repository of audio and audio-visual LLMs

SALMONN family: A suite of advanced multi-modal LLMs

1,536 stars125 forksUnknownApache-2.0

At a glance

What is it?
SALMONN is not one model but a family of research releases spread across branches, from the ICLR 2024 speech model to ELLSA and video-SALMONN 2. Here is what the repository actually contains, how to reach the code, and where it stops being the right choice.
Who is it for?
Adopt SALMONN if you are working on speech or audio-visual research and want published checkpoints and training code tied to named papers, and if you can accept that each model lives on a separate branch with its own setup. Do not adopt it if you need a single maintained package with a stable API, a pip install, or a support channel; the main branch is an index, not a library, and the README documents no rollback, no versioning policy and no release artifacts.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 37 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What SALMONN is, and what the main branch actually contains

The repository describes itself as the home of the SALMONN model family, which it calls a series of advanced multi-modal large language models. The important structural fact is that the main branch is an index page, not an implementation. Its top-level entries are .gitattributes, .gitignore, CODE_OF_CONDUCT.md, LICENSE, README.md, other_third-party_licenses/ and resource/. There is no model code, no inference script and no training script at the root. Everything of substance is on a branch or in a sibling repository.

The README lists the members: SALMONN 2, ELLSA, video-SALMONN 2 (in its own repository, bytedance/video-SALMONN-2), F-16, video-SALMONN-o1, a speech quality assessment line, video-SALMONN, and the original SALMONN. Several carry conference tags in the link text, including ICLR 2026 for ELLSA, ICML 2025 for F-16 and video-SALMONN-o1, ICASSP 2025 and ACL 2025 for the speech quality assessment work, ICML 2024 for video-SALMONN and ICLR 2024 for SALMONN. The audience is therefore research engineers and graduate students who already know which paper they are reproducing. If you arrived looking for one downloadable tool, the layout will be the first surprise.

How the family is organised across branches and sibling repositories

The mechanism is git branches plus separate repositories. The README's model list is a set of links, and most of them point at tree URLs on this same repository: salmonn2, ELLSA, video-salmonn-o1, speech_quality_assessment, videosalmonn and salmonn. Two entries leave the repository entirely: video-SALMONN 2 points to bytedance/video-SALMONN-2, and F-16 points to bytedance/F-16.

That means the data flow a reader should expect is: choose the paper, follow its link, land on a branch, and read that branch's own README for dependencies, checkpoints and commands. The top-level README does not carry a unified install path, and it does not describe a shared inference API across the family. Each release is self-contained by convention rather than by tooling. The News section is the closest thing to a changelog: entries run from the 2023-10-08 announcement of the SALMONN-13B checkpoint and inference code, through the 2024-04-07 release of training code, the 2024-05-28 release of annotations for the three-stage training of SALMONN (described as 600k SQA/AQA data and 50k audio-based storytelling data), the 2024-09-04 video-SALMONN release, the 2025-03-03 speech quality assessment scripts and checkpoints, the 2025-06-01 QualiSpeech dataset, the 2025-07-08 video-SALMONN 2 open-sourcing, and the 2026-04-20 ELLSA release. The last push to the repository was on 2026-08-24.

One design consequence is worth stating plainly: because the branches diverge, a fix or improvement on one branch does not reach the others. There is no shared core module the README points to. If you plan to run more than one member of the family, budget for more than one environment.

Getting to the code and running a first inference

The top-level README gives no install commands, so the honest answer is that installation instructions live on the branch or sibling repository for the model you want, and the README does not reproduce them. What the README does give is where the artifacts are. For the original SALMONN it points to the model checkpoint and inference code released on 2023-10-08, and the 7B version is published at tsinghua-ee/SALMONN-7B on Hugging Face, with a demo space at tsinghua-ee/SALMONN-7B-gradio.

So the first real step is to select the branch. If you want the ICLR 2024 model, the link is the salmonn branch; if you want the audio-visual model, it is videosalmonn; if you want the streaming full-duplex work, it is ELLSA. Cloning a single branch is the practical move, and it keeps the working tree small:

bash
git clone --branch salmonn --single-branch https://github.com/bytedance/SALMONN.git
cd SALMONN

After that, read the README on the branch you cloned. The top-level document does not tell you which Python version, which dependencies or which checkpoint layout that branch expects, and it would be wrong to guess. The one dataset the README does point to directly is QualiSpeech, for speech quality assessment, at huggingface.co/datasets/tsinghua-ee/QualiSpeech, which the README says can be used to develop your own audio LLM for speech quality assessment or to evaluate the low-level speech perception capabilities of existing audio LLMs.

Where SALMONN stops being the right tool

The clearest limitation is packaging. There is no pip package, no container image and no versioned release artifact mentioned anywhere in the README. Recent releases returned nothing. A team that needs to pin a dependency, reproduce a build in CI, or upgrade on a schedule has nothing to pin against except a commit hash on a branch.

The second limitation is scope. These are research models released alongside papers, and the repository is organised around that: the paper list is a BibTeX block, and the model list is a set of conference-tagged links. If your requirement is a speech recognition service with an SLA, a latency budget and a support contract, this is the wrong starting point regardless of model quality, because the README makes no claims about serving, throughput or deployment.

The third is discoverability. Nothing in the top-level README states which branch is current, which is superseded, or how the branches relate to each other beyond their paper titles. A reader who wants the newest audio-only model has to infer it from the News dates. That is a reasonable convention for a research group and an awkward one for anyone arriving cold.

How this differs from a general audio-language model toolkit

The natural comparison is with a toolkit that ships one model family behind a uniform interface, where you install once and switch checkpoints by name. SALMONN's approach is the opposite: each capability is a separate release with its own branch, its own README and, in two cases, its own repository. The difference shows up at the first upgrade. With a unified toolkit you change a model identifier; with SALMONN you change branch, and the surrounding code changes with it.

The trade-off is not obviously bad. A branch per paper keeps the code that produced a published result frozen and readable, which matters if you are reproducing a number or extending a specific architecture. The cost is that cross-model reuse is manual. If you want to compare video-SALMONN-o1 against video-SALMONN on the same inputs, the README offers no shared harness, so you would be writing the glue yourself. Choose SALMONN when the paper is the unit of work. Choose a unified toolkit when the pipeline is the unit of work.

Licence and the cost of keeping up

The repository is Apache-2.0, and there is an other_third-party_licenses/ directory at the root, which signals that components pulled in from elsewhere carry their own terms. Apache-2.0 is permissive and includes an explicit patent grant, but it also carries notice and attribution obligations, and the presence of third-party licence files means the effective terms for a given branch may be broader than the root LICENSE. Check the branch and that directory before redistributing anything. This is a description of what the repository states, not legal advice.

Upgrade cost is the other number to weigh. The News entries are dated and irregular: gaps between announcements run from weeks to most of a year. There is no deprecation policy, no migration guide and no compatibility statement between branches. The practical consequence is that adopting SALMONN is closer to vendoring a research codebase than to depending on a library. Plan to read the branch README each time you move, and expect to re-establish the environment rather than bump a version.

Editorial conclusion

Adopt SALMONN if you are working on speech or audio-visual research and want published checkpoints and training code tied to named papers, and if you can accept that each model lives on a separate branch with its own setup. Do not adopt it if you need a single maintained package with a stable API, a pip install, or a support channel; the main branch is an index, not a library, and the README documents no rollback, no versioning policy and no release artifacts. Before committing, verify three things yourself: which branch corresponds to the model you actually need, whether the checkpoint you want is on Hugging Face under the tsinghua-ee organisation, and whether the branch's own README supplies the dependency list, since the top-level README does not.

Frequently asked questions

What is bytedance/SALMONN?

It is the repository for the SALMONN model family, which the README describes as a series of advanced multi-modal large language models. The main branch is an index: the actual model code, checkpoints and training code sit on individual branches such as salmonn, videosalmonn and ELLSA, or in sibling repositories like bytedance/video-SALMONN-2.

Which branch of SALMONN should I use?

The README maps each release to a branch or repository and tags most of them with a conference: salmonn for the ICLR 2024 model, videosalmonn for ICML 2024, video-salmonn-o1 for ICML 2025, speech_quality_assessment for the ICASSP 2025 and ACL 2025 work, and ELLSA for ICLR 2026. Pick by the paper you are working from, since the README does not state which branch supersedes which.

Where are the SALMONN model checkpoints?

The README points to Hugging Face for the published weights: the 7B version is at tsinghua-ee/SALMONN-7B, and the QualiSpeech dataset is at tsinghua-ee/QualiSpeech. The 2023-10-08 news entry announces the SALMONN-13B checkpoint and inference code, and the 2025-03-03 entry announces finetuned checkpoints for speech quality assessment.

Does the bytedance/SALMONN repository have a pip package or container image?

The README does not mention a pip package, a container image or any versioned release artifact, and the repository's top-level entries are documentation, licence and resource files rather than a package. Installation instructions are expected on the individual branch you clone.

Official sources

  1. bytedance/SALMONN on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bytedance-salmonn.svg)](https://hysenlabs.com/projects/bytedance-salmonn)