Higgs Audio: the v2 repository that v3 tells you not to clone
Text-audio foundation model from Boson AI
At a glance
- What is it?
- The boson-ai/higgs-audio repo now opens with a banner telling you to skip it. Here is what still lives in the Python codebase, what moved to a hosted API and open weights, and who should pick which.
- Who is it for?
- Adopt the v3 hosted API or the bosonai/higgs-audio-v3-tts-4b weights if you want conversational TTS with zero-shot cloning and can accept the Boson Higgs Audio v3 Research and Non-Commercial License. Stay with the code in this repository only if you specifically need v2 or v2.5, whose documentation moved to README_V2.md.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 117 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the banner at the top of the repository actually changes
Open the repository today and the first thing you read is that Higgs Audio v3 is a standalone release and that you do not need this repo to use the latest model. That is an unusual thing for a project to say about itself, and it reframes the whole evaluation. This is no longer the place you go to run the current model. It is the home of the v2 and v2.5 line, the Python package that supported it, and a pointer to where v3 lives instead.
The split matters because the two halves have different licences and different installation stories. v3 ships as weights on Hugging Face under the Boson Higgs Audio v3 Research and Non-Commercial License, or as a rate-limited hosted API. The code in this repository is Apache-2.0, and the README also carries a third-party notice for boson_multimodal/audio_processing/, which is derived mainly from xcodec and has its own LICENSE file in that directory. If you are auditing a dependency for a commercial product, those three layers do not collapse into one answer.
The intended audience has narrowed. Someone who wants the newest speech model should follow the banner. Someone maintaining a pipeline built on the earlier generation, or reading the technical details and benchmarks that used to sit on the front page, is the actual reader of this repository now.
How the v2 codebase is laid out, and what the v3 split means for it
The top level is a normal Python package plus examples. boson_multimodal/ holds the library, examples/ holds runnable entry points, and README_V2.md holds the installation, examples, technical details and benchmarks that were moved off the front page. The example directory names describe the surfaces the project cared about: generation.py for synthesis, serve_engine/ and vllm/ for serving, voice_prompts/ and transcript/ and scene_prompts/ for conditioning inputs.
requirements.txt tells you more about the design than the README does. The pins are specific in places, transformers>=4.45.1,<4.47.0 and boto3==1.35.36, which means the project tracks a narrow window of the Hugging Face stack rather than following it. The presence of descript-audio-codec, vector_quantize_pytorch and dacite points at a codec-based audio pipeline with typed configuration, and s3fs plus boto3 suggest object storage is a first-class input path rather than an afterthought. torch and torchaudio are unpinned, so the CUDA build is your problem to match.
The v3 direction drops that package entirely. According to the README, v3 is served through SGLang-Omni or called over an OpenAI-compatible HTTP endpoint, so the architecture moves from an in-repo Python library to a serving stack you install separately. That is a real reduction in surface area for the maintainers and a real increase in the number of moving parts for anyone self-hosting.
Installing the v2 line from this repository
The README on the front page no longer carries v2 install steps; it says the full v2 and v2.5 documentation, including installation, has moved to README_V2.md. Treat that file as the source of truth and expect to follow it rather than the front page.
The repository is a standard setuptools project, so the package metadata is in pyproject.toml and setup.py, and dependencies are enumerated in requirements.txt. A typical checkout-and-install looks like this:
git clone https://github.com/boson-ai/higgs-audio.git
cd higgs-audio
pip install -r requirements.txt
pip install -e .After that, the runnable examples live under examples/. The README's own pointer is to README_V2.md for the working invocations, and examples/README.md for the example set. Because requirements.txt pins transformers below 4.47.0, run this in a virtual environment; installing it into a shared environment that already has a newer transformers will fight the pin.
If what you actually want is v3, do not install this package at all. The README's Option 1 is a hosted call:
export BOSON_API_KEY=bai-xxxx
curl https://api.boson.ai/v1/audio/speech \
-H "Authorization: Bearer $BOSON_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "higgs-audio-v3-tts", "input": "Hello, this is a test."}' \
--output out.mp3The API key comes from boson.ai/workspace, and the preview is described as free and rate-limited. The response body is written to out.mp3. Option 2 is self-hosting the weights, which the README recommends serving with SGLang-Omni:
export HF_TOKEN=hf_xxxxxxxxxxxxxxxx
hf download bosonai/higgs-audio-v3-tts-4b
sgl-omni serve --model-path bosonai/higgs-audio-v3-tts-4b --port 8000That downloads the weight set and starts a server on port 8000. Serving, voice-cloning and streaming recipes are in the model card and the SGLang-Omni cookbook, not in this repository.
The licence boundary is the first thing to check, not the last
The repository is Apache-2.0. The v3 weights are not. The README puts a note directly under the self-hosting instructions stating that Higgs Audio v3 is released under the Boson Higgs Audio v3 Research and Non-Commercial License, and that production, hosted or revenue-generating use requires a separate commercial license.
That is a clean, explicit boundary and it is worth taking at face value. It means the free hosted preview and the open weights are fine for research and evaluation, and a product built on either needs a conversation with Boson AI first. The Apache-2.0 licence on the code in this repository does not transfer to the v3 weights, and the xcodec-derived audio_processing directory carries its own LICENSE file, so a licence scan of a v2-based deployment has to read three files, not one.
I am not giving legal advice here, and the README does not spell out what counts as revenue-generating. If your use is anywhere near commercial, that ambiguity is the thing to resolve before you write code, because it determines whether the hosted API or the weights are even an option.
Where this is the wrong choice
The clearest failure mode is starting here for v3. The README says it in bold: do not clone this repository to use the latest model. A team that skims the description, sees a text-audio foundation model and clones, ends up installing a v2-era Python package with a pinned transformers range and a codec dependency chain, none of which the current model uses.
The second limitation is maintenance signal. The most recent push to this repository was on 2026-06-05. The front page has been repurposed as a redirect to a different release, and the substantive documentation for the models that remain here has been moved into README_V2.md. That is a repository in wind-down mode for its original purpose, and the README does not document a rollback path, a deprecation schedule, or how long the v2 line will keep receiving fixes.
The third is dependency fragility. Pinning transformers to a range below 4.47.0 while torch and torchaudio float means the environment is sensitive to whatever CUDA and PyTorch versions you bring. If your platform is already standardised on a newer transformers, this package will not slot in without isolation.
Finally, there is a scope question. The v3 banner describes conversational TTS across 100+ languages, zero-shot voice cloning and inline control of emotion, style and prosody. If your need is speech recognition, note that this material describes text-to-speech only; there is no STT pipeline documented here.
Alternatives and how the approach differs
The most direct alternative is the v3 path itself, and it is not really a competitor so much as the successor: instead of a Python library you import, you get an OpenAI-compatible HTTP endpoint or a weight set served by SGLang-Omni. The difference in approach is that inference moves out of your process. You gain a stable API surface and lose the ability to read and patch the model code in this repository.
The searches around this project repeatedly pair it with ElevenLabs, and the distinction is architectural rather than cosmetic. ElevenLabs is a hosted voice service with no open weight release; Higgs Audio v3 offers both a hosted preview and downloadable weights, which is the whole reason the licence note exists. If you need to run inference on your own hardware, that difference decides the choice. If you only need a voice in a product and do not care where it runs, the hosted option on either side is the simpler comparison, and the deciding factors become language coverage, cloning quality and price, none of which this material quantifies.
Within the open-weight space, the practical alternative is whatever TTS model your serving stack already supports. Because v3 is served through SGLang-Omni rather than a bespoke runtime, the switching cost is mostly the weight download and the prompt format, not a new deployment model.
What to verify before you commit
Check the licence that covers the exact artefact you plan to deploy. The repository is Apache-2.0, the v3 weights are non-commercial without a separate agreement, and boson_multimodal/audio_processing/LICENSE covers the xcodec-derived code. Three files, three answers.
Check the environment. If you are on the v2 path, confirm your transformers version can satisfy >=4.45.1,<4.47.0 and that your torch build matches your CUDA runtime, because neither torch nor torchaudio is pinned. If you are on the v3 path, confirm SGLang-Omni runs on your hardware before you plan around it.
Check the documentation you are actually reading. README_V2.md, not README.md, holds the v2 installation and benchmarks. The model card and the SGLang-Omni cookbook, not this repository, hold the v3 serving and cloning recipes. Support and contribution expectations are in SUPPORT_GUIDELINES.md, and the README does not describe a release cadence for either line.
Editorial conclusion
Adopt the v3 hosted API or the bosonai/higgs-audio-v3-tts-4b weights if you want conversational TTS with zero-shot cloning and can accept the Boson Higgs Audio v3 Research and Non-Commercial License. Stay with the code in this repository only if you specifically need v2 or v2.5, whose documentation moved to README_V2.md. Before committing, verify three things: which licence covers the model you would deploy, whether SGLang-Omni serves the weight set on your hardware, and whether your use is commercial, because the README states that production or revenue-generating use requires a separate commercial licence.
Frequently asked questions
What is Higgs Audio?
It is Boson AI's text-audio foundation model family, written in Python and licensed Apache-2.0 at the repository level. The current release is Higgs Audio v3, described as conversational TTS across 100+ languages with zero-shot voice cloning and inline control of emotion, style and prosody.
Is Higgs Audio v2 open source?
The code in this repository, including the v2 line, is Apache-2.0, and the README points to v2 and v2.5 models still available on Hugging Face. The v3 weights are a separate case: the README states they fall under the Boson Higgs Audio v3 Research and Non-Commercial License.
How do I install Higgs Audio v2?
The front-page README says the full v2 and v2.5 installation documentation has moved to README_V2.md, so that file is where the steps live. The repository itself is a setuptools project with dependencies in requirements.txt and metadata in pyproject.toml.
How do I use Higgs Audio?
For v3, the README gives two routes: call the OpenAI-compatible endpoint at api.boson.ai/v1/audio/speech with a BOSON_API_KEY, or download bosonai/higgs-audio-v3-tts-4b and serve it with SGLang-Omni on port 8000. For v2, the runnable examples are under examples/ and the working invocations are in README_V2.md.
How do I use Higgs Audio v2?
The README directs you to README_V2.md for the full v2 and v2.5 documentation, including installation and examples. The repository also ships an examples/ directory with generation.py and dedicated folders for serving, voice prompts and transcripts.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/boson-ai-higgs-audio)