Model or dataset
suno-ai/bark avatar
suno-ai/bark

Bark, Suno's text-to-audio model, and the places it stops being text-to-speech

🔊 Text-Prompted Generative Audio Model

39,274 stars4,667 forksJupyter NotebookMIT

At a glance

What is it?
Bark is a transformer text-to-audio model released by Suno under the MIT licence, with 100+ speaker presets and no voice cloning. It is happy to render your text as music, it tops out near 13 seconds by default, and the repository has not been pushed to since 2024-08-19.
Who is it for?
Use Bark if you need a voice that is not a corporate preset, you are writing a prompt in a language you want heard with its own accent, and you can accept output you have to listen to before shipping. Walk away if you need a named voice cloned, if a fixed duration matters, or if a command line entry point would carry your pipeline, because pyproject.toml declares no [project.scripts] and the distribution is still version 0.0.1a.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Probably not. The repository last received commits 25 months ago, on August 19, 2024.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

generate_audio is built for about 13 seconds of speech

The default call is documented as working well with around 13 seconds of spoken text. That number is the first thing to check if your output stops mid-sentence, because nothing in the import surface takes a duration and nothing raises an error when you exceed it. The audio simply ends where the model ends it.

Anything longer lives in notebooks/long_form_generation.ipynb, and that path is not a small detour. The repository's primary language is recorded as Jupyter Notebook, and the top level holds .gitignore, LICENSE, README.md, bark/, model-card.md, notebooks/, pyproject.toml and setup.py. The long-form work, the voice consistency enhancements and the worked examples are all in that notebooks directory, not in the packaged API. The 2023.05.01 update entry says as much when it announces a new notebooks section.

So the shape of the tool is: a short, well served call for lines and sentences, and a notebook you adapt yourself for anything that runs past a paragraph. If your use case is an audiobook or a podcast, that boundary is the whole decision.

history_prompt takes a preset name, and there is no voice cloning

Bark ships 100+ speaker presets and reaches them by name, not by sample:

python
text_prompt = """
    I have a silky smooth voice, and today I will tell you about
    the exercise regimen of the common sloth.
"""
audio_array = generate_audio(text_prompt, history_prompt="v2/en_speaker_1")

The string is a path inside the package rather than a label. pyproject.toml declares package data as assets/prompts/*.npz and assets/prompts/v2/*.npz for the bark package, which is where those presets live, and the v2/ prefix in the name is a directory in that list. The presets are NumPy archives shipped with the wheel, not a lookup against a hosted service.

What a preset gets you is an attempt at tone, pitch, emotion and prosody. The documentation is explicit that Bark does not currently support custom voice cloning, and the same note says the model tries to preserve music, ambient noise and other incidental audio. If you need one specific human voice, this is the wrong tool, and no amount of prompt engineering changes that, because the input that would select a voice is a shipped file.

Sometimes the model answers with music instead of speech

Bark can generate all types of audio and, in principle, does not see a difference between speech and music. That is the model's stated position and it is also its most reliable failure mode: sometimes it chooses to generate your text as music, and the documented workaround is to help it out by adding music notes around your lyrics. Wrapping the words in musical notation pushes the whole prompt toward the singing register rather than the speaking one.

The disclaimer explains why this matters more than it sounds. Bark was developed for research purposes, it is not a conventional text-to-speech model but a fully generative text-to-audio model, it can deviate in unexpected ways from the prompts you provide, and Suno does not take responsibility for any output generated. Put those together and the content type of a given generation is not guaranteed. A pipeline that assumes every call returns speakable narration will sometimes get a song, and nothing in the return value tells you which one you got, since generate_audio hands back a bare audio array.

For anything with a review step in front of a human, that is manageable. For batch generation into a player or a phone tree, you are listening to every call.

The language comes from your prompt text, so a German prompt can still speak English

There is no language parameter. Bark supports various languages out of the box and determines the language from the input text, then attempts the native accent for whatever it finds:

python
text_prompt = """
    Der Dreißigjährige Krieg (1618-1648) war ein verheerender Konflikt, der Europa stark geprägt hat.
    This is a beginning of the history. If you want to hear more, please continue.
"""
audio_array = generate_audio(text_prompt)

That example is the failure case and the feature at once, because the prompt is German history written mostly in German, and the documented outcome is English audio with a German accent. Mixing the two gets you neither pure German nor pure English. The same auto-detection cuts the other way, which is how a plain Korean paragraph needs no configuration at all:

python
text_prompt = """
    추석은 내가 가장 좋아하는 명절이다. 나는 며칠 동안 휴식을 취하고 친구 및 가족과 시간을 보낼 수 있습니다.
"""
audio_array = generate_audio(text_prompt)

Quality is uneven by design. English is described as best for the time being, with other languages expected to improve with scaling. Since detection is implicit, a misdetection gives you a plausible voice in the wrong language rather than an exception, so write prompts in a single language if the accent matters.

The wheel gives you an importable package and no command

pyproject.toml names the distribution suno-bark, declares requires-python of 3.8 or newer, and lists the runtime dependencies in full: boto3, encodec, funcy, huggingface-hub 0.14.1 or newer, numpy, scipy, tokenizers, torch, tqdm and transformers. setup.py is a one-line shim that calls setup(). The install section of the README is a heading of its own and the quick index links to it, and there is no [project.scripts] table anywhere in the build metadata, so installing gets you a library to import rather than a command to run.

The first run is three imported names and one call that fetches the weights:

python
from bark import SAMPLE_RATE, generate_audio, preload_models
from scipy.io.wavfile import write as write_wav
from IPython.display import Audio

# download and load all models
preload_models()

Two things are worth noticing in that snippet. SAMPLE_RATE, generate_audio and preload_models come from bark, while write_wav and Audio are scipy and IPython helpers the example borrows to save a file and play it back in a notebook. IPython is not in the runtime dependency list at all; it arrives with jupyter, which sits in the dev extra alongside pytest, flake8, mypy, pylint, hypothesis, bandit, isort and nbconvert. And preload_models is the step that downloads checkpoints, which is why a first run needs network access before it needs a GPU.

pyproject.toml still carries an Apache 2.0 comment above an MIT licence

The licence position is permissive, with one piece of metadata that will make a lawyer pause. The 2023.05.01 update entry states that Bark is now licensed under the MIT License, meaning it is now available for commercial use, and the model description says the pretrained checkpoints are ready for inference and available for commercial use. The repository licence is MIT.

Now look at the build metadata. Above the licence line in pyproject.toml there is a comment reading Apache 2.0, and the line below it points at the LICENSE file rather than naming a licence. So the file that decides your obligations is correct and the file that tooling reads for automated scanning is pointed at the right place, but a human skimming the build file sees a contradiction. That is worth raising with whoever maintains your dependency manifest rather than resolving yourself, and this is not a legal opinion, just a note that the metadata and the licence file disagree on the page.

The disclaimer sits alongside that permissiveness and is the part to keep in view. Use at your own risk, and please act responsibly, is the instruction the README gives, on a model that can deviate in unexpected ways.

Version 0.0.1a, no releases, and a last push on 2024-08-19

The version field in pyproject.toml is still 0.0.1a, the repository has no GitHub releases, and the last push was on 2024-08-19. The repository is not archived, so this is a project that stopped rather than one that was closed. The newest dated entry in the README's own update list is 2023.05.01, whose predecessor is the 2023.04.20 release announcement, so the documentation describes a 2023 state of the world and has not been revised since.

The dependency list is where that age shows up most sharply. Only huggingface-hub carries a lower bound, at 0.14.1 or newer. torch, transformers, tokenizers, encodec, numpy, scipy, tqdm and funcy are all unpinned, and those are exactly the packages that break between major versions. A fresh install in 2026 resolves whatever is current, against a codebase whose last commit is more than two years old. The dev extras are tighter, with hypothesis held above 6.14 and below 7 and isort above 5.0.0 and below 6, which tells you the project pinned what it tested and left the runtime loose.

Practically: vendor a lockfile, and treat the notebooks as the only current documentation of the advanced behaviour.

Replicate, a Hugging Face Space and a Colab notebook run it for you

The demos section points at three hosted routes, and for a first evaluation they are the cheapest option by a wide margin: a Space at huggingface.co/spaces/suno/bark, a Replicate page at replicate.com/suno-ai/bark, and a Google Colab notebook. Each of those is running the same weights somewhere you do not have to install torch, and the local cost of Bark is bounded by hardware rather than by code. The 2023.05.01 entry says you can use Bark with GPUs that have low VRAM, under 4GB, and that a smaller version of the model was added which offers additional speed-up with the trade-off of slightly lower quality.

That is the real comparison, local checkpoint against hosted checkpoint, and the difference is where your prompts and the resulting audio live. Locally, the text and the audio never leave the machine, you accept the VRAM and the roughly 13 second ceiling. Hosted, you accept that prompts and generated speech are transmitted to someone else's hardware and returned over a network, and you get whatever queueing that service applies. The README also notes an option for a smaller Bark with lower quality, which is the local answer to a machine that cannot hold the full checkpoint.

One naming trap to close with. The notice at the top of the README says that if you are looking for Suno's text-to-music models you should go to suno.ai, and this repository is the speech side only. Searches that send you here looking for a music generator are looking at the wrong model.

Editorial conclusion

Use Bark if you need a voice that is not a corporate preset, you are writing a prompt in a language you want heard with its own accent, and you can accept output you have to listen to before shipping. Walk away if you need a named voice cloned, if a fixed duration matters, or if a command line entry point would carry your pipeline, because pyproject.toml declares no [project.scripts] and the distribution is still version 0.0.1a. Before you build on it, read the LICENSE file rather than the pyproject.toml comment, which still says Apache 2.0, and pin torch and transformers yourself, since the runtime dependency list leaves both unconstrained.

Frequently asked questions

How do I install Bark AI?

The distribution is named suno-bark in pyproject.toml while the import is bark, and requires-python is 3.8 or newer. The runtime dependencies are boto3, encodec, funcy, huggingface-hub, numpy, scipy, tokenizers, torch, tqdm and transformers, so the install pulls a deep learning stack rather than a light package. There is no [project.scripts] entry, so nothing lands on your PATH.

Does Bark support voice cloning?

No. Bark tries to match the tone, pitch, emotion and prosody of one of its 100+ speaker presets and does not currently support custom voice cloning. The presets ship inside the package as NumPy archives under assets/prompts, including a v2 subdirectory, and you select one by name such as v2/en_speaker_1.

How much audio can Bark generate in one call?

By default generate_audio works well with around 13 seconds of spoken text. Longer output is worked through in notebooks/long_form_generation.ipynb, since the documented call takes a text prompt and an optional history_prompt but no length argument.

Which languages does Bark handle best?

Bark determines the language from the input text automatically and exposes no language parameter. English quality is described as best for the time being, with the expectation that other languages improve with scaling, and code-switched text produces the native accent of each language rather than a single blended voice.

Can I use Bark for commercial work?

Yes. The 2023.05.01 update states that Bark is licensed under the MIT License and is available for commercial use, and the model description says the pretrained checkpoints are ready for inference and available for commercial use. The same README still carries a research disclaimer saying Suno does not take responsibility for any output generated.

Official sources

  1. Official README
  2. Project repository