Model or dataset
stepfun-ai/Step-Audio2 avatar
stepfun-ai/Step-Audio2

Step-Audio 2: an audio model that reasons about how you sound, then answers out loud

Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.

1,519 stars114 forksPythonApache-2.0

At a glance

What is it?
StepFun open-sources three sizes of a talk-and-think audio model under Apache 2.0, plus a forked vLLM build with its own audio tokenizer, tool call parser and speech output parser. The interesting engineering is in that serving layer rather than in the weights.
Who is it for?
Step-Audio 2 is worth a look when a request is a conversation rather than a file to transcribe. The case for it is the audio-native design: paralinguistic reasoning, tool calling with a speech-shaped parser, and a model that answers in audio rather than handing you text to read aloud.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Three open weight sizes that answer with speech, not just text

Step-Audio 2 is described as an end-to-end multimodal large language model built for audio understanding and speech conversation. The distinguishing claim in the README is not transcription accuracy. It is that the model reads paralinguistic information, reasoning about cues such as the speaker's apparent age and emotional state, and that it responds in kind rather than transcribing first and reasoning second.

That design shows up in what ships. Three weights are released on Hugging Face and mirrored on ModelScope: Step-Audio 2 mini, Step-Audio 2 mini Base, and Step-Audio 2 mini Think. The plain `mini` is the instruction-following chat model, `mini Base` is the unaligned starting point, and `mini Think` is the reasoning variant that arrived with `examples-think.py` on 2025-09-15. All three are Apache 2.0, with a LICENSE file at the repository root matching what the README claims, which is one of the cleaner licensing setups in this kind of release.

The repository tree also carries `cosyvoice2/`, `flashcosyvoice/`, `token2wav.py` and a `--audio-parser` flag whose value ends in `tts_ta4`. That is the speech side of the system, and it means the audio stack in this project is not limited to listening.

The plain Transformers path needs a conda env and a pinned transformers

Setup is a fixed dependency set on Python 3.10 or newer, with PyTorch built against CUDA 12.1 and the CUDA Toolkit listed as a separate requirement. The README does not offer a CPU fallback or a pip package to install instead.

bash
conda create -n stepaudio2 python=3.10
conda activate stepaudio2
pip install transformers==4.49.0 torchaudio librosa onnxruntime s3tokenizer diffusers hyperpyyaml

git clone https://github.com/stepfun-ai/Step-Audio2.git
cd Step-Audio2
git lfs install
git clone https://huggingface.co/stepfun-ai/Step-Audio-2-mini

Two details in that sequence are load bearing. The `transformers==4.49.0` pin is exact, which tells you the audio model class depends on internals that move between releases. And the weights come from a separate `git clone` of the Hugging Face repository rather than a download step, run after `git lfs install` so that the large tensors actually arrive.

The audio stack is more than torch: `s3tokenizer` handles speech tokenization, `onnxruntime` pulls in an auxiliary runtime, `librosa` handles feature extraction, and `diffusers` and `hyperpyyaml` come along for the generation side. With the environment ready, the README's inference scripts are one line each:

bash
python examples.py
# python examples-base.py
# python examples-vllm.py
# python examples-think.py

Four scripts, four entry points: standard chat, the base weights, the vLLM path, and the thinking variant. A local Gradio interface is two more commands, `pip install gradio` then `python web_demo.py`, with `web_demo_vllm.py` in the tree for the served version.

Why the vLLM path needs a fork with its own tokenizer and parser modes

The README recommends the vLLM backend for streaming and multi-GPU deployment, and the reason it needs a fork is visible in the Dockerfile. It starts from `vllm/vllm-openai:v0.10.1`, uninstalls that vLLM, and compiles a different one from a branch of StepFun's own vLLM repository:

dockerfile
FROM vllm/vllm-openai:v0.10.1
RUN pip uninstall vllm -y
RUN pip install librosa
RUN git clone -b step-audio2-mini --depth 1 https://github.com/stepfun-ai/vllm.git /tmp/vllm

That is the whole customization. No other patches. So the branch name is a promise that the fork's delta is scoped to this model, and the shallow clone keeps the build small. The README warns that building the image yourself is very slow and needs 32GiB of memory, which is why a published image exists.

The serve command is where the audio support becomes visible. Three flags do not exist upstream: `--tool-call-parser step_audio_2`, `--tokenizer-mode step_audio_2`, and `--audio-parser step_audio_2_tts_ta4`, plus `--chat_template_content_format string` to tell the template that content arrives as a string rather than as a typed message array.

bash
docker run --rm -ti --gpus all \
    -v Step-Audio-2-mini:/Step-Audio-2-mini \
    -p 8000:8000 \
    stepfun2025/vllm:step-audio-2-v20250909 \
    -- vllm serve /Step-Audio-2-mini \
    --served-model-name step-audio-2-mini \
    --max-model-len 16384 \
    --max-num-seqs 32 \
    --tensor-parallel-size 1 \
    --enable-auto-tool-choice \
    --tool-call-parser step_audio_2 \
    --tokenizer-mode step_audio_2 \
    --audio-parser step_audio_2_tts_ta4 \
    --trust-remote-code

`--enable-auto-tool-choice` pairs with the tool call parser, and that combination is what the README's tool calling and multimodal RAG claim depends on. `--tensor-parallel-size 1` is the single GPU default, raised when a model does not fit. The audio output path is served through the OpenAI compatible endpoint on port 8000, with `examples-vllm.py` and the streaming variant `examples-vllm-stream.py` showing the client side.

Tool calling and RAG with the voice of retrieved speech

The two capabilities the README argues hardest for are tool calling and multimodal retrieval, and there is an unusual detail attached to the retrieval half. The model can access real-world knowledge from text and from acoustic material, and it can switch timbres based on retrieved speech. In practice that means a retrieved clip can supply not just content but a voice, which is a different thing from a text RAG pipeline that returns a string.

Tool calling matters here for a specific reason: an audio model that answers with a retrieved sentence has to decide what to say, and the tool call parser gives it a structured channel for reaching outside itself. The README ties this to fewer hallucinations for diverse scenarios. StepFun also published the evaluation sets behind these claims as datasets, `StepEval-Audio-Toolcall` and `StepEval-Audio-Paralinguistic`, both dated 2025-07-23, which is more useful to you than a headline number because you can inspect what the tasks actually ask for.

The paralinguistic benchmark is the one to look at first if you are evaluating this at all. Age and emotion inference from voice has obvious misuse potential, and knowing the task format tells you whether the model is inferring affect for a support triage case or something else entirely. The technical report on arXiv at 2507.16632 carries the full detail, and the README updates it as the work progresses rather than freezing one version at launch.

What this is not: a transcription API or a general TTS engine

Two comparisons are worth making explicitly, because the project name invites the wrong one.

Against Whisper-style ASR, Step-Audio 2 is not a competitor. Whisper takes audio and gives you text, and it is small, fast and boring in the way you want. Step-Audio 2 takes audio and gives you a conversational turn that accounts for how the audio sounded. If your requirement is a transcript with timestamps, this is the wrong tool and you will pay GPU time for reasoning you do not need.

Against text-first multimodal models with audio input, the difference is architectural rather than a benchmark delta. Those models encode audio into tokens and reason in the same channel as text, so paralinguistic detail tends to get compressed along the way. Step-Audio 2 keeps speech in the loop from tokenizer through parser to output, which is why the fork exists at all.

The honest limitation is operational. There is no pip install, no CPU path, an exact `transformers` pin, and a serving stack that only works on the fork. The repository also has no GitHub releases, so there is no version number to pin beyond the dated Docker image tag and the Hugging Face revision you clone.

Reading the repository as a research release rather than a product

What this repository is, in practice, is the code that accompanied a paper plus the weights to run it. The release history in the README reads like a lab notebook: a technical report on 2025-07-23, demonstration videos the next day, the mini and mini Base weights with `examples.py` on 2025-08-29, the vLLM backend on 2025-09-03, and the thinking variant with `examples-think.py` on 2025-09-15. Each entry pairs an artifact with the script that exercises it.

Two things follow from that shape. First, the code is written to be read alongside the report, so expect experiment scripts rather than a maintained library API. There is no package entry point, no versioning scheme and no changelog. Second, the README stops right after the local web demonstration section, and what a production evaluation would need next, batch inference, latency numbers, a quantization story, a licensing note on the weights beyond Apache 2.0, is not written down here.

On cadence, the last push to the repository was on 2026-03-16 and the most recent weights and examples date from September 2025. The Step-Audio 2 documentation page carries the demonstration material, and the fork branch is the other moving part to watch, since a vLLM upgrade upstream would eventually pull at it. Anyone planning to build on this should treat the pinned `vllm/vllm-openai:v0.10.1` base and the `step-audio2-mini` branch as the load-bearing version facts, because those are what everything else is written against.

Editorial conclusion

Step-Audio 2 is worth a look when a request is a conversation rather than a file to transcribe. The case for it is the audio-native design: paralinguistic reasoning, tool calling with a speech-shaped parser, and a model that answers in audio rather than handing you text to read aloud. The cost of that design is the serving layer, because the documented path runs on a fork of vLLM pinned to branch step-audio2-mini, built from source with MAX_JOBS=2, and the stock image will not understand the tokenizer or parser flags without it. Start with `python examples.py` on the plain Transformers path to find out whether the model quality is what you need, then budget the vLLM build separately. Check the hardware floor first: the dependency list pins PyTorch against cu121 and asks for the CUDA Toolkit, so there is no CPU story in this repository, and no commit has landed on the repository since 2026-03-16.

Frequently asked questions

What is Step-Audio 2 and what does it do that ordinary speech recognition does not?

Step-Audio 2 is an end-to-end multimodal model for audio understanding and speech conversation, released in mini, mini Base and mini Think sizes under Apache 2.0. Unlike a transcription model it reasons over paralinguistic cues such as apparent emotion and age, and it answers in speech rather than returning text for you to read aloud.

How do you run Step-Audio 2 on your own machine?

The README documents a conda environment on Python 3.10 or newer with `transformers==4.49.0`, then `python examples.py` for the plain Transformers path or `python web_demo.py` after `pip install gradio` for a local interface. Both paths need PyTorch built against CUDA 12.1, since no CPU route is described.

Why does Step-Audio 2 require a custom vLLM build instead of stock vLLM?

The serving flags are the reason. Upstream vLLM has no `--tool-call-parser step_audio_2`, `--tokenizer-mode step_audio_2` or `--audio-parser step_audio_2_tts_ta4` modes, so the Dockerfile starts from `vllm/vllm-openai:v0.10.1`, removes that package and compiles the `step-audio2-mini` branch of StepFun's own vLLM fork instead.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. stepfun-ai/Step-Audio2 on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/stepfun-ai-step-audio2.svg)](https://hysenlabs.com/projects/stepfun-ai-step-audio2)