Qwen3-ASR: an Apache-2.0 speech recognition family with a forced aligner
Qwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilingual speech/music/song recognition, language detection and timestamp prediction.
At a glance
- What is it?
- Qwen3-ASR ships two all-in-one recognition models (0.6B and 1.7B) covering 52 languages and dialects, plus a separate non-autoregressive aligner for timestamps. Here is what the repository documents, and where it stops.
- Who is it for?
- Adopt Qwen3-ASR if you need multilingual recognition with a single model handling both offline and streaming input, and you accept pinning transformers==4.57.6 and, for batch serving, vllm==0.14.0. Skip it if your pipeline depends on word-level timestamps from the ASR model itself: those come from the separate Qwen3-ForcedAligner-0.6B, which covers 11 languages rather than 52.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 96 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Qwen3-ASR covers, and the gap it targets
Most open speech recognition stacks force a choice. You either run a small model that is fast but brittle on accents and background music, or a large one that costs throughput. Qwen3-ASR approaches this by releasing two sizes of the same architecture: Qwen3-ASR-0.6B and Qwen3-ASR-1.7B. Both are described as all-in-one, meaning language identification and transcription come from one model rather than a detection step feeding a recognizer.
The README states the family supports 30 languages and 22 Chinese dialects, and the released-models table names them: Chinese, English, Cantonese, Arabic, German, French, Spanish, Portuguese, Indonesian, Italian, Korean, Russian, Thai, Vietnamese, Japanese, Turkish, Hindi, Malay, Dutch, Swedish, Danish, Finnish, Polish, Czech, Filipino, Persian, Greek, Hungarian, Macedonian and Romanian, alongside dialects such as Sichuan, Wu and Minnan. The audio types column lists speech, singing voice and songs with background music, which is a wider input claim than most ASR releases make.
The intended user is someone building a transcription service or a self-hosted pipeline who cannot send audio to a commercial API, or who wants to avoid per-minute pricing. The package is pure Python and the licence is Apache-2.0, so the weights and the inference toolkit carry the same permissive terms.
Two models, one aligner: how the family splits the work
The architecture is built on Qwen3-Omni, according to the introduction, which is the reason a single model can emit both a language tag and a transcript. The 1.7B checkpoint is the accuracy tier; the README claims it is competitive with the strongest proprietary commercial APIs. The 0.6B checkpoint is the throughput tier, and the README states it reaches 2000 times throughput at a concurrency of 128. That figure is a claim from the project, not something this article measured, and the README does not say which hardware or batch configuration produces it.
Timestamps are deliberately not part of the ASR models. Qwen3-ForcedAligner-0.6B is a separate non-autoregressive model that aligns text to speech and predicts timestamps for arbitrary units within up to 5 minutes of audio, in 11 languages: Chinese, English, Cantonese, French, German, Italian, Japanese, Korean, Portuguese, Russian and Spanish. The README says its timestamp accuracy surpasses end-to-end forced-alignment models. The practical consequence is that a subtitle pipeline needs two checkpoints loaded, not one, and the language coverage of the aligner is a subset of what the recognizers handle.
Both ASR models support offline and streaming inference from the same weights, and the README states they support transcribing long audio. Streaming is not a separate model or a distilled variant, which simplifies deployment but means the streaming path shares the same memory footprint as the offline path.
Installing qwen-asr and running a first transcription
The project is distributed as a Python package named qwen-asr, defined in pyproject.toml at version 0.0.6, requiring Python 3.9 or newer. Its dependency list is tightly pinned: transformers==4.57.6, accelerate==1.12.0, nagisa==0.2.11, soynlp==0.0.493, plus qwen-omni-utils, librosa, soundfile, sox, gradio, flask and pytz. vLLM is an optional extra pinned at vllm==0.14.0. Those exact pins matter if you already have a different transformers version in the same environment.
If your runtime cannot download weights during execution, the README gives manual download commands. ModelScope is recommended for users in Mainland China:
pip install -U modelscope
modelscope download --model Qwen/Qwen3-ASR-1.7B --local_dir ./Qwen3-ASR-1.7B
modelscope download --model Qwen/Qwen3-ASR-0.6B --local_dir ./Qwen3-ASR-0.6B
modelscope download --model Qwen/Qwen3-ForcedAligner-0.6B --local_dir ./Qwen3-ForcedAligner-0.6BThe Hugging Face route uses the huggingface_hub CLI instead. The README truncates the last command, so only the first two are reproduced here exactly as shown:
pip install -U "huggingface_hub[cli]"
huggingface-cli download Qwen/Qwen3-ASR-1.7B --local-dir ./Qwen3-ASR-1.7B
huggingface-cli download Qwen/Qwen3-ASR-0.6B --local-dir ./Qwen3-ASR-0.6BFor a first run, the repository ships four example scripts under examples/: example_qwen3_asr_transformers.py, example_qwen3_asr_vllm.py, example_qwen3_asr_vllm_streaming.py and example_qwen3_forced_aligner.py. The README points to the Hugging Face model cards for native Transformers usage, noting that torch.compile support was added on 2026-06-26.
The package also installs three console entry points, which is the fastest way to see output without writing code:
qwen-asr-demo
qwen-asr-demo-streaming
qwen-asr-serveThe first two launch Gradio demos (local Web UI), and qwen-asr-serve starts the serving process. The README does not document the flags these commands accept, so check the CLI modules in qwen_asr/cli/ before wiring them into a service.
Deployment paths: vLLM, Docker and the DashScope API
There are three ways to run inference, and they are not equivalent. The plain Transformers path is the simplest and works with the pinned transformers==4.57.6. The vLLM path is the one to choose for batch or concurrent serving; the README describes the inference framework as supporting vLLM-based batch inference, asynchronous serving and streaming inference. The third option is DashScope, Alibaba Cloud's hosted API, which the README links to separately. That last path is not self-hosting: audio leaves your infrastructure, and the Apache-2.0 licence on the weights does not govern the API service.
The repository contains a docker/ directory and a Docker section in the README, so containerized deployment is a supported path rather than an afterthought. The README does not state a published image name or tag, so treat the docker/ directory as the source of truth for how the image is built.
One deployment detail worth flagging: the required dependency list includes gradio and flask unconditionally, even for headless serving. If you install qwen-asr into a lean production image, you are pulling a web UI framework and a web server whether or not you use the demos. That is a packaging choice, not a bug, but it affects image size and your dependency audit.
Where Qwen3-ASR is the wrong tool
The aligner's language coverage is the clearest boundary. If you transcribe Thai, Vietnamese or Hindi and you also need per-word timestamps, Qwen3-ForcedAligner-0.6B does not list those languages. You would need a different alignment approach for that portion of your catalog, which reintroduces the second toolchain the family was meant to remove.
The 5-minute limit on alignment is a second hard edge. The README states the aligner handles timestamp prediction for arbitrary units within up to 5 minutes of speech. Long-form audio must be segmented before alignment, and the README does not document how the toolkit chunks it.
Version pinning is the third. transformers==4.57.6 and accelerate==1.12.0 are exact pins, not ranges. If your application also depends on a newer or older Transformers release, qwen-asr and that application cannot share one virtual environment without conflict. The README does not document a supported upgrade path or a compatibility matrix beyond the pins themselves.
Finally, on the streaming claim: the README says streaming and offline inference are unified in a single model, but it does not state a streaming latency figure or a chunk size. If your product needs sub-second partial transcripts, the documentation does not give you a number to plan against.
How it differs from Whisper-family and Paraformer-family setups
The comparison people search for is against Whisper Large V3, and the difference is structural rather than a score. Whisper models are trained for transcription and translation, and word-level timestamps in that ecosystem typically come from a separate alignment step, often a cross-attention heuristic or a dedicated aligner. Qwen3-ASR splits the same problem explicitly: the recognizer does not produce timestamps, and Qwen3-ForcedAligner-0.6B does, as a non-autoregressive model. The two-model design is more honest about the task boundary, but it also means two downloads and two loads.
Against faster-whisper, the trade is runtime philosophy. faster-whisper is a reimplementation of Whisper on a faster inference engine; Qwen3-ASR instead ships its own toolkit with a vLLM backend as an optional extra. If you already run vLLM for other models, Qwen3-ASR fits that infrastructure. If you do not, you are adding a serving stack.
Against FunASR and Paraformer, which come from the same broad ecosystem of Chinese speech tooling, the distinction is the foundation model. Qwen3-ASR is built on Qwen3-Omni and inherits its audio understanding, and the README frames the 1.7B model's strength as coming from that base. The released-models table also lists songs with background music as an audio type, which is a narrower set of systems than pure speech recognizers. The README makes no direct benchmark comparison against any of these projects, so any accuracy ranking would be speculation.
Maintenance, licence and what an upgrade costs you
The repository is not archived, and the last push was on 2026-06-26. The most recent news item is dated the same day and announces native Transformers support with torch.compile. That is the state of the project as documented: a release cadence tied to model publications rather than a continuous stream of patches.
The package version is 0.0.6, which signals pre-1.0 API stability. Combined with exact dependency pins, an upgrade is not a version bump you can take casually. Moving to a newer qwen-asr release may require moving transformers, accelerate and vllm together, and the README does not publish a changelog or a deprecation policy. There are no retrieved releases, so there is no release-notes trail to read before upgrading.
The licence is Apache-2.0 for the repository, and pyproject.toml declares license = { text = "Apache-2.0" }. That covers the code and the packaging. Model weights are distributed through Hugging Face and ModelScope collections, and this article cannot confirm whether the weights carry the same terms as the repository. If you are shipping a product, check the model card licence on the specific checkpoint you download rather than assuming the repository licence applies. This is not legal advice; it is a pointer to the one place the documentation leaves open.
Editorial conclusion
Adopt Qwen3-ASR if you need multilingual recognition with a single model handling both offline and streaming input, and you accept pinning transformers==4.57.6 and, for batch serving, vllm==0.14.0. Skip it if your pipeline depends on word-level timestamps from the ASR model itself: those come from the separate Qwen3-ForcedAligner-0.6B, which covers 11 languages rather than 52. Before committing, verify two things in your own environment: that the 0.6B and 1.7B checkpoints load under your Python version (the package requires 3.9 or newer), and that the language you need appears in the released-models table rather than only in the headline count.
Frequently asked questions
What is Qwen3-ASR?
It is an open-source speech recognition family from the Qwen team at Alibaba Cloud, consisting of Qwen3-ASR-1.7B and Qwen3-ASR-0.6B for language identification and transcription, plus Qwen3-ForcedAligner-0.6B for timestamp prediction. The README states the family covers 30 languages and 22 Chinese dialects.
How do I use Qwen3-ASR?
Install the qwen-asr Python package, then either follow the example scripts in examples/ (example_qwen3_asr_transformers.py, example_qwen3_asr_vllm.py, example_qwen3_asr_vllm_streaming.py, example_qwen3_forced_aligner.py) or run the installed console commands qwen-asr-demo, qwen-asr-demo-streaming and qwen-asr-serve. The README points to the Hugging Face model cards for native Transformers usage examples.
Is Qwen3-ASR open source?
Yes. The repository is licensed Apache-2.0, and pyproject.toml declares license = { text = "Apache-2.0" }. The weights are published in the Qwen Hugging Face and ModelScope collections; check the model card for the specific checkpoint you download.
What are the key differences between Qwen3-ASR and Whisper Large V3?
The structural difference is that Qwen3-ASR does not produce timestamps from the recognition model; timestamp prediction is handled by the separate non-autoregressive Qwen3-ForcedAligner-0.6B, which covers 11 languages. Qwen3-ASR also states support for streaming and offline inference from a single model and lists songs with background music among its audio types. The README does not publish a benchmark comparison against Whisper Large V3.
Does Qwen3-ASR work with vLLM?
Yes. vLLM is an optional dependency pinned at vllm==0.14.0 in pyproject.toml, and the README describes the inference toolkit as supporting vLLM-based batch inference, asynchronous serving and streaming inference. The repository includes examples/example_qwen3_asr_vllm.py and examples/example_qwen3_asr_vllm_streaming.py.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/qwenlm-qwen3-asr)