SenseVoice: ASR that also returns emotion and event tags
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
At a glance
- What is it?
- QwenAudio/SenseVoice is an MIT licensed speech foundation model that does recognition, language identification, emotion recognition and audio event detection in one pass. The released SenseVoiceSmall checkpoint covers five languages, not the fifty the wider research reports, and speaker diarization is a separate composed pipeline.
- Who is it for?
- Adopt SenseVoice if your audio is Mandarin, Cantonese, English, Japanese or Korean and you want transcription, language identification, emotion and audio event tags from one pass, particularly where throughput matters, since the README reports it running over five times faster than Whisper-Small on a non-autoregressive architecture.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 19 days ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 21, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Four tasks from one model
SenseVoice is a speech foundation model with several understanding capabilities at once: automatic speech recognition, spoken language identification, speech emotion recognition and audio event detection.
The transcription alone would put it in a crowded field. What makes it unusual is that the other three come out of the same pass rather than as separate models you run afterwards. A call centre analysing recordings typically wants the words, the language, whether the caller sounded angry, and whether there was background music or a cough in the segment. Getting those from one model instead of four changes the engineering considerably.
The README says the model supports detection of common human-computer interaction events including background music, applause, laughter, crying, coughing and sneezing, and claims emotion recognition that achieves and surpasses the current best emotion recognition models on its test data.
The released checkpoint is SenseVoiceSmall. The README is careful to note that speaker diarization, the task of deciding who spoke when, is a composed FunASR pipeline using separate FSMN-VAD and CAM++ models, and is not an output of the SenseVoiceSmall checkpoint itself.
Five languages, not fifty
This is the clarification to read before anything else. The README separates research scope from released checkpoint: the broader SenseVoice work reports training on more than 400,000 hours and support for more than 50 languages, while the released SenseVoiceSmall checkpoint supports Mandarin, Cantonese, English, Japanese and Korean.
That is a large gap and the README states it rather than letting the marketing imply the wider number. If your users speak Portuguese or Hindi, this is not your model today.
The five it does cover are a deliberate set. Mandarin and Cantonese are the pair most multilingual models handle badly, since Cantonese is often treated as Mandarin or not covered at all, and the README's benchmark section says SenseVoice-Small has advantages in Chinese and Cantonese recognition when compared against Whisper on AISHELL-1, AISHELL-2, Wenetspeech, LibriSpeech and Common Voice.
The benchmark images are in the repository rather than reproduced as numbers in the README, so anyone making a decision on accuracy should read the figures in image/asr_results1.png and asr_results2.png, or run their own evaluation.
Non-autoregressive inference
The architecture choice is what produces the speed claim. The README says SenseVoiceSmall uses a non-autoregressive end-to-end framework for low-latency inference, and states that at a similar parameter count it runs more than 5 times faster than Whisper-Small and 15 times faster than Whisper-Large, in the benchmark setup shown in the repository.
Autoregressive decoders generate one token at a time and each step depends on the last, so long utterances cost proportionally more. A non-autoregressive model produces its output in parallel, which is where that multiple comes from.
The trade is the usual one for the approach. Parallel decoding generally gives up some accuracy on difficult or ambiguous audio, because the model cannot revise an earlier decision in light of a later one. The README does not discuss that trade, and it does not publish an error rate alongside the speed figure.
For a workload where latency dominates, such as transcribing long call recordings in bulk, the speed matters more. For a workload where the transcript is the product, measure accuracy on your own audio before believing either number.
Installing and running it
Dependencies come from the repository's requirements file:
pip install -r requirements.txtThe README adds a version constraint that is easy to trip over: SenseVoiceSmall examples and the composed FunASR diarization path require funasr 1.3.26 or newer, and anyone who installed this repository earlier should run pip install -U "funasr>=1.3.26" before retrying the demos. The current deployment path given is funasr 1.4.14, covering Python, an OpenAI-compatible service and container workflows.
The repository ships several entry points rather than one CLI. There is demo1.py, demo2.py, demo_libtorch.py, demo_onnx.py, a webui.py and an api.py, plus model.py, export.py and export_meta.py for getting the model into other formats.
Long audio needs care. The README says long recordings must be segmented before reaching the encoder and that the inference example uses FSMN-VAD for that segmentation. There is also long_audio_no_vad.py, which the README describes as using bounded overlapping windows and preserving raw chunk outputs, so hour-scale recordings do not require one unbounded GPU allocation.
Fine-tuning is supported through finetune.sh and a deepspeed_conf directory, aimed at long-tail cases in a specific domain.
Emotion and event detection, with their limits
The README is more careful here than most model cards, and it should be read as guidance rather than as a capability claim.
On emotion, it says that without fine-tuning on the target data, SenseVoice was able to achieve and exceed the performance of the current best speech emotion recognition models, across test sets in Chinese and English covering performances, films and natural conversations. It also notes that SenseVoice-Large achieved the best performance on nearly all datasets, and that Small surpassed other open source models on most of them. The repository includes a reproducible evaluation contract under benchmarks/ser for a zero-shot CASIA or RAVDESS rerun, which reads the raw emotion tag and reports unweighted and weighted accuracy rather than parsing formatted transcription text.
On events, the README is candid about the gap. It says that although trained exclusively on speech data, SenseVoice can function as a standalone event detection model, and that on ESC-50 it achieved commendable results against BEATS and PANN, but that due to limitations in training data and methodology its event classification performance has gaps compared to specialised AED models.
That is the right way to use it: as a cheap signal alongside transcription, not as a replacement for a dedicated acoustic event classifier.
The llama.cpp runtime and its GPU problems
For people who want to run this without PyTorch, the project publishes prebuilt binaries for a FunASR llama.cpp and GGUF runtime, covering SenseVoice, Fun-ASR-Nano and Paraformer with built-in FSMN-VAD.
The most recent release, runtime-llamacpp-v0.2.1 from 2026-08-27, is informative about what is still rough. It fixes Vulkan selection when a requested device is reported as an integrated GPU, preferring a matching discrete GPU and falling back to integrated. It provides nine desktop archives for Linux, macOS and Windows including Vulkan and Windows CUDA variants. It states that Android and Mali are not an official prebuilt or validated target.
It also says the Windows Vulkan package is ready for Radeon 780M retesting, and that the selector fix does not claim to resolve a separate Radeon RX 9070 XT crash with code 0xC0000005 during initialisation. The Windows CUDA asset targets architecture 86, so other GPU architectures require building from source.
An earlier release, v0.1.9, moved the source snapshot onto current main to include a yaml.SafeLoader security fix, because GitHub-generated source archives were still shipping the older unsafe YAML loader.
That is a runtime in active development with real hardware gaps. Check the asset list against your GPU before assuming a binary exists for you.
Whisper as the alternative
The comparison the README invites is Whisper, and the two make opposite bets.
Whisper covers far more languages, has the largest tooling ecosystem of any open speech model, and is autoregressive, so it is slower per unit of audio at comparable size. It gives you a transcript and, depending on the variant, timestamps and translation. It does not give you emotion or audio event tags.
SenseVoice covers five languages, is non-autoregressive and therefore much faster per the README's own benchmark, and returns emotion and event tags in the same output. Its README claims an advantage in Chinese and Cantonese recognition specifically.
Choose Whisper if language coverage or ecosystem maturity decides it, or if you need translation. Choose SenseVoice if your audio is Mandarin or Cantonese, if throughput matters more than the last point of accuracy, or if the emotion and event tags are the reason you are transcribing at all. Running both on a sample of your own audio is a day's work and will tell you more than either README.
Licence and maintenance
SenseVoice is MIT licensed, with a LICENSE file at the root. That is permissive, allows commercial use, and carries no copyleft obligation, which is notable for a model of this size.
The last push to the repository was on 2026-09-10. Development is spread across two tracks: the model and Python code here, and the runtime binaries published through releases, with the newest being runtime-llamacpp-v0.2.1 on 2026-08-27.
The paper is on arXiv as 2407.04051, and checkpoints are published on both ModelScope and Hugging Face under FunAudioLLM, with online demos on both platforms.
Maintenance cost concentrates in two places. The funasr version floor moves, so upgrades need attention, and the runtime binaries are hardware-specific enough that a GPU change can mean building from source. Pinning funasr and testing the runtime asset on your actual hardware are the two checks worth doing before you commit.
Editorial conclusion
Adopt SenseVoice if your audio is Mandarin, Cantonese, English, Japanese or Korean and you want transcription, language identification, emotion and audio event tags from one pass, particularly where throughput matters, since the README reports it running over five times faster than Whisper-Small on a non-autoregressive architecture. Do not adopt it for broader language coverage, because the released checkpoint is five languages and not the fifty the research reports, and do not treat the event tags as a substitute for a dedicated classifier, which the README says still outperforms it on ESC-50. Speaker diarization needs a separate composed pipeline of FSMN-VAD and CAM++. Verify on your own audio before switching, and if you plan to use the llama.cpp runtime, check the release asset list against your GPU first, because the Windows CUDA build targets architecture 86 and a Radeon RX 9070 XT initialisation crash is still open.
Frequently asked questions
Which languages does SenseVoice support?
The released SenseVoiceSmall checkpoint supports Mandarin, Cantonese, English, Japanese and Korean. The README notes the broader research reports over 50 languages, but that is not what the published checkpoint covers.
Does SenseVoice do speaker diarization?
Not on its own. The README says diarization is a composed FunASR pipeline using separate FSMN-VAD and CAM++ models, and is not an output of the SenseVoiceSmall checkpoint.
What audio events can SenseVoice detect?
The README lists background music, applause, laughter, crying, coughing and sneezing among common human-computer interaction events, and says its event classification has gaps compared to specialised audio event detection models.
What version of funasr does SenseVoice need?
The README requires funasr 1.3.26 or newer for the SenseVoiceSmall examples and the diarization path, and gives funasr 1.4.14 as the current deployment path for Python, the OpenAI-compatible service and containers.
How does SenseVoice handle long recordings?
The README says long recordings must be segmented before the encoder, with FSMN-VAD used in the example. There is also long_audio_no_vad.py, which uses bounded overlapping windows so hour-scale audio does not need one unbounded GPU allocation.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/qwenaudio-sensevoice)