# ASR-LLM-TTS: a Python pipeline that chains SenseVoice, Qwen2.5 and three TTS engines

> ABexit/ASR-LLM-TTS wires speech recognition, a small language model and speech synthesis into runnable scripts, with real-time interruption, voiceprint registration and custom wake words. It is a demo collection, not a packaged library.

**ABexit/ASR-LLM-TTS** — This is a speech interaction system built on an open-source model, integrating ASR, LLM, and TTS in sequence. The ASR model is SenceVoice, the LLM models are QWen2.5-0.5B/1.5B, and there are three TTS models: CosyVoice, Edge-TTS, and pyttsx3

- Repository: https://github.com/ABexit/ASR-LLM-TTS
- Stars: 1,279 · Forks: 207
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/abexit-asr-llm-tts

## What ASR-LLM-TTS actually solves, and for whom

Most voice assistant demos stop at a transcription endpoint. This repository goes one step further and closes the loop: microphone audio in, recognized text out, a language model reply, and synthesized speech back. The three stages are SenseVoice for ASR, Qwen2.5-0.5B or Qwen2.5-1.5B for the dialogue turn, and a choice of CosyVoice, Edge-TTS or pyttsx3 for synthesis. The README describes the project as built on the CosyVoice codebase, with SenseVoice and Qwen2.5 added on top.

The intended reader is someone with a GPU who wants to see the whole chain run locally and then modify it. The repository is a flat set of numbered scripts, not an installable package. There is no setup.py, no pyproject.toml in the top-level listing, and no published release. That shape tells you the audience: people who read Python files and edit them, not people who want pip install and a stable import path.

A second audience is narrower. The 241130 update adds voiceprint registration and custom wake words, which points at hobbyist or research prototypes of a named, always-listening assistant. If you want a wake-word product, the mechanism here (pinyin matching on top of SenseVoice output) is worth reading even if you build elsewhere.

## The data flow from microphone to speaker

The chain is sequential and each stage hands text or audio to the next. SenseVoice transcribes a captured buffer. The transcript becomes the user turn in a Qwen2.5 prompt. The model's text output is passed to whichever TTS backend the script selected, and the resulting audio is played back, typically through pygame or sounddevice.

The real-time scripts add a gate in front of the ASR stage. According to the README, 13_SenceVoice_QWen2.5_edgeTTS_realTime.py uses webrtcvad for live voice activity detection with a detection window of 0.5 seconds, an effective speech activation rate of 40 percent, and a chunk size of 20 milliseconds. The README spells out the arithmetic: 500 ms divided by 20 ms gives 25 detection segments, and 10 active segments out of 25 crosses the threshold, so that half second is treated as speech and appended to the buffer. The README itself flags the weakness: a model-based VAD would suppress noise better.

The 241130 update adds two more mechanisms. Voiceprint recognition uses Alibaba's CAM++ model, trained on 3D-Speaker Chinese data, with a fixed directory for enrollment audio; if that directory is empty the system enters registration mode, and the README notes that longer enrollment audio generally gives a more stable voiceprint. Wake words are matched by converting SenseVoice's recognized Chinese characters to pinyin and comparing against a configured string, so the wake phrase is defined as pinyin. The default in 15.0_SenceVoice_kws_CAM++.py is 'ni hao xiao qian'; in 15.1 it is 'zhan qi lai'.

History memory is a queue of user and system turns. Each new round reads the history first and concatenates the new instruction, with a maximum history length defaulting to 512. The README maps scripts to features: 15.0 has no history memory, 15.1 does.

## Installing ASR-LLM-TTS and running a first real-time conversation

The README assumes Anaconda and ffmpeg are already installed and tells you to find those tutorials yourself. It then creates a Python 3.10 environment:

```bash
conda create -n chatAudio python=3.10
conda activate chatAudio
```

PyTorch is installed separately with a CUDA build. The README's tested combination is torch 2.3.1 with CUDA 11.8, and it notes that any version above 2.0 worked in local testing:

```bash
pip install torch==2.3.1 torchvision==0.18.1 torchaudio==2.3.1 --index-url https://download.pytorch.org/whl/cu118
```

The light path skips CosyVoice entirely and installs a shorter dependency list. The README's command pins each package:

```bash
pip install edge-tts==6.1.17 funasr==1.1.12 ffmpeg==1.4 opencv-python==4.10.0.84 transformers==4.45.2 webrtcvad==2.0.10 qwen-vl-utils==0.0.8 pygame==2.6.1 langid==1.1.6 langdetect==1.0.9 accelerate==0.33.0 PyAudio==0.2.14
```

Models are downloaded automatically by default. The README points at line 215 for the ASR model, where model_dir = "iic/SenseVoiceSmall" triggers the download, and line 220 for the language model, where model_name = "Qwen/Qwen2.5-1.5B-Instruct" pulls from Hugging Face. The README warns that the Hugging Face route needs a working connection and offers ModelScope mirrors for both models.

With the light dependencies in place, the README gives the verification command:

```bash
python 13_SenceVoice_QWen2.5_edgeTTS_realTime.py
```

If that script starts and responds to speech, the ASR, LLM and TTS stages are all wired correctly. CosyVoice is a separate, heavier path. The README lists pynini and WeTextProcessing as the dependencies users report the most trouble with, and gives this sequence:

```bash
conda install -c conda-forge pynini=2.1.6
pip install WeTextProcessing --no-deps
```

After the remaining CosyVoice packages, the README's second verification target is 10_SenceVoice_QWen2.5_cosyVoice.py. The repository also ships a requirements.txt with a broader pinned set, including vllm 0.6.4.post1, gradio 3.43.2 and modelscope 1.15.0, plus commented-out optional packages for GPU attention and quantization. Note that requirements.txt installs opencv-python-headless while the README's light command installs opencv-python; do not mix the two blindly.

## Where the pipeline breaks or is the wrong choice

The README is direct about the central trade-off: CosyVoice inference is slow and seriously damages conversational real-time behavior. That is why pyttsx3 and Edge-TTS were added as alternatives. If your goal is natural-sounding local synthesis, you pay in latency; if your goal is responsiveness, you accept a cloud call or a robotic voice.

The Edge-TTS path has its own history. The README states that EdgeTTS produced connection errors during testing and that upgrading to version 6.1.17 fixed it, with no VPN required. That is a version-sensitive dependency on a Microsoft service, and the fix is documented as a specific pin rather than a configuration option. Treat the pin as load-bearing.

The VAD gate is threshold-based. With a 0.5 second window and a 40 percent activation rule, short utterances and quiet speech can be missed, and the README's own improvement note says noise interference remains. There is no documented calibration procedure for different microphones or rooms.

Finally, the repository is a collection of scripts rather than a service. There is no documented rollback, no release history, and no upgrade path beyond pulling the branch. The last push to master was on 2026-06-03, so the code has been static for a while. If you need a maintained dependency with a compatibility promise, this is the wrong shape of project. It is also the wrong choice if you cannot supply a GPU and a working PyTorch CUDA install, since the README's tested configuration is CUDA 11.8.

## How this differs from a general voice assistant stack

The obvious comparison is Home Assistant's voice pipeline, which also chains wake word, speech-to-text, intent handling and text-to-speech, but does so as a managed integration with pluggable components and a configuration UI. The difference in approach is where the language model sits. Home Assistant routes transcripts through intent matching by default and treats a language model as an optional conversation agent. Here the LLM is the intent layer: Qwen2.5 receives the transcript directly and its free-form reply is spoken. That gives you open-ended conversation instead of command matching, and it gives you no structured intent object to act on.

A second comparison is FunASR's own examples, since SenseVoice comes from that project. FunASR ships inference utilities and model wrappers; this repository is the integration layer around them, plus the VAD gate, the pinyin wake-word matcher and the CAM++ voiceprint step. If you only need transcription, FunASR is the smaller dependency. If you need the full loop with interruption, this repository is the assembled version.

The third axis is the TTS choice itself. CosyVoice is the local, higher-quality option inherited from the CosyVoice codebase. Edge-TTS is a network call to Microsoft's service with no key required in the documented setup. pyttsx3 is offline and uses the platform speech engine. Three backends in one repository, selected by which script you run, is a pragmatic way to trade quality against latency without rewriting the loop.

## Licence, maintenance and what an upgrade costs

The repository is licensed Apache-2.0, and a LICENSE file sits at the top level. That is a permissive licence, but the repository also carries a .gitmodules entry and a third_party directory, and the README states the main code comes from the CosyVoice project. Dependencies pulled in that way can carry their own terms, and model weights downloaded from ModelScope or Hugging Face have their own licences separate from the code. Check each model card before commercial use; this is not legal advice, only a pointer to where the obligations live.

Maintenance is the weaker signal. There are no retrieved releases, so there is no versioned artifact to pin against. The last push to master was on 2026-06-03. Upgrading means pulling the branch and re-reading the numbered scripts, because the README's changelog is organized by date (241027, 241123, 241130) and features are added as new script files rather than as options in one entry point. A feature you rely on may live in 15.1 while its predecessor sits in 15.0, and the only documented difference between them is history memory.

The dependency pins add real upgrade cost. torch 2.3.1 with CUDA 11.8, transformers 4.45.2, funasr 1.1.12 and edge-tts 6.1.17 are all fixed in the README. Moving any one of them forward means re-testing the VAD gate and the TTS connection, since the README records that a stale edge-tts version was the cause of a connection failure.

## Conclusion

Adopt ASR-LLM-TTS if you want a working reference for a local speech loop and you are willing to run numbered scripts rather than import a library: start with 13_SenceVoice_QWen2.5_edgeTTS_realTime.py, which needs no CosyVoice dependencies. Do not adopt it if you need a supported package with semantic versioning, a stable API surface, or documented rollback; there are no releases and the README does not document upgrade or downgrade paths. Before committing, verify that the SenseVoiceSmall and Qwen2.5 model files resolve on your machine, that the pinned torch 2.3.1 plus CUDA 11.8 wheel matches your driver, and that edge-tts 6.1.17 is the version you install, since the README states earlier versions hit a connection error.

## FAQ

### What is ASR, LLM and TTS in the ASR-LLM-TTS pipeline?

ASR is speech recognition, here SenseVoice, which turns microphone audio into text. The LLM is Qwen2.5-0.5B or 1.5B, which produces the reply text, and TTS is the synthesis stage, available as CosyVoice, Edge-TTS or pyttsx3. The repository chains the three in that order in each script.

### What is the difference between ASR and TTS in ASR-LLM-TTS?

In this project ASR is the input side, where SenseVoice transcribes recorded audio into text for the language model. TTS is the output side, where CosyVoice, Edge-TTS or pyttsx3 converts the model's reply back into audio. They are separate stages with separate dependencies in the README's install steps.

### Is ASR-LLM-TTS considered AI?

The repository is built entirely from machine learning models: SenseVoice for recognition, Qwen2.5 for dialogue and CAM++ for voiceprints, all described in the README as open-source models. The surrounding code is ordinary Python that moves audio and text between them.

### Which TTS model should I use with ASR-LLM-TTS?

The README states that CosyVoice inference is slow and hurts conversational real-time behavior, which is why pyttsx3 and Edge-TTS were added. Edge-TTS requires version 6.1.17 because earlier versions produced connection errors, and the README says no VPN is needed at that version. The README does not rank them beyond that latency and reliability trade-off.

## Sources

- [ABexit/ASR-LLM-TTS on GitHub](https://github.com/ABexit/ASR-LLM-TTS)
- [Issues](https://github.com/ABexit/ASR-LLM-TTS/issues)
- [License: Apache-2.0](https://github.com/ABexit/ASR-LLM-TTS/blob/master/LICENSE)
- [README](https://github.com/ABexit/ASR-LLM-TTS/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/abexit-asr-llm-tts
