Model or dataset
wwbin2017/bailing avatar
wwbin2017/bailing

Bailing (百聆): a low-latency voice assistant that runs without a GPU

百聆 是一个类似GPT-4o的语音对话机器人,通过ASR+LLM+TTS实现,集成DeepSeek R1等优秀大模型,接入openClaw,真正的个人语音助手,时延低至800ms,Mac等低配置也可运行,支持打断

1,774 stars305 forksPythonMIT

At a glance

What is it?
Bailing chains FunASR, silero-vad, a configurable LLM and several TTS engines into a speech conversation loop, with OpenClaw wired in as the tool-execution layer. The README claims 800ms end-to-end latency, but also warns it is for personal study rather than production.
Who is it for?
Adopt Bailing if you want a readable Python reference for a full speech loop and you accept the README's own warning that it is for personal study, not production. Skip it if you need a supported product with a rollback story: the disclaimer states there is no technical support or warranty, and the README documents no upgrade or rollback procedure.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 178 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Bailing actually assembles, and for whom

Bailing is a Python application that turns a microphone into a conversation. The README describes it as an open source voice assistant built from four stages: speech recognition, voice activity detection, a large language model, and speech synthesis. The stated goal is a GPT-4o-like experience on hardware that has no GPU, which is the part that separates it from most voice-assistant demos. The repository targets people who want to run the whole loop on a laptop or a small server and read the code while doing it.

The README names the components without ambiguity. FunASR handles recognition, silero-vad decides which audio is worth processing, DeepSeek is the default language model, and text-to-speech is delegated to edge-tts, Kokoro-82M, ChatTTS, or the macOS say command. The project also ships a plugin layer with functions such as get_weather, schedule_task, open_application and web_search, plus a general aigc entry point that routes through OpenClaw. That is a broader scope than a transcription tool: Bailing wants to be an assistant that acts, and the README says so directly when it calls OpenClaw the layer that moves the project from chatbot toward an action-taking assistant.

The Robot loop: how audio becomes a reply and how interruption works

The core of the design is a component the README calls Robot, which owns task management, memory management and the coordination between the other modules. The README's flowchart image is the primary description of the data path, and the accompanying table is the clearest piece of specification in the document. It defines four states from two booleans, player active and user speaking.

When the player is active and the user is silent, playback continues normally. When the player is active and the user speaks, that is the interruption case. When the player is idle and the user is silent, nothing happens. When the player is idle and the user speaks, silero-vad has judged the audio valid and it goes to ASR.

That table is worth more than the feature list because it tells you where the hard part lives. Interruption is not a flag on the TTS call; it is a state machine that has to decide, mid-playback, whether incoming audio is speech and whether to stop the current utterance. The README says interruption can be configured by keyword or by voice, which implies two different sensitivities: a keyword trigger is cheap and predictable, while voice-based interruption depends on the VAD threshold being tuned to your microphone. Neither threshold is documented in the README, so expect to tune them against your own room.

Installing Bailing and getting one spoken exchange working

The README states Python 3.12 or higher, pip, and the dependency sets for FunASR, silero-vad, DeepSeek and the TTS engines. Start by cloning and installing both requirement files. The second one lives under third_party/OpenManus and is easy to miss.

bash
git clone https://github.com/wwbin2017/bailing.git
cd bailing
pip install -r requirements.txt
pip install -r third_party/OpenManus/requirements.txt

Configuration is file-based. The README points at config/config.yaml for ASR and LLM settings, and at the DeepSeek platform for an API key, noting that OpenAI, Qwen, Gemini and 01yi are alternatives. The recognition model is not bundled: you download SenseVoiceSmall into models/SenseVoiceSmall from the Hugging Face link in the README. If you intend to use the general AIGC path, the README also asks for model, base_url and api_key in third_party/OpenManus/config/config.toml, and marks that path as under test.

bash
cd server
python server.py

The README says server.py starts the backend and that this step is optional for local use. The local entry point is main.py from the repository root. For server mode, which the README recommends, it gives an openssl command to generate a self-signed certificate for development and then says to run python server.py without changing directory, after which you open http://localhost:8000 and press the start button. Note the inconsistency: the local instructions change into server/ first, the server instructions do not. Follow the server section if you are using the browser interface, and expect the certificate warning that a self-signed cert produces. The README does not document what port the certificate is bound to or how the two server modes differ beyond the working directory.

Where Bailing breaks down or is the wrong pick

The disclaimer is the most important paragraph in the repository. It states that Bailing is for personal learning and research, is not suitable for commercial use or production environments, may cause data loss or system faults, and comes with no technical support or warranty. That is unusually blunt, and it should govern the decision more than the feature list does.

Several concrete constraints follow. The default language model is a hosted API, so the offline story is partial: recognition and synthesis can run locally, but the reasoning step needs a key and a network path unless you point the configuration at a local endpoint. The AIGC path through OpenManus is described in the README as under test, with a suggestion to fall back to the v0.0.1 or v0.0.2 tags if it does not work. That is a documented escape hatch, and it also tells you the current default branch may not be the stable one for that feature.

Tool execution is the other risk surface. The plugin table includes open_application, which launches applications on macOS by name, and schedule_task, which creates timed reminders. A voice-triggered path into those functions means a misrecognized phrase can start an application or create a task. The README does not describe a confirmation step or a permission model for plugin calls. If you are evaluating Bailing for anything shared or multi-user, that gap matters more than latency.

How Bailing differs from a hosted voice API

The obvious alternative is a hosted speech stack such as OpenAI's realtime voice API, where recognition, reasoning and synthesis arrive as one managed service. The difference in approach is not quality, it is where the seams are. A hosted API hides the VAD threshold, the interruption policy and the model choice behind one endpoint; Bailing exposes all of them as separate modules you can replace. The README leans on this: ASR, VAD, LLM and TTS are described as independent and swappable, and the TTS list alone offers four engines, including macOS say as a zero-dependency fallback.

That modularity cuts both ways. You can run Kokoro-82M locally and keep audio on the machine, which a hosted API will not do. You also inherit the integration work: model weights to download, a YAML file to edit, an .env for OpenClaw credentials, and two requirement files that must agree. For a single developer on a Mac who wants to understand the loop, the trade is reasonable. For a team that needs an SLA, the hosted option wins on the axis that matters, and Bailing's own disclaimer concedes the point.

Maintenance, releases and what the MIT licence does not cover

The last push to the default branch was on 2026-04-06. The repository is not archived. Releases are sparse: v0.0.1 in October 2024, v0.0.2 in March 2025, and v0.0.3 in May 2025, which the release notes describe as adding AIGC capability. There is no changelog file in the top-level entries, so the release notes and the git history are the only upgrade record. The README does not document a rollback procedure; the closest thing is the advice to switch to the v0.0.1 or v0.0.2 tags when the AIGC configuration does not work.

Upgrade cost is dominated by the dependency pins. requirements.txt pins funasr==1.1.6, chattts==0.1.1, transformers==4.41.1 and numpy==1.26.4, while torch and torchaudio are unpinned. That combination means a fresh install can resolve a newer torch than the pinned transformers expects, and the README offers no guidance on that. The pins are a stability choice with a maintenance bill attached.

The MIT licence lets you use, modify and distribute the code provided you keep the licence notice. It does not change the disclaimer, and it does not grant rights to the model weights you download separately: SenseVoiceSmall, Kokoro-82M and ChatTTS each carry their own terms, and the README links out rather than restating them. Check those licences for your use case; the project's MIT grant does not speak for them.

Editorial conclusion

Adopt Bailing if you want a readable Python reference for a full speech loop and you accept the README's own warning that it is for personal study, not production. Skip it if you need a supported product with a rollback story: the disclaimer states there is no technical support or warranty, and the README documents no upgrade or rollback procedure. Before committing, check that config/config.yaml exposes the ASR and LLM keys you actually need, confirm the SenseVoiceSmall weights are in models/SenseVoiceSmall, and verify that the OpenClaw credentials in config/.env match the tool surface you intend to expose to voice input.

Frequently asked questions

What is Bailing (百聆)?

It is an open source voice conversation assistant that combines ASR, VAD, an LLM and TTS. The README describes it as a GPT-4o-like voice robot built to run without a GPU, with OpenClaw integrated for tool calls.

How do I install Bailing?

Clone the repository, run pip install -r requirements.txt and pip install -r third_party/OpenManus/requirements.txt, configure config/config.yaml, download SenseVoiceSmall into models/SenseVoiceSmall, and set an API key. The README states Python 3.12 or higher is required.

Does Bailing need a GPU?

The README states the project aims to deliver a GPT-4o-like experience without a GPU and to run on edge devices and low-resource environments. The reasoning step still uses a hosted model API by default, so a network path and an API key are needed.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Releases
  5. wwbin2017/bailing on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/wwbin2017-bailing.svg)](https://hysenlabs.com/projects/wwbin2017-bailing)