Model or dataset
opendilab/CleanS2S avatar
opendilab/CleanS2S

CleanS2S puts a full-duplex speech agent in one file, and the samples stop mid-sentence

High-quality and streaming Speech-to-Speech interactive agent in a single file. 只用一个文件实现的流式全双工语音交互原型智能体!

541 stars54 forksPythonApache-2.0

At a glance

What is it?
A prototype from OpenDILab where ASR, LLM, and TTS are joined by two WebSockets, the receive side carrying voice activity detection, with queues and threads keeping every stage non-blocking. It is built to be read and modified rather than deployed. Two things to know before you spend an evening on it: every demo is a muted Chinese-language video, and the sample outputs in the comparison table are cut off by the compute limit the project admits to.
Who is it for?
Read it if you want the shape of a speech-to-speech pipeline in front of you at once, since a single file with the stages named and the queues visible is worth more as a reference than as a dependency. Expect to supply your own models, your own machine, and your own judgement about latency, because nothing in the repository states which speech recognition, language, or speech synthesis model it wires in.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Single-file describes the pipeline, not the repository

The claim at the top is that every detail of an agent pipeline is put into a single standalone file, with no extra burden to configure dependencies or understand the project file structure. That claim is about the pipeline, and the repository around it is not one file. The tree carries a backend directory and a frontend_nextjs directory, and the document outline splits its getting-started section into a Backend (Server) half and a Frontend (Client) half, so there is a client to write and a server to run. The single-file framing is the point rather than an accident: the stated purpose is to let someone glance at the S2S pipeline and validate ideas on top of it, with every implementation easy to modify, a swappable model such as the LLM, new components, and a customisable pipeline. It is a reference implementation for researchers, and the framing is what makes the surrounding structure a fair trade.

ASR, LLM, and TTS joined by two WebSockets, with VAD on the receive side

The pipeline is named in one sentence: automatic speech recognition, a large language model, and text to speech, connected by two WebSocket components, a Receiver that contains voice activity detection and a Sender. Every piece of audio and text is streamed in both directions over that socket, which is what makes the conversation feel live rather than request and response. Blocking is handled explicitly rather than by luck: multi-threading and queueing keep the stream moving, every component is written as asynchronous and non-blocking, each one reads from an input queue and writes its result into another. That structure is where to look if you are adapting this, since the alternative to queue-per-stage is a stage that stalls the microphone. It is also why the file is worth reading end to end rather than skimming for the model calls.

An interruption becomes part of the context, not a discarded utterance

Full duplex means the user can speak and listen at the same time, and interruption is a first-class event rather than a cancellation. When you talk over the agent, it stops what it is doing and starts on the new input, carrying the context of the previous conversation and of the interruptions themselves. That second clause is the interesting one: what you interrupted with becomes something the agent can reason about later, which is what turns a barge-in from an audio artefact into conversational state. This rides on the WebSocket mechanism rather than on a separate protocol, and the project credits the library for making it practical. It is also the hardest part to evaluate from outside, because a demo that plays start to finish tells you nothing about how the agent behaves when cut off mid-sentence, and none of the linked conversations are labelled with interruption points.

Turn-based chat is named as the design flaw to avoid

The project argues this out rather than asserting it. Its position is that the assistant-style, turn-based response used in ordinary chatbots is one of the most important drawbacks for human-like conversation, and that the strategies it adds are meant to make exchanges more interactive and engaging. That framing matters for reading the rest of the design, because a turn-based assumption would be baked into the queues: one utterance in, one reply out, nothing in between. What the project does not offer is measurement. There are no latency figures, no word error rate, no interruption timing, and no comparison against a turn-based build of the same models, so the argument for full duplex rests on the demos and on the architecture rather than on anything you can check. Subjective Action Judgement is the other structural addition, described as enhancing the agent's ability to initiate actions proactively during a conversation, which is the piece that would let it speak first.

The RAG handler changes the register of the answer, not only its facts

Two optional components sit beside the language model. The WebSearchHelper class runs online searches from the user's query or to gather material relevant to the conversation, so the agent can reach outside its weights. The RAG class does retrieval from a database first and then generates, with the stated intent that replies are grounded in relevant, factual data. The comparison table for the two paths is more revealing than the descriptions. Given a foundation product, the plain LanguageModelHandler produces a structured product write-up with feature sections and ingredient lists, while RAGLanguageModelHelper produces a first-person reaction in conversational Chinese, enthusiastic and colloquial, referring to what the user previously struggled with. Both are truncated by the same limit, and both cells keep literal newline escape sequences rather than rendered breaks. The honest reading is that retrieval changed voice and specificity, not accuracy, and neither sample is long enough to show whether the grounding held up.

Output length is capped by compute, and the samples end mid-sentence

There is a note attached to the output examples stating that, because of computing resource limitations, the maximum token output is limited to a small size. Every cell in the comparison table stops in the middle of a clause, in the structured column and the retrieved column alike, which means the examples demonstrate the opening of a response rather than a completed one. That is a reasonable constraint for a prototype and a poor basis for judging quality, since truncation is exactly where a language model's weaknesses show. The same caveat applies to the conversation demos, which are hosted video attachments with a note asking you to unmute first. All five labelled topics are Chinese-language: two on investing, one on mood, one on gaokao volunteer choices, and one on stomach medicine. A reader who does not read Chinese cannot evaluate the speech, the prosody, or the turn-taking from anything in this repository.

The feature announcement links to a Chinese document, and the last commit is from April

Two small signals about maintenance sit next to each other. The Subjective Action Judgement feature is announced with a link into backend/README.zh.md, so following the project's newest capability from the English README lands in the Simplified Chinese documentation, and the English README itself links its Chinese counterpart in the same way. There is a README.zh.md at the top level, so the two-language setup is deliberate, but the anchor points are not consistently mirrored. On cadence, the last push was on 2026-04-07 and the repository publishes no GitHub releases, so there is no tag to compare against and nothing to install by version. The numbers around it are 541 stars, 54 forks, and 4 open issues, with Apache-2.0 as the licence and a NOTICE file at the root, which is the right shape for a project that assembles other people's models.

Editorial conclusion

Read it if you want the shape of a speech-to-speech pipeline in front of you at once, since a single file with the stages named and the queues visible is worth more as a reference than as a dependency. Expect to supply your own models, your own machine, and your own judgement about latency, because nothing in the repository states which speech recognition, language, or speech synthesis model it wires in. Do not judge output quality from the demos or the sample table: both are Chinese-language, both are muted or truncated, and the project itself says its token output is capped by computing resources.

Frequently asked questions

What is CleanS2S built from?

Automatic speech recognition, a large language model, and text to speech, connected by two WebSocket components: a Receiver that contains voice activity detection and a Sender. All audio and text streams in both directions over that socket, and each stage reads from an input queue and writes into another.

Does CleanS2S support interruption and full duplex?

Yes. The user can speak and listen at the same time, and interrupting the agent with new speech stops current processing and starts on the new input, carrying the context of the previous conversation and of the interruptions.

Does CleanS2S do web search or retrieval?

Both, as optional additions. A WebSearchHelper class conducts online searches from the user query or to gather conversation-relevant information, and an RAG class retrieves from a database first and then generates a response from what it retrieved.

Which models does CleanS2S use for speech recognition?

The repository describes the pipeline by role, naming ASR, LLM, and TTS stages, and states that the user can quickly change the model they like, for example the LLM. It does not name specific speech recognition or speech synthesis models anywhere in the README.

Is CleanS2S a production service?

It describes itself as a prototype and a reference implementation for researchers who want to glance at the S2S pipeline and validate ideas on top of it. The repository publishes no releases, the last push was on 2026-04-07, and the sample outputs are cut short by a stated computing resource limit on maximum token output.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. opendilab/CleanS2S on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/opendilab-cleans2s.svg)](https://hysenlabs.com/projects/opendilab-cleans2s)