# Qwen Audio Agent: a realtime voice runtime that keeps a backend coding agent in the conversation

> Qwen Audio Agent wraps a full-duplex voice frontend around ACP-compatible coding agents, so tasks run in the background while the conversation continues. It installs from npm, needs Node 22.22.2+, and depends on a DashScope key unless you point it at a local speech-to-speech service.

**QwenAudio/qwen-audio-agent** — A realtime voice runtime that keeps Agents talking, working, and present.  Real-time Voice Runtime for AI Agents

- Repository: https://github.com/QwenAudio/qwen-audio-agent
- Website: https://qwenaudio.github.io/qwen-audio-agent/
- Stars: 2,820 · Forks: 277
- Language: JavaScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/qwenaudio-qwen-audio-agent

## What Qwen Audio Agent actually solves

Most voice assistants answer one sentence and then go quiet, or they stall the moment a tool call starts. The README frames the project around that gap: conversation should not leave you waiting, and it should not halt because the agent is looking something up. Qwen Audio Agent is a voice runtime that keeps a full-duplex audio channel open while a backend agent does work, then speaks the result back into the same conversation.

The audience is narrower than "anyone who wants voice AI". It is for people who already have a coding or task agent they trust, and who want to talk to it instead of typing. The Agent Support table lists Qwen Code, OpenCode, OpenClaw, Qoder, MiniMax Code, Kimi Code, Hermes, CodeBuddy, Codex, Claude Code, DeepSeek and Pi, plus a generic ACP option and a frontend-only mode with no backend at all. If none of those names mean anything to you, the project has little to offer.

## Frontend answers, backend works: the two-track design

The architecture separates the realtime voice frontend from the task-executing backend agent, and the README describes the split plainly: questions that can be answered directly are answered immediately, and when tools or sustained processing are needed the task is delegated to the backend. The user always faces the same assistant.

That delegation is what makes interruption and parallel work possible. A task gets an independent lifecycle, the frontend can report progress or cancel it, and the result returns to the current conversation so you can ask follow-up questions or request modifications. Backends are unified under ACP, and the repository distinguishes native ACP integrations from external ACP adapters (Codex, Claude Code and Pi are listed as adapters, which means an extra component in the install path). The Gateway is the piece that owns the realtime model and the agent process. The .env.example notes that changing the Gateway's realtime model requires a restart, which tells you the Gateway is a long-lived process rather than something spawned per utterance.

## Installing Qwen Audio Agent and running a first session

The README gives one recommended install path: a global npm install. Node.js 22.22.2+ or 24.15.0+ and npm 10+ are required, and the package is published as qwen-audio-agent.

```bash
npm install -g qwen-audio-agent
```

After that, the CLI generates a config file. The README's quick start shows a single command, and the resulting file is dotenv-style.

```bash
qwenaudio config
```

The first key you fill in is the DashScope credential, because the default frontend is Qwen Audio Realtime. The .env.example also shows how to select the realtime model and voice, and how to switch the backend agent.

```dotenv
DASHSCOPE_API_KEY=your-key
QWEN_AUDIO_REALTIME_MODEL=qwen-audio-3.0-realtime-plus
AGENT_PROTOCOL=openclaw
```

If you want to skip DashScope entirely, the .env.example points the frontend at a speech-to-speech service you install and run yourself, in which case the Gateway needs no DashScope key and that service manages STT, LLM, TTS and voice selection.

```dotenv
QWEN_AUDIO_REALTIME_PROVIDER=speech-to-speech
SPEECH_TO_SPEECH_REALTIME_URL=ws://127.0.0.1:8765/v1/realtime
```

Three frontends exist: WebUI, terminal TUI, and a desktop floating orb for macOS, Windows and Linux. The README's build-from-source and GitHub-install routes live in a separate installation guide rather than the main file, so anyone who cannot use a global npm install is reading docs the README does not inline.

## Where the runtime gets in your way

The dependency on DashScope is the largest constraint. The default voice frontend is a hosted Qwen model, so latency and availability are not yours to control, and the .env.example treats the API key as mandatory for the default path. The local speech-to-speech route removes that dependency but adds a second service you must install, configure and keep running, with its own LLM authentication.

Model and voice selection is asymmetric in a way worth knowing before you build a config. The .env.example states that Audio and Omni voice preferences are stored separately and that switching model families does not rewrite the other family's preference, so a voice you set for one model will not carry over. The Omni models can accept negotiated realtime JPEG frames from the WebUI, but the same file says Desktop and TUI do not capture video at all. If your use case is visual, the frontend you pick decides whether it is possible.

Permission handling is another edge. The .env.example describes two modes, native (the backend asks) and full (highest privilege), and notes that Pi has no permission approval mechanism, so it always runs in full regardless of configuration. That is a documented boundary, not a bug, but it means the permission setting is not uniformly meaningful across backends. The Agent Support table also carries its own caveat: four-star ratings indicate active development or integrations not yet fully verified, which covers MiniMax Code, Hermes, CodeBuddy, Codex, Claude Code, DeepSeek and Pi.

## How it differs from wiring a voice model to a tool loop yourself

The obvious alternative is a speech-to-speech pipeline you assemble from a VAD, an STT model, an LLM and a TTS model, which is exactly what the Hugging Face speech-to-speech project provides and what Qwen Audio Agent added frontend integration for in v1.3.0. The difference is where the agent lives. A hand-built pipeline gives you full control over every stage and every model, and no vendor key is required if you host the components. Qwen Audio Agent instead assumes the intelligence sits in an existing ACP agent and treats voice as the transport into it.

That trade favors reuse over control. You inherit the backend agent's model configuration, tools, MCP servers, Skills and authentication rather than rebuilding them, and you get task lifecycle, progress reporting and result return as runtime behavior. What you give up is the ability to swap the language model independently of the agent, and you accept the Gateway as a process you run and restart. Notably, the project does not treat these as exclusive: you can point the frontend at a local speech-to-speech service and still keep an ACP backend, which is the configuration to consider if the DashScope requirement is the blocker.

## Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-09-10. Releases have been frequent: v1.10.0 on 2026-08-13, v1.10.1 on 2026-08-15, v1.11.0 on 2026-08-20, and the README's news section announces v2.0.0 as in development, with work on agent architecture, task lifecycle, multimodal input, memory and extensibility. A major version in flight is a real upgrade consideration: the release notes describe ongoing changes to the task lifecycle and agent architecture, which are the parts you would build against.

The package is licensed Apache-2.0, and the repository ships LICENSE, NOTICE and THIRD_PARTY_NOTICES.md, which is the usual arrangement for a project that bundles dependencies. Apache-2.0 permits commercial use and modification with the standard attribution and notice obligations; whether those obligations apply to your distribution is a question for your own counsel, not something this article can settle. If you embed the Gateway or the Realtime Provider extensions added in v1.11.0, check the exported entry points in package.json (./gateway-protocol, ./gateway-client-sdk, ./realtime-events and others) rather than importing internal paths, since those exports define the supported surface.

## Verifying a Qwen Audio Agent install before you standardize on it

The first thing to confirm is the backend integration type for your agent. The table separates native ACP from external ACP adapters, and the adapters require installing a base plus an adapter, which is more moving parts than the one-click native path. The second is your frontend choice against your input needs: if you need the Omni model's realtime image input, the .env.example limits that to the WebUI. The third is the model and voice pairing, since Audio and Omni voices are tracked separately and a model switch will not carry your preference across.

Start with AGENT_PROTOCOL set to none if you want to hear the voice frontend before committing to a backend. The Agent Support table lists that as a frontend-only mode with no config needed, and it isolates whether your audio path works from whether your agent integration works. Once voice is stable, set AGENT_PROTOCOL to your backend and re-check the permission mode, remembering that Pi ignores it. Keep the Gateway restart requirement in mind for any realtime model change, because an edit that appears to do nothing is usually a Gateway that has not been restarted.

## Conclusion

Adopt it if you already run an ACP-compatible backend agent and want voice as the front end rather than a separate chat app: the one-click install path reuses that agent's model configuration, tools, MCP, Skills and authentication. Skip it if you want a hosted, no-key voice assistant, since the default frontend runs on DashScope and the .env.example makes DASHSCOPE_API_KEY the first thing you fill in. Before committing, verify two things in your own setup: that your chosen backend appears in the Agent Support table with the integration type you need (native ACP versus an external adapter), and that your Node version satisfies the 22.22.2+ or 24.15.0+ requirement, because the package declares type module and ships ESM exports.

## FAQ

### Can Qwen Audio Agent process audio?

Yes. The project is a realtime voice runtime with a full-duplex frontend, and the README lists natural interruption and sustained multi-turn conversation among its core features. Audio is handled by the Qwen Audio Realtime model by default, or by a speech-to-speech service you run yourself.

### Does Qwen Audio Agent have an agent?

It runs your agent rather than shipping one. Backends are unified under ACP, and the Agent Support table lists Qwen Code, OpenCode, OpenClaw, Qoder, Kimi Code, Codex, Claude Code and others, plus a frontend-only mode with no backend.

### Is Qwen Audio Agent free?

The package is published under Apache-2.0, so the code itself is free to use. The default voice frontend calls DashScope and the .env.example requires a DASHSCOPE_API_KEY, so running it has a service cost unless you switch to a self-hosted speech-to-speech service.

## Sources

- [License: Apache-2.0](https://github.com/QwenAudio/qwen-audio-agent/blob/main/LICENSE)
- [Project website](https://qwenaudio.github.io/qwen-audio-agent/)
- [QwenAudio/qwen-audio-agent on GitHub](https://github.com/QwenAudio/qwen-audio-agent)
- [README](https://github.com/QwenAudio/qwen-audio-agent/blob/main/README.md)
- [Releases](https://github.com/QwenAudio/qwen-audio-agent/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/qwenaudio-qwen-audio-agent
