# py-xiaozhi: A Python Client for the Xiaozhi Voice Assistant Stack

> py-xiaozhi is a Python async client and MCP tool host that connects microphones, cameras and GPIO hardware to Xiaozhi AI services. It is the desktop and edge counterpart to the ESP32 firmware, and its dependency list is as much a constraint as a feature.

**huangjunsen0406/py-xiaozhi** — Open-source AI assistant ecosystem with MCP integrations, multimodal workflows, IoT support, and cross-platform voice interaction.

- Repository: https://github.com/huangjunsen0406/py-xiaozhi
- Website: https://huangjunsen0406.github.io/py-xiaozhi/
- Stars: 3,487 · Forks: 742
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/huangjunsen0406-py-xiaozhi

## What py-xiaozhi Solves, and Who It Is Actually For

The xiaozhi-esp32 firmware runs the Xiaozhi voice assistant on ESP32-S3 boards. That is a good fit for a speaker-sized device and a poor fit for anything with a camera, a screen, a GPIO rig or a full Linux userspace. py-xiaozhi exists to move the client half of that stack onto hardware that already runs Python: Windows 10+, macOS 10.15+, Linux x86_64 and ARM boards including the Raspberry Pi, Horizon Robotics RDK and Jetson Nano. The README describes it as evolving from the xiaozhi-esp32 firmware project, and notes it is used as an upstream dependency by D-Robotics for xiaozhi-in-rdk.

The intended user is someone building an embodied or edge assistant rather than a chat app. The feature list leans that way: Opus codec handling with RFC 6716 TOC frame parsing, camera capture feeding a vision-language model, a JSON-RPC 2.0 tool server, GPIO access for actuation, and Sherpa-ONNX wake word detection running on the device. If your goal is a text chatbot in a browser tab, this project is solving a problem you do not have, and the audio and platform dependencies will be dead weight.

## Event Loop, Protocol Layer, and the MCP Tool Server

The architecture is described in the README as event-driven on top of the asyncio event loop, with an application layer, a protocol layer and a UI layer kept separate, and component lifecycle managed through a bootstrap container with dependency injection. That is the shape of the code, not a marketing claim: the top-level layout has main.py, src/, libs/, models/, examples/mcp_plugins/ and a tests/ directory.

The transport is dual-protocol. The README lists WebSocket and MQTT with WSS/TLS encryption and auto-reconnection, which matches the dependency set: aiohttp and websockets for the socket path, paho-mqtt for the broker path. Audio arrives as Opus, and the README states the client performs auto frame detection by parsing the RFC 6716 TOC byte rather than assuming a fixed frame size. That detail matters on embedded boards, where resampling and buffering decisions determine whether voice feels responsive.

The MCP side is where the project diverges from a plain voice client. Tools are exposed as a modular JSON-RPC 2.0 server, and the README names music player, camera, screenshot, app management, weather and volume control as built-ins. The examples/mcp_plugins/ directory is the extension point. A camera tool and a GPIO tool sit behind the same interface as a weather lookup, so a voice command can end in a hardware action. The README does not document a permission model for those tools, which is worth noting before you wire a screenshot tool to a wake word on a shared machine.

## Installing py-xiaozhi and Running the Client Once

The README points to the project documentation site for startup tutorials and file descriptions, and requires Python 3.10 to 3.12. Two dependency files exist: pyproject.toml, which is the source of truth, and requirements.txt, whose header recommends uv. The GUI mode is an optional extra, so a headless install skips PySide6 and qasync.

A headless install with uv, matching the extra name declared in pyproject.toml:

```bash
uv sync
```

If you want the PySide6 and qasync GUI stack, the pyproject comment gives the extra form:

```bash
uv sync --extra gui
```

The requirements.txt header also allows a plain pip path. Note that requirements.txt lists PySide6 unconditionally while pyproject.toml keeps it optional, so the two files are not equivalent:

```bash
pip install -r requirements.txt
```

With dependencies in place, the entry point is main.py at the repository root. The README does not print a full invocation, so run the module and follow the documentation site for configuration of the AI endpoint, wake word model and UI mode:

```bash
python main.py
```

Expect the first run to fail on missing configuration rather than on missing code. The README states that voice wake-up requires downloading Sherpa-ONNX speech recognition models separately, and that camera features require a camera device and OpenCV support. Those are not pulled in by the sync step. It also states that after each update you should manually reinstall pip dependencies to pick up new ones, which is a real operational cost on a device you have already deployed.

## Where py-xiaozhi Breaks or Is the Wrong Choice

The dependency list is the first hard constraint. Python is pinned to 3.10 through 3.12, and several packages carry platform markers: pycaw, pywin32 and comtypes only on win32; applescript and two pyobjc frameworks only on darwin; gpiozero and lgpio only on linux. On Linux, the GPIO dependencies install whether or not you intend to use GPIO mode, and lgpio on a non-Pi x86_64 machine is at best inert. The README recommends a CPU with AVX support and at least 4GB of RAM, 8GB preferred, plus a 16kHz-capable audio device. A 1GB single-board computer will not run this comfortably.

The second constraint is the service dependency. The README describes network connection as a basic requirement, needed for AI services and online features. There is no documented local-only mode, so if your deployment site has no reliable internet, the voice path degrades to whatever offline pieces remain, which the README limits to wake word detection. That is a meaningful gap for robotics and industrial settings where the project otherwise looks attractive.

The third is documentation asymmetry. The README is strong on feature inventory and thin on operational detail. It does not document rollback after a bad update, does not describe how MCP tool permissions are enforced, and does not give a full main.py invocation with flags. The repository does contain deep-audit.md and risk-analysis.md at the top level, which suggests the maintainers have looked at failure modes themselves, but the README does not summarise their conclusions.

## py-xiaozhi Against xiaozhi-esp32 and xiaozhi-desktop

The nearest alternative is the project it came from. xiaozhi-esp32 is firmware for ESP32-S3 hardware: you flash a board and get a dedicated voice device. py-xiaozhi is a Python process on a general-purpose operating system. The practical difference is what you can attach. Firmware gives you a predictable, low-power device with tight audio timing. Python gives you OpenCV, a filesystem, a package manager, GPIO libraries and the ability to add an MCP tool in a few lines. If your assistant needs to read a camera frame and act on it, the firmware path means building that elsewhere.

xiaozhi-desktop is the other sibling, listed in the README as an Electron desktop client with AEC echo cancellation, Live2D and floating window modes, distributed as Windows and macOS installers. The split is straightforward: xiaozhi-desktop targets people who want an installable desktop app with an animated avatar and echo cancellation handled for them. py-xiaozhi targets people who want to run the client on a Raspberry Pi, an RDK board or a Jetson, script it, and extend it with their own tools. If you want the desktop experience specifically, the Electron client is the shorter path, and the README does not claim py-xiaozhi matches its AEC behaviour.

## Licence, Release Cadence and the Upgrade Bill

The licence is MIT, declared in both the LICENSE file and the pyproject.toml license field. That permits commercial use and modification with attribution and no warranty. It also means there is no contributor licence agreement or copyleft obligation to check, which simplifies redistribution on a shipped device. This is a description of the licence text, not legal advice; if you are embedding the client in a product, have counsel review the bundled third-party dependencies separately, since those carry their own terms.

The release history shows v2.0.8 on 2026-07-18, v2.1.0 on 2026-07-25 and v2.1.1 on 2026-07-26, and the last push to the repository was on 2026-08-17. Two patch releases inside eight days followed by a quieter month is a normal pattern for a project in this state, but it also means the API surface is still moving. The README's instruction to reinstall pip dependencies manually after each update is the concrete upgrade cost: there is no documented migration step, no version-pinned lockfile guarantee for downstream consumers, and uv.lock exists in the repository but the README does not describe how to use it for a reproducible deployment. If you deploy to a fleet of boards, budget for a rebuild-and-retest cycle per release rather than an in-place update.

## Conclusion

Adopt py-xiaozhi if you already use Xiaozhi services and want that stack running on a Raspberry Pi, an RDK board or a desktop instead of ESP32 hardware, and if you are comfortable pinning Python 3.10 to 3.12 and reinstalling dependencies after every update. Skip it if you need a standalone assistant with no cloud endpoint, if you cannot supply 4GB of RAM and a 16kHz audio device, or if your board is not in the supported list. Before committing, verify that your target platform is covered by the sys_platform markers in pyproject.toml, that your microphone actually opens at 16kHz, and that the main branch dependency set still resolves under your Python version, since the README states dependencies must be reinstalled manually after each update.

## FAQ

### What is Xiaozhi AI?

In this material, Xiaozhi refers to the assistant ecosystem around the xiaozhi-esp32 firmware project, which py-xiaozhi evolved from. py-xiaozhi is the Python client side of that stack, handling voice streaming, vision tasks and IoT control.

### Can I build an AI assistant using Python with py-xiaozhi?

Yes. py-xiaozhi is a Python async framework requiring Python 3.10 to 3.12, with a main.py entry point and an MCP plugin directory at examples/mcp_plugins/ for adding your own tools. You still need a network connection for the AI services, since the README lists that as a basic requirement.

### What is the difference between py-xiaozhi and xiaozhi-esp32?

xiaozhi-esp32 is firmware for ESP32 boards, while py-xiaozhi is a Python client that runs on Windows, macOS, Linux and ARM boards such as the Raspberry Pi, RDK and Jetson Nano. py-xiaozhi states it evolved from the firmware project and is used as an upstream dependency by D-Robotics.

### Does py-xiaozhi work without an internet connection?

The README lists a stable network connection as a basic requirement for AI services and online features. The only offline capability it documents is Sherpa-ONNX wake word detection, which requires downloading speech recognition models separately.

### How do I install py-xiaozhi on Linux?

The pyproject.toml and requirements.txt both support installation through uv or pip, and the GUI mode is an optional extra enabled with uv sync --extra gui. On Linux the dependency set includes gpiozero and lgpio for GPIO mode, and the README recommends at least 4GB of RAM and a 16kHz audio device.

## Sources

- [huangjunsen0406/py-xiaozhi on GitHub](https://github.com/huangjunsen0406/py-xiaozhi)
- [License: MIT](https://github.com/huangjunsen0406/py-xiaozhi/blob/main/LICENSE)
- [Project website](https://huangjunsen0406.github.io/py-xiaozhi/)
- [README](https://github.com/huangjunsen0406/py-xiaozhi/blob/main/README.md)
- [Releases](https://github.com/huangjunsen0406/py-xiaozhi/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/huangjunsen0406-py-xiaozhi
