XiaoZhi ESP32: an MCP-based voice chatbot for hardware boards
Project brief: An MCP-based chatbot | MCP . The previous 157-variant baseline was validated on ESP-IDF v6.0.1; the current matrix contains 171 variants, of which 170 support IDF 6.0.x and the ESP32-S31 variant requires IDF 6.1 or later.
At a glance
- What is it?
- XiaoZhi ESP32 turns ESP32-family boards into voice assistants that talk to large models and control devices over MCP. The firmware targets ESP-IDF 6.0.2, but board support and the cloud backend are two separate decisions.
- Who is it for?
- Adopt XiaoZhi ESP32 if you have a supported ESP32-S3 or ESP32-P4 board, are comfortable with ESP-IDF 6.0.2, and want device-side MCP control plus a choice between WebSocket and MQTT + UDP transports.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What XiaoZhi ESP32 solves, and for whom
Building a voice assistant on a microcontroller normally means assembling four separate things: a wake-word engine, an audio codec pipeline, a transport to a cloud model, and some way for the model to act on the physical device. XiaoZhi ESP32 packages all four into one firmware tree. The README describes it as a "voice interaction entry" that uses large models such as Qwen and DeepSeek and reaches multi-terminal control through the MCP protocol.
The audience is narrower than the feature list suggests. This is firmware for people who own or are willing to buy a supported board. The repository lists 138 board directories and 171 release variants, spanning Espressif's own ESP32-S3-BOX-3, M5Stack CoreS3 and AtomS3R, Waveshare's ESP32-S3-Touch-AMOLED-1.8, LILYGO's T-Circle-S3, Seeed's SenseCAP Watcher, and smaller community designs such as the XiaGe Mini C3 and the ESP-HI robot dog. If your hardware is on that list, the project has already solved the pin mapping, the audio routing and the display driver for you. If it is not, you are writing a custom board definition before you hear anything.
There is also a beginner path that skips the toolchain entirely. The README points to a firmware flashing guide for people who do not want to set up a development environment, and notes that the firmware connects to the official xiaozhi.me server by default, where personal users can register an account and use the Qwen real-time model for free. That is the fastest way to evaluate whether the interaction model suits you, and it costs nothing but a flash.
How the audio pipeline and MCP control actually fit together
The data flow starts on-device. Offline wake-word detection runs through ESP-SR, and the README states wake words are customizable, so the wake stage does not require a network round trip. Once awake, audio is encoded as Opus and streamed out. The README describes two communication transports: WebSocket, documented in docs/websocket.md, and MQTT plus UDP, documented in docs/mqtt-udp.md. That split matters. WebSocket is the simpler of the two and keeps one connection open. MQTT plus UDP separates signalling from media, which is the shape you want when the network is unreliable, at the cost of running two protocols and validating packets on both. The release notes for this cycle mention MQTT and UDP packet validation being hardened, which suggests that surface has seen real bugs.
On the server side, two pipeline shapes are supported: conventional streaming ASR into an LLM into TTS, and Realtime end-to-end voice models. The README notes that boards with AEC-capable hardware support realtime full-duplex interaction. That is a hardware-gated feature, not a firmware setting. A board without acoustic echo cancellation will fall back to the half-duplex pattern where the device stops listening while it speaks.
MCP is the part that distinguishes this from a plain voice assistant. Device-side MCP exposes local capabilities such as Speaker, LED, Servo and GPIO as tools the model can call. Cloud-side MCP extends the model outward to smart home control, PC desktop operation, knowledge search and email. The practical consequence is that a voice command can end in a GPIO toggle rather than a spoken answer, and the tool surface is defined by the firmware rather than hardcoded into the model prompt. Speaker recognition via 3D Speaker is also present, so the device can identify who is speaking.
Setting up the toolchain and a first flash
The mainline targets ESP-IDF v6.0 or later, with v6.0.2 named as the preferred stable SDK. ESP-IDF v5.5.2 is retained only for legacy board compatibility, so a new setup should not start there. The README recommends Cursor or VSCode with the ESP-IDF plugin, and notes that Linux compiles faster and produces fewer driver issues than Windows.
After installing the plugin and selecting ESP-IDF v6.0.2, the build follows the project's standard flow. The repository carries per-chip defaults files at the top level, including sdkconfig.defaults.esp32s3 and sdkconfig.defaults.esp32p4, which the build picks up according to the target you set. The set-target step selects the chip family and the matching sdkconfig.defaults file. The build produces the firmware image, and the flash step writes it and opens the serial console, where you should see the boot log and, after provisioning, the wake-word engine starting. The README does not spell out the individual idf.py invocations, so follow the ESP-IDF plugin's own build and flash commands rather than a snippet copied from here. If you are on a board other than the generic S3 target, the board directory under the repository's board tree is what determines pin assignments and peripherals, so confirm yours exists before assuming a clean build means a working device.
Network setup comes next. The README lists Wi-Fi provisioning through a hotspot or BluFi, plus wired Ethernet, USB RNDIS, and ML307/EC801E or NT26 Cat.1 4G as alternatives, with supported boards able to switch between Wi-Fi and 4G. Whichever path you use, the device needs to reach a server. Out of the box that is xiaozhi.me.
Where XiaoZhi ESP32 is the wrong choice
The largest constraint is the board matrix itself. 171 variants sounds generous until you look for your own hardware and find nothing. A custom board is possible, and docs/custom-board.md is the document for it, but that is a firmware porting task, not a configuration change. Budget for it accordingly.
The second constraint is the SDK version split. The README states that the previous 157-variant baseline was validated on ESP-IDF v6.0.1, and that the current matrix contains 171 variants, of which 170 support IDF 6.0.x while the ESP32-S31 variant requires IDF 6.1 or later. So one board in the set cannot be built on the preferred SDK. If you are pinning to v6.0.2 for the rest of your fleet, that board is a separate maintenance track. The migration guide at docs/esp-idf-6-migration.md is where the compatibility and board-validation details live, and the README defers to it rather than summarizing it.
Third, the firmware is not the whole product. The README describes the device connecting to the official server by default and points at a separate server project for self-hosting. Nothing in the firmware tree replaces the ASR, LLM and TTS backend. If your requirement is a fully offline assistant, this architecture does not meet it: wake-word detection is offline, but the conversation is not.
Finally, the repository is a large, fast-moving C++ codebase with a documented Google C++ code style requirement for contributions. The last push was on 2026-08-06, and the release cadence this year shows v2.2.6 in April, v2.4.0 in July and v2.4.2 in August. That pace is good for users and expensive for anyone maintaining a fork.
How XiaoZhi ESP32 differs from ESPHome voice assistants
The closest alternative in practice is the ESPHome voice assistant stack, which people also search for alongside this project. The difference is architectural rather than cosmetic.
ESPHome is a YAML-driven firmware generator. You describe components in configuration, the toolchain produces firmware, and the device integrates with Home Assistant as its control plane. Voice handling is assembled from those components. XiaoZhi ESP32 takes the opposite approach: it is a C++ application with a fixed audio pipeline and a protocol-level abstraction, MCP, for exposing device capabilities as tools. Control does not route through a home automation hub by default; it routes through the model, which calls MCP tools that the firmware declares.
That gives XiaoZhi ESP32 a shorter path from voice to physical action and a much wider set of possible backends, including the Qwen real-time model on xiaozhi.me. It also means you inherit the project's board support, its SDK version policy and its transport choices. ESPHome's advantage is that the configuration is declarative and the Home Assistant integration is the product, so upgrades are more predictable. If your goal is a Home Assistant satellite with local control, ESPHome is the more direct route. If your goal is a standalone voice device where the model can drive servos, LEDs and GPIO directly, and you are willing to track ESP-IDF 6, XiaoZhi ESP32 is the one that has already built that path.
Licence, maintenance and what an upgrade costs
The project is MIT-licensed, which is permissive: it allows commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is the standard reading of MIT and not legal advice; if you ship a product built on this firmware, have your own counsel confirm notice placement in your distribution.
The practical licence question is not the firmware but the services it talks to. The firmware is MIT. The default backend at xiaozhi.me is a separate service with its own terms, and the README's note that personal users can register and use the Qwen real-time model for free says nothing about commercial usage terms. Treat the firmware licence and the service terms as two separate reviews.
Upgrade cost is dominated by the ESP-IDF version, not by the application code. The README records that MQTT and BluFi cryptographic code migrated to PSA Crypto, and that IDF 6 component splits and third-party dependency compatibility had to be addressed. Those are exactly the changes that break a fork. If you maintain a custom board, expect to re-validate it against each IDF 6.x point release rather than treating the build as stable across them. The per-chip sdkconfig.defaults files at the repository root are the surface where those differences land.
The last push was on 2026-08-06, matching the v2.4.2 release. The repository is not archived.
Editorial conclusion
Adopt XiaoZhi ESP32 if you have a supported ESP32-S3 or ESP32-P4 board, are comfortable with ESP-IDF 6.0.2, and want device-side MCP control plus a choice between WebSocket and MQTT + UDP transports. Do not adopt it if your board is not in the 138 board directories, if you need ESP-IDF 5.5 as your mainline, or if you expect the repository to supply the language model, ASR and TTS backend, because the firmware connects to xiaozhi.me by default and a self-hosted server is a separate component. Before flashing, verify three things: your exact board directory exists, your board is not the ESP32-S31 variant that requires IDF 6.1 or later, and the wake word and language you need are covered by the 39 interface languages. The repository's own migration guide at docs/esp-idf-6-migration.md is the document to read first, not the README.
Frequently asked questions
What can XiaoZhi AI do?
It provides offline voice wake-up through ESP-SR, streams Opus audio to large models such as Qwen and DeepSeek, and supports both streaming ASR plus LLM plus TTS pipelines and Realtime end-to-end voice models. Device-side MCP exposes Speaker, LED, Servo and GPIO as tools the model can call, and cloud-side MCP extends it to smart home control, PC desktop operation, knowledge search and email.
Can ESP32 run AI?
XiaoZhi ESP32 runs the wake-word model on-device through ESP-SR and handles audio capture, encoding and display locally, while the language model itself runs on a server that the device reaches over WebSocket or MQTT plus UDP. The README also lists speaker recognition and camera vision input on supported boards.
Is there a XiaoZhi ESP32 alternative?
The ESPHome voice assistant stack is the closest alternative, and it works the other way around: ESPHome generates firmware from YAML configuration and integrates with Home Assistant as the control plane, while XiaoZhi ESP32 is a C++ application that exposes device capabilities as MCP tools for the model to call directly.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/78-xiaozhi-esp32)