Open-source project
xinnan-tech/xiaozhi-esp32-server avatar
xinnan-tech/xiaozhi-esp32-server

xiaozhi-esp32-server: a self-hosted backend for xiaozhi-esp32 voice devices

xiaozhi-esp32 ESP32 Backend service for xiaozhi-esp32, helps you quickly build an ESP32 device control server.

10,681 stars3,641 forksJavaScriptMIT

At a glance

What is it?
The xiaozhi-esp32-server project is the Python, Java and Vue backend that the open xiaozhi-esp32 firmware talks to. It is aimed at people who already own the hardware and want their own server instead of someone else's, and the README is explicit that it should not run in production.
Who is it for?
Adopt xiaozhi-esp32-server if you own an ESP32 device flashed with xiaozhi-esp32 firmware and want the conversation pipeline, memory and tool calls under your own control. Do not adopt it as a production voice platform: the README states the project is incomplete and has not passed a network security assessment.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 9 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What xiaozhi-esp32-server actually replaces

The xiaozhi-esp32 firmware is an open smart-hardware project. Out of the box its devices are usually pointed at a backend someone else runs, and the README frames the target audience precisely: people who have bought ESP32 hardware, have already connected it to a hosted backend, and now want to stand up their own. That is the whole pitch. This is not a general voice assistant framework, and it is not useful without a device on the other end.

The server owns the conversation chain. Speech recognition, a language model, speech synthesis, voice activity detection, optional voiceprint identification, memory and tool invocation all live behind it, and the device is mostly a microphone, a speaker and a network stack. The README lists MQTT plus UDP, WebSocket and HTTP as transports, and describes a console for managing users, agents and MCP commands pushed down to devices.

So the practical question this project answers is: can I keep the audio, the API keys and the conversation history on hardware I control? Yes, with the caveat that the project itself tells you not to expose it carelessly.

Two deployment shapes, and the database decision behind them

The README splits installation into a minimal mode and a full-module mode, and the split is really about where state lives. Minimal mode keeps data in configuration files and supports single-agent conversation; the README sizes it at 2 cores and 4 GB if you run FunASR locally, or 2 cores and 2 GB if every component is a remote API. Full-module mode adds multi-user and multi-agent management plus the console UI, stores data in a database, and the README asks for 4 cores and 8 GB with local FunASR or 2 cores and 4 GB with all APIs.

That is a meaningful fork. Choosing minimal mode means no database to back up, but it also means the multi-user console is out of reach. Choosing full mode means a database becomes part of your operational surface. The README does not document a migration path between the two, so the choice is better made before you accumulate data.

The README also distinguishes an all-free configuration (FunASR locally, glm-4-flash, EdgeTTS) from a streaming configuration that swaps in XunfeiStreamASR, qwen-flash and HuoshanDoubleStreamTTS. It states that streaming support arrived in version 0.5.2 and improved response time by roughly 2.5 seconds compared with earlier versions. Treat that as a claim from the project, not a measured result, and note that the free path trades latency for zero API spend.

Installing xiaozhi-esp32-server with Docker and running a first check

The README points at two deployment documents rather than inlining commands: docs/Deployment.md for the minimal install and docs/Deployment_all.md for the full-module install, each with a Docker route and a source route. The repository root carries Dockerfile-server, Dockerfile-server-base, Dockerfile-web and a docker-setup.sh script, which is consistent with the Docker path being the intended entry point. Because the README does not reproduce the compose file, the exact service names and ports are not something this article can quote; read the deployment document before typing anything.

What the repository does give you is a test entry point. The Makefile defines the fast test target and fans it out across the three components:

bash
make test-fast

That runs pytest in main/xiaozhi-server, mvn test in main/manager-api and npm test in main/manager-web, in that order. If you are deploying from source rather than from a prebuilt image, running it first tells you whether your Python, Java and Node toolchains are complete before you spend time on configuration.

For a first real use, the README describes two verification tools shipped in the tree. The audio interaction tool lives at main/digital-human/index.html and is started from that directory:

bash
cd main/digital-human
python start.py

The README says to then open http://127.0.0.1:8006/index.html, where you can check that audio is played and received correctly. The second tool measures component latency:

bash
cd main/xiaozhi-server
python performance_tester.py

According to the README, the performance tester exercises ASR, LLM, VLLM and TTS, and only tests models for which you have configured keys. That detail matters: a run with no keys configured will not tell you the pipeline is broken, only that it had nothing to call.

The README also lists a public test platform with an OTA endpoint and a WebSocket endpoint, and notes it is cleared daily and sized for six concurrent sessions. It is a way to see the thing working, not a place to keep anything.

The README's own warning is the main limitation

The warning section is unusually direct for an open source README. It states that the project's functionality is incomplete and that it has not passed a network security assessment, and it asks people not to use it in production. If you deploy it on a public network for learning, it says to apply the necessary protections. Take that at face value. A server that accepts device connections, holds third-party API keys for speech and language services, and exposes a management console is exactly the kind of surface you do not want reachable from the open internet without a reverse proxy and authentication in front of it.

The second warning concerns money. The project states it has no commercial relationship with any third-party API provider and does not guarantee their service quality or the safety of funds, does not host account keys, does not participate in payment flows and does not bear losses from topped-up balances. In practice this means every paid component you configure is a direct relationship between you and that provider, with your key sitting in your own configuration.

The third limitation is scope. This is a backend for one firmware family. If your device does not speak the xiaozhi protocol, none of the voice pipeline, memory or MCP tooling here applies to you.

How it differs from running the upstream xiaozhi-esp32 stack as shipped

The obvious alternative is the xiaozhi-esp32 firmware repository's own default arrangement, where devices connect to a hosted backend rather than one you operate. The difference is not features, it is who holds the keys and the data. With the hosted route you get a working device with no server to run, no database, no Docker host and no API accounts to manage, and you accept that your audio and conversation history pass through someone else's infrastructure under their terms.

With xiaozhi-esp32-server you take on the operational work the README describes: choosing minimal or full mode, provisioning the CPU and memory the README specifies, wiring ASR, LLM, VLLM and TTS providers, and securing the console. In exchange, the API keys are yours, the memory store is yours, and the RAGFlow knowledge base or MCP tools you attach are ones you chose.

There is a middle option worth naming: the README's all-free configuration, which pairs local FunASR with glm-4-flash and EdgeTTS. It keeps the self-hosted property while removing per-call cost, at the price of running speech recognition on your own CPU and accepting the latency that comes with it. The streaming configuration moves in the opposite direction, spending on hosted streaming services for faster turnarounds.

Maintenance cost, licensing and what the release cadence implies

The repository is not archived. Its last push was on 2026-07-24, which is the same date as the v0.9.6 release, and the two prior releases landed on 2026-06-30 and 2026-06-03. That is a roughly monthly release rhythm across the visible window, and the version numbers sit below 1.0, which is consistent with the README's statement that functionality is incomplete. Plan for upgrades rather than a frozen install.

Upgrade cost is uneven across the three components. The README's full-module deployment document includes a source deployment guide with automatic updates, which suggests the maintainers expect source deployments to be refreshed rather than rebuilt. The Makefile shows the test suite spans Python, Java and Node, so a source upgrade can require all three toolchains to be present even if you only care about the server. Docker users avoid that, at the cost of depending on the published images.

The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is a permissive licence, and it says nothing about the third-party services you connect to; their terms govern your use of them. This is not legal advice, and the README's own warning about provider relationships is the part worth reading twice.

Editorial conclusion

Adopt xiaozhi-esp32-server if you own an ESP32 device flashed with xiaozhi-esp32 firmware and want the conversation pipeline, memory and tool calls under your own control. Do not adopt it as a production voice platform: the README states the project is incomplete and has not passed a network security assessment. Before committing, verify that your device firmware speaks the same protocol as your chosen deployment mode, and check whether you need the database-backed multi-user console or the file-based minimal install.

Frequently asked questions

Does xiaozhi-esp32-server work without an ESP32 device?

The README states the project must be used together with ESP32 hardware and is aimed at people who already own a device and have connected it to a backend before. The audio interaction test tool at main/digital-human/index.html lets you check audio playback and reception without a device, but it is a diagnostic, not a substitute for one.

Should I choose the minimal install or the full-module install of xiaozhi-esp32-server?

The README maps minimal mode to smart conversation and single-agent management with data stored in configuration files and no database, and full-module mode to multi-user and multi-agent management with the console UI and data in a database. Minimal mode is sized at 2 cores and 4 GB with local FunASR, while full mode is sized at 4 cores and 8 GB.

Is xiaozhi-esp32-server safe to expose on the public internet?

The README explicitly asks users not to deploy it in production, noting the functionality is incomplete and the project has not passed a network security assessment. It advises applying necessary protections if you deploy it on a public network for learning purposes.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xinnan-tech-xiaozhi-esp32-server.svg)](https://hysenlabs.com/projects/xinnan-tech-xiaozhi-esp32-server)