Xiaozhi ESP32 Server Java: a two-process backend for ESP32 voice hardware
小智ESP32的Java企业级管理平台,提供设备监控、音色定制、角色切换和对话记录管理的前后端及服务端一体化解决方案
At a glance
- What is it?
- The Java rewrite of the Xiaozhi ESP32 server splits a Spring Boot admin platform from a dialogue worker that handles WebSocket and MQTT audio. Here is what the repository documents, how to bring it up, and where it stops.
- Who is it for?
- Adopt it if you already have Xiaozhi ESP32 hardware and want a Java stack you can read and extend, with MySQL 8.0 and Redis 7 already in your infrastructure. Do not adopt it if you expect a fully open feature set: the README's comparison image marks part of the functionality as commercial, and the documentation does not say which parts.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the Java port adds over the original Xiaozhi server
The upstream Xiaozhi ESP32 project provides firmware for ESP32 boards that listen, stream audio to a server, and speak back. xiaozhi-esp32-server-java is a server implementation for that firmware, written in Java, and its stated audience is people who already own ESP32 hardware and need a management platform around it. The README lists four groups: owners of ESP32 hardware, teams that want enterprise stability and scaling, individual developers who want a fast path to a working setup, and deployments expecting many concurrent devices.
The distinguishing choice is that this is a Java and Spring Boot codebase rather than a Python one. The stack table names Spring Boot, Spring MVC, MyBatis-Plus, Flyway and WebSocket on the backend, Vue.js and Ant Design on the frontend, MySQL 8.0 and Redis 7 underneath. For a team whose operations already run JVM services, that removes a runtime from the picture. Flyway handles schema creation, so the database does not need hand-written DDL before the first start.
The two-process split: xiaozhi-server on 8091 and xiaozhi-dialogue on 8092
The README describes a multi-module, dual-process architecture. Two independent processes share the same MySQL and Redis instances and can be deployed and scaled separately. xiaozhi-server listens on port 8091 and serves the admin side: REST API, user, device and role management, and OTA upgrade delivery. xiaozhi-dialogue listens on port 8092 and carries the real-time work: WebSocket and MQTT audio streams and the AI conversation pipeline.
The scaling story is the interesting part. The README states that dialogue supports horizontal expansion, that new instances register themselves with server automatically, and that load is balanced through device OTA. Put plainly, the admin process acts as the registry and the firmware is told where to connect. That is a real design decision rather than a deployment detail, because it means the admin process is on the critical path for every new device session, and the repository does not document what a dialogue instance does when server is unreachable.
On the dialogue side, the pipeline covers speech recognition, a large language model, and speech synthesis. The README lists Vosk, FunASR, Alibaba Cloud, Tencent Cloud and iFlytek for STT, and sherpa-onnx locally plus Volcano Engine, Alibaba Cloud and Edge TTS for TTS. LLM providers named are OpenAI, Zhipu, iFlytek Spark, Ollama, Dify and Coze. MCP tool protocol and Function Call are listed as extension mechanisms, alongside a RAG knowledge base and voice cloning.
Installing from source and running a first device session
The README's quick start assumes Linux and a source checkout. The first command clones the repository and enters it. The second is the one people miss: models/ and lib/ are not tracked in Git, so the download script has to run before anything else, and it fetches both the model files and the native libraries.
git clone https://github.com/joey-zhou/xiaozhi-esp32-server-java
cd xiaozhi-esp32-server-java
./scripts/download_models.sh # download models and native libraries (required on first run)
bin/all.sh start # build and start both server and dialogue
bin/all.sh status # check statusIf you plan to use a third-party STT or TTS provider instead of the local models, the README gives a lighter option: ./scripts/download_base.sh downloads only the VAD model and the native libraries.
For containers, the repository ships a docker-compose.yml with four services: mysql, node, server and dialogue. The node service builds from ./web and receives API_URL=http://server:8091. The server service maps port 8091 and receives SPRING_DATASOURCE_URL, SPRING_DATASOURCE_USERNAME and SPRING_DATASOURCE_PASSWORD, with Redis pointed at host redis. The compose file also builds mysql from Dockerfile-mysql on port 3306 and gates server and node on a mysqladmin healthcheck using user xiaozhi.
docker compose up -dOnce both processes are up, the README says xiaozhi-server prints an OTA address and xiaozhi-dialogue prints a WebSocket connection address. Those are what you feed into the firmware build, following the firmware document, so the device knows where to fetch configuration and where to stream audio.
Where the setup breaks: models, ports and the closed portion
The failure mode most likely to bite first is a checkout without models. Because models/ and lib/ live outside the repository, a plain clone followed by bin/all.sh start will not have what the local STT and TTS paths need. The README is explicit that the download script is required on first deployment, and that the base script is enough only when you are relying on third-party speech services.
Port collisions are the second. The compose file binds 3306, 8084, 8091 and 8092, and its own comment on the MySQL mapping tells you to change it to another unused port if 3306 is taken. The same logic applies to the other three, and the README does not describe a configuration path for moving 8091 or 8092.
The third is scope. The README carries a feature comparison image labelled open source versus commercial, with a note that some functionality is not open sourced and that commercial interest should go through the contact shown. The README does not enumerate which capabilities sit on the closed side, so anyone planning around a specific feature should confirm it exists in the open repository before designing around it. That is a genuine planning risk, not a footnote.
Finally, the resource envelope. The compose file caps the server service at 2 CPUs and 2G of memory, with reservations of 0.5 CPU and 512M. The README's own benchmark table, run on an 8-core 8G Tencent Cloud instance with 100 concurrent devices, reports memory between 1.8G idle and 1.96G peak and CPU reaching 80%. Those are the project's published numbers for its own test tool, not an independent measurement, but they do suggest that a 2G cap is tight once concurrency rises.
How it compares with the Python Xiaozhi server
The obvious alternative is the Python server that the Xiaozhi ESP32 ecosystem grew up around. The practical difference is not capability lists, it is what your team can operate. A Python deployment typically means managing a virtual environment, Python-level dependencies and a separate process supervisor, and the audio and model libraries are Python packages. This project replaces that with a Maven build, a Spring Boot jar and Flyway migrations, at the cost of a heavier runtime and a JVM memory floor.
There is a second, less obvious difference. The Python ecosystem's speech tooling is broader and changes faster, so if you want to experiment with a new STT or TTS model the week it appears, the Python side will usually have it first. This project's provider list is enumerated in the README and is a fixed set: Vosk, FunASR, Alibaba Cloud, Tencent Cloud, iFlytek, sherpa-onnx, Volcano Engine and Edge TTS. Adding a provider means writing Java against the project's interfaces.
If your constraint is a small single-board deployment with no JVM, or you want the largest pool of community examples to copy from, the Python server is the better fit. If your constraint is that everything you run in production is Java, Spring Boot and MySQL, this is the version that fits.
Licence, maintenance and what an upgrade costs
The repository is MIT licensed, and the LICENSE file sits at the top level. MIT is permissive: you can use, modify and redistribute the code, including commercially, provided the copyright notice and permission notice are retained. The README's disclaimer is separate from the licence and worth reading on its own terms. It states that the project supplies technical implementation code and no media content, that users are responsible for holding the rights to anything they play through it, and that example resources come from the network or user submissions and are for demonstration and testing. That is a content-liability statement aimed at whoever deploys the server, and it is the kind of thing that should go to whoever handles your legal review rather than being treated as boilerplate.
The commercial feature split noted above is a licensing-adjacent question too: the README does not state a separate licence for any closed component, because it does not describe those components at all.
On maintenance, the last push to the default branch was on 2026-09-03, and the most recent tagged release in the list is v5.1.0 from 2026-04-21, preceded by v5.0.0 on 2026-04-14 and v4.1.0 on 2026-02-21. The gap between the last push and the last release is worth noting if you depend on tagged artifacts rather than the default branch. The repository is not archived. There is a CHANGELOG.md at the top level, so upgrade notes live there. For upgrade cost specifically, the thing to check before pulling a new version is db/ and the Flyway migrations, since schema changes arrive through that path and the README does not document a rollback procedure.
Editorial conclusion
Adopt it if you already have Xiaozhi ESP32 hardware and want a Java stack you can read and extend, with MySQL 8.0 and Redis 7 already in your infrastructure. Do not adopt it if you expect a fully open feature set: the README's comparison image marks part of the functionality as commercial, and the documentation does not say which parts. Before committing, run ./scripts/download_models.sh and confirm the models/ and lib/ directories populate, then check bin/all.sh status reports both xiaozhi-server on 8091 and xiaozhi-dialogue on 8092.
Frequently asked questions
What is xiaozhi-esp32-server-java used for?
It is a Java server and admin platform for ESP32 hardware running the Xiaozhi firmware. It provides device monitoring, voice customization, role switching, conversation history, OTA upgrades and the real-time audio pipeline that connects the device to STT, an LLM and TTS.
How do I install xiaozhi-esp32-server-java?
Clone the repository, run ./scripts/download_models.sh to fetch the models and native libraries that are not stored in Git, then run bin/all.sh start to build and launch both processes. bin/all.sh status reports whether they came up.
Which ports does xiaozhi-esp32-server-java use?
The README assigns port 8091 to xiaozhi-server for the admin REST API and OTA, and 8092 to xiaozhi-dialogue for WebSocket and MQTT audio. The bundled docker-compose.yml also publishes 3306 for MySQL and 8084 for the Vue frontend.
Does xiaozhi-esp32-server-java include all features in the open source repository?
No. The README shows an open source versus commercial feature comparison and states that some functionality is not open sourced, without listing which parts. Confirm any specific capability in the open repository before planning around it.
What database and cache does xiaozhi-esp32-server-java require?
MySQL 8.0 and Redis 7. Both processes share the same MySQL and Redis instances, and Flyway creates the schema automatically on startup.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/joey-zhou-xiaozhi-esp32-server-java)