Self-hosted service
xerrors/Yuxi avatar
xerrors/Yuxi

Yuxi: a self-hosted multi-tenant knowledge agent platform that ties RAG, Neo4j and LangGraph together

可私有部署的多租户知识智能体平台:统一 RAG、知识图谱、多智能体、MCP/Skills、沙盒与权限管理。Self-hosted knowledge agent platform for RAG, knowledge graphs and multi-agent workflows.

7,235 stars1,149 forksPythonMIT

At a glance

What is it?
Yuxi bundles retrieval, knowledge graph extraction, LangGraph agent orchestration, MCP/Skills and sandboxed file execution behind one multi-tenant workspace. It is aimed at teams that want to run the whole stack themselves, and the Docker Compose install is where the real cost starts.
Who is it for?
Adopt Yuxi if you need retrieval, graph extraction, agent orchestration and per-department permissions inside your own network, and you can carry roughly ten containers plus a model API key. Do not adopt it if you only need a document chatbot, or if you cannot schedule a maintenance window: the README states that upgrading from v0.7.1 or v0.7.2 to the current release cannot be done with a plain docker compose up.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The team that owns the data but not the model

Yuxi targets a specific position in the stack. The README states it is for teams that need to control their own data, models and permissions. That sentence rules out two common buyers. It is not for someone who wants a hosted chatbot with a monthly seat price, and it is not for a solo developer who wants a library import. The unit of adoption is an organisation with departments, because permissions are modelled per user and per department, and because knowledge bases, agents, Skills and models are all scoped inside that permission system.

The problem it addresses is fragmentation. A team that wants a document question answering service today usually assembles a vector store, a parser, an orchestration framework, a file sandbox and an admin console, then writes the glue and the access control. Yuxi ships those as one Compose project. The trade-off is that you inherit its choices: PostgreSQL, Redis, MinIO, Milvus and Neo4j all arrive as required services, and the README lists them as the storage layer rather than optional add-ons.

How the pieces connect: Milvus chunks in, Neo4j edges out

The data flow visible in the README runs in one direction. Documents are uploaded, parsed by MinerU, PaddleX or RapidOCR into text, images and tables, then split into chunks and embedded into Milvus. That is the retrieval layer.

The graph layer is built on top of the chunks rather than beside them. The README says entities and relations are extracted from the document chunks of a Milvus knowledge base and written into Neo4j, where they participate in retrieval. So the graph is derived state: if you rebuild the vector index, expect to rebuild the graph too.

Above both sits the agent layer, which the README describes as LangGraph multi-agent orchestration with SubAgents, Skills, MCP and Tools. Long tasks run asynchronously through an ARQ worker; the interface shows step plans, subtask status and tool call logs, and pauses on a confirmation card when an action modifies files or calls a high-risk external interface. File-producing tasks land in a sandbox with an isolated filesystem, which the Compose file configures through SANDBOX_PROVIDER, SANDBOX_PROVISIONER_URL and a virtual path prefix of /home/gem/user-data.

Installing Yuxi from the v0.7.3 tag

The README pins the current repository configuration to v0.7.3 and gives a shallow clone of that tag as the first step. Do this on the host that will run the stack, not inside another container.

bash
git clone --branch v0.7.3 --depth 1 https://github.com/xerrors/Yuxi.git
cd Yuxi
./scripts/init.sh

On Windows PowerShell the README gives .\scripts\init.ps1 instead. The script creates .env, reads a SiliconFlow API key, and derives separate secrets for JWT, API key derivation and the sandbox provisioner. You can also copy .env.template by hand, but the Compose file marks two variables as mandatory: API_KEY_DERIVATION_SECRET and SANDBOX_PROVISIONER_TOKEN both fail the compose run with a message telling you to set them in .env or rerun the init script.

bash
docker compose up --build -d
docker compose ps
curl --fail http://localhost:5050/api/system/ready

The readiness endpoint is the honest signal here. When it returns status ready, open http://localhost:5173, create the super administrator from the page prompt and log in. The API documentation is served at http://localhost:5050/docs. If the curl call fails, the API container is not up yet; the Makefile uses the same idea in its reset target, polling docker compose exec -T api true before seeding users.

First real use: create a knowledge base, upload a PDF, wait for parsing and chunking to finish, then open the retrieval test workbench and run a query. The README states you can see the embedding pre-filter scores, the hybrid retrieval results and the reranked scores side by side. That screen is the fastest way to find out whether your embedding and rerank configuration is worth anything before you wire an agent to it.

The upgrade path is not docker compose up

This is the sharpest constraint in the README, and it is worth quoting rather than paraphrasing: upgrading from v0.7.1 or v0.7.2 to the current version cannot be done by running docker compose up directly. The README directs you to docs/advanced/deployment.md and says to complete backup and migration inside a downtime window.

The reason is visible in the Compose file. API_KEY_DERIVATION_SECRET carries a [v0.7.2+] marker and SANDBOX_PROVISIONER_TOKEN carries [v0.7.1+]. Both are required, both are generated by the init script, and neither exists in an older .env. An operator who pulls a newer image without reading the deployment document gets a container that refuses to start, with the error text pointing back at scripts/init.sh.

There is a second failure mode in the Makefile. Its reset target deletes docker/volumes and refuses to run at all if YUXI_STATE_DIR is set, either in the environment or in .env. That guard exists because the state directory can live outside the Compose volume tree. Treat make reset as a local development command only.

Where Yuxi is the wrong tool

If your requirement is a retrieval API inside an existing application, Yuxi is heavier than the job. You would be running PostgreSQL, Redis, MinIO, Milvus, Neo4j, the API, the worker, the web frontend, the sandbox provisioner and at least one document parsing service to answer questions over a few hundred PDFs. A vector store plus an embedding call gets you most of the way there with a fraction of the operational surface.

The knowledge graph is also not free. Extraction runs over document chunks during parsing, so graph quality is bounded by chunk quality, and the README does not describe a rollback path if an extraction pass produces a bad graph. The material is silent on incremental rebuild and on how extraction cost scales with corpus size.

Finally, the documentation is Chinese-first. The README links an English version at README.en.md, and the project homepage and quick-start pages are the authoritative sources, but the Compose comments and the Makefile comments are written in Chinese. A team without a Chinese reader will be slower on the deployment document that the upgrade path depends on.

What the sandbox and MCP layer actually buys you

The sandbox is the part that separates Yuxi from a chat interface with citations. Tasks that produce files run in an isolated filesystem, and the results are previewed and downloaded from the conversation. The Compose defaults show the shape of it: SANDBOX_EXEC_TIMEOUT_SECONDS defaults to 180, SANDBOX_MAX_OUTPUT_BYTES to 262144, and SANDBOX_PROVISIONER_DELETE_TIMEOUT_SECONDS to 120. TASKER_DEFAULT_TIMEOUT_SECONDS is far larger at 21600, which fits the long-running asynchronous worker model described in the README.

Those numbers are the tuning surface. A team running data analysis tasks will hit the 180 second execution ceiling before it hits the task ceiling, and raising it means accepting longer-lived sandbox processes. The README does not publish guidance on which values are safe, so treat the defaults as a starting point and measure against your own task mix.

MCP and Skills are the extension path for everything the platform does not ship. The README lists them alongside SubAgents and Tools as the multi-agent extension ecosystem, and permissions for Skills are scoped through the same user and department model as knowledge bases. That consistency is the design argument for adopting the whole platform instead of assembling parts.

Licence and the cost of staying current

Yuxi is MIT licensed. That permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. It says nothing about the services the Compose file pulls in: Milvus, Neo4j, PostgreSQL, Redis and MinIO each carry their own licence, and Neo4j in particular has editions with different terms. Check those separately before a production rollout. This is not legal advice.

The maintenance picture is current rather than frozen. The repository is not archived, and the last push was on 2026-09-10. Releases are frequent: v0.7.2 on 2026-09-02, v0.7.2.beta1 on 2026-08-23 and v0.7.1 on 2026-07-17. That cadence is the upgrade cost. A team that tracks every release should budget for reading docs/advanced/deployment.md each time, because the last two minor versions each introduced a required environment variable. A team that pins a tag and upgrades twice a year will spend a full downtime window per upgrade instead. Either way, the .env file is the artifact to back up first, followed by docker/volumes or whatever YUXI_STATE_DIR points at.

Editorial conclusion

Adopt Yuxi if you need retrieval, graph extraction, agent orchestration and per-department permissions inside your own network, and you can carry roughly ten containers plus a model API key. Do not adopt it if you only need a document chatbot, or if you cannot schedule a maintenance window: the README states that upgrading from v0.7.1 or v0.7.2 to the current release cannot be done with a plain docker compose up. Before committing, run ./scripts/init.sh on a staging host, confirm curl --fail http://localhost:5050/api/system/ready returns status ready, and read docs/advanced/deployment.md end to end.

Frequently asked questions

How do I install Yuxi on my own server?

Install Docker Engine and Docker Compose, clone the v0.7.3 tag, run ./scripts/init.sh to create .env, then run docker compose up --build -d. When curl --fail http://localhost:5050/api/system/ready returns status ready, open http://localhost:5173 and initialise the super administrator.

Can I upgrade Yuxi by just running docker compose up?

No. The README states that upgrading from v0.7.1 or v0.7.2 to the current version cannot be done with a direct docker compose up, and directs you to docs/advanced/deployment.md to do backup and migration in a downtime window. The Compose file requires API_KEY_DERIVATION_SECRET and SANDBOX_PROVISIONER_TOKEN, which older .env files do not contain.

How does Yuxi build its knowledge graph?

It extracts entities and relations from the document chunks of a Milvus knowledge base and writes them into Neo4j, where they participate in retrieval. Extraction happens during document parsing, and the interface shows entity totals, relation edge counts and build progress.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. xerrors/Yuxi on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xerrors-yuxi.svg)](https://hysenlabs.com/projects/xerrors-yuxi)