GustoBot: a LangGraph multi-agent recipe assistant you can retarget
五星大厨:全面Multi-Agent 的客服机器人,基于langraph实现,txt2sql ,txt2cypher, lightrag, 多模态 等
At a glance
- What is it?
- GustoBot is an Apache-2.0 Python project that wires LangGraph subgraphs around Neo4j, MySQL, Milvus, PostgreSQL pgvector and LightRAG to answer Chinese recipe questions. It is a template for vertical-domain customer service, not a drop-in product.
- Who is it for?
- Adopt GustoBot if you want a working reference for a LangGraph router plus tool subgraphs and you are willing to run Neo4j, MySQL, Milvus, PostgreSQL and Redis yourself. Do not adopt it if you need a single-binary chatbot or you cannot supply an LLM and embedding endpoint.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 32 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem GustoBot is aimed at: recipe questions that need four different stores
A recipe question is rarely one kind of question. "How many Sichuan dishes are in the collection?" is a count. "What goes with chicken and peanuts?" is a relationship. "Where does kung pao chicken come from?" is a passage of prose. "What does it look like?" is an image. A single vector index handles the third case and does badly at the first two. GustoBot's answer is to keep each store in its lane and let a router pick. The README frames this as a multi-layer knowledge graph over recipe names, ingredients, cooking steps, nutrition, regional cuisines and historical anecdotes, and the target audience is explicit: teams who want a "transferable, extensible, vertical-domain customer service template" rather than a finished recipe app. The README lists Pokemon encyclopedias, traditional Chinese medicine, legal consultation and government services as domains you could move to by swapping the knowledge source and graph schema.
Three layers, one router, and a fallback chain that decides the answer
The architecture is three tiers. L1 is the main routing layer: `analyze_and_route_query` combines heuristic keyword matching with an LLM structured output, and `route_query` then sends the request to one of six branches (general-query, additional-query, graphrag-query, text2sql-query, kb-query, image-query, file-query). L2 is a set of LangGraph subgraphs. The GraphRAG subgraph runs a Planner node that decomposes a question into subtasks, calls tools in parallel, and then a Summarize node merges the results. L3 is the atomic tool layer: Neo4j for Cypher, MySQL for Text2SQL aggregation, Milvus and PostgreSQL pgvector for retrieval, and an external search API in hybrid mode.
The retrieval order is the part worth reading closely. The README describes a PostgreSQL-first fallback: structured data from pgvector first, Milvus semantic search if that returns nothing, external search last. A reranker with two thresholds gates the pgvector path, and the README shows the comparison as `similarity >= KB_POSTGRES_SIMILARITY_THRESHOLD` and `rerank_score >= KB_POSTGRES_RERANK_THRESHOLD`, with Cohere, Jina, Voyage and BGE listed as supported rerankers. Two thresholds rather than one is a real design decision: a passage can be semantically close but still not answer the question, and the rerank score is what filters that out. The cost is two model calls per candidate set, and the README does not say what happens when the thresholds disagree or how they were chosen.
A Guardrails node sits inside the subgraphs and rejects queries outside the recipe domain, and the README also mentions Neo4j schema validation before Cypher execution. Conversation state goes through a LangGraph checkpointer backed by Redis.
Running GustoBot locally: Docker Compose, then the Vite frontend
The README requires Python 3.10, Node.js 16 or newer, and Docker with Docker Compose. The first step is cloning and creating the environment file. The comment in the README says to edit `.env` and configure the necessary API keys, so expect to fill in an LLM key and an embedding endpoint before anything starts.
git clone https://github.com/skygazer42/GustoBot.git
cd GustoBot
cp .env.example .envAfter editing `.env`, start the backend with Docker Compose. The README shows this as a single command, and the backend listens on port 8000.
docker-compose up -dThe frontend is not containerized in the README's instructions. You install and run it locally from the `web` directory, and the dev server listens on port 5173, which the README notes can be changed through `VITE_PORT`.
cd web
npm install
npm run devWhat the Compose stack actually starts
`docker-compose up -d` builds the backend image from the `server` target in the Dockerfile and starts it on port 8000. The Compose file also defines the dependencies the code expects: Redis on 6379 with `REDIS_HOST=redis`, Neo4j at `bolt://neo4j:7687`, Milvus on 19530, and MySQL with the database `recipe_db`. The backend container sets DNS to 8.8.8.8 and 8.8.4.4, which matters on hosts where the default resolver cannot reach the model endpoints. The Dockerfile installs from `requirements.txt` through the USTC PyPI mirror and switches apt to the Tsinghua Debian mirror, so a build outside mainland China will pull from those mirrors regardless of your local configuration.
The LightRAG build step is the sharp edge in the install
The Dockerfile takes `INIT_LIGHTRAG_ON_BUILD` as a build argument and, when true, initializes LightRAG data during the image build, persisting `LLM_BASE_URL`, `LLM_MODEL`, `EMBEDDING_MODEL`, `EMBEDDING_BASE_URL` and `EMBEDDING_DIMENSION` as environment variables. The defaults point at `https://dashscope.aliyuncs.com/compatible-mode/v1` with `qwen3-max` and at an embedding service on `http://139.224.116.116:3000/v1` with dimension 4096. The Compose file defaults `INIT_LIGHTRAG_ON_BUILD` to true and passes those values as build args. That means an unmodified `docker-compose up -d` will attempt to reach a third-party embedding host during the build, and the build will fail or hang if that host is unreachable. The Dockerfile comments warn against passing secrets through build args because they leak into build logs, yet `LLM_API_KEY` and `EMBEDDING_API_KEY` are declared as build args. Set `INIT_LIGHTRAG_ON_BUILD=false` and point `EMBEDDING_BASE_URL` and `EMBEDDING_DIMENSION` at your own service before building if you want a predictable first run.
Where GustoBot is the wrong tool
The dependency surface is the first limitation. A working instance needs Neo4j, MySQL, Milvus, PostgreSQL with pgvector, and Redis, plus an LLM endpoint and an embedding endpoint. That is six stateful services before the first question is answered. For a small recipe site, or for anyone who wants to embed a chatbot in a static page, a single Postgres with pgvector and one LLM call would cover the same ground with far less to operate. GustoBot is justified when you already run those stores or when the router-plus-subgraph pattern is the thing you are evaluating.
The second limitation is data provenance. The README states that the built-in recipe master data comes from OpenKG RecipeGraph, and that some historical food texts and eight pgvector sample records have no verifiable upstream source in the original commit, so they are only suitable for pipeline demonstration. A team that swaps in its own corpus inherits the pipeline but also inherits the job of documenting where that corpus came from. The README points to docs/DATA_SOURCES.md for lineage, use and licence notes.
The third is scope. The Guardrails node deliberately rejects questions outside the recipe domain. If your use case is general assistant behaviour, that node is working against you and has to be rewritten, not configured.
How this differs from a plain RAG chatbot or from LangChain's own SQL agent
The closest comparison is a single-store RAG pipeline: chunk documents, embed them, retrieve the top k, generate. That approach answers the anecdote question and fails the count question, because no amount of similarity search produces an accurate `COUNT` over a relational table. GustoBot's Text2SQL path exists precisely for that, and the README's example questions include both "how do you make kung pao chicken" and "how many Sichuan dishes are there".
The second comparison is LangChain's SQL agent, which also generates queries from natural language. The difference is the routing layer. A SQL agent assumes every question is a database question; GustoBot decides first whether the question is a graph question, a SQL question, a retrieval question, an image question or out of scope, and the subgraph only sees work it was chosen for. That extra hop costs a routing call and adds a failure mode: a misrouted question reaches a subgraph that cannot answer it, and the README does not describe a retry when that happens.
Licence, maintenance and what an upgrade costs
GustoBot is Apache-2.0, which permits commercial use and modification provided you keep the licence and notices, and it includes a patent grant. The repository also carries a SECURITY.md. Apache-2.0 covers the code in this repository; it does not cover the OpenKG RecipeGraph data or any corpus you ingest, and those carry their own terms. That is a question for your own legal review, not something the licence file answers.
Maintenance: the repository is not archived, and the last push was on 2026-08-29. The default branch is `develop`, not `main`, so the tip of the repository is a development branch. Releases are sparse: v0.1 in November 2025 and v0.1.2 in April 2026, while `pyproject.toml` still declares version 0.1.2. The gap between the last push and the last release means the `develop` branch contains work that no release captures. If you pin to a tag you get the released state; if you track `develop` you get the current state and no release notes for it.
Upgrade cost is dominated by the pinned dependencies. `requirements.txt` pins langchain 0.3.7, langgraph 0.2.60, pymilvus 2.3.7, neo4j 5.27.0 and `numpy<2.0.0`, with `langchain-core>=0.3.39,<0.4.0`. LangChain and LangGraph both move quickly, and a jump to a newer LangGraph release will touch every graph definition in the project. `pyproject.toml` sets `requires-python = ">=3.9"` and mypy to Python 3.9 while the README and Dockerfile both target Python 3.10, so the declared floor and the tested floor are not the same.
Editorial conclusion
Adopt GustoBot if you want a working reference for a LangGraph router plus tool subgraphs and you are willing to run Neo4j, MySQL, Milvus, PostgreSQL and Redis yourself. Do not adopt it if you need a single-binary chatbot or you cannot supply an LLM and embedding endpoint. Before committing, verify that the three-tier routing in gustobot/ actually matches the diagram in the README, check the data lineage notes in docs/DATA_SOURCES.md, and confirm that the LightRAG build step in the Dockerfile works against your own embedding service.
Frequently asked questions
How do I install GustoBot?
Clone the repository, copy .env.example to .env and fill in the API keys, then run docker-compose up -d for the backend. The frontend is started separately from the web directory with npm install and npm run dev. The README lists Python 3.10, Node.js 16+ and Docker Compose as requirements.
Which databases does GustoBot need to run?
The Compose file configures Redis, Neo4j on bolt://neo4j:7687, Milvus on 19530 and MySQL with the recipe_db database, and the README describes PostgreSQL pgvector as the first retrieval store. That is five stateful services before the LLM and embedding endpoints are counted.
Can GustoBot answer questions outside recipes?
The README states that a Guardrails node checks every query against the service scope and that out-of-scope questions are explicitly refused and redirected. The general-query route answers without calling external tools, but the domain boundary is enforced rather than configurable in the code shown.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/skygazer42-gustobot)