# pingcap/autoflow: a Graph RAG knowledge base on TiDB Serverless

> AutoFlow is a TypeScript application that turns documentation sites into a conversational search endpoint, with TiDB Serverless holding vectors, JSON and chat history. It ships as Docker Compose images, and the README still calls it early-stage software.

**pingcap/autoflow** — pingcap/autoflow is a Graph RAG based and conversational knowledge base tool built with TiDB Serverless Vector Storage. Demo: https://tidb.ai

- Repository: https://github.com/pingcap/autoflow
- Website: https://tidb.ai
- Stars: 2,975 · Forks: 192
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/pingcap-autoflow

## The documentation-site search problem AutoFlow targets

Most product teams have a docs site, a sitemap and no good way to answer "where is the setting for X" without a human. AutoFlow's stated purpose is a conversational knowledge base: you point its built-in crawler at official and documentation sites, it scrapes sitemap URLs, and the result is a Perplexity-style search page plus an embeddable JavaScript widget that sits in the bottom right corner of your own site. The README describes the crawler as designed to navigate official and documentation sites, and the widget as a way to give instant responses to product-related queries.

The audience is narrow on purpose. This is not a general chatbot builder. It is for teams whose corpus is public documentation and whose users ask product questions. If your knowledge lives in tickets, PDFs or a private wiki, the crawler-first design is the wrong entry point.

## How the Graph RAG pipeline is put together

The stack list is explicit: TiDB stores chat history, vectors, JSON and analytics; LlamaIndex is the RAG framework; DSPy handles the programming of foundation models; Next.js and Tailwind with shadcn/ui form the frontend. The repository layout matches that split, with backend/, core/, frontend/, docs/ and e2e/ as separate top-level directories, and Docker images published as tidbai/backend and tidbai/frontend.

The graph part is what separates it from a plain vector store. AutoFlow is described as graph rag, meaning a knowledge graph is built alongside embeddings, so retrieval can follow relationships between chunks rather than only nearest-neighbour similarity. The README does not document the graph schema, the extraction prompt, or how entities are merged across documents. That is a real gap: you can run the system, but you cannot audit how a given edge was created from the documentation provided.

Data flow in the Compose file is straightforward. The backend exposes port 80 inside the container, mapped to 8000 on the host, and reads configuration from .env. The frontend depends on the backend and receives BASE_URL: http://backend, so the browser talks to Next.js and Next.js talks to the API. A separate background service runs the same backend image with /usr/bin/supervisord on port 5555, which is where crawling and indexing work is likely executed. Redis sits under both as a dependency.

## Installing AutoFlow with Docker Compose

The README points deployment at the Docker Compose route, documented at autoflow.tidb.ai/deploy-with-docker, and notes the host should have 4 CPU cores and 8GB RAM. The repository ships docker-compose.yml, docker-compose.dev.yml and docker-compose-cn.yml, plus a .env.example to copy.

Start by creating your environment file. The template requires a TiDB connection and a secret key.

```bash
cp .env.example .env
python3 -c "import secrets; print(secrets.token_urlsafe(32))"
```

The second command is the one the .env.example itself suggests for generating SECRET_KEY. Paste the output into SECRET_KEY; the file states it must be greater than or equal to 32 characters and should not be shared.

Then fill in the TiDB block. The example uses a TiDB Serverless host and recommends creating a cluster from tidbcloud.com.

```bash
TIDB_HOST=xxxxx.prod.aws.tidbcloud.com
TIDB_USER=
TIDB_PASSWORD=
TIDB_DATABASE=
TIDB_SSL=true
```

Bring the stack up with the published images. The Compose file pins tidbai/backend:0.4.0 and tidbai/frontend:0.4.0, and publishes the UI on port 3000 and the API on port 8000.

```bash
docker compose up -d
```

After the containers report running, open http://localhost:3000. That is the frontend service port from the Compose file. If you intend to embed the JavaScript widget on another origin, set BACKEND_CORS_ORIGINS to your domain, which the .env.example describes as required for JS widgets.

One optional service deserves attention. local-embedding-reranker runs image tidbai/local-embedding-reranker:v4-with-cache on port 5001 and sits behind the profile local-embedding-reranker, so it does not start with a plain docker compose up. Its environment preloads the default embedding model and leaves PRE_LOAD_DEFAULT_RERANKER_MODEL set to false; the file notes you can switch it to true to preload the reranker as well. GPU support is commented out and would require uncommenting a deploy block with the nvidia driver.

## Where AutoFlow is the wrong tool

The README carries its own warning: AutoFlow is still in the early stages of development, and the next move is to make it a Python package so it becomes a RAG solution installed as pip install autoflow-ai. Treat that as a statement about the current shape of the project. There is no pip package today, so the supported path is containers.

The release history reinforces the point. The newest release listed is 0.4.0 from 2025-01-03, preceded by 0.4.0rc1 and v0.3.0 in December 2024. Anyone who needs a versioned upgrade path with frequent tagged releases should look elsewhere or accept running from main. The last push to the repository was on 2026-04-27, so the code has moved since the last tagged release, but tags are what most operators track.

There are operational limits too. The Compose file expects a reachable TiDB cluster; there is no bundled database service, so the stack is not self-contained. The embedding and reranking service is optional and profile-gated, which means the default deployment depends on whatever embedding configuration your .env supplies, and the README does not spell out which providers are supported. Finally, the crawler is sitemap-driven. A site with no sitemap, or one behind authentication, is outside what the documentation describes.

## AutoFlow compared with a plain LlamaIndex pipeline

The closest alternative is building the same thing directly on LlamaIndex, which AutoFlow already uses as its RAG framework. The difference is scope. A hand-rolled LlamaIndex service gives you control over chunking, the vector store and the query engine, and you can swap components without touching a UI. AutoFlow gives you less control and more assembled product: a crawler, a chat page, an admin surface, a widget and a background worker, all wired to TiDB for vectors, JSON and analytics in one database.

The trade-off is visible in the dependencies. Choosing AutoFlow means adopting TiDB Serverless as the vector store, DSPy as the prompt-programming layer, and Next.js as the frontend, because those are the components the images ship. If your organisation has standardised on Postgres with pgvector, AutoFlow is not the path of least resistance; the Compose file has no Postgres service and the .env.example only carries TIDB_* keys. Conversely, if you want the graph layer and the widget without writing retrieval code, assembling that yourself on LlamaIndex is a multi-week project. The honest framing is that AutoFlow trades flexibility for a working default, and the default is TiDB.

## Maintenance, upgrade cost and the Apache-2.0 licence

Upgrades are image-tag changes. The Compose file pins tidbai/backend:0.4.0 and tidbai/frontend:0.4.0, so moving to a newer build means editing those tags and recreating the containers. Because backend and background run the same image, both need the same tag or the API and the worker will diverge. The frontend image is separate and versioned independently, which the README's badge labels as tidbai/frontend and tidbai/backend respectively.

State lives in two places. TiDB holds chat history, vectors and JSON, and the Compose file mounts ./data into /shared/data for both backend and background, with Redis persisting to ./redis-data. Back up the TiDB cluster and the ./data directory together; restoring only one leaves the graph and its source files out of step.

On licensing, the README states AutoFlow is open source under the Apache License, Version 2.0, with the text in LICENSE.txt. That is a permissive licence, but the practical question for most teams is not the licence text: it is that the deployment depends on TiDB Serverless, a hosted service with its own terms, and on model providers you configure yourself. Review those service terms separately from the repository licence.

## Conclusion

Adopt AutoFlow if you already run TiDB Serverless or want a crawler, chat UI and embeddable widget in one Compose file, and you can accept early-stage software whose README says a pip package is the next move. Do not adopt it if you need a stable release cadence (the newest release listed is 0.4.0 from 2025-01-03) or you cannot supply a TiDB connection and a SECRET_KEY of at least 32 characters. Verify first that the crawler reaches your sitemap, that the embedding service you configure is reachable from the backend container, and that the Apache-2.0 licence terms fit how you plan to redistribute the frontend and backend images.

## FAQ

### What is the purpose of pingcap/autoflow?

It is an open source graph RAG knowledge base tool. The README describes it as built on TiDB Vector, LlamaIndex and DSPy, with a Perplexity-style conversational search page and an embeddable JavaScript widget for answering product-related queries.

### What is pingcap/autoflow software?

A TypeScript application whose stack is TiDB for chat history, vectors, JSON and analytics, LlamaIndex as the RAG framework, DSPy for programming foundation models, and Next.js with Tailwind and shadcn/ui on the front end.

### Is pingcap/autoflow free to use?

The README states AutoFlow is open source under the Apache License, Version 2.0, with the licence text in LICENSE.txt. The deployment itself needs a TiDB cluster, which the .env.example recommends creating through TiDB Serverless and which carries its own service terms.

### How do you use pingcap/autoflow?

The README points to Docker Compose deployment with 4 CPU cores and 8GB RAM. You copy .env.example to .env, supply SECRET_KEY of at least 32 characters and the TIDB_* connection values, then run docker compose up -d and open the frontend on port 3000.

## Sources

- [License: Apache-2.0](https://github.com/pingcap/autoflow/blob/main/LICENSE)
- [pingcap/autoflow on GitHub](https://github.com/pingcap/autoflow)
- [Project website](https://tidb.ai)
- [README](https://github.com/pingcap/autoflow/blob/main/README.md)
- [Releases](https://github.com/pingcap/autoflow/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pingcap-autoflow
