# SQLBot: a self-hosted text-to-SQL system built on LLMs and RAG

> SQLBot is a DataEase project that turns natural-language questions into SQL against your own databases, packaged as a Docker image with a web UI. The interesting part is the retrieval layer and the permission model; the awkward part is the licence and the model configuration.

**dataease/SQLBot** — 🔥 基于大模型和 RAG 的智能问数系统，对话式数据分析神器。Text-to-SQL Generation via LLMs using RAG.

- Repository: https://github.com/dataease/SQLBot
- Website: https://sqlbot.org/
- Stars: 6,861 · Forks: 889
- Language: JavaScript
- License: NOASSERTION
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/dataease-sqlbot

## What SQLBot actually solves, and for whom

Writing SQL is a translation problem. Someone knows the question ("which regions missed quota last quarter") and someone else knows the schema, and the two are rarely the same person. SQLBot targets that gap by accepting a natural-language question, generating SQL against a connected data source, and returning the result plus a chart.

The audience is narrower than "anyone who wants analytics". SQLBot expects you to have a database, a place to run a container, and an LLM endpoint. The README frames the product around ChatBI, conversational analytics, and lists integration paths into n8n, Dify, MaxKB and DataEase. That list tells you who the project is really for: teams already inside the FIT2CLOUD ecosystem, or teams assembling an internal tool stack from open source pieces and wanting a question-answering layer on top of their warehouse.

It is not a BI tool that happens to have a chat box. The output is SQL and a chart, not a governed semantic layer with certified metrics. If your organisation needs a single agreed definition of "active customer", SQLBot will happily generate three different queries that each define it differently.

## The RAG layer is what separates SQLBot from a prompt wrapper

Sending a schema dump to an LLM and asking for SQL is a solved demo and an unsolved product. The failure mode is well known: the model invents column names, joins tables that share no key, and produces syntactically valid SQL that returns the wrong number.

SQLBot's answer is retrieval-augmented generation over schema and business context. The README describes the mechanism in terms of a RAG pipeline feeding the model, and the "越问越准" section names the levers explicitly: custom prompts, a terminology library, and maintained SQL examples used to calibrate logic. Those three artifacts are the retrieval corpus. The terminology library maps business vocabulary to physical columns; the SQL examples act as few-shot anchors for the patterns your organisation actually uses.

The practical consequence is that SQLBot's accuracy is a function of curation effort, not of the model alone. A fresh install pointed at a database with cryptic column names and no terminology entries will perform noticeably worse than the same install after someone has spent a week writing examples. The README says the effect improves with use, which is true, but it undersells that the improvement is manual work, not automatic learning.

Data flow, as far as the repository shows: question enters the frontend, backend assembles retrieval context from the configured sources, calls the configured LLM provider, receives SQL, executes it against the data source, and returns rows to the frontend for rendering. The g2-ssr directory in the repository root indicates chart rendering happens server-side through a separate build stage in the Dockerfile, which is why the image build is multi-stage and pulls a base image from an Aliyun registry.

## Installing SQLBot with Docker and running a first question

The README's quick start assumes a Linux server with Docker installed. The one-liner below is copied from it. It publishes ports 8000 and 8001, mounts four data directories plus the PostgreSQL data volume, and runs privileged, which the container needs for its bundled database.

```bash
docker run -d \
  --name sqlbot \
  --restart unless-stopped \
  -p 8000:8000 \
  -p 8001:8001 \
  -v ./data/sqlbot/excel:/opt/sqlbot/data/excel \
  -v ./data/sqlbot/file:/opt/sqlbot/data/file \
  -v ./data/sqlbot/images:/opt/sqlbot/images \
  -v ./data/sqlbot/logs:/opt/sqlbot/app/logs \
  -v ./data/postgresql:/var/lib/postgresql/data \
  --privileged=true \
  dataease/sqlbot
```

After the container starts, the README says to open http://<your server IP>:8000/ and log in as admin with the password SQLBot@123456. Change that password immediately; it is published in the README and in the docker-compose.yaml as DEFAULT_PWD.

If you prefer Compose, the repository ships a docker-compose.yaml. It is stricter than the docker run form: the SECRET_KEY variable uses the ${SECRET_KEY:?...} syntax, so Compose refuses to start until you create a .env file containing it. The relevant environment block looks like this.

```yaml
environment:
  POSTGRES_SERVER: localhost
  POSTGRES_PORT: 5432
  POSTGRES_DB: sqlbot
  POSTGRES_USER: root
  POSTGRES_PASSWORD: Password123@pg
  PROJECT_NAME: "SQLBot"
  DEFAULT_PWD: "SQLBot@123456"
  SECRET_KEY: ${SECRET_KEY:?Set a unique SECRET_KEY in .env before starting SQLBot}
  LOG_LEVEL: "INFO"
  SQL_DEBUG: False
```

Once logged in, the first real task is connecting a data source and an LLM provider. The README lists supported providers: Alibaba Cloud Bailian, Qianfan, DeepSeek, Tencent Hunyuan, iFlytek Spark, Gemini, OpenAI, Kimi, Tencent Cloud, Volcano Engine, MiniMax, and a generic OpenAI-compatible option. Everything except OpenAI itself is reached through OpenAI-compatible APIs, which means an internal gateway exposing that interface should work under the generic entry. Set SQL_DEBUG to True if you want generated SQL logged while you tune terminology and examples.

## Where SQLBot breaks down

The licence is the first constraint most teams will hit. The repository's licence field reads NOASSERTION, and the README states the project follows the FIT2CLOUD Open Source License, described there as GPLv3 with additional restrictions: you may not replace or modify the SQLBot logo or copyright information, and derivative works must comply with GPLv3. That is a copyleft licence with a branding clause. If you plan to embed SQLBot inside a closed product, or to white-label it, the branding restriction is the clause to read carefully. This is not legal advice; the LICENSE file is the authority.

Second, the privileged container. Both the README command and docker-compose.yaml set privileged: true. That grants the container broad access to the host. It is there because SQLBot bundles PostgreSQL and manages its own data directory, but it means you should not run this on a host shared with workloads you do not control.

Third, generated SQL is not guaranteed correct. Nothing in the README claims otherwise, and the entire terminology-and-examples apparatus exists because raw generation is unreliable. Treat SQLBot as a drafting tool. A wrong aggregate that looks plausible is worse than an error message, and SQLBot has no documented verification step between generation and execution.

Finally, the README does not document rollback, backup procedures, or an upgrade path between versions. Releases exist (v1.10.1 on 2026-08-27, v1.10.0 on 2026-07-16, v1.9.0 on 2026-06-09) but the README says nothing about how to move between them or what happens to your terminology library and SQL examples when you do. The last push to the repository was on 2026-09-09.

## SQLBot against building the pipeline yourself

The obvious alternative is not another product; it is assembling the same thing from parts. A retrieval store for schema and examples, a prompt template, an OpenAI-compatible client, and a small web frontend is perhaps a few hundred lines. Teams with strong Python or TypeScript engineers often take this route because they want the retrieval corpus to live in their own git repository and the permission model to match their existing identity provider.

The difference in approach is where the work sits. A homegrown pipeline gives you total control over what goes into the context window and how results are cached, and it costs you the frontend, the chart rendering, the workspace isolation model, and the integration hooks. SQLBot ships those already. The README describes workspace-level resource isolation and fine-grained data permissions, plus embedding options (web embed, popup embed, MCP invocation) for dropping the assistant into another application. Rebuilding the permission layer alone is more work than most teams estimate.

If your interest is specifically MCP, note that the docker-compose.yaml carries an MCP-related setting, SERVER_IMAGE_HOST, defaulted to a placeholder http://YOUR_SERVE_IP:MCP_PORT/images/. It needs a real value before MCP image references resolve. The README lists MCP as a supported integration path but does not document the setup in the excerpt available here.

## Who should adopt SQLBot, and what to check first

Adopt it if you have a database, a capable LLM endpoint, and a container host, and you want conversational querying without writing the retrieval and permission layers. The Docker path is genuinely short: one command, a browser, a login. Teams already running DataEase, MaxKB, 1Panel or n8n have the clearest fit, since the README names those as integration targets and the project shares an organisation with them.

Do not adopt it if a copyleft licence with a branding restriction is a blocker, if you cannot run a privileged container, or if you expect the tool to produce audited numbers without human review. It is also the wrong choice if your question is "what is our revenue" rather than "what does the revenue table contain". SQLBot answers questions about data, not about definitions.

Verify before you commit: that SECRET_KEY is set in .env (Compose will not start without it), that the default admin password has been changed, that your model provider is in the supported list or exposes an OpenAI-compatible API, and that you have a plan for the terminology library and SQL examples, because those files are what make the output usable. The README does not describe how to export or version them, so decide that yourself before the corpus grows.

## Maintenance, upgrades and licence cost

The repository was last pushed on 2026-09-09 and is not archived. Releases have arrived at a steady cadence through 2026: v1.9.0 in June, v1.10.0 in July, v1.10.1 in August. That is a real release rhythm, but the README does not document an upgrade procedure, so the practical upgrade cost is unknown from the repository files. The docker-compose.yaml mounts ./data/postgresql as a host volume, which at least means your data survives a container replacement, provided you keep that directory.

Operationally, the bundled PostgreSQL and the privileged flag mean SQLBot is closer to an appliance than a library. You are maintaining a container, a database, and a model endpoint, not just a dependency. The LOG_LEVEL and SQL_DEBUG environment variables are the only documented observability controls; the logs volume maps to /opt/sqlbot/app/logs.

On licensing, the README's summary is that the FIT2CLOUD Open Source License is GPLv3 plus restrictions on logo and copyright removal, with GPLv3 obligations for derivative works. If you distribute a modified SQLBot, those obligations travel with it. If you only run it internally, the practical effect is smaller, but the branding clause still applies to anything you show users. Read the LICENSE file rather than the README summary.

## Conclusion

Adopt SQLBot if you already run a database you can point it at and you want a chat interface over it without building the retrieval and permission layers yourself. Skip it if you need a permissive licence, if you cannot host a capable LLM endpoint, or if you expect the generated SQL to be correct without review. Before committing, verify three things: that the SQLBot@123456 default password is changed, that SECRET_KEY is set in .env as docker-compose.yaml demands, and that your intended model provider appears in the supported list or speaks the OpenAI-compatible API.

## FAQ

### What exactly is SQL used for in a tool like SQLBot?

In SQLBot, SQL is the intermediate output: the system takes a natural-language question, generates SQL against a connected data source, executes it, and returns the rows plus a chart. The README describes this as a ChatBI workflow built on large language models and RAG.

### Which AI is better for SQL?

SQLBot does not rank providers by SQL quality. The README lists supported providers (Alibaba Cloud Bailian, Qianfan, DeepSeek, Tencent Hunyuan, iFlytek Spark, Gemini, OpenAI, Kimi, Tencent Cloud, Volcano Engine, MiniMax, and a generic OpenAI-compatible entry) and states that accuracy improves through custom prompts, a terminology library and maintained SQL examples.

### Is SQL still worth learning in 2026 if SQLBot generates it?

SQLBot generates SQL but does not verify it. The README frames the tool around custom prompts, a terminology library and SQL examples used to calibrate logic, all of which require someone who can read the generated queries. The README makes no claim that generated SQL is correct without review.

### Which is harder, SQL or Excel?

SQLBot does not compare the two. It targets the SQL side of the problem: the README describes it as a system for conversational data analysis that generates SQL from natural-language questions against a connected data source.

## Sources

- [dataease/SQLBot on GitHub](https://github.com/dataease/SQLBot)
- [Issues](https://github.com/dataease/SQLBot/issues)
- [Project website](https://sqlbot.org/)
- [README](https://github.com/dataease/SQLBot/blob/main/README.md)
- [Releases](https://github.com/dataease/SQLBot/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dataease-sqlbot
