Model or dataset
spring-ai-alibaba/DataAgent avatar
spring-ai-alibaba/DataAgent

Spring AI Alibaba DataAgent: a Java analyst that writes SQL, runs Python, and ships a report

Spring AI Alibaba DataAgent

2,669 stars638 forksJavaApache-2.0

At a glance

What is it?
DataAgent is a Spring Boot service that turns a question into SQL, executes generated Python in a sandbox, and renders an HTML report with ECharts. It is for Java teams that already run Spring AI Alibaba and want an analyst agent they can self-host.
Who is it for?
Adopt DataAgent if your team is already on Spring Boot 3.4.8+ and Spring AI Alibaba 1.1.2.2, and you want the analyst loop (SQL, Python, report) inside your own infrastructure with an OpenAI-compatible model endpoint. Do not adopt it as a drop-in replacement for a BI tool or if you cannot run Docker, because the Python analysis path depends on the sandbox.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap DataAgent fills between a chat model and a BI dashboard

A chat model can write a SELECT statement. It cannot run it, notice that the result set is wrong, write a follow-up query, plot the numbers, and hand back a document. DataAgent is built for that longer loop. The README describes it as an "enterprise-grade intelligent data analyst" that goes beyond Text-to-SQL tools, and the feature table lists four capabilities that map onto four stages: intent understanding over multiple turns, SQL generation against a schema, Python execution for analysis the SQL layer cannot express, and report generation as HTML or Markdown with ECharts charts.

The audience is narrow and specific. This is a Java project, built on Spring AI Alibaba Graph, requiring Spring Boot 3.4.8+ and JDK 17+. If your data platform is a Python shop, the surrounding stack (Maven wrapper, Spring Boot service, Nuxt frontend) is friction you would pay for no reason. If your team already runs Spring Boot services and wants an analyst agent inside the same deployment boundary, the fit is much closer. The README also notes compatibility with any model that follows the OpenAI interface, naming Qwen and Deepseek as examples, so the model choice is not locked to one vendor.

StateGraph, the sandbox, and where the Python actually runs

The orchestration layer is a StateGraph, the graph abstraction from Spring AI Alibaba. The README says the Text-to-SQL conversion is "based on StateGraph" and that the system supports multi-table queries and multi-turn intent understanding. That means the analyst is not a single prompt but a graph of nodes with shared state, which is what makes the human-in-the-loop feature possible: the README describes a mechanism where the user can intervene and adjust the plan during the plan generation stage, before execution continues.

The part that deserves attention is Python execution. The README states that generated code runs in task-level containers through Spring AI Alibaba Sandbox, with PEP 723 dynamic dependencies, resource limits, retry on failure, and automatic cleanup. PEP 723 is the inline script metadata convention, so a generated script can declare its own dependencies and the sandbox resolves them per task rather than from a shared environment. Task-level containers plus automatic cleanup is the right shape for untrusted generated code, and the resource limits are the constraint that keeps a runaway analysis from consuming the host. Docker is listed as required when the workflow needs to execute Python steps, which tells you the sandbox is container-backed rather than a subprocess.

RAG sits alongside this. The README says a vector database is integrated for semantic retrieval over business metadata and terminology, with the stated goal of improving SQL generation accuracy. The knowledge documentation distinguishes semantic models, business knowledge, and agent knowledge, so retrieval is aimed at schema and vocabulary rather than at documents in general. A model registry allows switching LLM and Embedding models at runtime, which matters because the embedding model and the vector store have to agree on dimensions.

Installing DataAgent and running the first agent

The README points to docs/QUICK_START.md for the full guide and gives a three-step path in the repository root. Prerequisites are JDK 17+, MySQL 5.7+, Node.js 22+, pnpm 11+, and Docker when the workflow executes Python steps.

The first command imports the schema. It runs from the repository root and reads the SQL file bundled with the management module:

bash
mysql -u root -p < data-agent-management/src/main/resources/sql/schema.sql

You should see MySQL prompt for the root password and then create the tables without error. If the import fails, check the MySQL version first, since the README states 5.7+ and the schema is the authority on what it actually needs.

The backend starts through the Maven wrapper, scoped to the management module only:

bash
./mvnw -pl data-agent-management spring-boot:run

The -pl flag limits the build to data-agent-management, which is the Spring Boot service. Expect the usual Spring Boot startup log and the service listening on its configured port; the README does not print that port, so take it from the application configuration in that module.

The frontend is a Nuxt app in its own directory, installed and started with pnpm:

bash
cd data-agent-frontend-nuxt
pnpm install && pnpm dev

After that, the README says to open http://localhost:3000 in a browser and create your first data agent. That port is the Nuxt dev server, not the backend. Configuration details for models, the vector store, and API keys live in docs/DEVELOPER_GUIDE.md and docs/ADVANCED_FEATURES.md, which the README links but does not reproduce.

What the setup asks of you before the first useful answer

The three commands above get the system running. They do not get it answering questions about your data. Between those two points sit the things the README lists as features but does not walk through: an OpenAI-compatible chat model and embedding model have to be configured, the vector database has to be mounted if you want RAG, and the semantic model and business knowledge have to be described before SQL generation has anything to retrieve. The README's knowledge documentation exists precisely because that configuration is where accuracy comes from.

The API Key layer is worth noting for a different reason. The README describes API Key lifecycle management with fine-grained permission control, and the advanced features document covers calling the system through an API Key. That suggests DataAgent is designed to be embedded in other applications, not only driven from the Nuxt UI. If you only need the UI, that machinery is weight you carry without using.

The Makefile at the repository root is not an install path. It forwards every target to CI/make/common.mk, java.mk, linter.mk, and tools.mk, so make targets belong to the project's own build and lint pipeline rather than to deployment. Do not expect a make install here.

Where DataAgent is the wrong choice

Docker is a hard dependency for the analysis path. The README lists it as required when the workflow needs to execute Python steps. In an environment where containers cannot run, or where generated code must not leave the host process, the Python analysis feature is unavailable, and the remaining Text-to-SQL path is a much smaller product than the feature table suggests. That is a deployment constraint, not a bug, but it changes what you are adopting.

The release history is the second thing to weigh. The most recent release listed is 1.0.0-rc7, published on 2026-07-29, following 1.0.0-rc6 on 2026-07-17 and 1.0.0-rc5 on 2026-03-12. These are release candidates. The repository is not archived and the last push was on 2026-09-20, so work is ongoing, but there is no 1.0.0 final in the release list. Teams that require a stable version number before adopting a dependency should treat that as a gap to close with the maintainers rather than something to assume away.

Finally, the README does not document rollback, backup, or migration procedures for the schema it asks you to import. For a tool that reads from your operational database and writes reports, that silence is a real operational question, and the README is the source that is silent on it.

DataAgent against a Python-first analyst stack and against MCP clients

The closest alternative in kind is a Python-based analyst agent built on a framework such as LangChain or LlamaIndex, where the agent, the analysis code, and the plotting libraries all live in one language. The difference is not capability, it is the boundary. In a Python stack, the generated analysis code runs in the same ecosystem as the orchestration, so there is no cross-language handoff and no separate sandbox component to operate. In DataAgent, the orchestration is Java and the analysis is Python, and the Spring AI Alibaba Sandbox is the bridge. You pay for that bridge in an extra runtime dependency, and you get a service that fits a Spring Boot deployment, exposes API Key management, and can be extended with Java-side plugins through the documented extension points for vector stores and models.

A second comparison is not a substitute but an overlap. DataAgent implements MCP and can act as a tool server, per the README, exposing NL2SQL and agent management to MCP-capable clients such as Claude Desktop. If your actual need is "let my existing assistant query this database," running the full DataAgent stack with its Nuxt frontend and API Key layer may be more surface than the problem requires. The MCP server capability is the smaller, more targeted option, and it is worth deciding which of the two you are adopting before you install anything.

Licence, upgrade cost, and what a version bump touches

DataAgent is licensed under Apache License 2.0, and the LICENSE file is at the repository root. The source headers in the Makefile carry the standard Apache 2.0 notice. Apache 2.0 permits commercial use and modification and includes a patent grant, but it also carries notice and attribution obligations, and the licence text itself is the authority on what those are. This is a description of the licence, not legal advice; if your organisation has a policy on Apache 2.0 dependencies, the LICENSE file is what your review should read.

The upgrade surface spans three runtimes. A version bump can move the Spring AI Alibaba dependency (the README badges 1.1.2.2), the Spring Boot baseline (3.4.8+), and the frontend toolchain (Node.js 22+, pnpm 11+) independently, and the schema.sql file is versioned with the backend. That means an upgrade is not a single artifact swap: the database schema, the Java service, and the Nuxt frontend move together, and the release candidates listed above are the units you would be tracking. Budget for reading the release notes for each candidate rather than assuming a patch-level change.

Editorial conclusion

Adopt DataAgent if your team is already on Spring Boot 3.4.8+ and Spring AI Alibaba 1.1.2.2, and you want the analyst loop (SQL, Python, report) inside your own infrastructure with an OpenAI-compatible model endpoint. Do not adopt it as a drop-in replacement for a BI tool or if you cannot run Docker, because the Python analysis path depends on the sandbox. Before committing, verify that the schema.sql import matches your MySQL version, that your model endpoint follows the OpenAI interface, and that a vector store is optional in your deployment rather than assumed.

Frequently asked questions

What is Spring AI Alibaba DataAgent?

It is an enterprise-grade intelligent data analyst built on Spring AI Alibaba Graph. The README describes it as going beyond Text-to-SQL to execute Python analysis and produce HTML or Markdown reports with ECharts charts.

How do I install and start Spring AI Alibaba DataAgent?

Import the schema with mysql, start the backend with ./mvnw -pl data-agent-management spring-boot:run, then run pnpm install && pnpm dev in data-agent-frontend-nuxt and open http://localhost:3000. The README points to docs/QUICK_START.md for the full configuration guide.

Does Spring AI Alibaba DataAgent require Docker?

Docker is listed as a prerequisite when the workflow needs to execute Python steps. The README states that generated code runs in task-level containers through Spring AI Alibaba Sandbox, so the Python analysis path depends on a container runtime being available.

Which models and vector databases does Spring AI Alibaba DataAgent support?

The README says the system is compatible with chat and embedding models that follow the OpenAI interface, naming Qwen and Deepseek as examples, and that it supports mounting any vector database. A built-in model registry allows switching LLM and Embedding models at runtime.

What is the difference between Spring AI Alibaba DataAgent and a plain Text-to-SQL tool?

The README positions DataAgent as going beyond Text-to-SQL: it adds Python deep analysis in a sandbox, automatic report generation with ECharts charts, a human-in-the-loop plan review step, and RAG retrieval over business metadata and terminology.

Under what licence is Spring AI Alibaba DataAgent released?

The project uses the Apache License 2.0, and the LICENSE file is at the repository root. The README states the licence in its own section as well.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. spring-ai-alibaba/DataAgent on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/spring-ai-alibaba-dataagent.svg)](https://hysenlabs.com/projects/spring-ai-alibaba-dataagent)