Model or dataset
Zafer-Liu/Data-Analysis-Agent avatar
Zafer-Liu/Data-Analysis-Agent

Data-Analysis-Agent: Conversational SQL, Chart Generation, and Business Insights from Natural Language

🚀你的私人数据分析助手。通过对话式交互,自动生成可视化报表与商业洞察,让数据决策变得像聊天一样简单。 🚀 Your personal data analysis assistant. Say goodbye to complex SQL and Excel formulas. An LLM-powered data analysis agent. Chat with your data to instantly generate visualizations and business insights. Making data-driven decisions has never been easier.

2,603 stars226 forksJavaScriptNOASSERTION

At a glance

What is it?
Zafer-Liu/Data-Analysis-Agent is a Python agent that accepts natural language questions about your data and automatically generates SQL, selects from 43 chart types, and produces business summaries. It connects to uploaded Excel or CSV files and to MySQL, PostgreSQL, SQLite, and SQL Server databases, with a web UI and optional Windows installer.
Who is it for?
With release 1.4.0 from 2026-09-19, the project is under active development. The CC BY-NC 4.0 license prohibits commercial use without written permission from the author, limiting adoption to internal analytics, research, and personal use.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 12 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the Agent Does and Who It Targets

Data-Analysis-Agent addresses a specific bottleneck in business intelligence: non-technical users who need insights from structured data are blocked by the requirement to write SQL or use complex BI tools. The agent accepts questions in plain language, derives the query from the question, executes it against the connected data source, selects the most appropriate visualization, and presents the result with a natural language business summary.

The README describes the target user as a non-technical business analyst who wants to interact with data by chatting rather than querying. The examples given in the README illustrate the scope:

text
最近三个月销售额趋势如何?
哪个地区利润最高?
帮我生成用户增长图

These translate to: what is the sales trend over the last three months, which region has the highest profit, and generate a user growth chart. The agent handles all steps from schema inspection through chart rendering without requiring the user to specify the query structure or chart type.

How the Analysis Pipeline Works

The analysis process follows a documented four-step sequence that is surfaced to the user through SSE (Server-Sent Events) streaming output:

text
[1/4] 正在读取数据结构...
[2/4] 正在生成 SQL...
[3/4] 正在执行查询...
[4/4] 正在生成图表与洞察...

In step one, the agent reads the schema of the connected data source: table names, column names, and data types. In step two, it generates a SQL query based on the natural language question and the schema context. In step three, it executes the query and retrieves the result set. In step four, it selects a chart type from the library and generates the visualization alongside a business insight summary.

The streaming output is visible in the UI as each step completes, which the README describes as more transparent and interactive than traditional BI tools that show only the final result. The analysis process uses the LLM for the natural language understanding and SQL generation steps; the chart rendering and execution steps run as local code against the connected database.

Installing and Running the Agent

The agent supports three installation paths. The simplest option for Windows is the installer package, available from GitHub Releases. The README specifies the requirements as Python 3.10 or later and Windows 10 or 11, 64-bit. The installer creates a desktop shortcut for Business Analytics Agent.

For cross-platform use, a zip archive with platform-specific start scripts is available. On Windows, double-clicking start.bat launches the agent. On macOS, start.command performs the same function.

For Docker deployment, the Dockerfile produces a Python 3.11-based image:

dockerfile
FROM python:3.11-slim

The Dockerfile installs system packages for SQL Server support (unixodbc, libodbc2), OpenMP libraries for numpy and scikit-learn (libgomp1), and matplotlib rendering dependencies (libgl1, libglib2.0-0). After installing from requirements.txt, the container starts with:

dockerfile
CMD ["python", "-u", "app.py"]

Two environment variables are documented in the Dockerfile: BAA_HOST (set to 0.0.0.0 for network accessibility) and BAA_SKIP_DEPENDENCY_CHECK (set to 1 for faster startup in the cloud deployment). The default port is 5001.

Chart System and Export Formats

The agent automatically selects from 43 chart types organized into six categories. The categories and their members are documented in the README with specific chart names:

Comparison charts include Marimekko (absolute and percentage), Bar, Grouped Bar, Stacked Bar, Diverging Bar, Dot Plot, Waffle, Bullet, Sankey, Heatmap, and Waterfall. Time series charts include Line, Circular Line, Slope, Sparkline, Bump, Cycle, Area, Stacked Area, Horizon, and Connected Scatter. Distribution charts include Histogram with Pareto, Pyramid, Error Bar, Box-and-Whisker, Violin, Ridgeline, Beeswarm, and Stem-and-Leaf. Geographic charts include Flow Map, Dot Density Map, and Choropleth Map. Relationship charts include Scatter, Bubble, Radar, Chord, Arc, Network Diagram, and Parallel Coordinates. Proportion charts include Treemap, Sunburst, Nightingale, and Pie.

The agent also supports generating statistical analyses including outlier handling (truncation and winsorization), decile grouping, K-Means clustering, and decision tree modeling.

Export formats include cleaned Excel spreadsheets, DOCX-format reports, and PPT presentations in a built-in style. The requirements.txt lists python-docx and python-pptx as the export dependencies.

Data Source Support and LLM Configuration

The agent connects to four categories of data sources. File uploads support Excel (handled via openpyxl, xlrd, and python-calamine) and CSV. Database connections support SQLite, MySQL (via pymysql), PostgreSQL (via psycopg2-binary), and SQL Server (via pyodbc). The README marks DuckDB and Spark as planned future additions.

In addition to the standard relational databases, version 1.3.0 added Feishu (Lark) multi-dimensional table support. A Feishu robot integration reads from Feishu multidimensional tables, routes the data through DuckDB for in-memory SQL analysis, and returns results to the group chat.

The agent supports multiple LLM providers. The README lists DeepSeek, OpenAI, and AtlasCloud as supported providers, along with any OpenAI SDK-compatible API endpoint through custom base_url, model, and api_key configuration. The default models are deepseek-v4-flash for DeepSeek, gpt-4o-mini for OpenAI, and deepseek-v4-pro for AtlasCloud. MCP (Model Context Protocol) extension is also supported for connecting external tools.

The agent also supports MCP (Model Context Protocol) extension for connecting to local or remote MCP servers to expand the agent's available tools. A separate knowledge base input feature allows uploading domain-specific business documents: the agent uses this context to better interpret questions that use company-specific terminology or naming conventions. Tutorial documents for both features are in the Information/ directory of the repository.

A companion tool, Agent Manager (Zafer-Liu/Agent_Manager), provides a desktop management layer for the agent. Adding the agent to Agent Manager allows one-click start and stop, real-time log and port monitoring, opening the web UI directly from the desktop application, and generating a temporary public-facing share URL for demonstrations.

PandasAI is a comparable Python library that adds natural language querying to pandas DataFrames. It focuses on DataFrame-level analysis in Python code rather than providing a web UI, does not include a built-in chart library of 43 types, and does not connect to SQL databases directly. The key difference is the interface: Data-Analysis-Agent provides a web UI and report export, while PandasAI is a Python library that returns results to a Python environment.

License and Maintenance Status

The repository carries a CC BY-NC 4.0 license for the software itself. The README states explicitly that commercial use is prohibited without authorization from the author, and that the author has applied for a Chinese software copyright. The non-commercial restriction applies to use, redistribution, and derivative works. Internal enterprise analytics teams that use the software within their organization without external distribution may not be subject to the commercial use clause, but organizations with any commercial element should contact the author before deployment.

The project also has a NOASSERTION value in the GitHub license field, which is a separate indicator that the machine-readable license metadata is not fully configured. The CC BY-NC 4.0 text in the README governs actual usage.

The latest release is 1.4.0 from 2026-09-19, with the previous releases v1.3.1 on 2026-09-03 and v1.3.0 on 2026-08-23. The rapid release cadence and the last push on 2026-09-17 indicate active development. The v1.3.0 changelog documented improvements to long-term memory: fixing memory extraction failures caused by thinking model output and JSON format issues, adding 24-hour automatic memory consolidation, session recovery protection, and an option to disable long-term memory in the general settings. A SECURITY.md is present in the repository for reporting security issues. An install.sh script in the repository root provides a guided setup for Linux and macOS users who prefer a script-based installation over the manual steps. The requirements.txt includes statistical analysis libraries including statsmodels and pmdarima for time series analysis, and scikit-learn for the K-Means clustering and decision tree features. The pyproject.toml configures ruff for linting, targeting Python 3.10 with checks for syntax errors, undefined names, and common bug patterns via the E9, F, and B rule sets. For SQL Server connectivity, the Dockerfile documents that the unixodbc and libodbc2 packages must be present on the host, and the requirements.txt specifies pyodbc as the driver. Teams deploying on systems without these libraries will need to install them separately before the SQL Server connection will function.

Editorial conclusion

With release 1.4.0 from 2026-09-19, the project is under active development. The CC BY-NC 4.0 license prohibits commercial use without written permission from the author, limiting adoption to internal analytics, research, and personal use. Teams on Windows can use the packaged installer. Linux or Docker deployments should start from the Dockerfile, which sets BAA_HOST=0.0.0.0 and BAA_SKIP_DEPENDENCY_CHECK=1 as the required environment variables. Verify that the connected database user has read access to the schemas you intend to query before running the agent.

Frequently asked questions

What is a data analysis agent and what does this one do specifically?

A data analysis agent accepts natural language questions about structured data and automates the query, visualization, and insight steps. Data-Analysis-Agent specifically generates SQL from plain language questions, selects from 43 chart types, and produces business summaries with SSE streaming output visible during the analysis process.

What databases does Data-Analysis-Agent support?

The agent connects to SQLite, MySQL, PostgreSQL, and SQL Server for database sources, and also accepts uploaded Excel and CSV files. Feishu multi-dimensional tables are supported since version 1.3.0. DuckDB and Spark support are listed as planned future additions.

Can Data-Analysis-Agent be used commercially?

The README applies a CC BY-NC 4.0 license and explicitly states that commercial use is prohibited without written permission from the author. Internal business analytics use may be permitted depending on the interpretation, but any external commercial deployment requires contacting the author.

Official sources

  1. Issues
  2. README
  3. Releases
  4. Zafer-Liu/Data-Analysis-Agent on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zafer-liu-data-analysis-agent.svg)](https://hysenlabs.com/projects/zafer-liu-data-analysis-agent)