# DeepAnalyze: an 8B agentic LLM that runs the data science pipeline end to end

> DeepAnalyze is an open source, 8B-parameter agentic model from Renmin University of China and Tsinghua that plans, writes and executes analysis code across structured and unstructured files. It is a research release with a real deployment burden, and the README does not cover rollback or failure recovery.

**ruc-datalab/DeepAnalyze** — DeepAnalyze is the first agentic LLM for autonomous data science. 🎈你的AI数据分析师，自动分析大量数据，一键生成专业分析报告！

- Repository: https://github.com/ruc-datalab/DeepAnalyze
- Website: https://ruc-deepanalyze.github.io
- Stars: 4,665 · Forks: 737
- Language: Python
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/ruc-datalab-deepanalyze

## The gap DeepAnalyze targets: analysis that stops at the code cell

Most LLM data analysis tools generate a pandas snippet and hand it back. The user still has to run it, read the error, fix the column name, run it again, and decide what to plot. DeepAnalyze is built to close that loop. The README describes it as "the first agentic LLM for autonomous data science" and lists the full pipeline as its scope: data preparation, analysis, modeling, visualization, and report generation, with no human intervention in between.

The intended user is not a beginner who cannot read a dataframe. It is someone who already knows what a good analysis looks like and wants the mechanical part done: locating the right files among many, writing the exploratory code, and assembling the result into something presentable. The repository also points at a benchmark called CoDA-Bench, released in June 2026, which evaluates whether code agents can handle data-intensive analytical tasks in a Linux sandbox with hundreds of data files. That is a fair description of the setting DeepAnalyze is designed for.

## How the agent works: file discovery, code execution, report assembly

The README frames the mechanism as open-ended data research over mixed sources. The model is expected to handle structured data (databases, CSV, Excel), semi-structured data (JSON, XML, YAML), and unstructured data (TXT, Markdown), then produce what the project calls analyst-grade research reports.

The repository layout shows how that is packaged. The root holds deepanalyze.py and a deepanalyze/ package, plus run.py as an entry point. A demo/ directory contains separate front ends: demo/chat, demo/chat_v2, demo/cli, demo/jupyter, demo/deepanalyze_general, and demo/mock_vllm. The presence of a mock vLLM demo is telling. It suggests the team expects people to develop against a fake backend before committing GPU time, which is a sensible choice for a project whose real inference path is expensive.

The sandbox is the load-bearing part. The March 2026 WebUI v2 update added Docker-based sandboxed code execution. An agent that writes and runs its own Python needs that isolation; without it, generated code runs with whatever access the host process has. The README does not describe what the sandbox restricts, only that it exists.

## Installing DeepAnalyze and running a first analysis

The repository ships a requirements.txt at the root. It lists numpy, pandas, openpyxl, scikit-learn, seaborn, torch, transformers, matplotlib, plotly, and statsmodels for the analysis stack, and requests, websockets, python-multipart, uvicorn, fastapi, openai, and pypandoc for the web UI. Note the first line: vllm is commented out, so it is not installed by default.

The README gives no install command, so the only installable artifact named in the repository is the requirements file itself. Install it with the standard pip form.

```bash
pip install -r requirements.txt
```

If you intend to serve the model yourself, vllm has to be installed separately, because the line in requirements.txt is commented out. The file records only vllm>=0.8.5, which is a floor, not a tested pin.

The model weights are on Hugging Face at RUC-DataLab/DeepAnalyze-8B. The project also offers API keys, announced in December 2025, with an application form and a usage guide at docs/DeepAnalyze_API_Key_Usage_Guide.md. If you go the API route, you skip the weights entirely.

For a first run, the CLI demo is the least ceremony. The README notes that as of November 2025 DeepAnalyze supports OpenAI-style API endpoints and is reachable through a command line terminal UI, contributed by a community member. The README does not document the exact invocation, so read demo/cli before starting anything.

Expect the agent to take a question in natural language, explore the files you point it at, and emit code and a report. What you should watch for on the first run is whether the sandbox is active. If code executes without a container boundary, stop and configure the Docker path in demo/chat_v2 before pointing the agent at anything sensitive.

## The deployment cost is the real constraint, not the model size

An 8B model is small by current standards, but the surrounding system is not. You need a vLLM server, a web UI process, and a Docker sandbox, and the agent will generate and execute many code cells per task. Token consumption scales with the number of exploratory steps, not with the length of the question. A vague prompt about a directory of CSVs can turn into dozens of execution rounds.

The README does not publish latency, throughput, or cost figures, so there is no way to size this from the documentation alone. That is a genuine gap for anyone planning a deployment. The API key path exists precisely because self-hosting is heavy; the December 2025 announcement routes users through a Google Form or a Feishu Form, which means access is granted rather than open.

The second constraint is trust. The agent writes code and runs it. If the sandbox configuration is wrong, or if you disable it for convenience, you have given a language model execution rights on your machine. The README states that Docker-based sandboxed execution exists but does not document its boundaries, so the guarantee is only as good as the container configuration you write yourself.

## Where DeepAnalyze is the wrong tool

If your task is a fixed query against a known schema, DeepAnalyze is overkill. A SQL statement or a ten-line pandas script is faster, cheaper, and reproducible. An agent that explores is valuable when the data layout is unknown; it is a liability when the layout is known and the answer must be identical every time.

It is also the wrong choice when you need auditable, deterministic output. The README describes autonomous completion of tasks, and autonomy means the path taken can vary between runs. For regulated reporting, that variance is the problem, not the feature.

Finally, the documentation is thin on failure. The README does not document rollback, does not describe what happens when generated code fails repeatedly, and does not specify a retry or abort policy. Before depending on this for anything scheduled, you would need to read the source in deepanalyze/ rather than the README.

## How it differs from a general coding agent

The obvious comparison is a general-purpose coding agent pointed at a Jupyter notebook. The difference is the training data and the target. DeepAnalyze ships with DataScience-Instruct-500K on Hugging Face, a dataset of data science instructions, and the model was tuned for that domain rather than for software engineering in general. The project also maintains CoDA-Bench specifically to measure agents on data-intensive analytical tasks.

That specialization shows in the file-type coverage the README claims: databases, CSV, Excel, JSON, XML, YAML, TXT, and Markdown in one workflow. A general coding agent can read those formats, but it has no particular reason to treat Excel and YAML as first-class inputs to the same analysis.

The trade-off runs the other way too. A general agent benefits from a much larger body of community tooling, plugins, and documentation. DeepAnalyze's ecosystem is the demo/ directory and the API guide. The related project DeepPrep, announced for release in July 2026, is a data-preparation companion that turns raw tables into analysis-ready data, which suggests the authors know that preparation is a separate problem worth its own system.

## Maintenance, licensing, and what the release cadence tells you

The last push to the repository was on 2026-08-30. The repository is not archived. The README shows a steady stream of dated updates through 2026: WebUI v2 in March, CoDA-Bench in June, DeepPrep announced in July. There are no retrieved releases, so distribution is through the repository and Hugging Face rather than a versioned package. That means upgrading is a git pull, and there is no changelog to read between commits.

The licence is MIT. That is permissive: it allows commercial use, modification, and redistribution with the licence text retained. It says nothing about the model weights, which live on Hugging Face under the RUC-DataLab organization and may carry their own terms. It also says nothing about the API service, which is a separate arrangement governed by the application form. If you plan to build a product on the weights, check the model card separately from the LICENSE file. This is not legal advice; the point is that MIT on the code does not automatically cover the checkpoint or the hosted endpoint.

Upgrade cost is the practical concern. With no releases and a commented-out vLLM dependency, a pull can change the inference path without warning. Pinning a commit hash is the only reliable way to keep a working deployment stable.

## Conclusion

Adopt DeepAnalyze if you have GPU capacity, a sandbox for generated code, and analysis tasks where a written report is the deliverable rather than a number you will act on immediately. Do not adopt it if you need a stable hosted service, if you cannot isolate code execution, or if your data cannot leave your network and you have no hardware to run the 8B weights. Verify three things first: that the 8B checkpoint fits your GPU memory, that the sandbox in demo/chat_v2 works in your environment, and that the API key application path in docs/DeepAnalyze_API_Key_Usage_Guide.md is still open.

## FAQ

### What is DeepAnalyze used for?

It is an agentic LLM for autonomous data science. The README lists data preparation, analysis, modeling, visualization, and report generation, and says it can research structured, semi-structured, and unstructured data sources to produce analyst-grade reports.

### How accurate is AI data analysis with DeepAnalyze?

The README does not publish accuracy figures for DeepAnalyze. The project points to CoDA-Bench, a benchmark released in June 2026 for evaluating whether code agents can handle data-intensive analytical tasks, which is the setting the model targets.

### How can deep learning be used for data analysis in DeepAnalyze?

DeepAnalyze uses an 8B language model that plans and executes code rather than applying a trained network to the data. The model weights are on Hugging Face at RUC-DataLab/DeepAnalyze-8B, and the repository includes a Docker-based sandbox for the code the agent runs.

### What does the phrase deep analysis mean in the context of this project?

In this repository the name refers to an agentic LLM for autonomous data science, not to a generic analytical technique. The README ties it to open-ended data research across structured, semi-structured, and unstructured files.

## Sources

- [Issues](https://github.com/ruc-datalab/DeepAnalyze/issues)
- [License: MIT](https://github.com/ruc-datalab/DeepAnalyze/blob/main/LICENSE)
- [Project website](https://ruc-deepanalyze.github.io)
- [README](https://github.com/ruc-datalab/DeepAnalyze/blob/main/README.md)
- [ruc-datalab/DeepAnalyze on GitHub](https://github.com/ruc-datalab/DeepAnalyze)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ruc-datalab-deepanalyze
