Model or dataset
zi-yue-1129/DATAGEN avatar
zi-yue-1129/DATAGEN

DATAGEN: a LangGraph multi-agent pipeline for hypothesis generation, analysis and report writing

DATAGEN: AI-driven multi-agent research assistant automating hypothesis generation, data analysis, and report writing.

1,810 stars251 forksPythonMIT

At a glance

What is it?
DATAGEN is an MIT-licensed Python project that wires eight LLM agents into a LangGraph state machine, pauses for a human to approve a hypothesis, then runs data analysis, search and report writing. Its configuration and dependency choices, not its marketing copy, are what decide whether it fits your workflow.
Who is it for?
Adopt DATAGEN if you already hold API keys for at least one of OpenAI, Anthropic or Google, you are comfortable editing YAML to assign a different model to each agent, and you want a human checkpoint between hypothesis generation and the rest of the pipeline.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What DATAGEN actually automates, and who it is built for

The project frames itself as a research assistant: you point it at a data file, describe the analysis in plain language, and it produces hypotheses, code, charts and a written report. The README's own example is a single string pasted into main.py that names a CSV and asks for machine learning analysis plus graphical reports. That is the intended entry point, and it tells you who this is for. It is for a data analyst or researcher who can read Python, has API keys, and wants the boring middle of an analysis (writing the plotting code, drafting the narrative) handled by agents rather than by hand. It is not a no-code tool. There is a frontend/ directory in the repository root, but the README documents only the script path, so anyone expecting a hosted interface will be working from the source tree. The eight named agents map cleanly onto the phases of a short research project: hypothesis_agent proposes, process_agent supervises, code_agent writes analysis code, visualization_agent draws, searcher_agent looks things up, report_agent writes, quality_review_agent checks, and note_agent records the run. The decomposition is the product. If you only wanted an LLM to write pandas code, you would not need a graph.

The LangGraph state machine and the human checkpoint in the middle

The README states the workflow explicitly: hypothesis generation, then a human choice to continue or regenerate, then processing (analysis, visualization, search, report writing), then quality review, then revision as needed. That ordering matters more than the agent list. The human step sits early, before any code is written or any file is read, which means the expensive part of the run does not start until you have approved a research direction. Regeneration is a loop back to the hypothesis agent rather than a restart, so a bad first hypothesis costs you one model call, not a full pipeline. The note_agent is described as recording the research process, and the README calls it a state-tracking mechanism for context retention across phases. In practice this is the component that keeps the report agent from contradicting the visualization agent. The trade-off is that state accumulates. There is no documented compaction step, so a long run with many revisions carries the whole history forward into each subsequent model call. Whether that becomes a cost problem depends on your context window and on how often you ask for a regenerated hypothesis.

Installing DATAGEN and running a first analysis from main.py

The README gives a Conda-based install and requires Python 3.10 or higher. Clone the repository, create the environment, and install the pinned dependency set. Note that the clone URL in the README points at starpig1129/DATAGEN rather than the zi-yue-1129 path, so use whichever remote you actually have access to.

bash
git clone https://github.com/starpig1129/DATAGEN.git
conda create -n datagen python=3.10
conda activate datagen
pip install -r requirements.txt

Once dependencies are in place, the README says to place your data file in the data directory, edit the user_input variable in the main() function of main.py, and run the script. The example string it gives names the file and states the task.

python
user_input = '''
datapath:YourDataName.csv
Use machine learning to perform data analysis and write complete graphical reports
'''

Then start the run from the activated environment. The graph should reach the hypothesis step and stop for your choice before any analysis code executes.

bash
python main.py

Environment variables you must fill before anything runs

The repository ships a file literally named `.env Example`, and the README instructs you to rename it to `.env` and fill in the values. Four settings are marked required: WORKING_DIRECTORY, CONDA_ENV, CHROMEDRIVER_PATH, and the working directory is also consumed by the filesystem MCP server. CONFIG_DIRECTORY is optional and defaults to config/. The API keys for OpenAI, Anthropic and Google are all marked optional, which only makes sense in combination with the agent_models.yaml file: if you assign every agent to a single provider, the other keys stay empty. FIRECRAWL_API_KEY is optional but the README warns that query capabilities may be reduced without it. The ChromeDriver requirement is the one that catches people. Selenium 4.37.0 is pinned in requirements.txt, and the default path expects a Linux binary at ./chromedriver-linux64/chromedriver, so a macOS or Windows user must change that value and supply a matching driver.

sh
# Your data storage path (required)
WORKING_DIRECTORY = ./data/

# Configuration directory path (optional)
CONFIG_DIRECTORY = config

# Conda environment name (required)
CONDA_ENV = datagen

# ChromeDriver executable path (required)
CHROMEDRIVER_PATH = ./chromedriver-linux64/chromedriver

Assigning a different model to each agent through agent_models.yaml

This is DATAGEN's most interesting design choice. Rather than one model for the whole pipeline, each agent gets its own provider and model entry in agent_models.yaml inside CONFIG_DIRECTORY. The README's example mixes three providers in a single config: gpt-5-nano for hypothesis generation, gemini-2.5-pro for note taking, and claude-haiku-4-5 for code. Temperature is per-agent and ranges from 0.0 to 2.0. The stated purpose is environment switching: point CONFIG_DIRECTORY at a different folder and you get a different set of models, with config_local suggested for local development because it is already in .gitignore. That is a clean pattern for separating a cheap dev setup from an expensive production one. The cost is that model names live in YAML and are not validated at import time, so a typo or a retired model surfaces as a runtime failure partway through a graph run, after you have already paid for the hypothesis step.

yaml
agents:
  hypothesis_agent:
    provider: openai
    model_config:
      model: gpt-5-nano
      temperature: 1.0
  note_agent:
    provider: google
    model_config:
      model: gemini-2.5-pro
      temperature: 1.0
  code_agent:
    provider: anthropic
    model_config:
      model: claude-haiku-4-5
      temperature: 1.0

Where DATAGEN breaks down or is the wrong tool

The dependency pins are aggressive and will fight you. requirements.txt pins langchain 1.0.2, langchain-core 1.0.0, langgraph 1.0.1 and langchain-community 0.4 in the same file, alongside langchain-openai 1.0.1, langchain-anthropic 1.0.0 and langchain-google-genai 3.0.0. If you already have a LangChain stack in the same environment, expect to isolate DATAGEN in its own Conda env, which is exactly why CONDA_ENV is a required setting. The second failure mode is the browser dependency. Selenium is pinned and a ChromeDriver path is required, so headless server deployments need a working Chrome or Chromium install that the README does not walk through. Third, there is no release history. The repository has no retrieved releases, so you are tracking main and there is no version to pin against or changelog to read before upgrading. If your team requires a tagged artifact before anything reaches production, DATAGEN is the wrong shape today. Finally, the agents are only as good as their prompts, and the README describes the agent configuration system as Progressive Disclosure without documenting the prompt files themselves, so tuning behaviour means reading config/agents/ directly.

How DATAGEN differs from a plain LangChain agent or AutoGen

A single LangChain agent with a code-execution tool can already write pandas code and plot a chart. The difference here is the fixed graph. DATAGEN commits to a research shape in advance: hypothesis, human approval, execution, review, revision. That shape is enforced by LangGraph rather than negotiated between agents at runtime, which makes runs more predictable and makes the human checkpoint a first-class node instead of something you bolt on. Compared with conversational multi-agent frameworks where agents message each other freely, DATAGEN's agents are roles in a pipeline with a supervisor, not peers in a debate. If your task does not look like a research project (for example, a recurring ETL job that needs the same transformation every night), the hypothesis and review stages are pure overhead. If your task does look like one, the fixed graph is the reason to pick this over assembling the same loop yourself.

Licence, maintenance and what an upgrade costs you

DATAGEN is MIT-licensed, and the LICENSE file sits at the repository root. MIT permits commercial use and modification provided the copyright notice and permission notice are retained; it offers no patent grant and no warranty, so if your organisation cares about patent language you should read the text rather than assume. The last push to the repository was on 2026-08-16, so the code is recent, but there are no retrieved releases and therefore no versioned upgrade path. Upgrading means pulling main and re-reading requirements.txt for changed pins, then re-checking agent_models.yaml against whatever model names your providers still serve. Because the LangChain and LangGraph pins are exact, a single bump in that file can ripple across every agent, and there is no changelog in the repository to tell you what changed. Budget for reading diffs on main rather than for a versioned release process. Note also that the README refers to the project as previously named AI-Data-Analysis-MultiAgent, so older issues and forks may be filed under the former name.

Editorial conclusion

Adopt DATAGEN if you already hold API keys for at least one of OpenAI, Anthropic or Google, you are comfortable editing YAML to assign a different model to each agent, and you want a human checkpoint between hypothesis generation and the rest of the pipeline. Do not adopt it if you need a packaged service with a versioned release, if you cannot supply a ChromeDriver binary, or if your data cannot leave your machine and you have not verified that the ollama provider path works for every agent. Before committing, run main.py once on a small CSV and confirm three things: that the human choice step actually blocks for input, that the note_agent writes to the directory named by WORKING_DIRECTORY, and that the model names in config/agent_models.yaml still resolve on your provider accounts.

Frequently asked questions

What does DATAGEN do?

It runs a multi-agent research pipeline: a hypothesis agent proposes a direction, you choose whether to continue, and then code, visualization, search and report agents carry out the analysis under a process supervisor. The README describes the output as data analysis, visualizations and a written report.

What are data generation tools?

This question is about a different category of software, and DATAGEN is not one of them. DATAGEN analyses data you already have; it takes a file such as a CSV placed in WORKING_DIRECTORY and produces analysis code, charts and a report rather than new synthetic records.

What is DATAGEN?

DATAGEN is an MIT-licensed Python project, previously named AI-Data-Analysis-MultiAgent, that builds a LangGraph state machine over eight named agents. It requires Python 3.10 or higher and a Conda environment named by the CONDA_ENV setting.

Why was DataGen scrapped?

The README does not describe the project being discontinued or scrapped. It documents an install path, a configuration system and a running workflow, and the last push to the repository was on 2026-08-16.

What arg is datagen from?

The README does not mention an alternate reality game or any related fiction. DATAGEN here is a Python multi-agent data analysis project that uses LangChain, LangGraph and a set of LLM providers.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. zi-yue-1129/DATAGEN on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zi-yue-1129-datagen.svg)](https://hysenlabs.com/projects/zi-yue-1129-datagen)