# Agentic Data Scientist: a multi-agent framework that plans before it codes

> K-Dense's Python framework splits data science work between a planning agent on OpenRouter and a coding agent on Claude Code, then reviews the result. Here is what it installs, what it costs to run, and where it breaks.

**K-Dense-AI/agentic-data-scientist** — An end-to-end Data Scientist

- Repository: https://github.com/K-Dense-AI/agentic-data-scientist
- Website: https://k-dense.ai
- Stars: 825 · Forks: 122
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/k-dense-ai-agentic-data-scientist

## What Agentic Data Scientist actually does with a research question

The project takes a natural language request, for example a differential expression analysis, and runs it through a multi-agent workflow instead of a single prompt. The README describes the goal as separating planning from execution, validating work continuously, and adapting the approach based on progress. That framing matters more than the feature list: the target user is someone who has a dataset and a question but does not want to hand-write every exploratory script, and who is willing to trade determinism for iteration.

The intended audience is narrower than "data scientists" in general. Because the coding agent runs through Claude Code, and the planning and review agents run through OpenRouter, the tool assumes you already hold accounts with both providers. A team that has standardized on a single vendor, or that cannot send data descriptions to a third-party router, is not the audience. The README also points to a hosted product, K-Dense Web, for users who need more capability than the open-source framework provides, which tells you where the maintainers see the ceiling of this repository.

## The plan, execute, validate loop and where each agent runs

The architecture is split by role rather than by data. Planning and review agents are ADK agents backed by OpenRouter; the coding agent is a Claude Code agent backed by Anthropic. The .env.example file makes this explicit with separate model settings: DEFAULT_MODEL and REVIEW_MODEL default to a Gemini model routed through OpenRouter, while CODING_MODEL is commented as needing to be an Anthropic model available in Claude Code.

Two mechanisms are worth calling out. First, the framework claims to create an analysis plan with success criteria before any code runs, then track progress against those criteria at every step, and finally revise the plan when execution turns up something unexpected. Second, tool access is partly external: the README lists MCP integration and names Context7 as a server for live library documentation, with an optional CONTEXT7_API_KEY in the environment file. That is a sensible place to put a dependency, because library APIs drift faster than agent prompts.

The trade-off is visible in the same design. A self-correcting plan is harder to audit than a fixed sequence of steps. The README explains why the loop exists (catch issues early, validate both plans and implementations) but does not describe how a revised plan is versioned or how a reviewer would reconstruct why the run changed direction.

## Installing Agentic Data Scientist and running a first analysis

There are two prerequisites before the Python package itself. The README requires the Claude Code CLI, installed globally through npm, and two API keys. Install the CLI first:

```bash
npm install -g @anthropic-ai/claude-code
```

Then install the framework. The README gives two routes: a persistent tool install, or a one-off run through uvx with no installation. The uvx form is useful for a first look because it does not touch your environment:

```bash
uv tool install agentic-data-scientist
uvx agentic-data-scientist --mode simple "your query here"
```

Both keys must be present, either exported in the shell or written to a .env file in the project directory. The .env.example shows the exact names:

```bash
OPENROUTER_API_KEY=your_key_here
ANTHROPIC_API_KEY=your_key_here
```

Now a real run. The README stresses that --mode is mandatory, and that this is deliberate so you are aware of the complexity and API cost. Orchestrated mode is the full multi-agent workflow; simple mode is direct coding with no planning stage. To analyze a CSV with the full loop, passing the file with -f and sending output to a named directory:

```bash
agentic-data-scientist "Perform differential expression analysis" --mode orchestrated --files data.csv --working-dir ./my_analysis
```

By default the run writes into ./agentic_output/ in the current directory and preserves those files afterward. Adding --temp-dir switches to temporary storage with automatic cleanup, which is what you want for a throwaway exploration. --verbose turns up logging, and --log-file redirects it somewhere specific. If you want the agents to stay off the network entirely, set DISABLE_NETWORK_ACCESS=true, which the README says removes the fetch_url tool from the ADK agents and the WebFetch and WebSearch tools from the Claude Code agent.

## Cost, keys and the two-vendor dependency

The most concrete constraint is that a single run bills against two providers at once. Planning and review traffic goes to OpenRouter; coding traffic goes to Anthropic. The README warns that choosing a mode is about being aware of API costs, and the difference between modes is real: simple mode skips the planning stage entirely, so it spends fewer tokens on the same request. If you are evaluating the tool, run the same prompt in simple mode first and only move to orchestrated once you know the planning stage earns its spend.

The dependency list adds a second cost that is easy to miss. google-adk is pinned to exactly 2.5.0, while claude-agent-sdk, google-genai, litellm and mcp are floor-pinned with >=. That asymmetry means an ADK upgrade is a deliberate edit to pyproject.toml, but a transitive upgrade of the SDKs can happen on any fresh install. Python is constrained to >=3.12,<3.13, so a 3.13 environment will refuse to install. The optional dev extra pulls pytest, pytest-asyncio and ruff if you intend to run the test suite in tests/.

## Where this framework is the wrong choice

Two failure modes stand out from the documentation as written. The first is reproducibility. A workflow that revises its own plan based on intermediate findings is a poor fit for anything that has to produce the same output twice, such as a regulated reporting pipeline or a model that must be retrained identically on a schedule. The README does not document rollback, plan versioning, or a way to freeze a run, so if your requirement is a fixed DAG, a workflow orchestrator is the better tool and this is not it.

The second is the network default. Agents have web search and URL fetching enabled unless you set DISABLE_NETWORK_ACCESS=true. For anyone working with data that cannot leave the environment, that default is the wrong one, and it is worth checking before the first run rather than after. The README also does not state what happens to uploaded files, how long they persist in the working directory, or whether agent traffic is logged by the providers. Those are questions to raise with the vendor, not gaps this review can fill.

## How it differs from a notebook assistant or a plain script generator

The obvious alternative is asking a chat assistant to write the analysis code, then running it yourself. The difference is the review stage. A chat assistant produces one draft and stops; this framework routes the draft through a separate review agent, compares it against success criteria set during planning, and can send it back. That extra pass is the whole product. It is also why the tool costs more and takes longer than a single completion.

A second alternative is a general workflow engine with an LLM node, where you define the steps and the model fills in one of them. That keeps the control flow in your hands and the model in a box. Agentic Data Scientist inverts it: the model proposes the control flow, and you supply the goal. Neither approach is strictly better. If your analysis steps are already known and stable, the workflow engine is cheaper and auditable. If the analysis steps are exactly what you are trying to discover, the adaptive loop is the point.

## Licence, maintenance and what an upgrade actually involves

The repository is MIT licensed, with the licence declared in pyproject.toml and a LICENSE file at the top level. MIT is permissive: it allows commercial use and modification, and it places no copyleft obligation on your own code. It also comes with no warranty, and it grants no rights over the third-party services the framework calls. Your use of OpenRouter, Anthropic and Context7 is governed by their terms, not by this licence. That is a factual boundary, not legal advice; check the provider terms yourself if the data is sensitive.

On maintenance, the last push was on 2026-08-18 and the repository is not archived. The release history shows v0.2.3 on 2026-05-29, with v0.2.2 and v0.2.1 both landing on 2025-11-21. That is a modest cadence with a gap between the November 2025 pair and the May 2026 release, so plan for the possibility that a fix you need arrives on the maintainers' schedule rather than yours. Upgrading the package is the easy part; the harder part is the pinned google-adk==2.5.0, which will hold you back if a newer ADK becomes a requirement elsewhere in your stack. Read docs/ and CONTRIBUTING.md before assuming a behaviour is documented somewhere in the README, because the README is a quick-start and not a reference.

## Conclusion

Adopt it if you already pay for Anthropic and OpenRouter access and want a repeatable plan-execute-review loop around messy analysis work; skip it if you need a deterministic pipeline, since the planning agent can revise the plan mid-run and the README does not document rollback. Before committing, verify the two API keys resolve, confirm Python 3.12 is the interpreter your tool runner picks, and check that the pinned google-adk==2.5.0 installs cleanly on your platform.

## FAQ

### What is Agentic Data Scientist?

It is an MIT-licensed Python framework from K-Dense that runs data science tasks through a multi-agent workflow, separating planning from execution and validating results as it goes. It is built on Google's Agent Development Kit and the Claude Agent SDK, and it is installed from PyPI as agentic-data-scientist.

### How do I install Agentic Data Scientist?

Install the Claude Code CLI with npm install -g @anthropic-ai/claude-code, then either run uv tool install agentic-data-scientist or call it directly with uvx agentic-data-scientist. Python must be 3.12, since pyproject.toml constrains requires-python to >=3.12,<3.13.

### Which API keys does Agentic Data Scientist need?

Two. OPENROUTER_API_KEY covers the planning and review agents, and ANTHROPIC_API_KEY covers the coding agent. Both can be exported as environment variables or placed in a .env file, as shown in .env.example.

### What is the difference between orchestrated and simple mode in Agentic Data Scientist?

Orchestrated mode runs the full multi-agent workflow with planning, execution and validation. Simple mode does direct coding with no planning stage, which the README presents as the faster option for quick tasks and question answering.

### Can I run Agentic Data Scientist without giving the agents network access?

Yes. Setting DISABLE_NETWORK_ACCESS to true or 1 disables the fetch_url tool for the ADK agents and the WebFetch and WebSearch tools for the Claude Code agent. Network access is enabled by default, so this has to be set explicitly.

## Sources

- [K-Dense-AI/agentic-data-scientist on GitHub](https://github.com/K-Dense-AI/agentic-data-scientist)
- [License: MIT](https://github.com/K-Dense-AI/agentic-data-scientist/blob/main/LICENSE)
- [Project website](https://k-dense.ai)
- [README](https://github.com/K-Dense-AI/agentic-data-scientist/blob/main/README.md)
- [Releases](https://github.com/K-Dense-AI/agentic-data-scientist/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/k-dense-ai-agentic-data-scientist
