Model or dataset
K-Dense-AI/karpathy avatar
K-Dense-AI/karpathy

Karpathy: An Agentic ML Engineer for Local Research Workflows

An agentic Machine Learning Engineer

1,555 stars177 forksPythonMIT

At a glance

What is it?
K-Dense-AI/karpathy is an open-source Python project that combines the Claude Agent SDK with Google ADK to create a conversational ML engineer, using a local sandbox to run training pipelines. It targets individual researchers who want to drive real ML toolchains from a chat interface, without writing agent scaffolding from scratch.
Who is it for?
Karpathy suits ML researchers comfortable with Claude Code who want to drive training and experimentation through conversation rather than building agent scaffolding themselves. It is a poor fit for teams needing reproducibility guarantees, remote compute, or a production serving layer: the sandbox is local, temporary, and requires manual data loading.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 42 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A Conversational Interface for ML Training Pipelines

K-Dense-AI/karpathy positions itself as an agentic ML engineer: a Python application that accepts natural language instructions and translates them into real machine learning workflows. The target user is an individual researcher or engineer who already works with Claude Code and wants to offload repetitive parts of model training, such as data preprocessing, hyperparameter search, and evaluation scripting, to a conversational agent rather than writing orchestration code by hand.

The project is not a notebook replacement or a GUI dashboard. It requires the user to supply datasets manually, state goals in plain text, and then monitor the outputs the agent writes to a local sandbox directory. The README is direct: it calls the implementation simple and frames the project as a demonstration of what Scientific Agent Skills can do in an ML context. There is no claim of production readiness, and the project has no GitHub releases as of the last push on 2026-08-18.

A detail worth noting: the project is named after Andrej Karpathy but is not endorsed by or affiliated with him. The README states this explicitly.

How the Sandbox Model Works

The central design choice in karpathy is the sandbox directory. When `start.py` runs, it creates a `sandbox` subdirectory, copies the `.env` file into it, sets up a Python virtual environment inside the sandbox with ML packages including PyTorch, transformers, and scikit-learn, and pulls down skills from the Scientific Agent Skills library. The agent operates inside that sandbox: it reads files, writes scripts, runs training loops, and stores outputs there.

The agent itself is built on two frameworks layered together. The Claude Agent SDK handles the agentic loop, tool calls, and skill execution. Google ADK (Agent Development Kit) provides the web interface and the session routing. When the ADK web server starts, it exposes a local interface at http://localhost:8000 where the user selects the karpathy agent and types instructions. The agent then dispatches to whichever Scientific Agent Skills match the task, and those skills run Python code inside the sandbox.

The OpenRouter API sits between the agent and the underlying language model. The `AGENT_MODEL` environment variable controls which model the agent uses. The agent does not call any model API directly; it routes all inference through OpenRouter, which means the model choice is decoupled from the code but also means every inference call goes over the network to a third-party service.

Setup: Clone, Configure, and Run start.py

The project requires Python 3.13 or higher and the `uv` package manager. The README does not document a Docker path, so local Python is the only documented route.

Clone the repository and move into it:

bash
git clone https://github.com/K-Dense-AI/karpathy.git
cd karpathy

Install dependencies with uv:

bash
uv sync

Create a `.env` file in the `karpathy` directory with your API keys:

bash
OPENROUTER_API_KEY=your_openrouter_api_key_here
AGENT_MODEL=your_model_name_here

The `OPENROUTER_API_KEY` is required. The `AGENT_MODEL` variable sets which model OpenRouter will route to. Both values are copied into the sandbox when `start.py` runs, so agents inside the sandbox can use them.

Then run the startup script:

bash
python start.py

This creates the sandbox, installs ML packages into it, copies the `.env`, and starts the ADK web interface. The README says to navigate to http://localhost:8000 and select `karpathy` in the top-left dropdown. All agent outputs appear in the `sandbox` directory, and any datasets or input files must be placed there manually before the agent can use them.

To set up the sandbox without starting the web interface, the README documents a separate command:

bash
python -m karpathy.utils

To start the ADK web interface independently afterward:

bash
adk web

Scientific Agent Skills: Where the ML Knowledge Lives

The capabilities of the karpathy agent come entirely from the Scientific Agent Skills library, a separate repository maintained by K-Dense-AI. The README describes it as a collection of 130+ ready-to-use agent skills covering research, science, engineering, analysis, finance, and writing, built on the open Agent Skills standard. The `start.py` script pulls a subset of those skills into the sandbox so the agent has them available at runtime.

The karpathy agent is designed for ML specifically, so the relevant skills are those for PyTorch model training, scikit-learn workflows, data transformation, and evaluation. The README does not enumerate which specific skills are included in the default sandbox; the full catalog lives in the upstream Scientific Agent Skills repository at github.com/K-Dense-AI/scientific-agent-skills.

This design separates concerns cleanly. The karpathy repository is thin: it provides the agent configuration, the startup script, and the sandbox wiring. The actual ML knowledge sits in the skills library. Updates to the skills, such as a new training loop pattern or a new evaluation helper, propagate to karpathy through the upstream repository without requiring a code change here.

Limitations: Local Only, Manual Data Loading, No Compute Control

Several real constraints follow from the sandbox design. First, the sandbox is local: all model training happens on the machine running the script. The README lists Modal sandbox integration as an upcoming feature, with the ability to choose compute type, but as of the last push on 2026-08-18 that integration is not present in the codebase. Users needing GPU-accelerated training on remote infrastructure must provision that themselves.

Second, the agent cannot load data autonomously. The README states plainly that any files the agent should use must be manually added to the `sandbox` directory. There is no file upload mechanism in the web interface and no automated data fetching.

Third, the project has no GitHub releases. The version in `pyproject.toml` is `0.1.0`, indicating early-stage software. There is no changelog and no migration guide. The `pyproject.toml` pins minimum versions rather than exact ones, so a `uv sync` on a fresh clone may resolve different transitive dependency versions over time.

Fourth, `claude-agent-sdk>=0.2.87` and `google-adk>=2.1.0` are both relatively new packages with their own release cadences. A breaking change upstream, particularly in the Claude Agent SDK, could silently alter agent behavior.

Comparing Karpathy to a Direct LangChain Agent Setup

The most comparable alternative is building a custom agent with LangChain, which also runs locally, supports tool calling, and can invoke Python subprocess calls for ML training. The difference is in what each approach provides out of the box.

With LangChain you write the tool definitions, manage the context window, decide how to structure the agent loop, and integrate ML libraries yourself. The flexibility is high but the initial work is real, particularly for researchers who want to focus on the science rather than the scaffolding.

With karpathy, the scaffolding is pre-built around the Claude Agent SDK and Google ADK, and the Scientific Agent Skills library supplies ML-specific tools without requiring the user to write them. The trade-off is that you are tied to the K-Dense-AI skills library and the OpenRouter inference path. You cannot swap in a locally served model without modifying the agent configuration, and the skills themselves are not under your direct control.

For a researcher who already uses Claude Code and wants fast iteration on ML experiments, karpathy removes the scaffolding burden. For a team that needs to audit every component of the agent stack or run inference on-premises, a LangChain setup with full custom tooling is the more appropriate path.

Maintenance Record and Licensing

The repository is not archived. The last push was on 2026-08-18. The project lists upcoming features: Modal sandbox integration and the possibility of exposing K-Dense Web features in the open-source version. These signals suggest active development intent, but the project has no releases and no version history beyond `0.1.0`. There is no CHANGELOG and no roadmap document beyond the two bullet points in the README.

The disclaimer at the bottom of the README notes explicitly that the project is not endorsed by or affiliated with Andrej Karpathy, which is relevant for anyone who might assume community or organizational backing tied to that name.

The license is MIT. The core karpathy code can be used, modified, and redistributed freely. The Scientific Agent Skills library, which the agent depends on at runtime, is a separate repository with its own licensing terms; anyone deploying karpathy in a production or commercial context should review the upstream skills repository's license independently.

The dependency on `claude-agent-sdk` also carries the Anthropic API usage terms, and the OpenRouter key implies a separate commercial API relationship. Neither of these is part of the MIT license on this repository, so the effective cost and terms of running the agent extend beyond the open-source code itself.

Editorial conclusion

Karpathy suits ML researchers comfortable with Claude Code who want to drive training and experimentation through conversation rather than building agent scaffolding themselves. It is a poor fit for teams needing reproducibility guarantees, remote compute, or a production serving layer: the sandbox is local, temporary, and requires manual data loading. Before adopting it, verify that the OpenRouter model you intend to use is available, since all inference routes through that API and a missing model blocks the agent entirely.

Frequently asked questions

How do you use Karpathy skills?

The skills come from the Scientific Agent Skills library and are installed automatically into the sandbox when you run `python start.py`. You then interact with the agent through the ADK web interface at http://localhost:8000, and the agent dispatches to the relevant skills based on your instructions.

How do you use the Karpathy agent?

Run `python start.py` from the project root after placing a `.env` file with your `OPENROUTER_API_KEY` and `AGENT_MODEL`. The script sets up the sandbox and starts the ADK web interface; navigate to http://localhost:8000, select the karpathy agent, and type your ML instructions. Any files the agent needs must be placed in the `sandbox` directory manually.

Does the Karpathy agent require a paid K-Dense subscription?

The open-source repository runs using an OpenRouter API key, which is a separate paid service with its own pricing. The README notes that a K-Dense web subscription provides substantially more powerful multi-agent capabilities, but the basic open-source version operates independently of a K-Dense account.

Official sources

  1. Issues
  2. K-Dense-AI/karpathy on GitHub
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/k-dense-ai-karpathy.svg)](https://hysenlabs.com/projects/k-dense-ai-karpathy)