# AIDE ML: the reference implementation of a tree-search ML engineering agent

> AIDE ML is Weco AI's open source reference build of the AIDE algorithm, a Python agent that writes and benchmarks machine learning code against a metric you describe in plain English. It is a research package, not a managed service, and the README is honest about that split.

**WecoAI/aideml** — AIDE: an LLM agent for machine learning engineering - the research Weco grew out of. Referenced in OpenAI MLE-bench.

- Repository: https://github.com/WecoAI/aideml
- Website: https://weco.ai
- Stars: 1,553 · Forks: 235
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/wecoai-aideml

## The problem AIDE ML addresses: metric-driven code search, not model training

Most ML tooling assumes you already know which model to fit. AIDE ML starts one step earlier. You hand it a directory of data, a goal sentence and an evaluation metric, and it writes candidate Python scripts, runs them, reads the resulting score, and keeps the branches that score well. The README frames the package as the open-source reference build of the AIDE algorithm, with three layers named explicitly: the algorithm described in the paper, this repository as a lean implementation for experimentation, and the Weco product for teams that want tracking and broader code-optimization scope. That layering is the clearest signal of intended audience. Agent-architecture researchers are told they can swap search heuristics, evaluators or LLM back-ends. ML practitioners are told they can build a pipeline quickly given a dataset. Nobody is told this is a deployment target. The interface is deliberately thin: no YAML grid, no wrapper class around your model. The README gives the example `aide data_dir=… goal="Predict churn" eval="AUROC"`, which is the whole contract.

## How the agentic tree search works, and what each node holds

The mechanism is a tree over code. Each generated Python script becomes a node in a solution tree. The LLM produces patches that spawn children from an existing node, and the metric returned by running a child decides whether that branch is kept and expanded or pruned. The README describes this as iterative agentic tree search and ships an HTML visualiser so you can inspect the full tree and the code attached to each node. Two configuration keys control the shape of that search: `agent.steps` sets the number of improvement iterations and defaults to 20, while `agent.search.num_drafts` sets drafts per step and defaults to 5. The coding model is chosen with `agent.code.model` and defaults to `gpt-4-turbo`. The README states the plumbing is model-neutral across OpenAI, Anthropic, Gemini, or any local LLM that speaks the OpenAI API, and requirements.txt carries both `openai>=1.69.0` and `anthropic>=0.20.0`. The claim that motivates the tree structure comes from outside the repository: the README reports that OpenAI's MLE-bench, covering 75 Kaggle competitions, found AIDE's tree search wins four times more medals than the best linear agent, OpenHands. That is a cited result, not something this package measures for you.

## Installing AIDE ML and running a first optimisation

The package requires Python 3.10 or later, per setup.py and the README badge. Install it from PyPI, then export an OpenAI key, then run the CLI against one of the bundled example tasks. The README's quick start is three commands, and the third one is the actual optimisation.

```bash
pip install -U aideml
export OPENAI_API_KEY=<your-key>
aide data_dir="example_tasks/house_prices" \
     goal="Predict the sales price for each house" \
     eval="RMSE between log-prices"
```

When the run finishes, the README says you will find `logs/<id>/best_solution.py` with the best code found and `logs/<id>/tree_plot.html`, which you open in a browser to inspect the solution tree. If you would rather drive it from Python, the README gives an `aide.Experiment` object with `data_dir`, `goal` and `eval` arguments and a `run(steps=...)` method that returns a best solution exposing `valid_metric` and `code`.

```python
import aide

exp = aide.Experiment(
    data_dir="example_tasks/bitcoin_price",
    goal="Build a time series forecasting model for bitcoin close price.",
    eval="RMSLE"
)
best_solution = exp.run(steps=2)
print(best_solution.valid_metric)
```

There is also a Streamlit UI, but it is not in the installed package. The README instructs you to clone the repository, run `pip install -e .` to pull in Streamlit, then `cd aide/webui` and `streamlit run app.py`. In the sidebar you paste an API key, upload data, set Goal and Metric, and press Run AIDE. The repository also ships a Dockerfile and a Makefile whose `docker-run` target mounts logs, workspaces and the example tasks directory and passes `OPENAI_API_KEY` through as an environment variable.

## Where AIDE ML stops: cost, metric brittleness and the missing production path

Every step in the search is an LLM call that produces code, followed by a real execution of that code. The defaults of 20 steps and 5 drafts per step describe a search that generates on the order of a hundred candidate scripts, each one run against your dataset. The README does not publish a token budget or a wall-clock estimate, and the repository does not ship a cost calculator, so the practical cost is whatever your provider charges for that volume plus the compute to execute the scripts. That is a real constraint for anyone pointing this at a large dataset. The second limitation is the metric interface. `eval` is a single string, and the agent has to turn it into something it can compute and compare. A metric that needs custom aggregation, a held-out split you control, or a business rule that is not expressible in one sentence is a poor fit. Third, this is the reference build by design. The README routes production users to the Weco platform for experiment tracking and enhanced user control, which means the open repository does not promise those features. If you need run comparison across teams, audit trails or a support contract, the package is the wrong layer. Finally, requirements.txt pins a wide and heavy set of agent-side packages, including torch, torchvision, torchaudio and torchtext, so the install footprint is not small even when your own task is tabular.

## AIDE ML compared with a fixed AutoML pipeline

The nearest thing to compare against is a conventional AutoML library that searches a predefined space of models and hyperparameters. The difference is what gets searched. A conventional pipeline searches over estimator choices and numeric ranges that its authors wrote down in advance; the space is bounded by the library's catalogue, and the output is a fitted model plus a score. AIDE ML searches over Python source. The candidate space is whatever the LLM can write, which includes feature engineering, preprocessing and validation logic, not just model selection. That is why the README describes each script as a node rather than each configuration as a point. The trade-off runs the other way too. A fixed AutoML pipeline is deterministic given a seed, its runtime is predictable, and its search space is documented. AIDE ML's search depends on the model you configure and on the code the model happens to generate, so two runs are not the same experiment. The README's own framing supports this reading: it calls the repository lean and aimed at experimentation and extension, and points at the Weco product for the broader, tracked version. For a tabular problem where gradient boosting is obviously the answer, a fixed pipeline is cheaper and more reproducible. AIDE ML earns its cost when the pipeline itself is the unknown.

## Licence, maintenance and the cost of upgrading

The repository is MIT licensed, and the README badge and setup.py both carry that. MIT is permissive: you can use, modify and redistribute the code, including commercially, provided the copyright notice and permission notice are preserved. That is the licence text, not legal advice, and if you are redistributing the package inside a product you should read the LICENSE file at the repository root rather than take this summary as sufficient. On maintenance, the last push to the default branch was on 2026-09-03, and the most recent release listed is v0.2.2 from 2025-11-05, with v0.2.0 before it in January 2025 and v0.1.4 in April 2024. The gap between releases is measured in months, so treat upgrades as occasional events rather than a stream. The upgrade cost that matters most is the dependency set in requirements.txt. It pins exact versions for numpy, pandas, scikit-learn, scipy and others, and pulls torch and a long tail of agent-side libraries. Any environment that already has its own pinned versions of those packages will conflict, so the Dockerfile, which builds the environment in a separate stage and copies a virtualenv into a slim runtime image running as a non-root user, is the cleaner isolation path than installing into an existing environment.

## Conclusion

Adopt AIDE ML if you are an agent-architecture researcher or an ML practitioner who wants a inspectable tree of candidate scripts and is willing to supply an LLM key and a dataset directory. Do not adopt it if you need experiment tracking, multi-user control or a supported production path; the README points those users at the hosted Weco product instead. Before committing, verify that your metric can be expressed as a single string the agent can parse, that your data fits the data_dir layout the examples use, and that you are comfortable with the pinned dependency set in requirements.txt, which includes torch and a long list of agent-side packages.

## FAQ

### What is AIDE ML and what does it do?

AIDE ML is the open-source reference build of the AIDE algorithm, an LLM-driven agent that writes, evaluates and improves machine learning code. It runs an agentic tree search over generated Python scripts, using your evaluation metric to prune and guide the search.

### How do I install AIDE ML?

Install it from PyPI with pip install -U aideml, which requires Python 3.10 or later. The README then has you export an LLM key such as OPENAI_API_KEY before running the aide command.

### Which LLM models can AIDE ML use for writing code?

The README states the plumbing is model-neutral across OpenAI, Anthropic, Gemini, or any local LLM that speaks the OpenAI API. The coding model is selected with the agent.code.model flag, which defaults to gpt-4-turbo.

### Where does AIDE ML put the best code it finds?

After a run finishes, the README says the best code is written to logs/<id>/best_solution.py and the solution tree is rendered to logs/<id>/tree_plot.html.

## Sources

- [License: MIT](https://github.com/WecoAI/aideml/blob/main/LICENSE)
- [Project website](https://weco.ai)
- [README](https://github.com/WecoAI/aideml/blob/main/README.md)
- [Releases](https://github.com/WecoAI/aideml/releases)
- [WecoAI/aideml on GitHub](https://github.com/WecoAI/aideml)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/wecoai-aideml
