Model or dataset
openai/mle-bench avatar
openai/mle-bench

MLE-bench: OpenAI's Kaggle-Based Benchmark for Measuring AI Agent ML Engineering Ability

MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering

1,759 stars260 forksPythonNOASSERTION

At a glance

What is it?
MLE-bench is a benchmark from OpenAI that tests AI agents on 75 Kaggle competitions spanning three difficulty tiers, grading each run against the same medals a human competitor could earn, with a public leaderboard that was paused in April 2026 while the submission process is redesigned.
Who is it for?
MLE-bench is the right tool for researchers evaluating AI agents on real ML engineering tasks. The grading logic, dataset construction code, and agent scaffolding are all open source and runnable.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 159 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What MLE-bench Measures and Why It Differs from Code Benchmarks

Software engineering benchmarks such as SWE-bench test whether an agent can resolve a GitHub issue: read the issue, modify the code, and pass the existing test suite. MLE-bench asks a different question. Given a machine learning competition dataset, can an agent write, debug, and iterate on ML code to produce a submission that scores above a stated threshold?

The benchmark uses 75 Kaggle competitions as its task set. For each competition, the benchmark defines a grading threshold corresponding to the bronze medal standard: a submission that outperforms at least 40 percent of human competitors. Runs are graded against this threshold, not against arbitrary scores.

Three difficulty tiers partition the 75 competitions. Low (also called Lite) holds the most accessible competitions. Medium holds competitions that require more domain knowledge or iteration. High holds the hardest competitions in the set. The leaderboard reports separate scores for each tier as well as an aggregate over all competitions, allowing comparison across difficulty levels rather than only on the aggregate.

The paper accompanying the benchmark is available at arxiv.org/abs/2410.07095. The benchmark was introduced by OpenAI to measure how well AI coding agents handle the full ML engineering cycle, from data loading and preprocessing through model training, hyperparameter search, and submission formatting.

The 75-Competition Dataset and Grading Logic

The 75 competitions in MLE-bench were selected and curated from the Kaggle platform. The benchmark repository includes the code used to construct the dataset, which covers how competitions were selected, how evaluation code was ported, and how grading thresholds were determined.

Each competition in the benchmark has a corresponding grading script that reads the agent's submission file and produces a numeric score. The grading scripts are included in the repository's `mlebench/competitions/` directory, which is packaged as part of the `mlebench` distribution. This means the grading logic is auditable: researchers can inspect exactly how each competition is scored before running an evaluation.

The benchmark's evaluation mechanism requires Docker. Each agent run is isolated in a container to prevent cross-contamination between runs. The `docker>=7.1` dependency in `pyproject.toml` is a hard requirement.

The Kaggle API is used to download competition datasets during evaluation. The `kaggle>=1.6,<1.7` dependency pins this. Kaggle requires authenticated API access and acceptance of competition-specific terms of service before data can be downloaded. These legal requirements mean that automated bulk downloading of all 75 competitions for a first run requires active management.

Running the Benchmark with the mlebench CLI

The project is a Python package named `mlebench`, as defined in `pyproject.toml` with Python 3.11 or newer required. The `pyproject.toml` registers a `mlebench` command-line entry point that maps to `mlebench.cli:main`. This CLI is the primary interface for preparing competition data, running agents, and grading submissions.

The `agents/` directory contains the agent implementations evaluated for the paper. The `run_agent.py` script at the top level is the entry point for running an agent against a competition set. The `examples/` directory includes a `README.md` documenting basic usage patterns and an example for rule violation detection in the `examples/rule_violation_detection/` subdirectory.

Grading reports from previous runs are stored in the `runs/` directory, organized by run group. To produce a leaderboard score from grading reports, the reports must be organized in the `runs/` folder as specified in the README. The benchmark outputs scores for each difficulty tier separately and reports confidence intervals, which is relevant because the 75-competition set is small enough that variance between runs can be significant.

Reading the Leaderboard: Difficulty Tiers and Score Distributions

The leaderboard in the README is the most informative single artifact in the repository for understanding the current state of AI ML engineering capability. As of the data available in this repository, the top entry is Famou-Agent 2.0, which scores 80.3 percent on Low, 64.04 percent on Medium, 42.22 percent on High, and 64.44 percent overall, using Gemini-3-Pro-Preview with a 24-hour budget.

A clear pattern emerges: every agent performs substantially better on Low competitions than on High. For most entries, the High tier score is 20 to 30 percentage points below the Low tier score. This gap reflects the difference between competitions that can be solved by adapting a standard approach and competitions that require domain-specific insight or significant hyperparameter search.

The running time column shows that most entries use a 24-hour budget. Some entries use 12 hours and score lower. The relationship between time budget and score is not strictly linear, but the data consistently shows that agents given 12 hours score below their 24-hour counterparts.

The leaderboard also contains an additional submissions section for entries that are not directly comparable to the main leaderboard, including one entry that used test-set feedback during the run. These are separated to preserve the integrity of the main ranking.

Why the Leaderboard Was Paused in April 2026

The README includes an update dated 2026-04-24: "We are currently not taking any new submissions to the leaderboard while we develop an improved process for ensuring submissions are fair and comparable."

This pause followed the appearance of a submission labeled as using test-set feedback, which the additional submissions section separates from the main leaderboard. The concern is that agents with access to test-set information during a run cannot be directly compared to agents that operate without it.

The pause means that new entries cannot be added to the leaderboard under the current process. Researchers can still run the benchmark and compute their own scores using the provided grading logic, but those scores will not appear on the official leaderboard until the new submission process is defined and announced.

This matters for anyone planning to publish results using MLE-bench as evidence of their agent's capability: results from runs after 2026-04-24 cannot currently be cross-referenced with the published leaderboard entries.

Limitations: Kaggle License Constraints and Narrow Domain Coverage

Kaggle competitions come with individual terms of service that restrict how data can be used. Downloading competition data requires active Kaggle account credentials and acceptance of each competition's terms. This creates a practical barrier to running the full benchmark in fully automated CI environments and limits reproducibility for researchers who do not already have Kaggle access.

The 75 competitions in MLE-bench are drawn entirely from structured data tasks that Kaggle hosts. They do not cover other ML engineering domains such as fine-tuning language models, deploying inference services, writing distributed training code, or working with unstructured code repositories. A high MLE-bench score says specifically that an agent can do Kaggle-style tabular and image ML tasks well, not that it can do ML engineering broadly.

The benchmark also does not measure the quality of the code the agent writes, only the quality of the submission it produces. An agent that produces correct results through brittle or unreadable code scores identically to one that writes clean, maintainable code.

MLE-bench versus SWE-bench and the Scope of Each Benchmark

SWE-bench is the most widely cited software engineering benchmark for AI agents. It tests an agent's ability to resolve real GitHub issues: read the issue, change the code, pass the tests. The tasks are software engineering tasks: fixing bugs, implementing features, refactoring code.

MLE-bench tests ML engineering tasks: writing pipelines, training models, tuning for competition metrics. The two benchmarks cover adjacent but non-overlapping skill sets. An agent that scores well on SWE-bench has demonstrated it can modify existing codebases. An agent that scores well on MLE-bench has demonstrated it can build and iterate on ML pipelines from scratch against a quantitative performance target.

For research teams deciding which benchmark to run, the distinction is in what they want to measure. SWE-bench is the right benchmark for agents designed to help with software maintenance and code improvement. MLE-bench is the right benchmark for agents designed to assist with model development, data science competitions, or applied ML research.

The last push to the MLE-bench repository was on 2026-04-24. The repository uses a NOASSERTION license field, which means the license terms should be verified before redistributing the benchmark or the evaluation code.

Editorial conclusion

MLE-bench is the right tool for researchers evaluating AI agents on real ML engineering tasks. The grading logic, dataset construction code, and agent scaffolding are all open source and runnable. It is not suitable as a continuous integration benchmark: each run requires 12 to 24 hours of compute, Kaggle data licenses restrict automated downloads, and the public leaderboard is currently paused. Before running an evaluation, verify that the Kaggle API credentials are configured and that the target competitions' data can be legally downloaded under your intended use.

Frequently asked questions

What is MLE-bench?

MLE-bench is a benchmark from OpenAI that measures how well AI agents perform at machine learning engineering tasks. It uses 75 Kaggle competitions as its task set, organized into low, medium, and high difficulty tiers, and grades agent submissions against the bronze medal threshold from those competitions.

What difficulty levels does MLE-bench use?

MLE-bench organizes its 75 competitions into three tiers: Low (also called Lite), Medium, and High. The leaderboard reports separate scores for each tier, and most agents score substantially higher on Low than on High.

Is MLE-bench still accepting new leaderboard submissions?

No. As of 2026-04-24, the README states that new submissions to the leaderboard are paused while the team develops an improved process for ensuring submissions are fair and comparable. The benchmark code itself remains available to run locally.

Official sources

  1. Issues
  2. openai/mle-bench on GitHub
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openai-mle-bench.svg)](https://hysenlabs.com/projects/openai-mle-bench)