SWE-bench Pro: the default dataset quietly became a different 642 tasks
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
At a glance
- What is it?
- A long-horizon software engineering benchmark whose second version shrank the split, moved the default dataset configuration, and left the install and usage instructions describing the previous version. The news section also records removed unit tests, a capped earlier results table, and an open leaderboard issue.
- Who is it for?
- SWE-bench Pro is the more honest benchmark of the two lines the page names, and its V2 split is defensible: every task ships with a verifier, a reference solution and a public image, and one command demonstrates that the reference patch resolves all 642. The cost of that rigour is that the numbers are not comparable to the earlier ones without care.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 17 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The default dataset configuration changed from 731 tasks to 642
The single most consequential line in the news section is the one about which configuration is the default.
The overview section tells a reader to load the dataset with no configuration name at all:
from datasets import load_dataset
swebench = load_dataset('ScaleAI/SWE-bench_Pro', split='test')That call returned the original split. It now returns V2, because the news entry says the default configuration is V2 and the original is kept under an explicit name and a version tag.
So the same two lines of code return a different dataset than they did before the September release. The task count went from 731 to 642, which is 89 fewer tasks. The access snippet a reader is most likely to copy is also the one whose behaviour changed.
The V2 section shows all three explicitly:
from datasets import load_dataset
v2 = load_dataset('ScaleAI/SWE-bench_Pro', split='test') # V2, 642 tasks (default config)
hard = load_dataset('ScaleAI/SWE-bench_Pro', 'hard', split='test') # HARD-51
v1 = load_dataset('ScaleAI/SWE-bench_Pro', 'v1', split='test') # original 731 tasksNone of this is hidden, but it is in the news section rather than next to the snippet it affects.
Every install and usage step below the V2 section is for the previous pipeline
One sentence in the V2 section governs everything after it: the evaluation script, the run scripts, the dockerfiles and the container images are the V1 pipeline, kept for reproducing V1 numbers with the V1 configuration.
What follows that sentence is the whole manual. Installing dependencies, installing Docker, configuring Modal or local Docker, finding the correct container image, generating patches, gathering them into a single file, and evaluating them. None of those steps is V2.
V2 has its own instructions in a separate file, referenced but not reproduced, and it uses a different tool entirely: a task directory per instance plus a runner that takes a path, an execution backend, a parallelism figure and an oracle flag.
harbor run -p v2/tasks -e modal -n 50 -a oracle # reference patch resolves 642/642So the shape of the documentation is: a short section describing the current release and pointing elsewhere, followed by several screens of instructions for the superseded one, plus a compatibility sentence explaining which is which.
The default dataset and the default instructions are therefore mismatched by construction. A reader who loads the default split and then follows the usage section is assembling a V1 pipeline to evaluate V2 data.
That oracle line is also the strongest verification claim on the page, since it says the reference patch passes every task in the split.
The recommended path needs a cloud account and the free path is labelled beta
The configuration step is headed as recommended, with local Docker offered as a beta alternative.
The recommended route asks you to run a setup command that generates a token, then tells you to confirm the credentials in a configuration file under your home directory:
modal setup # Follow the prompts to generate your tokenThe alternative is described as needing no additional setup, with a flag to pass when running evaluations. So the comparison is between a hosted execution service that requires provisioning a token against a local container runtime that works out of the box.
Marking the local route beta and the cloud route recommended is a normal choice for a project whose compute needs are bursty, and the parallelism flag in the V2 command supports that reading, since parallel execution is the reason to reach for a hosted backend.
It does mean the recommended path is not the one that runs on your own hardware, which sits oddly against the Docker emphasis elsewhere on the page. Docker is called the mechanism for reproducible evaluations, and then the recommended way to use that Docker is to hand the work somewhere else.
Both paths evaluate the same tasks, so this is about where the containers run rather than what gets measured.
Modal and the Docker SDK are both unconditional requirements, and both are optional
The requirements file has seven entries in three commented groups, and the comments contradict the file's own structure.
The first group is core: a dataframe library, a progress bar and the datasets library. The second is described as a cloud evaluation backend that is optional if using local Docker. The third is a container SDK described as optional if using the cloud backend.
Both of the last two are listed as requirements rather than as extras, so a plain install of the requirements file pulls in a cloud SDK and a container SDK together, and the reader ends up with both even though the page says either one suffices.
The fourth group is a hub client for dataset management, listed unconditionally and correctly so, since the dataset has to be fetched either way.
No entry carries an upper bound. Combined with a core dataframe library floored at 1.5, that is a wide resolution range, and for a project whose entire purpose is producing comparable numbers across systems, the evaluation environment's own version drift is a larger source of variance than most benchmark authors discuss.
The install step for all of this is one line.
Unit tests were removed because some required the year to be 2025
The news section records a change to the benchmark's own tests, and the reason given is more interesting than the change.
Unit tests were removed that were outdated, with the example given being tests that required the year to be 2025, and tests that were not previously intended to be included.
A test that asserts the current year is a test with an expiry date. It passes when the suite is written, keeps passing through that year, and starts failing on the first of January regardless of whether any patch is correct.
For a benchmark, that is a scoring hazard rather than a maintenance nuisance. Every instance carrying such a test would become unsolvable by any model, including the reference patch, on a fixed date. The removal is the right call and the disclosure is welcome, since it changes the difficulty of the split.
The wider pattern is that benchmark validity decays with wall-clock time, and the usual defences are pinning fixtures or freezing the clock. A patch generator can freeze a clock; a benchmark's own test suite generally cannot.
Combined with the separate entry about an earlier results table being updated to remove a cap, the news section describes a project that has had to revise its own measurement instruments twice in a year.
An earlier leaderboard entry admits a cap, and another entry is still open
Two entries in the news section bear on whether a number from this benchmark can be cited.
The first says results were updated to remove a cap, and points at a results page for the uncapped version. That implies earlier published figures were capped, and a capped figure is not the same quantity as an uncapped one.
The second says issues with the leaderboard were identified and are being addressed. That entry sits between the September V2 release and the February test removal, so at some point after the V2 release the project's own record says its published leaderboard had a problem and the fix was still in progress.
Neither statement is retracted later on the page, and no closure notice appears in the entries that follow.
That is worth taking at face value rather than reading around. A benchmark leaderboard is a measurement instrument, and an instrument with a known unresolved issue produces numbers that are provisional regardless of how the tasks themselves were validated. The per-task validation is a separate property from the correctness of the aggregate.
The repository also ships a results page at its root, an error analysis directory, and a trajectories directory, so the intermediate artefacts behind published figures are in the tree. Whether they correspond to the current split is another question the page does not answer.
V1 container images live under a personal Docker Hub account
The container images are split by version in a way that is easy to miss.
For V2, every task has a public image in a GitHub container registry, named by instance id, and the dataset carries a column pointing at it. Pulling needs no login. That is a first-party registry under the project's own organisation.
For V1, described as legacy, prebuilt images are on Docker Hub under an account whose name is not the organisation's. The page links the repository directly.
That matters for two reasons. The V1 images are not under the same publisher as the V2 images, so a reader auditing where images come from has to look at two different accounts. And the example code for finding the right image hardcodes that same personal account name in the string it prints, so anyone copying the snippet inherits it.
The example itself is otherwise good practice. Rather than documenting a naming scheme, it reads the instance id and the image tag from two dataset columns and prints the fully qualified name, which means the snippet keeps working whatever the registry is called.
The V1 section also carries one operational note that would otherwise cost an afternoon: the images run bash by default, so invoking bash again inside them does not do what you would expect. That is filed as an issue reference rather than explained.
Two submodules, a results page at the root, and checked-in trajectories
The top-level layout explains what kind of repository this is.
There are two directories that are third-party agent frameworks, and a submodule configuration file at the root, so both are git submodules. The usage section flags the submodule requirement for the larger of the two and explains that its contents hold instructions for setup, for running against instances, and for configuring model parameters and turn limits.
The second submodule is not mentioned in the usage section at all, even though the news section announces it was added. A clone without submodule initialisation gets an empty directory in both cases, and only one of them is warned about.
Then there is an evaluation script, a helpers directory, run scripts, dockerfiles, an error analysis directory, a trajectories directory, and a results page at the repository root.
Trajectories and error analysis being in the tree is the most interesting of these. A benchmark that publishes the model runs that produced its numbers lets anyone re-read how a patch was generated, which is a stronger reproducibility claim than publishing the scores.
The V2 directory sits alongside all of it as a self-contained subtree with its own readme, its own task directories, a subset id list, a tooling directory and per-task images in a registry. The version boundary is enforced by directory layout rather than by branches, which is why both versions remain readable side by side.
Editorial conclusion
SWE-bench Pro is the more honest benchmark of the two lines the page names, and its V2 split is defensible: every task ships with a verifier, a reference solution and a public image, and one command demonstrates that the reference patch resolves all 642. The cost of that rigour is that the numbers are not comparable to the earlier ones without care. Pin the dataset configuration by name, not by default, and read which pipeline you are running, because the install and usage steps on the page are the previous version's. Treat the earlier leaderboard entries as capped rather than as scores, and check whether the leaderboard issue noted in the news has been closed before citing any of it.
Frequently asked questions
What is in the SWE-bench Pro dataset?
The V2 split is 642 validated tasks across 11 repositories, each shipped as a self-contained task directory with a verifier, a reference solution and a public container image. A HARD-51 subset is provided as a separate configuration, and the original 731-task split is kept under an explicit v1 configuration and tag.
Which pipeline do the SWE-bench Pro usage instructions describe?
The previous one. The page states that the evaluation script, run scripts, dockerfiles and Docker Hub images are the V1 pipeline, kept for reproducing V1 numbers with the v1 configuration. V2 uses a separate task-directory format and its own runner, with instructions in a separate readme.
Does running SWE-bench Pro require a cloud account?
Not strictly. Configuring the hosted execution backend is marked as the recommended path and requires a setup command that generates a token, while local container execution is offered as a beta alternative needing no extra setup and a flag when running evaluations.
Why did SWE-bench Pro remove some unit tests?
Because some were outdated, with the example given being tests that required the year to be 2025, and others were not previously intended to be included. A test that asserts the current year fails on a fixed date regardless of whether a patch is correct.
Where are the SWE-bench Pro container images published?
V2 images are public on a GitHub container registry, referenced by a dataset column, and need no login to pull. The legacy V1 images are on Docker Hub under an account name other than the project's organisation, and the example code that resolves an image hardcodes that name.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/scaleapi-swe-bench-pro-os)