CLI tool
SWE-bench/SWE-bench avatar
SWE-bench/SWE-bench

SWE-bench: A Docker-Based Harness for Grading Model-Generated Patches on Real GitHub Issues

SWE-bench: Can Language Models Resolve Real-world Github Issues?

5,946 stars994 forksPythonMIT

At a glance

What is it?
SWE-bench turns real GitHub issues into a patch-generation task and grades the result with containers. The v5 CLI needs a task repo, 120GB of disk and a run_id you never reuse.
Who is it for?
Adopt SWE-bench if you need a reproducible, containerized score on real GitHub issues and can give it an x86_64 host with 120GB of free storage, 16GB of RAM and 8 CPU cores. Do not adopt it as a fast regression suite or on an ARM laptop unless you pass --task-repo and let Docker Buildx build images locally.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The task SWE-bench defines: issue plus codebase, output is a patch

SWE-bench is a benchmark for evaluating large language models on real world software issues collected from GitHub. The framing is narrow on purpose. A model receives a codebase and an issue, and it must produce a patch that resolves the described problem. Nothing about conversation, tool use or repository navigation is specified by the benchmark itself; only the patch is graded.

The intended audience is people who report model scores and people who build the systems that produce them. If you are writing a paper, filling a leaderboard entry, or comparing two agent scaffolds on identical inputs, the harness is the point. If you want a general code assistant, this is not that. The repository also carries the code and data for two papers: SWE-bench at ICLR 2024 and SWE-bench Multimodal at ICLR 2025.

Several dataset variants exist behind one interface. The README lists aliases full, verified, multimodal and multilingual, and says DATASET also accepts a HuggingFace id or a local path, with anything else passed through as given. SWE-bench Verified is described in the news section as a subset of 500 problems that real software engineers have confirmed are solvable, produced with OpenAI Preparedness. Multimodal v2 is described as fully open source with 480 tasks available for local evaluation.

How the harness works: task repo, Docker images, cached results

The mechanism is containerized evaluation. Since June 2024 the project has used Docker for what the README calls more reproducible evaluations. The current v5 CLI builds images from a task repo rather than shipping prebuilt ones, which is why the setup step clones swe-bench-tasks and runs a check against that checkout.

A run produces docker build logs in logs/build_images and evaluation logs in logs/evaluation, and the run summary lands in logs/evaluation/<run_id>/results.json. The report command can re-grade saved logs without starting containers, which is the cheapest way to re-read a result you already have.

The design detail most likely to burn you is caching. The README states that the harness caches results by run_id and instance_id only. Run the same instance with the same run_id twice, even with a different prediction diff, and the harness reuses the cached result from the first run and does not re-evaluate. A new prediction requires a new run_id. This is a deliberate speed choice with a sharp edge: an unchanged run_id silently returns stale grades rather than erroring.

Installing SWE-bench and running a first gold evaluation

Install Docker first, following the Docker setup guide the README links. On Linux the README also points at the post-installation steps. Then build SWE-bench from source. The clone below is the documented path, and the task repo checkout is required for local image builds with the v5 CLI; the README notes the path shown is only an example and can be replaced with any local checkout location.

bash
git clone [email protected]:SWE-bench/SWE-bench.git
cd SWE-bench
pip install -e .

git clone --depth 1 https://github.com/SWE-bench/swe-bench-tasks.git ./swe-bench-tasks
swebench dataset check ./swe-bench-tasks

The dataset check is the gate. If it passes, test the installation with a gold run, which evaluates the reference patch for a single instance instead of a model prediction. The README gives this exact example, using sympy__sympy-20590 and the run id validate-gold.

bash
swebench eval verified --gold \
    -i sympy__sympy-20590 \
    --run-id validate-gold \
    --task-repo ./swe-bench-tasks

On an M-series Mac or another ARM-based system, the README says to use --task-repo so the images are built locally with Docker Buildx. Once that gold run completes, evaluating real predictions is the same command with a predictions path and worker count.

bash
swebench eval verified -p <path_to_predictions> --run-id <run_id> -j <num_workers>

To generate predictions rather than grade them, the README shows swebench infer verified -m gpt-5 -o preds -w 8, which uses mini-SWE-agent. Images can be built or pulled ahead of time with swebench images build verified -j 8, checked with swebench images check multilingual, and leftover containers removed with swebench images clean --run-id <run_id>. The older python -m swebench.harness.run_evaluation form still works and takes the same arguments.

The resource wall, the ARM caveat and the run_id trap

SWE-bench is heavy, and the README says so in a warning rather than a footnote. The recommendation is an x86_64 machine with at least 120GB of free storage, 16GB of RAM and 8 CPU cores, with fewer than min(0.75 * os.cpu_count(), 24) workers. Docker Desktop users are told to raise virtual disk space to roughly 120 free GB. A laptop with a full disk will not complete a full evaluation, and the failure will look like a build error rather than a capacity problem.

Caching is the second failure mode, described above: same run_id plus same instance equals a reused result. The third is platform. The v5 CLI builds images from a task repo, and the README singles out M-series Macs and other ARM systems as needing --task-repo for local Buildx builds. This is not the tool for a quick sanity check on a MacBook Air.

There is also a scope limit worth stating plainly. The benchmark grades a patch against tests. It does not measure whether a patch is maintainable, idiomatic, or something a reviewer would merge. A passing score and a good change are different claims, and the repository does not document a way to evaluate the second one.

sb-cli and Modal: the same evaluation without the local cluster

The most direct alternative to running the harness locally is sb-cli, released in January 2025 and described as the cloud-based evaluation tool for submitting runs to the SWE-bench leaderboards. The difference is where the containers live. Local swebench eval builds and runs images on your machine and leaves logs in logs/; sb-cli moves that step off your hardware and into a submission flow aimed at the leaderboard. If your goal is a published number rather than a private one, sb-cli removes the 120GB requirement and the worker tuning. If you need to inspect build logs or re-grade offline with swebench report, the local path is the one that gives you the artifacts.

The second option is Modal, credited in the news section for running evaluations entirely on the cloud, with a pointer to the evaluation documentation. That is closer to the local harness than sb-cli is, in that you keep control of the run while renting the machines. The trade-off is the same in both cases: you give up direct access to the local logs directory and the ability to re-grade from disk without a network round trip. None of these three paths changes the benchmark itself, only where the Docker images are built and run.

Licence, maintenance and what an upgrade actually costs

The repository is MIT licensed, and pyproject.toml declares the same with license = {file = "LICENSE"}. For most users that means the harness and its code can be reused with attribution, but the datasets are a separate question. The README points at HuggingFace for the multimodal dataset, and the dataset terms live there rather than in this repository, so check the dataset card before redistributing task data. That is a fact about where the terms are, not legal advice.

The last push to the repository was on 2026-09-18, and the repository is not archived. The project is still moving: the news section records Multimodal v2 opening up on September 1, 2026, the sb-cli release in January 2025, the Docker harness migration in June 2024, and SWE-bench Verified in August 2024. pyproject.toml requires Python 3.10 or newer, so a 3.8 environment will not install it despite what the badge in the README suggests.

Upgrade cost is concentrated in the CLI. The v5 CLI builds images from a task repo, and the older python -m swebench.harness.run_evaluation invocation still works with the same arguments. That compatibility note is the thing to lean on during an upgrade. What the README does not document is a rollback path for a run that produced bad images, beyond swebench images clean --run-id <run_id>, so keep the task repo checkout pinned if you need to reproduce an older score.

Editorial conclusion

Adopt SWE-bench if you need a reproducible, containerized score on real GitHub issues and can give it an x86_64 host with 120GB of free storage, 16GB of RAM and 8 CPU cores. Do not adopt it as a fast regression suite or on an ARM laptop unless you pass --task-repo and let Docker Buildx build images locally. Before trusting a number, verify that the task repo check passes, that each evaluation uses a fresh run_id, and that the results.json you are reading came from the run you think it did.

Frequently asked questions

What does SWE-bench mean?

It is the name of a benchmark for evaluating large language models on real world software issues collected from GitHub, where a model is given a codebase and an issue and must output a patch. The repository holds the code and data for the ICLR 2024 paper and the ICLR 2025 Multimodal paper.

What is SWE-bench Verified?

The README describes it as Part 2 of a collaboration with OpenAI Preparedness: a subset of 500 problems that real software engineers have confirmed are solvable. It is available as the verified alias in swebench eval.

What is SWE-bench Lite?

The README does not define a lite alias. It notes that DATASET accepts an alias, a HuggingFace id or a local path, and that anything else is passed through as given, so SWE-bench/SWE-bench_Lite works as a HuggingFace id.

What is SWE-bench Multimodal?

It is a variant integrated into the repository in January 2025, with its own paper and dataset. The news section states that Multimodal v2 is fully open source with 480 tasks available for local evaluation, and multimodal is one of the dataset aliases.

What is SWE-bench Multilingual?

Multilingual is listed among the dataset aliases accepted by swebench eval, alongside full, verified and multimodal. The README gives the example swebench images check multilingual for verifying images exist on the registry, but it does not describe the task contents.

Is SWE-bench open source?

Yes. The repository is MIT licensed, and pyproject.toml declares the same licence file. The task datasets are distributed separately, with the multimodal dataset hosted on HuggingFace.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. SWE-bench/SWE-bench on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/swe-bench-swe-bench.svg)](https://hysenlabs.com/projects/swe-bench-swe-bench)