SWE-smith: generating unlimited training tasks for software-engineering agents
[NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents
At a glance
- What is it?
- SWE-smith is an MIT-licensed toolkit from the SWE-bench group that turns any GitHub repository into a SWE-gym and creates unlimited program-repair and localization tasks for training SWE-agents. It requires Docker and Linux, and it is research infrastructure, not an end-user coding tool.
- Who is it for?
- Adopt SWE-smith if you build or fine-tune software-engineering agents and need training tasks at a scale curated benchmarks cannot provide, and you can run Docker on Linux. Do not expect it to help your own coding session or to run on Windows or macOS, which the maintainers do not support.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The data problem SWE-smith attacks
Training a model to act as a software-engineering agent needs many realistic tasks: bugs to fix, files to localize, tests that must pass. Curating those by hand does not scale. SWE-smith, from the SWE-bench group and a NeurIPS 2025 Datasets and Benchmarks spotlight, is a toolkit for producing them at volume. It turns any GitHub repository into a SWE-gym, creates unlimited tasks for that repository, spanning file localization, program repair and SWE-bench-style problems, and supports training a language model into a stronger SWE-agent. The audience is researchers and teams building or fine-tuning coding agents who need training data beyond the fixed public benchmarks. This is infrastructure for making agents, not an agent you run on your own code, and the README is clear about that framing.
From a repository to task instances
The mechanism is a pipeline that manufactures tasks from a repo. The README lays out the sequence: create an execution environment for the repository, synthesize task instances, keep the tasks that break one or more unit tests, and generate the issue text that describes each task. The unit-test signal is the core idea, a synthesized change that breaks tests is a candidate repair task with an automatic correctness check built in. The output is a dataset of task instances, and the project publishes a 52,000-instance dataset built this way, plus 250-plus prebuilt environment images, one Docker image per represented repository. That volume is the point: rather than a fixed benchmark, SWE-smith is a factory that can keep producing tasks from new repositories as long as their tests can be run.
Loading the dataset and training
For consumers, the fastest entry is the published dataset, and the README shows loading it and pulling a Docker container per task:
from swesmith.profiles import registry
from datasets import load_dataset
ds = load_dataset("SWE-bench/SWE-smith", split="train") # Loads all 52k task instances
for task in ds:
rp = registry.get_from_inst(task) # Get the RepoProfile for the task
container = rp.get_container(task) # Docker container with the task initializedEach task resolves to a RepoProfile and a Docker container with the task set up, which is why Docker is a hard requirement. To build environments from your own repositories rather than use the dataset, you install the package from source per the project's installation guide, then run the environment-construction and instance-synthesis steps. The README states SWE-smith was developed and tested on Ubuntu 22.04 and does not plan to support Windows or macOS.
Real constraints: Linux, Docker, and research scope
The limitations are stated plainly and matter. SWE-smith requires Docker to create execution environments, and the maintainers explicitly do not plan to support Windows or macOS, so it is a Linux tool in practice. It is research infrastructure for producing training data and training agents, not a tool that improves your own coding session, so the audience is narrow. The generated tasks are validated by whether they break unit tests, which is a strong automatic signal but bounded by the quality and coverage of the repository's tests. And the scale that makes it useful, 52,000 instances and hundreds of Docker images, is also a resource commitment: building and running environments at that volume needs real compute and disk. These are the right trade-offs for its purpose, but they rule it out as a lightweight utility.
SWE-smith versus SWE-bench and hand-built task sets
The natural comparisons are the group's own SWE-bench and manually curated task sets. SWE-bench is an evaluation benchmark: a fixed, curated set of real issues used to measure an agent, not to generate training data at will. Hand-built task sets give you control over every example at the cost of the labor to create them and a ceiling on how many you can make. SWE-smith's difference is generation: it scales task creation from arbitrary repositories, with an automatic test-based validity check, so you can train on far more and more varied tasks than a curated set provides. The README reports it was used to fine-tune a 32B coder model with a large jump on SWE-bench Verified, which is the intended loop, generate with SWE-smith, evaluate on SWE-bench. Use SWE-bench to measure, SWE-smith to produce the data you train on.
MIT license and project standing
SWE-smith is MIT-licensed, so the toolkit and its outputs are freely reusable under that license, though the repositories it builds environments from carry their own licenses you must respect. The last push was on 2026-09-14, and the project has the standing of a peer-reviewed release: a NeurIPS 2025 spotlight, a documented dataset and trajectories on Hugging Face, and a trained model card. That academic grounding is a reason to trust the method's description, and the committed dataset and environment images let you start without building anything. Treat it as active research infrastructure: begin from the published 52k dataset to learn the shape of the tasks, and move to source installation and environment construction only when you need tasks from repositories the dataset does not cover.
Editorial conclusion
Adopt SWE-smith if you build or fine-tune software-engineering agents and need training tasks at a scale curated benchmarks cannot provide, and you can run Docker on Linux. Do not expect it to help your own coding session or to run on Windows or macOS, which the maintainers do not support. Start from the published SWE-bench/SWE-smith dataset with load_dataset to understand the task format, and install from source to construct environments from your own repositories only when you need them.
Frequently asked questions
What is SWE-smith?
SWE-smith is an MIT-licensed toolkit from the SWE-bench group that turns any GitHub repository into a SWE-gym and generates unlimited tasks, such as program repair and file localization, for training software-engineering agents.
What does SWE-smith require to run?
Docker, to create per-task execution environments, and Linux in practice: it was developed and tested on Ubuntu 22.04, and the maintainers do not plan to support Windows or macOS.
How is it different from SWE-bench?
SWE-bench is a fixed evaluation benchmark used to measure an agent. SWE-smith generates training tasks at scale from arbitrary repositories, validated by whether a change breaks unit tests, so you train on data it produces and evaluate on SWE-bench.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/swe-bench-swe-smith)