# DeepSWE: A Coding Agent Benchmark Built on Original, Multi-Language Engineering Tasks

> DeepSWE is a benchmark from Datacurve that measures coding agents on 113 original software engineering tasks drawn from active open-source repositories, covering TypeScript, Go, Python, JavaScript, and Rust. Each task runs in an isolated Docker environment and is graded by a program-based verifier.

**datacurve-ai/deep-swe** — Measuring frontier coding agents on original, long-horizon engineering tasks

- Repository: https://github.com/datacurve-ai/deep-swe
- Website: https://deepswe.datacurve.ai/
- Stars: 1,781 · Forks: 118
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/datacurve-ai-deep-swe

## What DeepSWE Measures and Who Uses It

Most coding agent benchmarks draw tasks from GitHub issue histories or synthetic problem sets. DeepSWE takes a different approach: all 113 tasks are original, long-horizon engineering problems drawn from active open-source repositories, written by humans rather than generated automatically. The README describes them as tasks that require understanding a real codebase and making changes that are verified by real tests, rather than tasks extracted after the fact from resolved pull requests.

The benchmark covers five languages: TypeScript, Go, Python, JavaScript, and Rust. This multi-language scope distinguishes DeepSWE from benchmarks that focus exclusively on Python repositories. Tasks requiring different toolchains, build systems, and testing frameworks stress an agent's ability to work across ecosystems, which single-language benchmarks cannot measure.

The primary audience is researchers and engineers building or evaluating frontier coding agents who want a benchmark closer to real engineering work than synthetic or historical GitHub issues provide. The PROVENANCE.md file in the repository documents how each task in the corpus was sourced and validated by the benchmark authors.

## The Harbor Task Format and Isolated Environments

Each task in DeepSWE uses the Harbor task format, which organises task files into a consistent directory structure:

```text
task.toml         Metadata (repo, base commit, language, image, limits)
instruction.md    The prompt the agent sees
environment/      Dockerfile reproducing the prebuilt image
tests/            Verifier entry point, held-out tests, and grader config
solution/         Reference solution (held out from the agent)
```

The `task.toml` specifies the repository, the base commit, the Docker image, and resource limits for the run. This makes the environment fully reproducible: anyone with the task.toml can rebuild the exact environment the agent sees. The `instruction.md` is the only file the agent sees during evaluation; the tests, grader configuration, and reference solution are held out.

The verifier checks observable behaviour, not symbol names or internal structure. A solution passes if it produces the correct output under the held-out tests, regardless of how the agent implemented it. This design allows agents to refactor, rename, or restructure code as long as the externally observable test results match. The reference patch in `solution/` exists so reviewers can spot-check the correctness of each task offline, but it plays no role in automated grading.

## Installing Pier and Running the Benchmark

DeepSWE runs through Pier, a command-line tool from Datacurve that drives the Harbor evaluation framework. Install it with:

```bash
uv tool install datacurve-pier
```

Clone the benchmark repository first:

```bash
git clone https://github.com/datacurve-ai/deep-swe
```

To run the full benchmark with Claude Opus 4.8 as the model:

```bash
export ANTHROPIC_API_KEY=...
pier run -p deep-swe/tasks --agent mini-swe-agent --model anthropic/claude-opus-4-8
```

To run with GPT-5.5:

```bash
export OPENAI_API_KEY=...
pier run -p deep-swe/tasks --agent mini-swe-agent --model openai/gpt-5.5
```

To run a reproducible random subset of 10 tasks:

```bash
pier run -p deep-swe/tasks --agent mini-swe-agent --n-tasks 10 --sample-seed 0
```

To run a single task by its ID:

```bash
pier run -p deep-swe/tasks/<task-id> --agent mini-swe-agent
```

Pier also drives other agents directly: `claude-code`, `codex`, `gemini-cli`, and `opencode`. The `--env modal` flag runs tasks in parallel on Modal rather than locally.

## How the Verifier Scores Agent Runs

After each task run, the verifier produces output files in a `verifier/` subdirectory:

```text
verifier/
    reward.json      Structured scores (binary reward + pass fractions)
    ctrf.json        Machine-readable test report with failure messages
    test-stdout.txt  Raw suite output and a list of failure reasons
    run.log          Raw stdout/stderr captured during the run
    reports/         Framework-native report/log files from the grader
```

The `reward.json` contains the binary reward (pass or fail) and per-test pass fractions. The `ctrf.json` is a machine-readable test report compatible with the Common Test Results Format, which is used for aggregating results across runs and integrating with CI systems.

Since v1.1, grading uses Harbor's separate verifier environment. This requires Pier newer than version 0.3.0. The agent works in an isolated environment and commits its work upon completion. A `[[verifier.collect]]` hook in each `task.toml` then extracts these commits as a patch, which is applied and graded in a pristine container. This separation between the agent environment and the verifier environment ensures the grader cannot be influenced by files the agent left in its working directory.

## Pier vs. Harbor: What Pier Adds for Air-Gapped Evaluations

Harbor is the upstream framework for sandboxed coding agent evaluations, developed by the harborframework.com project. Pier began as a fork of Harbor to address a specific gap: Harbor blocks all outbound traffic in tasks configured with `allow_internet = false`. This network isolation blocks not only the task environment but also the dependency installation steps an agent might need, and it blocks the LLM API calls the agent itself makes.

Pier solves this by adding per-agent network allowlists. Each agent is granted only the network access it needs, while the task environment remains isolated from arbitrary outbound traffic. This design lets an agent call its LLM provider and install packages from a known registry without opening the task sandbox to all outbound connections.

Additional features Pier adds beyond Harbor: more complete trajectory metadata captured during each run, a built-in trajectory viewer for inspecting the full sequence of actions an agent took, and `pier critique run` for structured post-hoc analysis of agent trajectories. These tools help benchmark operators understand not just whether an agent solved a task but how it approached the problem and where it failed. All leaderboard scores on the DeepSWE website were produced using Pier running `mini-swe-agent` on Modal infrastructure.

## Scope Limits and Comparison with SWE-bench

DeepSWE has 113 tasks. SWE-bench, the most widely used coding agent benchmark, draws from GitHub issues across a larger number of Python repositories and has a verified variant (SWE-bench Verified) that uses human-validated issues. DeepSWE's smaller task count means confidence intervals on aggregate scores are wider. A difference of a few percentage points between two agents may not be statistically reliable at 113 tasks, whereas SWE-bench's larger corpus supports finer-grained comparisons.

DeepSWE's strength relative to SWE-bench is multi-language coverage and the use of original tasks rather than resolved GitHub issues. Resolved issues can appear in LLM training data if the agent's base model was trained on code repositories. Original tasks written specifically for the benchmark after the training cutoff avoid that contamination concern.

The repository includes a `tasks/` directory at the top level containing all task subdirectories. The `.gitignore` and `LICENSE` are the only other top-level files besides the tasks directory, README, and PROVENANCE.md. The README does not describe how tasks are distributed across languages, so it is not possible to determine from the README alone how many tasks fall in each language category. Operators wanting language-specific analysis would need to inspect the task.toml files directly or parse the PROVENANCE.md document, which the README mentions as the source of record for task provenance.

The last push to the main branch was on 2026-08-26. The repository has no GitHub releases.

## Conclusion

DeepSWE is most useful for teams building or evaluating coding agents who want a benchmark with original multi-language tasks and reproducible Docker-based grading. It is not a general-purpose development tool and requires Pier 0.3.0 or later. Before running a full evaluation, verify that your agent supports the Pier CLI interface and that the tasks directory for your target language subset includes enough tasks to produce a statistically meaningful sample. The leaderboard on the DeepSWE website shows scores for several frontier models and provides a baseline for comparison before you invest compute in a full run.

## FAQ

### What is DeepSWE?

DeepSWE is a benchmark for measuring coding agents on original, long-horizon software engineering tasks. It contains 113 tasks across TypeScript, Go, Python, JavaScript, and Rust, drawn from active open-source repositories. Each task runs in an isolated Docker environment and is graded by a program-based verifier that checks observable behaviour.

### What is the DeepSWE benchmark?

DeepSWE is a coding agent evaluation benchmark created by Datacurve. It uses the Harbor task format with isolated environments and held-out verifiers. Tasks are run using the Pier CLI. The benchmark scores agents on binary pass/fail reward and per-test pass fractions, with results exported in CTRF format.

### How do I run DeepSWE tasks against my own agent?

Install `datacurve-pier` with `uv tool install datacurve-pier`, clone the deep-swe repository, and use `pier run -p deep-swe/tasks --agent <your-agent> --model <model-id>`. Pier supports the `mini-swe-agent` built in, as well as `claude-code`, `codex`, `gemini-cli`, and `opencode`. Use `--n-tasks` and `--sample-seed` for reproducible subsets.

## Sources

- [datacurve-ai/deep-swe on GitHub](https://github.com/datacurve-ai/deep-swe)
- [Issues](https://github.com/datacurve-ai/deep-swe/issues)
- [License: Apache-2.0](https://github.com/datacurve-ai/deep-swe/blob/main/LICENSE)
- [Project website](https://deepswe.datacurve.ai/)
- [README](https://github.com/datacurve-ai/deep-swe/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/datacurve-ai-deep-swe
