# EdgeBench measures agents at 12 hours, not at one shot

> ByteDance Seed's EdgeBench puts agents into 134 executable tasks, lets them iterate for 12 hours or more, and fits a log-sigmoid curve to roughly 38,000 hours of interaction. Only 51 tasks are public, the harness ships as a package called sforge, and the two leaderboards do not agree on which model leads the Formal category.

**ByteDance-Seed/EdgeBench** — EdgeBench: Unveiling scaling laws of learning from real-world environments

- Repository: https://github.com/ByteDance-Seed/EdgeBench
- Website: https://edge-bench.org/
- Stars: 455 · Forks: 18
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/bytedance-seed-edgebench

## The header that is commented out and the numbers that are not

The README opens with a block of HTML that has been commented out: the tagline and a line of statistics claiming 134 real-world tasks, 51 open-source, 6 capability categories and 38,000+ hours of agent interaction. None of that renders on the repository page, and all of it reappears further down in the overview. So the headline figures exist twice in the source and once on screen. The substantive numbers are worth having in view: 134 executable tasks, a public release of 51, six categories named Scientific & ML, Systems & SE, Optimization, Knowledge, Formal and Games, and 12 or more hours of interaction per task. Supporting material sits alongside: a dataset on Hugging Face under ByteDance-Seed/EdgeBench, a documentation site, a tech report on arXiv, a Discord invite and a WeChat QR image.

## The full board and the open-source board disagree

Two leaderboards are published, and they are not two views of the same ranking. On the full 134-task set at 12 hours, Claude Opus 4.8 scores 51.3, GPT-5.5 48.4, GPT-5.4 39.3, GLM-5.1 37.4 and DS-V4-Pro 31.0. On the 51-task open-source subset the same order holds, but every score drops: 44.2, 43.1, 34.2, 30.4 and 25.7. The category tables move more than the totals do. Opus leads Formal on the full set at 55.0, while on the subset GPT-5.5 leads that category at 49.0 against Opus's 40.9. A conclusion about category strength therefore depends on which board you read.

## The log-sigmoid fit, and the tasks that stop moving

The central claim is that performance follows a log-sigmoid scaling law as a function of interaction time, with R squared of 0.998, derived from about 38,000 hours of agent interaction. A fit that tight says agents improve predictably rather than erratically. The per-task table shows the other half of that story: curves that stop. GLM-5.1 on `graph_node_classification` moves 49.4 to 52.3 by the four-hour mark and then reports 52.3 unchanged for the remaining eight hours. On `ffmpeg_swscale_reimplementation` the same model goes 0.3, 0.3, 0.4, 2.2, 2.2, 2.2. Extra hours buy nothing once a task stops yielding. Keep one caveat in view while reading the fit: it was computed over all 134 tasks, while the public release is 51, so the headline result rests mostly on runs the community cannot inspect or repeat.

## Empty cells appear as dashes in the per-task table

The per-task table states that missing valid results are shown as a dash, and there are notable gaps. `rust_multicrate_reconstruction` carries dashes for Opus 4.8 at every time budget while four other models report numbers. `dabic_gravity_inversion` starts at 12.7 for DS-V4-Pro with no two-hour value. Those gaps matter for reading an average: a model with missing cells is being compared on whatever subset of tasks it completed. Nothing in the visible documentation explains why a cell is empty, whether the run crashed or was never attempted, or how a leaderboard total treats the difference.

## The repository is EdgeBench, the package is sforge

The Python distribution does not share the repository's name. Its metadata declares:

```toml
[project]
name = "sforge"
version = "1.1.0"
requires-python = ">=3.10,<3.14"
license = { text = "Apache-2.0" }
```

The description calls it an evaluation harness for frontier agents, the build backend is hatchling, and the console script is `sforge`. Dependencies lean on Docker (7.0 or newer), FastAPI, uvicorn, huggingface-hub, rich and tqdm, which is what a harness that drives containers and serves results would need. There is no GitHub release for any of this: version 1.1.0 lives in `pyproject.toml`, and the classifier claims Production/Stable for a research harness with no published artifact. The repository carries no `tests/` directory either, although pytest is a dev dependency.

## One backend flag, three agent scaffolds, one effort dial

SForge 1.1.0 is where the operational choices live. `--backend e2b` runs everything in E2B Sandboxes so no Docker daemon or cluster is needed, with official templates published under the `edgebench` namespace on E2B; the extra is pinned as `e2b>=2.46.0,<3`. `--agent opencode` joins Claude Code and Codex as built-in scaffolds, and `--effort {low,medium,high,max}` sets reasoning effort uniformly across agents, which is the flag that makes a cross-model comparison mean something. Two example directories ship with the repository: `examples/all-tasks-k8s/` for running the set on Kubernetes and `examples/single-task-docker/` for one task locally. The huggingface-hub dependency is what pulls the published dataset into a run, so a local evaluation and a hosted dataset share one input path.

## What 12 hours per task actually costs

The design choice that makes EdgeBench different also makes it expensive. One model on one task is 12 hours of agent interaction, and the full board is 134 tasks across five models, so a complete reproduction is measured in thousands of hours of wall clock and provider spend. That is the reason the public subset exists, and the reason the E2B backend matters, but the trade does not shrink. The visible documentation does not describe the scoring function, the cost of a run, or how failures during a trajectory are handled, so treat the leaderboard as the output of a particular harness rather than a portable measurement of model quality. The licence is Apache-2.0 in both the repository metadata and the package metadata, which is the one piece of housekeeping here that lines up cleanly, and CONTRIBUTING.md and SECURITY.md are both present for a project with no releases.

## Conclusion

EdgeBench fits a team that wants to know how an agent improves with wall-clock time rather than how it does on the first try, and that can afford 12 hours per task per model. Do not use the two leaderboards interchangeably, since the 51-task subset reorders the Formal category and scores every model lower. Verify first that a per-task run exists for your stack, because some rows on the leaderboard are empty, and check the scoring function in the tech report before treating a score as a research result.

## FAQ

### How many tasks does EdgeBench contain and how many are public?

134 real-world tasks across six capability categories, with 51 released publicly along with the full evaluation framework. Each task allows 12 or more hours of interaction, and the analysis behind the scaling law covers about 38,000 hours of agent interaction in total.

### What does the EdgeBench scaling law claim?

That performance follows a log-sigmoid scaling law as a function of interaction time, with R squared of 0.998. The analysis is based on roughly 38,000 hours of agent interaction across all 134 tasks, and the details are in the tech report.

### How do I run EdgeBench tasks without Docker?

Pass --backend e2b to run everything inside E2B Sandboxes, which needs no local Docker daemon or cluster, using official templates published under the edgebench namespace. The package ships that backend as an optional extra pinned to e2b>=2.46.0,<3.

### Which coding agents can EdgeBench evaluate?

Claude Code and Codex are built in as scaffolds, and --agent opencode adds a third. The --effort flag accepts low, medium, high or max and applies the same reasoning effort setting across agents so the comparison is not confounded by defaults.

## Sources

- [ByteDance-Seed/EdgeBench on GitHub](https://github.com/ByteDance-Seed/EdgeBench)
- [Issues](https://github.com/ByteDance-Seed/EdgeBench/issues)
- [License: Apache-2.0](https://github.com/ByteDance-Seed/EdgeBench/blob/main/LICENSE)
- [Project website](https://edge-bench.org/)
- [README](https://github.com/ByteDance-Seed/EdgeBench/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bytedance-seed-edgebench
