Open-source project
ByteDance-Seed/EdgeBench avatar
ByteDance-Seed/EdgeBench

EdgeBench: Measuring How AI Agents Learn Over Extended Real-World Tasks

EdgeBench: Unveiling scaling laws of learning from real-world environments

455 stars18 forksPythonApache-2.0

At a glance

What is it?
EdgeBench is ByteDance-Seed's benchmark of 134 real-world tasks that tracks how autonomous AI agents improve over up to 12 hours of interaction with executable environments. It differs from one-shot benchmarks by measuring a full trajectory of improvement rather than a single pass/fail score.
Who is it for?
EdgeBench suits researchers and agent framework developers who want to measure not whether an agent solves a task, but how its performance changes as a function of interaction time. It is the wrong tool for teams who need a fast evaluation loop: running at 12 hours per task means a full evaluation against all 134 tasks takes weeks, not minutes.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 11 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What EdgeBench Measures and Who It Is For

EdgeBench addresses a gap in standard agent evaluation. Most benchmarks measure whether an agent can complete a task in a single attempt. EdgeBench places agents in executable task environments and lets them iterate for up to 12 hours, recording the score at 2-hour intervals. The result is a performance trajectory, not a single number.

The dataset covers 134 real-world tasks spread across six categories: Scientific and ML, Systems and SE, Optimization, Knowledge, Formal verification, and Games. ByteDance-Seed has publicly released 51 of those tasks; the remaining 83 are used for the private leaderboard evaluation and are not available for download. The public task set is hosted on HuggingFace at ByteDance-Seed/EdgeBench.

The primary audience is AI systems researchers studying agent learning dynamics and teams building or evaluating coding-capable models who need to understand performance over time rather than in a single shot. It is not suited for product teams running quick evaluations before a release; the time requirement makes it impractical for that use case.

The Log-Sigmoid Scaling Law Finding

The central finding in the EdgeBench technical report at arxiv.org/abs/2607.05155 is that performance across all 134 tasks follows a log-sigmoid scaling law as a function of interaction time. The reported R-squared value for this fit is 0.998, based on roughly 38,000 hours of agent interaction data.

The leaderboard table in the README shows scores for multiple models at 2-hour increments up to 12 hours. For the full 134-task benchmark at 12 hours, Claude Opus 4.8 reaches 51.3, GPT-5.5 reaches 48.4, GPT-5.4 reaches 39.3, GLM-5.1 reaches 37.4, and DS-V4-Pro reaches 31.0. The same models score lower on the 51-task public subset. At the 2-hour mark, scores are substantially lower across all models, illustrating the improvement curve the benchmark is designed to capture.

Category-level scores at 12 hours show uneven difficulty. Systems and SE tasks show the highest scores for most models, while Optimization tasks show the lowest. This structure lets researchers study which types of problems benefit most from extended iteration versus which approach a ceiling quickly.

Installing SForge and Configuring the Evaluation Harness

The evaluation harness is a Python package named sforge, separate from the EdgeBench dataset. The pyproject.toml in the repository identifies the package as sforge version 1.1.0, with a Python version constraint of 3.10 or later and up to 3.13. The package installs the sforge command-line binary.

The default execution backend requires Docker: the harness depends on the docker Python package at version 7.0 or later. Version 1.1.0 of SForge introduced an alternative E2B backend, selected with the flag --backend e2b, which runs tasks inside E2B sandboxes with no Docker or Kubernetes cluster required. Official E2B templates are published under the edgebench namespace on the E2B platform.

Three flags govern the main execution modes. The --agent flag selects the agent scaffold: Claude Code and Codex were built-in from earlier versions, and --agent opencode was added in 1.1.0. The --effort flag accepts low, medium, high, or max and sets reasoning effort uniformly across agent types. The --backend flag defaults to Docker and accepts e2b as the alternative.

Optional dependencies for the e2b backend require the e2b package at version 2.46 or later (below 3), which the pyproject.toml lists as a separate install target. Example configurations are available in the examples/single-task-docker/ directory for Docker-based single-task runs and examples/all-tasks-k8s/ for Kubernetes-based runs across the full task set.

The Six Task Categories and Public Task Structure

The 51 publicly released tasks span the same six categories as the full benchmark, giving researchers a representative subset. Scientific and ML tasks include physics inversion problems, graph node classification, and reinforcement learning. Systems and SE tasks include compiler work, vector search throughput, exchange core performance, and vulnerability analysis. The per-task score tables in the README show scores for each task at each 2-hour checkpoint for all five benchmarked models, so researchers can study individual task curves rather than aggregate numbers.

The README names specific tasks in the public set, including bipedalwalker_locomotion_rl, borden_source_inversion, ann_vector_search_qps, exchange_core_throughput, and juliet_vulnerability_analyzer. Each task has a scoring function embedded in the environment. The agent receives that score as feedback during its iteration. Without intermediate feedback, the log-sigmoid learning curve would not be observable, so the feedback mechanism is structurally central to what EdgeBench measures.

For Formal verification and Games categories, the tasks require logical reasoning and strategic play rather than code generation and debugging. Separating these categories in the reporting lets users see which capabilities improve over extended interaction and which do not show the same scaling behaviour.

Limitations: Undisclosed Tasks, Runtime Dependencies, and Evaluation Cost

The most direct limitation is that 83 of 134 tasks are not publicly released. Researchers who want to reproduce the full leaderboard scores cannot do so independently; they must use the hosted evaluation service at edge-bench.org. This means the full scaling law claim and leaderboard positions cannot be independently verified using only the open-source code and public tasks.

The harness requires Docker or E2B for every run. Teams without container infrastructure must set up that infrastructure before running any evaluation. The Kubernetes example in examples/all-tasks-k8s/ addresses the full-scale case, but it adds significant operational complexity for teams that want to add a new model to the leaderboard.

Each task runs for up to 12 hours. Evaluating the full public set of 51 tasks against a single model requires up to 612 compute-hours if tasks run sequentially. The Kubernetes configuration allows parallelism, but the compute cost of a thorough evaluation is high. EdgeBench is not designed for rapid iteration during model development.

There are no GitHub releases for the EdgeBench repository itself. The sforge Python package is at version 1.1.0 per pyproject.toml, but there is no matching tagged release in the repository.

EdgeBench vs. SWE-bench

SWE-bench is a benchmark that evaluates agents on resolving GitHub issues, measuring what fraction of issues an agent can close in a single attempt. It is the best-known benchmark for coding-capable agents and the nearest comparison to EdgeBench in purpose, but the methodological difference is fundamental.

SWE-bench measures one-shot performance: the agent either resolves the issue or it does not. EdgeBench measures improvement over time: the agent receives feedback from the environment, continues working, and the benchmark records the score at each 2-hour checkpoint. SWE-bench is suited for measuring whether a model can solve a class of task at all. EdgeBench is suited for measuring how a model's ability to learn within a task scales with time.

The two benchmarks are complementary. A team using SWE-bench to compare model capabilities for a product decision is answering a different question than a research team using EdgeBench to study agent learning dynamics. Choosing between them depends entirely on whether the question is about peak capability or about the shape of the performance curve over extended interaction.

Editorial conclusion

EdgeBench suits researchers and agent framework developers who want to measure not whether an agent solves a task, but how its performance changes as a function of interaction time. It is the wrong tool for teams who need a fast evaluation loop: running at 12 hours per task means a full evaluation against all 134 tasks takes weeks, not minutes. Before adopting it, verify that your agent runtime is compatible with Docker or E2B sandboxes, since the SForge harness requires one of those execution environments. The dataset on HuggingFace at ByteDance-Seed/EdgeBench provides the task definitions, and the tech report at arxiv.org/abs/2607.05155 documents the scaling law finding.

Frequently asked questions

What is EdgeBench?

EdgeBench is ByteDance-Seed's benchmark of 134 real-world tasks that measures how autonomous AI agents improve over up to 12 hours of interaction with executable environments. It tracks performance at 2-hour intervals to capture the full learning trajectory, not just a final score.

How many EdgeBench tasks are publicly available?

51 of the 134 tasks are publicly released via the HuggingFace dataset at ByteDance-Seed/EdgeBench. The remaining 83 tasks are used for the private hosted leaderboard at edge-bench.org and are not available for download.

Which agent runtimes does EdgeBench support?

The SForge harness has built-in support for Claude Code, Codex, and OpenCode as agent scaffolds. Tasks run inside Docker containers by default, or inside E2B sandboxes when using the --backend e2b flag added in SForge 1.1.0.

Official sources

  1. ByteDance-Seed/EdgeBench on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bytedance-seed-edgebench.svg)](https://hysenlabs.com/projects/bytedance-seed-edgebench)