Open-source project
Accio-org/CommerceAgentBench avatar
Accio-org/CommerceAgentBench

The harness moves Commerce Agent Bench scores by more than the models do

CommerceAgentBench: Benchmarking Long-Horizon Agents in High-Fidelity, Stateful, and Reproducible Replicas of Real Online Services

1,268 stars70 forksHTMLApache-2.0

At a glance

What is it?
Commerce Agent Bench runs 107 long-horizon commerce tasks in fresh containers against local replicas of real services, and it publishes the same thirteen models under three different harnesses, which turns out to be the most informative thing in the repository. The token columns move by a factor of two for identical models, and the readme says plainly that those columns are telemetry rather than efficiency scores.
Who is it for?
Commerce Agent Bench fits teams choosing between agent models for commerce work who want long-horizon tasks with state rather than question answering, and who will read the harness column as part of the result rather than as background.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 35 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The same model costs twice as much on a different harness

The reference results put the same thirteen model families through the same 107 tasks under three harnesses, and the file says every row compares directly across them. What that shows is that the harness is not a neutral container. Claude Opus 5 passes 65 of 107 tasks on one harness and 60 on another, while its token count moves from 2.05 million to 3.47 million. GPT-5.6 Sol passes 52 tasks on the first and 53 on the second, but its token total goes from 1.15 million to 2.09 million. The largest gap is a model that takes 79 steps on one harness and 138 steps with 11.04 million tokens on the other. The readme states the reason directly: tool granularity, runtime scheduling and provider usage accounting differ, so steps, time and tokens are descriptive telemetry and not normalized efficiency scores.

The published numbers came from endpoints you do not have

Reproducing the leaderboard is not the same as running it. The published scores were produced through evaluation endpoints managed by the maintainers, with one named model acting as the judge, while the public path in the repository expects bring-your-own credentials. The live leaderboard is named as the source of record and the tables in the file as a snapshot, which is the right arrangement but also means the numbers in the repository will drift from the site. There is a second uncontrolled variable. Every model ran with thinking enabled at its provider's default reasoning effort, and since that default differs between vendors, so does the effective harness tier. The readme flags this rather than hiding it, and it means the comparison is between vendor defaults, not between equal reasoning budgets.

The dependency list is two libraries, and the manifest explains why

For a harness that starts containers, drives browsers and shells out to graders, the manifest declares exactly two dependencies, and the comment above them is the most instructive comment in the repository. Verifier graders run on the host, through a shell script invoked by the command line entry point, not inside the workstation image the agent operates in. Two graders import third-party libraries, and if those are missing the grader emits a message about a missing import instead of the real error and the entire task family scores zero on a fresh machine that installed the package and nothing else. Both libraries are therefore pinned for that reason, with the two grader file paths written next to them, and the instruction is to keep the list tight and pin only what a grader uses today.

Apache for the code, Creative Commons for the tasks

The repository carries two licence files, and the split is deliberate and explained in the package metadata rather than left to be discovered. The distributed Python package is the Apache-2.0 half, and the task suite is licensed separately under Creative Commons Attribution 4.0 with its own licence file at the root. The classifier comment says the data directory is not packaged at all, so installing the harness does not redistribute the dataset. That is the correct handling for a benchmark, where the code is software and the tasks are a redistributable corpus that someone may want to reuse in their own evaluation. Around the two licences sit citation metadata, a contributors file, a contributing guide, a security policy and third-party notices, which is the expected set for something meant to be cited.

The previous name is still a directory

The project was previously called RealReplicaBench and was renamed to Commerce Agent Bench to better reflect its intended scope and leave room to expand, and the readme says so in a note near the top. The old name has not gone away, though: a directory carrying it is still at the repository root, beside the current harness directory and the task data directory. That is a reasonable way to avoid breaking anything that imports the old path, and it also means a search of the repository returns both names, which is confusing the first time. The package description carries the same drift in a milder form, naming one harness in a one line summary while the readme compares three.

Twenty two vision tasks against twenty eight browser tasks

The task set is counted twice and both counts total 107. By surface it is 53 command line tasks, 28 browser tasks, 16 file tasks and 10 API or MCP tasks. By capability it is 65 text-only, 20 browser-text-capable and 22 vision-required. Put side by side, the second split does not sit inside the first: 22 vision-required tasks cannot all be browser tasks when there are only 28 browser tasks in total, so vision tasks exist on at least one other surface, and the readme does not say which. That matters for anyone planning a vision-only evaluation, because the task data rather than the documentation is where they would have to find it. Every task runs in a fresh container and is graded by its own verifier, deterministic or language-model assisted depending on the task.

A hundred and four static pages, and the state lives in the runtime

The public gallery lets a visitor browse 104 rendered pages across eight user interface mock services, and it is explicit that this is a static visual tour: state-changing interactions happen inside the benchmark runtime instead. That distinction matters because the three named task surfaces are the stateful ones. Product publishing covers structured catalog and listing operations, freight booking covers multi-step logistics workflows, and storefront operations covers visual configuration and stateful editing. Behind them sit local mock services standing in for SaaS, commerce, messaging, document and operational systems, so an agent operates a real interface and changes real state without anyone needing a production account. What you can audit afterwards is the run record: resolved configuration, trajectory, verifier result, artifacts, logs and container metadata.

Editorial conclusion

Commerce Agent Bench fits teams choosing between agent models for commerce work who want long-horizon tasks with state rather than question answering, and who will read the harness column as part of the result rather than as background. Before quoting any number, check whether it came from the maintainers' evaluation endpoints or from your own credentials, remember that reasoning effort was left at each vendor's default rather than pinned, and read the two licence files, since the code and the task data are under different terms. Anyone wanting a reproducible, effort-controlled leaderboard should wait for that rather than take the snapshot tables at face value.

Frequently asked questions

How many tasks does Commerce Agent Bench contain?

107 tasks. They split into 53 command line, 28 browser, 16 file and 10 API or MCP tasks by surface, and separately into 65 text-only, 20 browser-text-capable and 22 vision-required tasks by capability.

How are Commerce Agent Bench tasks graded?

Each task runs in a fresh container and has its own verifier, which is either deterministic or language-model assisted. Every run preserves the resolved configuration, the trajectory, the verifier result, artifacts, logs and container metadata.

Can I reproduce the published Commerce Agent Bench scores?

Not exactly. The published scores were produced through evaluation endpoints managed by the maintainers with a named judge model, while the public path in the repository uses bring-your-own credentials. The live leaderboard is named as the source of record, and the tables in the file as a snapshot.

Why does Commerce Agent Bench depend on so few libraries?

It declares only openpyxl and PyYAML. A comment in the manifest explains that verifier graders run on the host rather than inside the workstation image, and that without these two libraries two task families fail with a missing-import message and score zero on a fresh install.

Does Commerce Agent Bench need real accounts for the services it tests?

No. Local mock services stand in for SaaS, commerce, messaging, document and operational systems, so tasks run against stateful replicas of those systems without any production account being required.

What was Commerce Agent Bench called before?

RealReplicaBench. It was renamed to reflect its intended scope and allow for future expansion, though a directory carrying the old name is still present at the repository root.

Official sources

  1. Accio-org/CommerceAgentBench on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/accio-org-commerceagentbench.svg)](https://hysenlabs.com/projects/accio-org-commerceagentbench)