Browser Arena measures four steps, and two of them are API calls
The Browser Arena. Comparing how cloud browser providers perform.
At a glance
- What is it?
- An open benchmark that times the full lifecycle of a cloud browser session across nine providers from fixed EC2 instances. Its scoring method, fixed anchors on reliability, latency and cost, is the part worth arguing with, because min-max scaling zeroes a provider for a 1.7 point gap.
- Who is it for?
- Browser Arena is worth reading if you are choosing between cloud browser providers and want one test applied identically to all of them, because the connect and goto numbers isolate browser behaviour from API design in a way a vendor benchmark cannot.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Four steps, two of which are API calls
The whole benchmark is one lifecycle, run identically against every provider:
Create session -> Connect via CDP -> Navigate page -> Release session
(API) (Playwright) (goto) (API)Written out as steps, the first action is creating a session through the provider's own API, the second is connecting Playwright over the Chrome DevTools Protocol, the third is navigating to a URL and waiting for the domcontentloaded event, and the fourth is releasing the session through the API again. Results are reported as medians across all successful runs rather than as averages, which is the right choice when a provider has a tail of slow sessions. There are two modes: sequential, which measures 100 sessions per provider one at a time, and concurrent, which measures 100 sessions per provider in batches of 10 parallel sessions. The concurrent mode is the one that answers whether a provider holds up under parallel load rather than whether it is fast when alone.
Two EC2 boxes, two Node versions, three documented minimums
Comparability comes from pinning the runners, and the table of test environments is small enough to memorise. The us-east-1 box is a t3.micro running linux x64 on Node v20.20.0, and the us-west-1 box is a t3.micro on linux x64 with Node v18.20.8. That second row is the interesting one, because Node 18 sits below the engines requirement in package.json, which asks for node greater than or equal to 20. The README meanwhile says the project requires Node.js greater than or equal to 18, and the container image starts from node:26-slim. So the repository states three different minimums for the same interpreter, and the region that runs Browserbase and Notte is the one on an older runtime than the manifest permits. Nobody is measuring browser performance here, so it does not corrupt the results, but it tells you how much of the documentation is maintained separately from the code.
Warm-up runs, one URL, and a 30 second pause on 429
The fairness section lists five rules, and each one closes a different hole. Ten warm-up runs happen before measurement so that cold start effects do not land inside the sample. The same URL is used for every provider, example.com, which removes any argument about page weight or origin speed. No provider specific tuning is applied and only default SDK settings are used, so nobody gets a hand-tuned client. Rate limits get an automatic 30 second backoff on a 429 error rather than a failed run. And the last rule is the one with the most weight, because most provider SDKs auto-retry transient errors, so the reported success rates are post-retry outcomes rather than first attempt outcomes. That is a defensible choice for a benchmark about what you will experience in production, and an unhelpful one if what you wanted to know was which provider has the flakiest network, because the two questions are silently merged into one number.
Round trip times already differ by six times
Before any measurement happens, the project records TCP and TLS round trip times with curl time_appconnect, taken as the median of ten runs, and publishes them per provider. Hyperbrowser at connect-us-east-1.hyperbrowser.ai comes in at 9 ms, Notte at 12 ms with its host redacted, Steel at connect.steel.dev and Kernel at api.onkernel.com both at 14 ms, Browser Use at cdp.browser-use.com at 29 ms, Anchor Browser at connect.anchorbrowser.io at 38 ms, and Browserbase at connect.usw2.browserbase.com at 62 ms. The runner column shows which EC2 box made each measurement. So a provider at 62 ms is being compared against one at 9 ms before a browser exists, which is a fact the leaderboard cannot remove. It also explains the placement of providers in the benchmark list: Notte and Browserbase run from us-west-1, Steel, Kernel, Hyperbrowser, Anchor Browser and Browser Use run from us-east-1, and Lightpanda runs from us-west-1.
Create and release measure the API, not the browser
The benchmark is candid about which of its own numbers mean what, and the admission appears in a section on what the stages actually measure. The create and release timings reflect provider API design, specifically whether session management is synchronous or asynchronous, more than they reflect browser speed. A provider whose API returns a session synchronously will look slower on create than one that queues the session and hands back a handle, even if both start a browser at the same moment. The connect plus goto timings are named as the best proxy for actual browser performance, since that pair measures how long it takes to get a real Chrome running and load a page. Read the leaderboard with that split in mind. A provider that wins on connect and goto and loses on create is not slower, it has a different API shape. The regression alert example in the same spirit shows a Notte session release going from 26 ms to 1.07 s, a 41.2 times change, which is an API layer regression rather than a browser one.
Fixed anchors instead of min-max, on purpose
Each provider gets a single score from 0 to 1 combining reliability, latency and cost, weighted equally at 33/33/33 by default, with presets such as Speed first or Budget first available on the leaderboard. Each dimension is normalised against fixed anchors rather than against the observed data. Reliability runs from 90 percent, which the README calls unusable in production, to 100 percent. Latency runs from 10,000 ms, described as a practical timeout threshold, down to 0 ms. Cost runs from 0.20 dollars an hour, about twice the most expensive current provider, down to zero. The final figure is a weighted sum of the three normalised values. Two alternatives are named and rejected: ratio to best, where a 6.8 times latency range against a 2.4 times cost range lets cost dominate at equal weights, and min-max on observed data, where 98.3 percent reliability becomes 0.00 when everyone else is at 100 percent.
Nine providers, nine keys, one cron line each
Adding a provider means adding a key and a row. The .env.example lists one variable per provider, including a project id for Browserbase, with Tilion marked as beta and its key requested from tilion.dev, and Lightpanda carrying a key plus a commented CDP URL because its sessions are created on CDP connect and it defaults to us-west. Daily runs are automated by scripts/run-and-publish.sh, which benchmarks one region's providers at concurrency 1 and 10, stages only that region's results directory and pushes to main. Two EC2 boxes share that script and differ only by a region flag:
0 3 * * * /home/ubuntu/browserarena/scripts/run-and-publish.sh --region us-east
0 4 * * * /home/ubuntu/browserarena/scripts/run-and-publish.sh --region us-westThe west job is staggered an hour behind the east job to avoid push races, since both write to the same branch. Each box clones the repository, gets the keys for its own providers in .env, and is given an SSH deploy key with write access, one key per box, with cron output going to a gitignored logs directory.
A health endpoint that answers 503 on regression
The alerting story is unusually complete for a benchmark project. A GET to https://www.browserarena.ai/api/health reports whether a provider has regressed against its own 28 day baseline, returns 200 when healthy, and returns 503 when degraded or when results have gone stale. Because the failure signal is an HTTP status code, an uptime monitor can watch it with no custom code, and Better Stack, Checkly or Cronitor are named as candidates. The endpoint is meant to be curled, not read:
$ curl -s https://www.browserarena.ai/api/health | jq -r .summary
Notte session release 26ms -> 1.07s (41.2x) since 2026-09-05Reproducing the numbers yourself takes a clone, an install, a copied .env and one command, and the results are queryable locally with DuckDB against the SQL files in queries/. A Dockerfile wraps the same entry point with the benchmark, run count, provider and concurrency exposed as environment variables, and railway.json is there for running it as a service. The background is worth knowing too: Steel built the first version of this benchmark as steel-dev/browserbench and tested five providers, and this project extends it with more providers, the scoring system, the public leaderboard and a structure for adding more benchmarks.
Editorial conclusion
Browser Arena is worth reading if you are choosing between cloud browser providers and want one test applied identically to all of them, because the connect and goto numbers isolate browser behaviour from API design in a way a vendor benchmark cannot. It is less useful if you need a latency comparison that ignores where the provider is hosted, since round trip times already differ by more than six times before a session starts and two providers run from the wrong region for their endpoints. Before trusting a ranking, read the anchors rather than the order, check which region each provider is tested in, and decide whether a 33/33/33 weighting matches your own use. If you run it yourself, expect to supply an API key per provider and to fix the Node version mismatch between the manifest and the documentation.
Frequently asked questions
How does Browser Arena measure a cloud browser provider?
It runs four steps for every provider: create a session through the provider API, connect Playwright over CDP, navigate to a URL and wait for domcontentloaded, then release the session through the API. Results are medians across successful runs, in either sequential mode or batches of 10 parallel sessions.
Which providers does Browser Arena rank?
Nine of them: Notte, Tilion in beta, Browserbase, Steel, Kernel, Hyperbrowser, Anchor Browser, Browser Use and Lightpanda. Each is tested from one of two EC2 t3.micro boxes, with Notte, Browserbase and Lightpanda on the us-west-1 runner and the rest on us-east-1.
How does the Browser Arena value score turn latency into a number?
Reliability, latency and cost are each normalised to a 0 to 1 scale against fixed anchors and combined with equal 33/33/33 weights by default. Latency scores 0 at 10,000 ms and 1 at 0 ms, reliability scores 0 at 90 percent and 1 at 100 percent, and cost scores 0 at 0.20 dollars an hour and 1 at zero. Presets such as Speed first shift the weights.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nottelabs-browserarena)