AgentLab: twelve browser benchmarks, and the timeout that only Ray enforces
AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reproducibility.
At a glance
- What is it?
- ServiceNow's AgentLab is a research harness for web agents built on BrowserGym and Ray, covering twelve benchmarks from 33 to 31586 tasks. It is opinionated about how it fails: a hung job can halt a small study entirely, and the automatic timeout exists only in the parallel backend. The package also classifies itself as Pre-Alpha.
- Who is it for?
- Use this if you are running reproducible web agent experiments and want the study directory, the parallel backend and the recovery path in one harness, since the two-call reload is the part that saves a long run. Skip it if you want a benchmark you can start without standing up the environment yourself, because AgentLab links out for that on every row.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 78 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The package classifies itself Pre-Alpha while the tag reads 0.4.2
The trove classifier in the package metadata says Development Status :: 2 - Pre-Alpha. That is the project's own assessment of itself, and it sits alongside a tagged history that reads v0.4.0 in February 2025, v0.4.1 in December 2025 and v0.4.2 in January 2026.
The gap between the last tag and the branch explains part of it. The last recorded push is 2026-07-17, about six months after v0.4.2 shipped, and the version field in the manifest is dynamic rather than written down, taken from the git tag by hatch-vcs. A build from main therefore has no version to report at all, which is also why nothing in the release list covers the work done since January.
The README opens with a warning rather than a pitch. It says AgentLab is meant to provide an open, easy to use and extensible framework to accelerate the field of web agent research, and that it is not meant to be a consumer product, to be used with caution. A research harness that will happily launch hundreds of browser automations against live sites is a reasonable thing to be cautious about.
The licence is stated in two places and missing in one. The manifest carries license = Apache-2.0 and a matching Apache Software License classifier, and the README badge links the Apache 2.0 text. The repository's own licence field is recorded as NOASSERTION, which is the GitHub side saying it could not classify the file rather than a claim that no licence applies.
Python 3.11 to 3.12, and three model SDKs in the base dependency set
The manifest requires Python 3.11 or 3.12, written as a lower bound of 3.11 and an upper bound of 3.13. The Makefile pins the development interpreter more narrowly still, syncing with uv against Python 3.12 specifically, and a uv.lock sits at the repository root so that resolution is committed rather than rediscovered.
The base dependency list is long and tells you what the project actually is. pydantic for configuration, dask and distributed for scheduling, browsergym at 0.7.1 or newer as the benchmark layer, joblib, openai, anthropic and litellm all in the base set, huggingface_hub, tiktoken, gymnasium, and torch at 2.2.2 or newer. Two entries shape the whole interface: ray with its default extras is the parallel backend, and gradio at 5.5 or newer is the UI assistant.
Three model SDKs in the required set, rather than one behind an extra, is the clearest signal of scope. The advertised unified LLM API covers OpenRouter, OpenAI, Azure and a self-hosted TGI deployment, and that breadth is paid for at install time by every user regardless of which one they pick.
Only two things are optional. The hint extra pulls sentence-transformers at 5.0.0 or newer, and the transformers extra pulls transformers at 4.38.2 or newer. Everything else, including torch and ray, is required to run a study at all.
One pip line, then a benchmark you have to stand up yourself
The documented install is two commands. Run pip install agentlab, then run playwright install to get the browser binaries. After that the instruction is to prepare the required benchmark using whatever the setup column for that row links to, because AgentLab does not ship the environments it evaluates against.
That last step is the bulk of the real work. The benchmark table points at per-benchmark setup guides for WebArena, WebArena Verified, WorkArena, VisualWebArena, AssistantBench, MiniWoB, OSWorld and TimeWarp, and two entries, GAIA and Mind2Web-live, have no setup link and are marked as coming soon.
Configuration is four environment variables, and the comments on them matter more than the names. AGENTLAB_EXP_ROOT is the root directory for experiment results and defaults to $HOME/agentlab_results. OPENAI_API_KEY is needed only if OpenAI models are used, OPENROUTER_API_KEY only if OpenRouter models are used, and the pair AZURE_OPENAI_API_KEY with AZURE_OPENAI_ENDPOINT only if Azure models are used. You supply credentials for one provider, not four.
There is also a single command that puts an assistant in front of a live page, invoked as agentlab-assistant with a start_url argument, and a second form that takes an agent_config pointing at a dotted module path so you can drive your own agent instead.
Twelve benchmarks whose task counts are not comparable
The benchmark table is the centre of the project, and it is worth reading the columns rather than just the names. Task counts run from 33 to 31586. WorkArena L1 has 33 tasks, MiniWoB 125, AssistantBench 214, WorkArena L2 and L3 341 each, OSWorld 369, WebArena and WebArena Verified 812 each, VisualWebArena 910, TimeWarp 1386, and WebLinx 31586.
That spread is not a difficulty ranking. WebLinx carries a max step of 1, meaning a single action per task, while WorkArena L2 and L3 allow 50 and everything else lands between 10 and 30. A success rate on WebLinx and a success rate on WorkArena measure different amounts of work, so quoting them side by side says nothing without the step budget beside them.
Three further columns vary in ways that change what a run costs. Seed diversity is High for the three WorkArena levels, Medium for MiniWoB and None everywhere else, so only those four give you repeated sampling. Multi-tab support is on for WebArena, WebArena Verified, VisualWebArena, AssistantBench and TimeWarp, and off for the rest. Hosted method ranges across self hosted docker, a demo instance, live web, a self hosted dataset and self hosted static files.
AssistantBench and the two coming-soon entries are the only ones marked live web, which means a benchmark that hits the public internet rather than a container you control.
Every cell in the leaderboard column reads soon
The last column of the benchmark table is the BrowserGym leaderboard entry, and it reads soon for all twelve rows without exception. That includes WebArena and MiniWoB, the two entries marked self hosted, one as a docker image and one as static files, both of which can be run without touching the public internet.
So the table is precise about everything you control and empty about the one thing you would want it for. Task counts, step budgets, seed diversity, multi-tab support and hosting method are all filled in. Comparative standing against other agents is uniformly unpopulated.
The leaderboard itself is not in this repository. It is a HuggingFace space under the ServiceNow organisation, browsergym-leaderboard, and the point of the soon markers is that results have not been submitted to it yet rather than that the space does not exist.
The research context is a paper rather than a repository. The README points at the BrowserGym ecosystem paper on arXiv, identifier 2412.05467, for the details of the benchmark layer that AgentLab builds on top of.
A hung job halts a small run and Ray is the escape, which inverts the usual reason to scale
The job timeout section is the most operationally useful paragraph in the README. It says the complexity of the wild web, Playwright and asyncio can sometimes cause jobs to hang, and that the effect of this is to disable workers until the study is terminated and relaunched.
The consequence depends on how you launched. If you are running jobs sequentially or with a small number of workers, a hang could halt your entire study until you manually kill and relaunch it. In the Ray parallel backend there is a system that automatically terminates jobs exceeding a specified timeout, and the section says this is particularly useful when task hanging limits your experiment.
Read together, the safe configuration is the large one. A single-worker sequential run has no timeout enforcement and one hung task takes the whole study down, while a many-worker Ray run recovers on its own. The usual argument for scaling up is throughput, and here it is also the argument for not losing a night of compute.
There is no equivalent protection for the interactive assistant, since a single page under agentlab-assistant has no worker pool behind it and nothing in the timeout section covers it.
Relaunching is two calls, and the agent config has to be importable to survive
Launching a study is one call. You build it with make_study, passing a benchmark name, a list of agent arguments and a comment, then run it with a job count:
# Import your agent configuration extending bgym.AgentArgs class
# Make sure this object is imported from a module accessible in PYTHONPATH to properly unpickle
from agentlab.agents.generic_agent import AGENT_4o_MINI
from agentlab.experiments.study import make_study
study = make_study(
benchmark="miniwob", # or "webarena", "workarena_l1" ...
agent_args=[AGENT_4o_MINI],
comment="My first study",
)
study.run(n_jobs=5)Recovering is two more, and it is the part people get wrong when they just rerun. You load the existing study from its directory with Study.load, select the tasks that never finished or that raised with find_incomplete(include_errors=True), and then call run() again. The include_errors flag is what makes it a recovery rather than a resume, since without it a task that crashed looks identical to one that was never scheduled.
The comment above the import is a real constraint rather than advice. The agent arguments object has to be imported from a module reachable on PYTHONPATH so that it can be unpickled, which matters because Ray moves that object between worker processes. A class defined in a notebook or a __main__ block will not survive distribution.
For anything beyond the defaults there is main.py at the repository root, which the README calls a lazy CLI that is actually more convenient, with the instruction to comment and uncomment the lines you need and to modify them at will, but not to push them to the repository.
The test harness strips the benchmark's own pins so both can be installed
The Makefile is where the awkward parts of benchmarking show up. The osworld target clones OSWorld and then edits its requirements.txt with sed, replacing the pinned numpy, torch, tqdm and pandas requirements with bare unversioned names, before installing those requirements and then OSWorld itself in editable mode.
That is not housekeeping, it is conflict resolution. AgentLab pins torch at 2.2.2 or newer and OSWorld pins its own versions, and a literal reading of both files cannot be satisfied. Removing the benchmark's pins lets the already-resolved versions satisfy both sides, at the cost of running the benchmark against dependency versions it was not written for.
The MiniWoB target does the opposite kind of work, pinning by commit rather than by version. It clones miniwob-plusplus, checks out a specific SHA, serves the html directory over a local HTTP server on port 8080, and records the process id in a file. A separate check target curls the URL and fails loudly if the server is not up.
The test target chains five steps, and the ordering matters when something breaks: setup, miniwob, check-miniwob, run-tests, stop-miniwob. A failing test run stops the chain before the server is torn down, so a red suite leaves a stray server and a stale pid file behind. The test command itself deselects anything marked not pricy, which is how tests needing paid model calls are kept out of the default run, and lint runs black in check mode plus darglint, which verifies that docstrings document the arguments their functions actually take.
Editorial conclusion
Use this if you are running reproducible web agent experiments and want the study directory, the parallel backend and the recovery path in one harness, since the two-call reload is the part that saves a long run. Skip it if you want a benchmark you can start without standing up the environment yourself, because AgentLab links out for that on every row. Before your first large run, read the job timeout section, decide on Ray rather than a few workers, and check the licence yourself, since the repository record says NOASSERTION while the package states Apache-2.0.
Frequently asked questions
How do I install AgentLab and run a first study?
Run pip install agentlab, then playwright install. The benchmark itself is not shipped, so you prepare it using the setup instructions linked from the benchmark table. Set AGENTLAB_EXP_ROOT, which defaults to $HOME/agentlab_results, plus the key for whichever provider you use: OPENAI_API_KEY, OPENROUTER_API_KEY, or the AZURE_OPENAI_API_KEY and AZURE_OPENAI_ENDPOINT pair.
Which Python version does AgentLab support?
The manifest requires Python 3.11 or 3.12, written as >=3.11,<3.13. The Makefile pins the development interpreter more narrowly by syncing with uv against Python 3.12, and a uv.lock at the repository root commits the resolution.
Can I resume an AgentLab study that failed?
Yes. Load the study directory with Study.load, call find_incomplete(include_errors=True) to select tasks that never finished or that raised, then call run() again. Hung jobs are a separate problem: they disable workers until the study is terminated and relaunched, and automatic timeout termination only exists in the Ray parallel backend.
What licence is AgentLab released under?
The package metadata states Apache-2.0 in both the license field and the classifier, and the README badge links the Apache 2.0 text. The repository licence field is recorded as NOASSERTION, so the machine-readable repository record is empty even though the package states the licence in two separate places.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/servicenow-agentlab)