Model or dataset
xlang-ai/OSWorld-V2 avatar
xlang-ai/OSWorld-V2

OSWorld-V2: a release-pinned benchmark for long-horizon computer use agents

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

344 stars52 forksPythonApache-2.0

At a glance

What is it?
OSWorld-V2 scores agents on multi-step desktop tasks, and its real cost is not the Python package but the discipline of keeping code, task files, assets and mocked websites on one release tag.
Who is it for?
Adopt OSWorld-V2 if you need a reproducible, release-pinned score for an agent that drives a real desktop, and you are willing to run the provider VMs and the mocked websites yourself. Do not adopt it if you want a single pip install that runs a leaderboard in one process, or if you only need short single-step GUI checks, where the release machinery is pure overhead.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OSWorld-V2 measures that a short GUI benchmark cannot

OSWorld-V2 exists for one narrow problem: scoring a computer use agent on tasks that take many steps to finish. The project describes itself as benchmarking computer use agents on long-horizon real-world tasks, and the repository organises around that claim. The unit of evaluation is not a screenshot classification or a single click target. It is a task the agent works through in a desktop environment, with the run recorded as a trajectory.

The audience is narrow too. This is research infrastructure: the pyproject classifiers list Development Status 4 - Beta and Intended Audience Science/Research, and the declared operating system is POSIX Linux only. If you are building a product that automates a browser flow, OSWorld-V2 is not a runtime you ship. It is the thing you point at your agent to argue that it works. The homepage, the paper and the trajectory viewer exist to make that argument legible to reviewers.

The design consequence is that everything is versioned together. A score is only meaningful relative to a named benchmark release, which is why the README repeats the warning not to mix releases and not to replace a pinned tag with main or latest.

How a run is assembled: code, tasks, assets, websites, provider images

The architecture is a set of separately versioned components that must agree. The README names five for the recommended release: the OSWorld-V2 code at xlang-ai/[email protected], Python task files at xlangai/[email protected], task assets at xlangai/[email protected], mocked websites either team-hosted under the suffix site.hku.icu or self-hosted from Task-Web/[email protected], and provider VM images that reuse the AWS and Docker artifacts pinned by the 2026.06.24 release.

Task files and task assets are separate downloads. The task files come from a Hugging Face dataset, the assets from a gated one, and the assets are pointed at through the OSWORLD_FILE_BASE_URL environment variable. That split is deliberate: the task definitions are small and text-like, the assets are the large binary payload the mocked websites and applications need.

The provider layer sits underneath. Docker is documented for Linux, and the README also refers to AWS image artifacts, so the environment can be local containers or cloud VMs. The repository carries desktop_env/, mm_agents/, hybrid_agents/, monitor/ and an evaluation_examples/ directory, which maps to the obvious split: environment, agent implementations, and scoring. The task_loader.py, run.py and lib_run_single.py entries at the top level are the execution path.

The mocked websites are the part people underestimate. Because the tasks target real-world workflows, the benchmark ships stand-in sites that behave like the originals. Those sites are versioned independently in a separate repository, so a benchmark release is only reproducible if the website deployment matches.

Installing OSWorld-V2 and running the download scripts

The README gives two setup paths: an agent-assisted one built around a skill called setup-osworld, and a manual one. The manual path starts with a clone and uv sync, and the project metadata requires Python >= 3.12 through pyproject.toml.

bash
git clone https://github.com/xlang-ai/OSWorld-V2
cd OSWorld-V2
uv sync

After that, the release-aware download scripts fetch the task files and the assets. Both take --benchmark-release, and the assets script also takes --target-dir and --clean.

bash
uv run scripts/tools/download_osworld_v2_tasks.py \
  --benchmark-release osworld-v2-2026.08.08
bash
uv run scripts/tools/download_osworld_v2_assets.py \
  --benchmark-release osworld-v2-2026.08.08 \
  --target-dir cache/osworld_v2_assets \
  --clean

export OSWORLD_FILE_BASE_URL="$(pwd)/cache/osworld_v2_assets"

Expect the first script to land Python task files and the second to populate the cache directory you named. The export is what makes the assets discoverable at run time, so it belongs in the same shell as the evaluation, not in a separate session.

If you use the team-hosted websites rather than self-hosting, the README says to change the website suffix:

bash
export WEBSITE_HOST_SUFFIX="site.hku.icu"

The dependency set also moved between releases. The 2026.08.08 notes state that the uv dependency files changed and now require imageio-ffmpeg>=0.6.0, with two ways to satisfy it.

bash
uv sync --frozen
# Or install only the new package in an existing host environment:
uv pip install "imageio-ffmpeg>=0.6.0"

One bootstrap detail matters. The README instructs using main only as a setup bootstrap for the release-aware download scripts, then switching the evaluation checkout to v2026.08.08. Running the benchmark itself from main is explicitly discouraged.

Where OSWorld-V2 breaks: release skew and dropped hosting

The failure mode the project documents most openly is version skew. Because five components carry their own tags, mixing them produces a number that looks fine and means nothing. The README states the rule plainly: do not mix releases or replace a pinned tag with main or latest. A team that clones main, downloads 2026.08.08 tasks and points at a website deployment from an older tag has built a benchmark that no one else can reproduce.

The second failure is hosting churn, and it has already happened. The 2026-09-08 update says the project no longer hosts the osworld-v2-2026.06.24 websites at web.hku.icu, and that using that version now requires self-hosting Task-Web/[email protected]. Anyone who ran the June release against the hosted sites and wants to rerun it later has to stand up the website repository themselves. The benchmark version matrix still lists 2026.06.24 as Active, but the hosted column says self-host. That is a real operational cost attached to a benchmark that presents itself as a versioned artifact.

The third constraint is scope. The project metadata declares POSIX Linux only, so Windows and macOS desktops are not the environment this evaluates. And optional heavy dependencies for larger agents or OCR and model stacks are gated behind uv sync --extra full, which the README ties to v1 tasks. If your agent needs those stacks for v2 work, the README does not promise them.

Finally, this is a benchmark, not a harness you extend casually. The migration document exists precisely because moving from OSWorld 1.0 touches dependency setup, task conversion, provider reuse, AWS notes, mocked websites, GitLab, agent migration and result comparability. Comparability is the last item for a reason.

OSWorld-V2 against OSWorld 1.0: reuse the providers, rebuild the rest

The obvious alternative is the project's own predecessor, xlang-ai/OSWorld. The README points to docs/MIGRATING_FROM_OSWORLD_V1.md and lists what the move involves: dependency setup, task conversion, provider reuse, AWS notes, mocked websites, GitLab, agent migration, and result comparability.

The difference in approach is not a rewrite of the agent loop. The provider layer is the part designed to carry over, and the 2026.08.08 release makes that concrete by reusing the AWS and Docker image artifacts pinned by 2026.06.24 rather than rebuilding them. What changes is everything around the environment: the task files are now versioned Hugging Face datasets rather than whatever shipped in the repository, the assets are a separate gated download, and the websites are a separately versioned deployment. Tasks and assets are decoupled in v2 in a way the v1 layout does not suggest.

The practical consequence for a team already on v1 is that scores are not comparable across the two. The migration guide treats result comparability as its own checklist item, and the README's release warnings reinforce that a v2 number is only meaningful within a named v2 release. If your published results are v1 numbers, migrating means re-running, not re-labelling.

Maintenance, licensing and what an upgrade actually costs

The repository is not archived, and the last push was on 2026-09-11, which is recent. Two releases are listed, v2026.08.08 and v2026.06.24, and the version matrix marks both Active while recommending 2026.08.08. The project explicitly warns that main carries development code and that reproducible evaluation should use a supported release.

Upgrade cost is where this benchmark differs from a library. The 2026.08.08 notes lay out five steps: re-download the Python task files, re-download the task assets with --clean, update the OSWorld-V2 checkout to v2026.08.08, update the self-hosted websites to Task-Web/[email protected] or change WEBSITE_HOST_SUFFIX for the team-hosted ones, and update the host environment because the uv dependency files changed and now require imageio-ffmpeg>=0.6.0. Provider VM images, by contrast, are unchanged between the two releases, so the expensive artifact is the one that did not move. Budget the upgrade as a re-download plus a dependency sync, not as a code refactor.

Licensing is Apache-2.0, declared in pyproject.toml with license-files = ["LICENSE"]. That covers the code in this repository. It does not automatically cover the task assets, which live in a separate gated Hugging Face dataset, or the mocked websites, which live in a different repository. If you plan to redistribute a task set or publish trajectories, check the terms on those components separately. This is a description of where the licences sit, not legal advice.

The dependency list is broad: it pulls in provider SDKs for OpenAI, Anthropic, Zhipu, Groq, Together and Google GenAI, plus pyautogui, playwright, pymupdf, librosa and opencv-python, among others. A host environment for this benchmark is not small, and that is a maintenance line item on its own.

Editorial conclusion

Adopt OSWorld-V2 if you need a reproducible, release-pinned score for an agent that drives a real desktop, and you are willing to run the provider VMs and the mocked websites yourself. Do not adopt it if you want a single pip install that runs a leaderboard in one process, or if you only need short single-step GUI checks, where the release machinery is pure overhead. Before you commit, verify three things on the 2026.08.08 manifest: that your provider images match the 2026.06.24 artifacts it reuses, that your website suffix points at a deployment you control, and that your checkout is on v2026.08.08 rather than main.

Frequently asked questions

What is OSWorld-V2?

It is a benchmark for computer use agents on long-horizon real-world tasks, published by xlang-ai under Apache-2.0 with Python as the primary language. A run combines release-pinned code, Hugging Face task files, gated task assets, mocked websites and provider VM images.

Does OSWorld-V2 have a verified variant?

The README does not describe a verified variant of OSWorld-V2. It names two benchmark releases, osworld-v2-2026.06.24 and osworld-v2-2026.08.08, and the version matrix marks both as Active while recommending 2026.08.08.

What is OSWorld SFT?

The README and pyproject.toml do not mention supervised fine-tuning or an SFT component. What the repository does document is evaluation: task files, task assets, mocked websites, provider images, and a trajectory viewer for inspecting runs.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. xlang-ai/OSWorld-V2 on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xlang-ai-osworld-v2.svg)](https://hysenlabs.com/projects/xlang-ai-osworld-v2)