Model or dataset
xlang-ai/OSWorld avatar
xlang-ai/OSWorld

OSWorld: a benchmark that grades computer-use agents inside real Ubuntu and Windows VMs

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

3,168 stars537 forksPythonApache-2.0

At a glance

What is it?
OSWorld runs multimodal agents against live desktops and scores them with executable checks rather than screenshots. The setup cost is a VM provider plus a large Python dependency tree, and the benchmark numbers changed with OSWorld-Verified.
Who is it for?
Adopt OSWorld if you are publishing agent results and need executable, environment-state checks instead of a curated demo. Skip it if you cannot give it a VM provider with KVM, VMware, VirtualBox, Modal, or a cloud account, because the environment is the benchmark.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 17 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OSWorld measures that a screenshot benchmark cannot

OSWorld is a benchmark for agents that operate a computer through the same surfaces a person uses: a desktop, a window manager, applications, files. The repository describes it as benchmarking multimodal agents for open-ended tasks in real computer environments, and the task data lives under evaluation_examples. The intended audience is agent researchers and model teams who need a score that survives the question "did it actually do the thing?" A chat transcript or a screenshot cannot answer that. An OSWorld task ends with a check on the state of the machine, so an agent that produces plausible narration but leaves the spreadsheet unchanged does not pass.

The scope is deliberately wide. Topics on the repository include GUI, RPA, code generation, reinforcement learning and vision-language models, and the dependency list includes openpyxl, python-docx, python-pptx, pypdf, pdfplumber and librosa, which is a reasonable hint at the file types the checks inspect. This is not a narrow app-automation suite. That breadth is also where the cost sits: a real desktop, real applications and real files have to exist before any measurement happens.

How the agent loop, VM provider and evaluator fit together

The architecture separates three things that are easy to conflate. The desktop_env package owns the machine: it boots a VM, exposes screenshots and input, and resets state between tasks. The mm_agents directory holds the agent side, the policy that reads observations and emits actions. The evaluator, driven by run.py and lib_run_single.py, takes the finished task and runs the check defined for it. The gemini_action_parser.py file at the repository root is a reminder that turning model output into a legal click or keystroke is its own parsing problem, and different model families need different handling.

Provider abstraction is the load-bearing design decision. The README states that the environment code was refactored to decompose VMware integration and to start supporting other platforms such as VirtualBox, AWS, Azure and Docker. In practice you choose a provider at construction time through provider_name, and the rest of the loop is unchanged. That is why the same quickstart can run on a laptop with VMware Fusion or on a server with KVM. It is also why failures are often environmental rather than agentic: a VM that will not boot looks, from the logs, like a run that never started.

Scoring is scripted against the environment, not judged by a model. The evaluation_examples data is what carries the per-task setup and the check, and show_result.py and lib_results_logger.py exist to read the run output back out. The README does not document rollback or resumption semantics for an interrupted run, so treat a killed job as a lost run rather than a resumable one.

Installing OSWorld on a desktop with VMware or VirtualBox

The README splits installation by host type. On a desktop, laptop or bare metal machine that is not itself virtualized, it points at VMware Workstation Pro (VMware Fusion on Apple chips) or VirtualBox, and it warns that parallelism and macOS on Apple chips may not be well supported under VirtualBox. Start by cloning and installing the Python dependencies; the README asks for Python 3.10 or newer and recommends Conda.

bash
git clone https://github.com/xlang-ai/OSWorld
cd OSWorld
conda create -n osworld python=3.10
conda activate osworld
pip install -r requirements.txt

If you only want the environment and not the benchmark tasks, the README gives a shorter path through the published package. Note that pyproject.toml declares requires-python >= 3.12 while the README and setup.py say 3.10 or newer, so pin your interpreter deliberately rather than assuming the two agree.

bash
pip install desktop-env

Next, install the hypervisor and make sure the vmrun command is on your path. The README offers this check and says a successful install prints the currently running virtual machines.

bash
vmrun -T ws list

After that the README states that the setup script downloads the necessary virtual machines and configures the environment automatically. The first real use is the quickstart entry point, which takes a provider name and a headless flag. The README shows this form for the Modal provider, and the same two arguments are what you change when you switch providers.

bash
python quickstart.py --provider_name modal --headless true

For Docker, the README says to pass provider_name docker and os_type Ubuntu or Windows when initializing DesktopEnv. Before that it asks you to confirm KVM support on Linux with the following command, where a value greater than zero means the processor should support KVM. macOS hosts generally do not support KVM, and the README directs those users back to VMware.

bash
egrep -c '(vmx|svm)' /proc/cpuinfo

Where OSWorld is the wrong tool

The heaviest constraint is the environment itself. OSWorld needs a real desktop VM, which means a hypervisor, a disk image, and a host that can run both. The README is explicit that macOS hosts generally do not support KVM and should use VMware instead, so a container-only CI runner without KVM is not a supported target. If your team's evaluation pipeline is a slim Linux container farm, OSWorld does not slot into it without adding virtualization.

Interrupted runs leave residue. The README warns that if an experiment is interrupted abnormally, residual Docker containers may remain and affect system performance over time, and it gives a cleanup command. That is a real operational hazard for anyone scheduling many runs.

bash
docker stop $(docker ps -q) && docker rm $(docker ps -a -q)

There is also a versioning trap. The 2025-07-28 update introduced OSWorld-Verified with fixed issues, more AWS support, and what the README calls more effective benchmark signals, and it asks readers to compare results against the new numbers on the official website. A score produced on an older checkout is not the same measurement as a score produced now. If you need a stable target for regression testing, this is a moving benchmark, not a frozen suite. Finally, the dependency list is long and pins some heavy packages, including torch and transformers, so expect a substantial install and a large image.

OSWorld versus OSWorld-Verified, and the alternatives worth naming

The most relevant comparison is internal. OSWorld-Verified is the same project after a major revision: the 2025-07-28 note lists fixed community-reported issues, more AWS support with parallelization that the README says can reduce evaluation time to within 1 hour, and changes intended to make the benchmark signals more effective. The README also states that new model results were run on the latest version and posted to the official website, and asks that your results be compared with those. So the practical choice is not "OSWorld or something else" but "which revision, and against which leaderboard". Running the old checkout and comparing to the new site numbers is the easiest way to produce a misleading result.

Against outside options, the meaningful difference is where the task lives. An API-level agent benchmark hands the model a function-calling interface and checks the return values; OSWorld hands the model a screen and checks the machine. That makes OSWorld slower and more fragile to set up, and it is the only one of the two that can catch an agent that clicks the wrong cell in a spreadsheet. A pure web benchmark sits in between, since a browser is easier to reset than a desktop but still exercises pixels. If your product is a browser agent, a desktop benchmark overstates your risk surface; if your product drives installed applications, a web benchmark understates it. The repository also lists RPA among its topics, which is the closest traditional analogue, but RPA tooling records deterministic scripts while OSWorld measures a model's ability to choose actions.

Maintenance, licence and what a run actually costs

The repository is not archived and the last push was on 2026-08-30, so the codebase is moving. That cuts both ways for adopters. You get fixes such as the ones folded into OSWorld-Verified, and you inherit churn in task definitions and scoring, which is exactly the part you want stable when you are comparing two models. Pin a commit for any result you intend to publish, and record it alongside the score.

Licensing is Apache-2.0, per the repository and the classifier in setup.py. That is permissive and includes a patent grant, which matters if you are embedding the harness in a commercial evaluation stack. This is a description of the licence file, not legal advice; the benchmark tasks and any bundled disk images may carry their own terms, and the README links to a Google Drive cache and a Hugging Face asset host for supporting files, which are separate from the code licence. Check those before redistributing anything.

Upgrade cost is dominated by the environment rather than the Python package. A version bump can mean a new VM image, a re-download of the pre-staged files the README points to, and re-running baselines because the signals changed. Budget for a full re-baseline after any upgrade that touches evaluation_examples, not just a pip install.

Editorial conclusion

Adopt OSWorld if you are publishing agent results and need executable, environment-state checks instead of a curated demo. Skip it if you cannot give it a VM provider with KVM, VMware, VirtualBox, Modal, or a cloud account, because the environment is the benchmark. Before quoting a number, confirm which release you ran: the 2025-07-28 OSWorld-Verified update changed tasks and signals, and the project asks that results be compared against the new website numbers. Also check the evaluation_examples JSON for the task you plan to cite, since that file is what defines the scoring script.

Frequently asked questions

Is OSWorld still active?

The repository is not archived and the last push was on 2026-08-30. The most recent release listed is v0.1.16 from 2024-06-26, so activity is visible in the main branch and in the 2025-07-28 OSWorld-Verified update rather than in tagged releases.

What is OSWorld?

It is a benchmark from xlang-ai for multimodal agents performing open-ended tasks in real computer environments, released with a NeurIPS 2024 paper. Tasks run against actual Ubuntu or Windows desktops, and the repository ships the environment, the agent side and the evaluation examples.

What is OSWorld-Verified?

It is the 2025-07-28 revision of the benchmark. The README lists fixed community-reported issues, more AWS support with parallelization that can reduce evaluation time to within 1 hour, and more effective benchmark signals, with new model results posted to the official website.

How does OSWorld compare with OSWorld-Verified?

They are the same project at different revisions rather than two products. The README asks that results from the latest version be compared against the new numbers on the official website, which implies older checkouts are not directly comparable.

Is OSWorld free to use?

The repository is licensed Apache-2.0, which permits commercial use. The benchmark also depends on external services and assets, including VM images, a Google Drive cache and cloud providers such as AWS, Azure and Modal, and those carry their own costs and terms.

What is the OSWorld benchmark?

It is a set of tasks in evaluation_examples that an agent attempts inside a live desktop VM, with the outcome checked against the resulting machine state rather than against a transcript. Providers include VMware, VirtualBox, Docker with KVM, Modal, AWS and Azure.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. xlang-ai/OSWorld on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xlang-ai-osworld.svg)](https://hysenlabs.com/projects/xlang-ai-osworld)