Open-source project
microsoft/WindowsAgentArena avatar
microsoft/WindowsAgentArena

Windows Agent Arena: a Windows 11 VM harness for benchmarking desktop AI agents

Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.

904 stars105 forksPythonMIT

At a glance

What is it?
WAA ships a Docker image, a 30GB Windows 11 golden snapshot and an Azure ML path so the same agent loop can be scored across hundreds of desktop tasks. The cost is a heavy setup and a Windows-only task surface.
Who is it for?
Adopt Windows Agent Arena if you are evaluating a computer-use agent on Windows-only desktop tasks and you can absorb the Docker plus Windows 11 VM setup, or if you already have Azure ML and want the parallel runner. Skip it if your target is a Linux or browser-only agent, or if you need a lightweight harness that installs in one command; OSWorld-style Linux environments cover that ground instead.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 170 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap WAA fills: scoring desktop agents on real Windows

Most agent benchmarks hand the model a clean API or a browser. A desktop agent instead has to look at pixels, decide where to click, and cope with whatever window is on screen. Windows Agent Arena targets exactly that setting. The README describes it as a scalable Windows AI agent platform for testing and benchmarking multi-modal, desktop AI agents, and the stated goal is a reproducible and realistic Windows OS environment where agentic workflows can be tested across a diverse range of tasks.

The audience is narrow and specific. This is research infrastructure for people who train or compare computer-use agents, not a runtime you would embed in a product. The repository topics (agentic, ai-benchmark, computer-use, desktop-agent, windows) point the same way. If you are shipping an automation script for your own desktop, the setup cost here is out of proportion to the problem.

One design decision worth noting: the environment is a full Windows 11 VM, not a mock. That is what makes the results credible, and also what makes the harness heavy.

How the harness works: Docker image, golden snapshot, command server

The architecture separates the agent from the machine it controls. On the host side, a Docker image built from the src directory runs the agent loop and talks to model endpoints. On the guest side, a Windows 11 VM runs the actual desktop. The README states that the 30GB snapshot represents a fully functional Windows 11 VM with all the programs needed to run the benchmark, and that this VM additionally hosts a Python server which receives and executes agent commands.

So the data flow is: agent code in the container observes the screen, produces an action, and sends it to the Python server inside the VM, which executes it and returns the new state. The README links separate architecture diagrams for the local and cloud deployments, which confirms the two paths share this split.

For scale, the README says WAA supports deployment of agents at scale using Azure ML cloud infrastructure, allowing parallel runs and delivering benchmark results for hundreds of tasks in minutes rather than days. The cloud path is the reason the repository depends on azure-ai-ml, azure-identity, azureml-core and azureml-dataset-runtime.

A difficulty switch changes how much the agent must do itself. Per the 2024-11-10 update, setting diff_lvl="hard" in src/win-arena-container/start_client.sh means that in many tasks the agent must initialize and set up the task, for example finding and opening the right application, instead of having the task configured for it.

Installing Windows Agent Arena and running a first local benchmark

The README lists three pre-requisites: a running Docker daemon (on Windows, Docker with WSL 2 is recommended), an OpenAI or Azure OpenAI API key, and Python 3.9, with Conda suggested for an ad hoc environment.

Start by cloning and installing the Python dependencies in a 3.9 environment. The README gives these commands:

bash
git clone https://github.com/microsoft/WindowsAgentArena.git
cd WindowsAgentArena
conda create -n winarena python=3.9
conda activate winarena
pip install -r requirements.txt

Next, create config.json at the root of the project. Only the keys for the endpoint you actually use are needed:

json
{
    "OPENAI_API_KEY": "<OPENAI_API_KEY>",
    "AZURE_API_KEY": "<AZURE_API_KEY>",
    "AZURE_ENDPOINT": "https://yourendpoint.openai.azure.com/"
}

Then pull the base image and build the WinArena image. The README says the base image contains the dependencies and models needed to run the code in src, and that the build produces windowsarena/winarena:latest from the current src code:

bash
docker pull windowsarena/winarena-base:latest
cd scripts
./build-container-image.sh

If Dockerfile-WinArena-Base changed, the README notes you can pass --build-base-image true; ./build-container-image.sh --help lists the other options.

The last preparation step is the VM. Download the Windows 11 Enterprise Evaluation ISO (90-day trial, English, United States, roughly 6GB) from the Microsoft Evaluation Center, rename it to setup.iso, and copy it into src/win-arena-container/vm/image. The README then describes an automatic setup that produces the WAA golden image. Expect a long wait and about 30GB of disk: this is a full Windows install, not a container start.

For a first scored run, the README's own example for the top performing mode in the paper is a local script invocation:

bash
./run-local.sh --som-origin mixed-omni --gpu-enabled true

That mode pairs the Navi agent with Omniparser, which Microsoft open-sourced on 2024-10-23. If you only want the default configuration, the README's earlier instruction is to change diff_lvl="normal" to diff_lvl="hard" inside src/win-arena-container/start_client.sh, which means the client script is the place to look for run parameters.

Where Windows Agent Arena gets in the way

The setup is the first limitation, and it is not a small one. You need a Docker daemon, Python 3.9 specifically, a model API key, a 6GB Windows ISO, and enough disk for a 30GB snapshot. That snapshot has to be built before any task runs, and the README does not document rollback or how to repair a golden image that boots into a bad state. If the automatic setup fails midway, the documented path restarts from the ISO rather than from a checkpoint.

The evaluation surface is Windows-only by construction. Any agent whose target is a Linux desktop, a mobile UI, or a pure browser session cannot be meaningfully scored here, and porting tasks means rebuilding the VM image rather than editing a task file.

Cloud scale is tied to Azure ML. The README presents Azure ML as the mechanism for parallel runs and fast results across hundreds of tasks, so the throughput claim is conditional on that subscription. The local path exists, but the README does not make a parallel-throughput claim for it.

Finally, the local path assumes WSL or Linux as the host. The README's local deployment section is titled Local deployment (WSL or Linux), and the Windows guidance is to run Docker with WSL 2. A native Windows host without WSL is not the documented configuration.

On maintenance: the last push to the repository was on 2026-04-13, and the most recent release listed is v0.0.4 from 2024-09-28. The codebase has moved since that release, but the tagged artifacts are old.

OSWorld and the difference between Linux and Windows agent evaluation

The closest comparison in this space is OSWorld, which evaluates agents in a Linux desktop environment. The difference is not cosmetic. A Windows agent has to deal with a different window manager, a different set of preinstalled applications, and Windows-specific dialogs; a Linux agent does not. If your deployment target is Windows, an OSWorld score tells you something about general computer-use ability but nothing about how the agent handles the Windows shell.

The trade runs the other way too. OSWorld-style Linux environments are generally lighter to stand up than a 30GB Windows snapshot, and they are not tied to Azure ML for parallel execution. If your agent is platform-agnostic and you want fast iteration, the Windows VM is overhead you are paying for realism you do not need.

A second reference point is the WAA paper itself, Evaluating Multi-Modal OS Agents at Scale (arXiv 2409.08264), which is where the task suite and the reported agent results live. The repository is the harness; the paper is the definition of what is being measured. Reading only the README will not tell you how tasks are scored.

Licence and the cost of keeping a fork alive

The repository is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is permissive enough for internal research and for shipping derived harness code, but the licence covers the code in this repository, not the Windows 11 Enterprise Evaluation ISO, the OpenAI or Azure OpenAI endpoints, or Azure ML compute. Those carry their own terms, and the ISO is explicitly a 90-day trial.

Upgrade cost is dominated by the golden image. The README's build script rebuilds windowsarena/winarena:latest from src, but changes to Dockerfile-WinArena-Base require the --build-base-image flag, and any change to the VM image means redoing the ISO-to-snapshot pipeline. Bumping the base image in isolation is cheap; touching the guest is not.

The Python side is small: requirements.txt lists five packages (azure-ai-ml, azure-identity, azureml-core, azureml-dataset-runtime, pandas), so dependency drift is not the maintenance burden here. The VM is.

Editorial conclusion

Adopt Windows Agent Arena if you are evaluating a computer-use agent on Windows-only desktop tasks and you can absorb the Docker plus Windows 11 VM setup, or if you already have Azure ML and want the parallel runner. Skip it if your target is a Linux or browser-only agent, or if you need a lightweight harness that installs in one command; OSWorld-style Linux environments cover that ground instead. Before committing, verify the golden image build path end to end: download the Windows 11 Enterprise Evaluation ISO, rename it to setup.iso, place it at src/win-arena-container/vm/image, and confirm the 30GB snapshot boots and its Python server accepts commands. Also check whether the diff_lvl setting you want is normal or hard, because the two produce different task setups.

Frequently asked questions

What does Windows Agent Arena do?

It is a Microsoft platform for testing and benchmarking multi-modal desktop AI agents in a reproducible Windows 11 environment, with a local Docker path and an Azure ML path for parallel runs.

What do I need before installing Windows Agent Arena?

The README lists a running Docker daemon (Docker with WSL 2 on Windows), an OpenAI or Azure OpenAI API key, and Python 3.9, with Conda suggested for the environment. The VM setup additionally requires the Windows 11 Enterprise Evaluation ISO.

How do I run the hardest Windows Agent Arena tasks?

The 2024-11-10 update says you change diff_lvl="normal" to diff_lvl="hard" in src/win-arena-container/start_client.sh, which makes agents initialize and set up many tasks themselves instead of receiving a prepared task configuration.

Does Windows Agent Arena require Azure?

No for a local run, which uses Docker and the Windows 11 VM. Yes for the scale story: the README ties parallel execution and fast results across hundreds of tasks to Azure ML, which is why the requirements file includes azure-ai-ml and azureml-core.

What licence does Windows Agent Arena use?

The repository is MIT licensed. That covers the code, not the Windows 11 Enterprise Evaluation ISO (a 90-day trial), the model endpoints, or Azure ML compute.

Official sources

  1. License: MIT
  2. microsoft/WindowsAgentArena on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-windowsagentarena.svg)](https://hysenlabs.com/projects/microsoft-windowsagentarena)