e2b-dev/open-computer-use: an open source computer use agent on an E2B Desktop Sandbox
AI computer use powered by open source LLMs and E2B Desktop Sandbox
At a glance
- What is it?
- Open Computer Use runs an LLM-driven agent inside a cloud Linux desktop from E2B, driving keyboard, mouse and shell. The model split in config.py is the interesting part, and the Gradio grounding dependency is the sharp edge.
- Who is it for?
- Adopt it if you want to read and modify the agent loop rather than consume a hosted computer use API, and if you are comfortable editing config.py to pick grounding, vision and action models. Do not adopt it if you need a packaged product, a documented release history, or a grounding path that works without Hugging Face Spaces.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 83 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What problem open-computer-use solves, and for whom
Most computer use demos are tied to one vendor's model and one vendor's virtual machine. Open Computer Use separates those two decisions. The desktop comes from the E2B Desktop Sandbox, a cloud Linux machine the README describes as secure, and the models come from a list of providers you edit in config.py. The README states the project supports 10+ LLMs plus OS-Atlas, ShowUI and any other model you want to integrate.
The audience is narrow but real: engineers who want to read an agent loop end to end, swap a model, and watch what changes. The README points to a design article, 'How I taught an AI to use a computer', which suggests the code is meant to be studied, not just run. If you want a hosted API that clicks buttons for you, this is the wrong shape of project. If you want to see how grounding, vision and action get wired together and then change one of them, it is the right one.
The three-model split in config.py
The mechanism that matters is that the agent does not use one model for everything. config.py assigns three roles: a grounding model, a vision model and an action model. The README's example is OSAtlasProvider for grounding, a Groq Llama 3.2 provider for vision, and a Groq Llama 3.3 provider for action.
That split is the design decision worth arguing about. Grounding is the task of turning a screenshot plus an instruction into a coordinate or an element reference. It is a different problem from describing a screen, and a different problem again from emitting the next action. Providers in providers.py are annotated with which roles they can fill: Llama 3.2 is vision only on Fireworks, OpenRouter and Llama API, and vision plus action on Groq. Llama 3.3 is action only. DeepSeek is action only. Gemini 2.0 Flash, GPT-4o, GPT-4o mini and Claude are listed as vision plus action. Moonshot and Mistral AI appear as providers, with Pixtral named for vision and Mistral Large for actions.
The consequence is that you cannot casually swap a single model string and expect everything to work. If you point the vision slot at an action-only model, nothing in the README says the code will catch that for you. The role annotations are documentation, not a type system.
Installing open-computer-use and running a first prompt
The README's prerequisites are Python 3.10 or later, git, an E2B API key and an API key for at least one LLM provider. The install instructions are written for a machine with Homebrew, which installs Poetry and ffmpeg:
brew install poetry ffmpegThen clone the repository and enter it:
git clone https://github.com/e2b-dev/open-computer-use/
cd open-computer-useCreate a .env file in the project directory. The only variable the README marks as required for the sandbox is the E2B key, and you add keys only for the providers you selected in config.py. Hugging Face Spaces do not require an API key, but the README says HF_TOKEN is required to bypass Gradio rate limits, which matters because OS-Atlas and ShowUI are served through Hugging Face Spaces.
E2B_API_KEY="your-e2b-api-key"
GROQ_API_KEY=...
HF_TOKEN=...Install dependencies and start the agent:
poetry install
poetry run startThe README states the agent opens and prompts for its first instruction, and that the display stream becomes visible a few seconds after the Python program starts. You can skip the prompt:
poetry run start --prompt "use the web browser to get the current weather in sf"What you should see is a window streaming the sandbox desktop while the agent works. The pyproject.toml entry point is start = "main:main", and pywebview with the qt extra plus pyqtwebengine are dependencies, so the display path is a desktop webview rather than a browser tab.
Where the grounding path is fragile
The README is explicit that HF_TOKEN is needed to bypass Gradio rate limits, and that OS-Atlas and ShowUI are hosted on Hugging Face Spaces. That makes an external, rate-limited service part of the default grounding path. The README does not document a retry policy, a timeout, a fallback grounding provider, or what the agent does when a grounding call fails. It also does not document rollback of an action the agent already took.
The honest reading is that the default configuration trades reliability for a free grounding model. You can avoid it by choosing a grounding-capable provider from providers.py, but the README does not map which providers can serve the grounding role the way it does for vision and action. That gap is the first thing to check in the source before you build on this.
There is a second limitation that is easy to miss. The README says the project uses Ubuntu but is designed to work with any operating system. 'Designed to' is not 'tested on'. Nothing in the README describes what changes when the sandbox image is not Ubuntu, so treat that as an intention rather than a supported matrix.
How it differs from Anthropic's computer use and from browser-only agents
Anthropic's computer use is a model capability exposed through the API: you send a screenshot and a tool definition, and Claude returns an action. Open Computer Use is the surrounding harness. It owns the sandbox lifecycle through e2b-desktop, owns the display stream, and owns the loop that decides which model handles which step. The README's provider list includes Anthropic as one vision-plus-action option, which makes the relationship concrete: you can run this project with Claude as the model, and what you are adding is the E2B desktop, the three-role split and the local UI.
Compared with browser-only agents, the scope here is the whole desktop. The README says the agent operates the computer via the keyboard, mouse and shell commands, and that the user can pause and prompt the agent at any time. Shell access is the meaningful difference. It also widens the blast radius, which is presumably why the sandbox is a cloud machine rather than your laptop.
Maintenance, licence and upgrade cost
The repository is not archived. The last push was on 2026-07-09. There are no retrieved releases, so version 0.1.0 in pyproject.toml is the only version identifier available, and there is no changelog to read before upgrading.
The dependency set is where upgrade cost lives. e2b is pinned at ^1.5.0, e2b-desktop at ^1.7.2, openai at ^1.59.5 and anthropic at ^0.44.0. The caret ranges mean Poetry will pull newer minor and patch versions within those majors, so an upgrade can move the sandbox SDK and the model SDKs at the same time. pywebview with the qt extra and pyqtwebengine 5.15.x add a compiled Qt dependency to the install, which is the part most likely to break on a machine that is not set up for it.
The licence is Apache-2.0, which permits commercial use and modification and includes a patent grant. It also requires that you keep the licence and notice files and state significant changes. That is a description of the terms, not legal advice; if you redistribute a modified version, have someone qualified read the LICENSE file.
What to check before you build on it
Read providers.py first. The README says providers are imported from there and lists which models are vision-only, action-only or both, but it does not say which can ground. That file is the source of truth for whether your intended configuration is legal.
Second, decide what happens when Hugging Face Spaces is slow or rate limited. The README gives you HF_TOKEN as the mitigation and nothing else. If you cannot tolerate that dependency, plan to replace the grounding model rather than patch around it.
Third, budget for the E2B key. Every run starts a cloud desktop, and the README treats the E2B API key as a prerequisite on the same level as Python. There is no local sandbox mode described. If your environment forbids sending screenshots to a hosted grounding endpoint, this project's default configuration does not fit, and you would need to run OS-Atlas or ShowUI yourself, which the README does not walk through.
Editorial conclusion
Adopt it if you want to read and modify the agent loop rather than consume a hosted computer use API, and if you are comfortable editing config.py to pick grounding, vision and action models. Do not adopt it if you need a packaged product, a documented release history, or a grounding path that works without Hugging Face Spaces. Before committing, verify three things: that your chosen providers are actually covered by the API keys in .env, that HF_TOKEN is set so the OS-Atlas and ShowUI grounding calls are not rate limited, and that your machine can run the pywebview Qt display stream, which is where the README's install path starts.
Frequently asked questions
How do I use open-computer-use?
Install Poetry and ffmpeg, clone the repository, create a .env file with E2B_API_KEY and the API keys for the providers you selected in config.py, then run poetry install and poetry run start. The README states the agent opens and prompts for its first instruction, and that you can pass --prompt to start it with a specific instruction instead.
What is OpenAI computer use, and how does open-computer-use relate to it?
The README lists OpenAI as one of the supported providers, with GPT-4o and GPT-4o mini marked as vision plus action. Open Computer Use is the harness around such models: it supplies the E2B Desktop Sandbox, the display stream and the three-role model split in config.py.
Which models can open-computer-use run?
The README states it supports 10+ LLMs plus OS-Atlas and ShowUI and any other model you integrate. providers.py lists Fireworks, OpenRouter, Llama API, Groq, DeepSeek, Google, OpenAI, Anthropic, HuggingFace Spaces, Moonshot and Mistral AI, with role annotations such as vision only or action only.
Does open-computer-use need an E2B API key?
Yes. The README lists an E2B API key as a prerequisite alongside Python 3.10 or later and git, and the .env example shows E2B_API_KEY as the required variable. The sandbox is a cloud desktop, so every run depends on that key.
Why does open-computer-use ask for an HF_TOKEN?
The README says the Hugging Face token is required to bypass Gradio rate limits. OS-Atlas and ShowUI, the two grounding options it names, are served through Hugging Face Spaces, so the default grounding path goes through that endpoint.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/e2b-dev-open-computer-use)