Model or dataset
simular-ai/Agent-S avatar
simular-ai/Agent-S

Agent S: a Python framework that drives a real desktop GUI with mouse and keyboard

Agent S: an open agentic framework that uses computers like a human

12,312 stars1,446 forksPythonApache-2.0

At a glance

What is it?
Agent S, published on PyPI as gui-agents, takes a natural-language task, reads the screen, and acts through clicks and keystrokes. This review covers what the repository actually ships, how to install it, and where it stops being the right tool.
Who is it for?
Adopt Agent S if you are researching OS-level agents or automating a desktop you control and can watch, and you accept that the safety story is yours to build. Do not adopt it if you need a hosted, production-grade agent with a support contract; the README points that audience at Simular Cloud and Sai instead.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 24 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Agent S fills: tasks with no API behind them

Most automation assumes an interface. A REST endpoint, a CLI, a library. A large share of the software people actually use every day has none of those. Legacy desktop clients, internal admin panels, vendor tools that ship a window and nothing else. Agent S targets exactly that gap. The README describes it as an open source computer use agent framework that "takes a natural-language task, looks at the screen, and completes the task by clicking, typing, and scrolling in ordinary desktop and web applications, with no API integration and no per-app scripting required."

The audience is narrower than the tagline suggests. This is a research framework first. The repository ships benchmark setup directories (osworld_setup/, WAA_setup.md, evaluation_sets/) alongside the library, which tells you the primary consumer is someone reproducing or extending published results. The secondary audience is engineers automating their own machine, where the cost of a wrong click is low and a human is nearby. If you want a finished product rather than a framework, the README is explicit: it points to Simular Cloud and to Sai, the hosted agent built on the same ideas.

How the agent turns a screenshot into a click

The architecture visible in the repository is a loop, not a pipeline. The agent captures the screen, a multimodal model reasons over what it sees plus the task description, and the framework emits a low-level action: a click at coordinates, a keystroke, a scroll. Then it captures again. The package name gui-agents and the topic tags (grounding, planning, memory, retrieval-augmented-generation, in-context-reinforcement-learning) describe the pieces the maintainers consider part of that loop.

The grounding problem is the interesting one. A model that says "click the Save button" has not solved anything until something maps that phrase to pixels. The dependency list shows the intended approach: paddleocr and pytesseract for text detection in the screenshot, scikit-learn for the matching machinery around it. That is a classical detection stack rather than a purely end-to-end vision model, and it shapes the failure modes. Text that OCR cannot read, custom-drawn icons with no label, or a canvas-rendered UI will not ground reliably.

Input is injected through pyautogui on every platform, with pyobjc on macOS and pywinauto plus pywin32 on Windows. Those platform-conditional entries in setup.py mean the same agent code path produces different native calls depending on where it runs. The console entry point agent_s maps to gui_agents.s3.cli_app:main, so the S3 generation of the agent is what the installed command launches.

Installing gui-agents and running the agent_s command

The distribution name and the project name differ, which trips people up. The repository is Agent-S; the package you install is gui-agents, as declared in setup.py. Python support is bounded: setup.py declares python_requires=">=3.9, <=3.12", so a 3.13 environment will refuse the install rather than fail later.

A plain pip install pulls the full dependency set, including paddleocr and paddlepaddle:

bash
pip install gui-agents

The repository also ships a requirements.txt at the top level with the same core list, which is the path to use if you are working from a clone rather than from PyPI. Note the platform markers in that file: pyobjc is pulled only on Darwin, pywinauto and pywin32 only on Windows.

The install provides a console script named agent_s. Running it with no arguments is the documented entry point into the CLI:

bash
agent_s

What you should see is the CLI starting up. The README does not reproduce the CLI's flags, so the reliable next step is the Usage section of the README rather than guessing at options. Expect the first run to be slow: paddleocr and paddlepaddle are large, and the agent needs a vision-capable model from OpenAI, Anthropic, or an open-weight provider before it can do anything at all. The README states that models from all three categories are supported, and the repository includes a models.md file for provider configuration.

Where the loop breaks: OCR, cost, and no rollback story

Three limitations are visible from the repository alone, and none of them are hidden.

First, grounding is only as good as the pixels. Because text extraction leans on paddleocr and pytesseract, the agent's competence is uneven across interfaces. A dense web form with standard widgets is a friendly case. A game, a CAD tool, or an app that draws its own controls is not. The dependency choice is reasonable, but it is a choice with a predictable ceiling.

Second, every step costs a model call. The loop captures the screen and reasons about it repeatedly, so a twenty-step task is twenty rounds of multimodal inference. The README's cost comparisons are made about Sai, the hosted product, not about running this framework yourself. Anyone budgeting for self-hosted use should measure their own spend rather than assume the hosted figures transfer.

Third, and most important: the README does not document rollback, sandboxing, or an undo mechanism. An agent with pyautogui access to your desktop can delete a file, send a message, or click through a confirmation dialog, and nothing in the repository provides a built-in guard for that. This is not a criticism of the code so much as a statement about scope. The framework gives you the loop; the containment is an exercise left to the operator.

Agent S against scripted automation and hosted computer use

The obvious alternative is Selenium, which is already in the dependency list. The difference in approach is fundamental. Selenium drives a browser through the DOM: you address elements by selector, and the automation is deterministic and fast. Agent S drives the operating system through pixels: you address elements by describing them, and the automation is probabilistic. If your target is a web page you control, Selenium is strictly better, cheaper, and easier to debug. Agent S earns its place when the target is a native window, a mixed desktop workflow, or a page whose DOM you cannot reach.

The other alternative is the hosted route, and the README makes this comparison for you. Sai, Simular's hosted agent, is described as reaching 73% on OSWorld 2.0, ahead of GPT-5.6 Sol at 62.57% as reported by OpenAI, and at lower cost. Those numbers are about the hosted product. The distinction that matters for adoption is operational: a hosted agent comes with someone else's infrastructure, monitoring, and update cadence. Self-hosting gui-agents gives you control over the model provider and the machine, and hands you the operational burden in exchange.

Licence, maintenance, and what an upgrade actually costs

Agent S is Apache-2.0, which is permissive: commercial use, modification, and redistribution are allowed, with the usual requirements around preserving notices and stating changes. That is a summary of the licence family, not legal advice, and if you are embedding this in a product you should read the LICENSE file in the repository root.

Maintenance is current. The last push to the default branch was on 2026-09-05, and the most recent tagged release is v0.3.2 from 2025-12-16. The release history shows a steady cadence: v0.3.0 and v0.3.1 landed a day apart in October 2025, then v0.3.2 in December. Version numbers in setup.py track the release tags, so pip install gui-agents==0.3.2 is a pinned, reproducible target.

The upgrade cost is the dependency surface. paddleocr, paddlepaddle, pyautogui, selenium, and three separate model SDKs (openai, anthropic, google-genai) all move independently of this project. A pinned version is the sane default for anything you depend on. The project's own version history also shows the API is not frozen: S1, S2, S2.5, and S3 are distinct generations, and the CLI entry point points specifically at gui_agents.s3, which means a future S4 could move it. Budget for reading release notes before bumping.

Editorial conclusion

Adopt Agent S if you are researching OS-level agents or automating a desktop you control and can watch, and you accept that the safety story is yours to build. Do not adopt it if you need a hosted, production-grade agent with a support contract; the README points that audience at Simular Cloud and Sai instead. Before committing, verify three things: that your Python version falls inside the 3.9 to 3.12 window declared in setup.py, that a vision-capable model key is available for the provider you intend to use, and that paddleocr and paddlepaddle install cleanly on your platform, since they dominate the dependency weight and are the most likely source of a broken first run.

Frequently asked questions

What is the difference between Agent S and the gui-agents package?

They are the same project under two names. The repository is simular-ai/Agent-S, while setup.py declares the distribution as gui-agents version 0.3.2. Installing gui-agents is what gives you the agent_s command.

Which operating systems does Agent S support?

The README states it runs on macOS, Windows, and Linux, and setup.py carries platform markers for pyobjc on Darwin and pywinauto plus pywin32 on Windows. The classifiers list Windows, Linux, and macOS.

What Python version do I need to install Agent S?

setup.py declares python_requires=">=3.9, <=3.12". Environments outside that range will not install the package.

What are the top 3 AI agents?

The repository does not rank AI agents in general. It only makes claims about itself and Simular's hosted product: the README states Agent S3 was the first computer use agent to surpass human performance on OSWorld at 72.60%, and that Sai reached 73% on OSWorld 2.0.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. simular-ai/Agent-S on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/simular-ai-agent-s.svg)](https://hysenlabs.com/projects/simular-ai-agent-s)