# The AskUI SDK gives an agent a screen and a keyboard, and its dependency pins read like a changelog

> A Python automation framework that drives desktops, phones and embedded displays by looking at pixels rather than reading selectors, with two ways to call it: single commands, or a sentence of intent. The most telling file in the repository is not the readme but the dependency list.

**askui/python-sdk** — Enable AI to control your desktop, mobile and HMI devices

- Repository: https://github.com/askui/python-sdk
- Website: https://docs.askui.com/
- Stars: 553 · Forks: 61
- Language: Python
- License: MIT
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/askui-python-sdk

## Selectors break, and that is the argument for the whole design

The case for a vision-based approach opens with the failure mode of everything else: a button moves, a label changes, a layout shifts, and the script breaks. What follows from that in the experience of most teams is a maintenance tax, with brittle selectors, conditional logic written for edge cases, and tests updated every time a designer touches a screen.

The alternative is described in two parts. One is finding interface elements by what they look like or say rather than by a selector in the page, which removes the coupling to layout. The other is giving the system a high-level instruction and letting the model work out the steps, which removes the coupling to the sequence.

Both parts come with the same honest caveat, which the documentation does not hide: there are two modes, single-step commands or agentic instructions, and they have very different properties. A single step is a lookup with a model in the middle of it. An instruction is a model deciding what to do with a screen. If you only need the first, the framework is a much more predictable tool, and the examples show both styles side by side because they are meant to be mixed.

## Two entry points, and the intent one is short enough to be tempting

The whole interface in the readme's first example is a context manager, an act method, and a get method. The act call takes a sentence of intent and does the multi-step work; the get call takes a question and returns text pulled off the screen.

```python
from askui import ComputerAgent

with ComputerAgent() as agent:
    # Complex multi-step instruction
    agent.act(
        "Open a browser, navigate to GitHub, search for 'askui vision-agent', "
        "and star the repository"
    )

    # Extract information from the screen
    balance = agent.get("What is my current account balance?")
    print(f"Balance: {balance}")

    # Combine both approaches
    agent.act("Find the login form")
    agent.type("user@example.com")
    agent.click("Next")
```

The last three lines are the part worth copying, because they show the two modes composed. An instruction to find the form, then a literal type, then a literal click. That is a realistic pattern for anything that touches credentials or money, and it is the pattern the documentation is steering you towards without quite saying so.

There is also a separate agent class for phones alongside the computer one, and a documentation set covering both. Mobile is a genuinely different target rather than a rebranded desktop, which is why it gets its own class and its own tooling.

## You pick the model, and you can bring your own

Running any of this needs a model behind it, and the project hosts its own models while supporting several third-party ones out of the box, with two named in the readme. The framework's own hosted option is the easiest starting point because it needs no provider account of your own.

That path needs two things from you: a workspace and a token, obtained from the project's hub, with the free trial called out as not requiring a card. Both go in environment variables, and the readme gives the syntax separately for Unix shells and for Windows PowerShell:

```shell
export ASKUI_WORKSPACE_ID=<your-workspace-id-here>
export ASKUI_TOKEN=<your-token-here>
```

If you would rather use your own account or your own cloud, that is a documented path with its own page, and the readme frames it as plugging in a provider rather than forking the library. This matters more than it sounds: an organisation with an existing agreement for one model has a policy reason not to route UI automation through someone else's inference, and being able to switch is what makes the tool adoptable inside a company that has opinions about that.

## The dependency pins are a written history of upstream breakage

The manifest is the most informative file in the repository, because the pins are not arbitrary. One dependency is capped below its next major release with a comment explaining that the newer version dropped a keyword argument the SDK still sends, which crashed every model call. Another is pinned to an exact version with a comment saying that listing tools inside the runner fails otherwise. A third carries a comment noting it needs an older Python because of a transitive scientific dependency.

There are also two upper bounds on Python itself, which means the package declares a range rather than a floor, and the readme mentions only the lower one. That mismatch is the kind of thing that turns into a confusing installation failure on a machine with the newest interpreter.

Reading the rest of the list tells you what the framework is actually made of: a client for its own device-side component, both vendor model libraries, a Google client, a protocol-buffer runtime, a gRPC stack, an MCP implementation, a templating engine, an analytics client, a retry library, image hashing, an image library, clipboard access, a machine identifier, a settings library, a FastAPI server, a Gradio client, an async helper library, and Android device tooling. That is a large surface for a UI automation library, and each entry is a place where an upstream release can break you.

## Optional extras exist for eight integrations, and the readme says to skip them

The install is one command, and the framework needs a recent Python:

```bash
pip install askui
```

The default install is deliberately broad and the extras are deliberately narrow. Document reading for office formats is one, routed through a converter library and intended for use inside the extraction method. Running a model vendor through their managed cloud services is two more, one per cloud, for organisations that route those models through an internal agreement rather than a direct account.

Tracing is a third, covering export over the standard protocol plus optional instrumentation of the HTTP client and the database layer, and the documentation explicitly says it is for production pipelines and not needed for a local script. There is a browser extra, a flag for continuous integration images, and an everything extra.

What to take from the table is the recommendation attached to the everything option: prefer picking individual extras when you know what you need, and the readme goes further and suggests starting minimal and adding as you need them. That is unusual advice for a package that could have sold you everything at once.

## Custom tools are passed in per call, not registered globally

Extending the agent is a matter of passing a list of tools into the call, and the tools come from a store with separate modules for computer-scoped and universal behaviour. The example saves a screenshot to a directory and prints to the console:

```python
from askui import ComputerAgent
from askui.tools.store.computer import ComputerSaveScreenshotTool
from askui.tools.store.universal import PrintToConsoleTool

with ComputerAgent() as agent:
    agent.act(
        "Take a screenshot of the current screen and save it, then confirm",
        tools=[
            ComputerSaveScreenshotTool(base_dir="./screenshots"),
            PrintToConsoleTool()
        ]
    )
```

Scoping tools to a single call rather than registering them on the agent object is a small choice with real consequences. It means an agent can have a screenshot-taking tool for one step and not for another, which is exactly the granularity you want when you are debugging a specific action and do not want every step writing a file.

The standard interface for third-party tools is the Model Context Protocol, so a tool written for another agent framework can be handed to this one. There are example files for model providers, for driving several machines from one agent, for reading documents, for handling secrets, and for a computer-only setup, which tells you the intended shape of a real installation pretty precisely.

## Caching, secrets, and the loop you have not thought about yet

Three documentation entries exist that most frameworks in this space do not have at all, and they are the ones that separate a demo from something running unattended.

The first is caching, which exists because vision-based steps are expensive: the framework's stated goal is to use a paid model call only when it is genuinely necessary. Without it, every act call is a billable request and a loop over a hundred steps is a hundred requests. The second is secrets, which lets an agent use a password without the value ever being exposed to the model, which is the single most important feature on this list for anything touching a login. The third is callbacks, for injecting your own logic into the control loop between steps.

Around those sit observability, reporting that turns agent logs into test reports, and structured extraction from screenshots and files. That combination, caching plus secret handling plus reporting, is what tells you this is aimed at scheduled runs rather than at a person watching a browser. It is also a fair signal of where the project is going, since a vision agent that can report, cache and hold secrets is halfway to being a test runner.

## Conclusion

Decide whether this is a testing tool or an automation tool before you install it, because the two produce very different code. If you are rewriting brittle selectors, the vision approach is genuinely better than what you have. If you need a test that fails when a specific button is missing, an agent that decides for itself what the screen means is the wrong instrument. Two operational points are worth reading before the first run: the hosted model needs a workspace and a token, and the whole approach bills model calls unless you set up caching, so a loop that runs a hundred steps is a hundred model calls unless you say otherwise.

## FAQ

### What is the AskUI Python SDK for?

It lets a Python script or an AI agent control desktop machines on Windows, macOS and Linux, mobile devices on Android and iOS, and embedded HMI displays. Elements are found by what they look like or say rather than by brittle XPath or CSS selectors, and calls can be single commands or high-level instructions.

### How do I install the askui package?

With a single pip install, and it needs a recent Python. The default install covers everyday automation with the smallest footprint. Optional extras add specific integrations one at a time, including office document reading, running models through managed cloud services, tracing for production pipelines, and browser automation, and the documentation recommends starting minimal and adding extras as needed.

### Do I need an AskUI account to use it?

You need a model behind the agent, and the easiest path is the hosted option, which requires signing up to get a workspace ID and an access token set as environment variables. If you would rather use your own provider, including your own cloud account, the project documents plugging in a model provider rather than forking the library.

### How do I keep agent costs down?

Caching is a documented feature whose stated purpose is to avoid expensive model calls when they are not needed. That matters because every vision-based step is a model request, so a long loop multiplies. Secrets are handled separately, so an agent can type a password without the value ever being sent to the model.

### Can I add my own tools to the agent?

Yes, by passing a list of tools into an individual call rather than registering them globally. Tools come from a store with separate modules for computer-scoped and universal behaviour, and the standard Model Context Protocol is supported, so tools built for another agent framework can be reused.

### What Python versions does askui support?

The readme states a minimum of Python 3.10. The package manifest declares a narrower range with both a lower and an upper bound, and the upper bound exists because one transitive dependency needs an older Python, so a machine with a very new interpreter will not be supported even though the readme only mentions the floor.

## Sources

- [askui/python-sdk on GitHub](https://github.com/askui/python-sdk)
- [License: MIT](https://github.com/askui/python-sdk/blob/main/LICENSE)
- [Project website](https://docs.askui.com/)
- [README](https://github.com/askui/python-sdk/blob/main/README.md)
- [Releases](https://github.com/askui/python-sdk/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/askui-python-sdk
