Model or dataset
showlab/ShowUI avatar
showlab/ShowUI

ShowUI: an end-to-end vision-language-action model for GUI agents

[CVPR 2025] Open-source, End-to-end, Vision-Language-Action model for GUI Agent & Computer Use.

1,909 stars142 forksPythonApache-2.0

At a glance

What is it?
ShowUI is a 2B-parameter vision-language-action model that turns a screenshot plus a query into a GUI action. It ships training code, a vLLM inference path and a Gradio-based API, and it is the model behind the Computer Use OOTB project.
Who is it for?
Adopt ShowUI if you are building a GUI agent and want a small, open checkpoint you can fine-tune on your own screenshots, or if you want to read training code for grounding and navigation tasks. Do not adopt it if you need a supported, versioned product with a documented rollback story, or if you are not prepared to read QUICK_START.md, GRADIO.md and TRAIN.md and resolve the dependency set in requirements.txt yourself.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 160 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem ShowUI targets: clicking the right pixel from a screenshot

A GUI agent has to look at a screen and decide where to click, type or drag. The hard part is grounding: mapping a natural-language instruction such as a query about a button to a coordinate in an image. ShowUI is presented as an end-to-end vision-language-action model for this job, which means the screenshot goes in and an action comes out, without a separate detection stage that finds UI elements first and a planner that consumes them second. The repository describes it as open-source, end-to-end and lightweight, and the released checkpoint is ShowUI-2B, hosted on Hugging Face under showlab/ShowUI-2B with a mirror on ModelScope. The audience is narrow and technical: researchers and engineers building computer-use agents, and people who want to fine-tune a GUI model on their own screenshots rather than call a closed API. The README links a paper (arXiv 2411.17465), a Hugging Face Spaces demo, slides, and a companion project, computer_use_ootb, for running the model locally against a real desktop. If your problem is document understanding or general visual question answering, this is not aimed at you.

How ShowUI works: UI-guided token selection and a Qwen2.5-VL base

The mechanism the README highlights is UI-guided token selection. A screenshot is cut into patches; the repository's own example images show a Chrome screenshot producing 1296 patches, while applying a UI graph reduces the count to 167 UI components. The point is to shrink the visual token budget before the language model has to reason over it, which is what makes a 2B model plausible for this task at all. The training code supports grounding and navigation training on Mind2Web, AITW and Miniwob, and it accepts self-customized models including ShowUI, Qwen2VL and Qwen2.5VL. The update log records support for fine-tuning and inference of Qwen2.5-VL as of 2025.3.2, and vLLM inference as of 2025.2.13. Training features listed in the README include DeepSpeed, BF16, QLoRA, SDPA or FlashAttention2, Liger-Kernel, mixed dataset training, interleaved data streaming, random image resize with crop and pad, Wandb monitoring, and multi-GPU and multi-node runs. There is also an iterative refinement mode for grounding accuracy, available in the Spaces demo, and a separate ShowUI-pi project for GUI dragging. Note the shape of the repository: it is a research codebase with notebooks and scripts, not a packaged library. main/, model/, prepare/, utils/ and ds_configs/ are directories you read and adapt, not modules you import from a stable API.

Installing ShowUI and running a first inference

The README does not put install steps in the main file. It points to QUICK_START.md for local model usage and GRADIO.md for the local Gradio interface, so those two files are where the real commands live. What the repository does give you at the top level is requirements.txt, which pins accelerate==0.30.1, deepspeed==0.13.1, bitsandbytes==0.43.1, selenium==4.12.0, transformers>=4.47.0, plus qwen-vl-utils, vllm, peft, datasets and others. Because several of those pins are exact and transformers is a floor, expect to spend time on a clean environment rather than installing into an existing one.

The lightest first use does not need a GPU at all. The README states that api.py calls the model through the Hugging Face Gradio client, so you provide a screenshot and a query and the inference happens remotely.

bash
python3 api.py

Run that from the repository root and you should get a client that accepts an image path and a query, then returns the model's action output. The README's wording is that you do not need a GPU to deploy the model locally.

For local inference on your own hardware, the repository provides a notebook rather than a script. The README says to see inference_vllm.ipynb for vLLM inference, and notes that the gpu_num parameter can be adjusted to use multiple GPUs for faster inference.

python
# inference_vllm.ipynb
# adjust gpu_num to spread inference across several GPUs
gpu_num = 1

If you want the Gradio interface instead of the notebook, GRADIO.md is the installation reference. The training path is separate: TRAIN.md covers setup, and train.py plus the ds_configs/ directory hold the DeepSpeed configuration. The README's training section lists QLoRA and BF16 as options, which matters because the 2B checkpoint is small enough that quantized fine-tuning on a single card is a realistic thing to attempt.

Where ShowUI gets in your way

The documentation is uneven. The README is a link hub: it tells you that QUICK_START.md, GRADIO.md and TRAIN.md exist, but the top-level file does not restate their contents, so you cannot judge setup difficulty without opening them. There are no retrieved releases for this repository, which means there is no versioned artifact to pin and no changelog to diff against when something breaks. Upgrades therefore mean tracking the main branch. The last push was on 2026-04-24, so the code is not stale, but a moving main branch with no releases is a different kind of dependency than a tagged package.

The dependency set is the second friction point. requirements.txt mixes training, evaluation and demo concerns in one file: deepspeed and bitsandbytes sit next to selenium and opencv-python, and vllm is unpinned. Installing all of it to run one notebook pulls in far more than inference needs. The repository does not document a minimal install for inference only, and it does not document rollback if a dependency upgrade breaks your environment.

Third, the model is 2B parameters. That is the selling point and the constraint. The README describes iterative refinement as a way to improve grounding accuracy, which implies single-pass grounding is not always right. If your task needs high accuracy on unusual interfaces, plan for the refinement loop or for fine-tuning, not for a one-shot answer. And if you need a supported product with an SLA, a hosted endpoint and a compatibility guarantee, this research repository is the wrong tool.

ShowUI versus calling a hosted computer-use API

The obvious alternative is a closed computer-use model behind an API. The difference is not accuracy, which this material does not let you compare. The difference is where the model runs and what you can change. A hosted API gives you a stable endpoint, no GPU bill and no environment to maintain, but you cannot fine-tune it on your own screenshots, you cannot inspect the token selection step, and every screenshot leaves your machine. ShowUI runs locally, ships its training code, and exposes the UI-guided token selection implementation in the showui module as of 2024.12.23, so you can read and modify the part that decides which visual tokens survive. You pay for that with setup work, dependency resolution and the absence of releases. Within the open-source space, the README's own acknowledgement points to SeeClick for codes and datasets, and the update log points to ShowUI-π for dragging and ShowUI-Aloha for human demonstration workflows; those are sibling projects from the same lab with narrower scopes, not drop-in replacements. Pick ShowUI when you need to own the model and the data path; pick a hosted API when you need an endpoint this quarter and your screenshots are not sensitive.

Licence, maintenance and what an upgrade costs you

ShowUI is Apache-2.0. That is a permissive licence, and it is the same licence family as many of the dependencies in requirements.txt, but the file does not enumerate them and this article is not legal advice. Check the licences of the checkpoint you download and of qwen-vl-utils and the Qwen base models you fine-tune from, since your obligations follow the weights and the base model, not only the repository. The README asks that you cite the paper if you find the work helpful, which is a request, not a licence term.

On maintenance, the facts are these: the repository is not archived, and the last push was on 2026-04-24. The update log runs from 2024.11.16 through 2026.2.21 and shows continuing work across inference backends, training support and sibling projects. What it does not show is a release cadence. With no retrieved releases, an upgrade means pulling main and re-reading TRAIN.md and QUICK_START.md, because a change in the training entry point or the vLLM integration will not announce itself in a version number. Budget for that: pin your own fork or commit hash, and keep a working environment snapshot, since the repository documents no rollback procedure.

Editorial conclusion

Adopt ShowUI if you are building a GUI agent and want a small, open checkpoint you can fine-tune on your own screenshots, or if you want to read training code for grounding and navigation tasks. Do not adopt it if you need a supported, versioned product with a documented rollback story, or if you are not prepared to read QUICK_START.md, GRADIO.md and TRAIN.md and resolve the dependency set in requirements.txt yourself. Verify first that your GPU can hold the 2B model in the precision you plan to use, and check the inference_vllm.ipynb notebook against your CUDA and vllm versions before you build anything on top of the API.

Frequently asked questions

How do I install ShowUI locally?

The README does not put install steps in the main file; it directs you to QUICK_START.md for local model usage and GRADIO.md for the local Gradio interface. The top-level requirements.txt lists the dependencies, including accelerate==0.30.1, deepspeed==0.13.1, bitsandbytes==0.43.1 and transformers>=4.47.0.

Does ShowUI need a GPU to run?

Not for the API path. The README states that api.py is based on the Hugging Face Gradio client, so you do not need a GPU to deploy the model locally. Local inference is a separate route, covered by inference_vllm.ipynb, where the gpu_num parameter controls how many GPUs are used.

Can I fine-tune ShowUI on my own screenshots?

The repository includes training code, with TRAIN.md as the setup reference and train.py as the entry point. The README lists grounding and navigation training on Mind2Web, AITW and Miniwob, self-customized models including ShowUI, Qwen2VL and Qwen2.5VL, and efficient training options such as DeepSpeed, BF16 and QLoRA.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Project website
  4. README
  5. showlab/ShowUI on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/showlab-showui.svg)](https://hysenlabs.com/projects/showlab-showui)