clickclickclick: two model roles, four interfaces and a placeholder clone URL
Autonomous Android and computer use using any LLM (local or remote)
At a glance
- What is it?
- An experimental framework for Android and desktop automation driven by LLMs, split into a planner and a finder with an executor behind them. It runs from a CLI, a Gradio interface, a Python API and a REST endpoint, works with Ollama, Gemini and OpenAI models, and its documentation still carries a template clone URL.
- Who is it for?
- clickclickclick fits a developer experimenting with agent-driven device automation who wants planner and finder to be different models, since that separation is the whole design and it is the cheapest way to trade cost against reliability. It does not fit production automation, and the README says as much in one sentence before the install steps.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Three components, and two model roles you can swap
The design is stated as three components: a planner, a finder and an executor. What you configure is two of them, because the planner decides the steps and the finder locates the elements on screen, and both are chosen separately from the command line or the setup wizard.
That separation is the interesting part. A planner that understands intent and a finder that reads a screenshot are different jobs, and the README is explicit that the split has a cost: a small local model such as `qwen3.5:4b` works as planner, with slow basic navigation, but is not reliable as finder because its UI element detection is poor.
The framework targets two platforms, `android` as the default and `osx`, and lists Ollama models including Llama 3.2-vision and qwen3.5:4b alongside Gemini and GPT 4o. Its own recommendation is specific about the best current pairing: Gemini 3.1 Flash-Lite as both planner and finder, which is the opposite of the default.
For macOS the dependency list carries a Python package for AppleScript, and for Android the README states the hard prerequisite plainly: adb has to be installed on the machine where the code runs.
The default pair is asymmetric: OpenAI plans, Gemini finds
The documented defaults are not the recommended configuration, and that gap is worth knowing before you draw conclusions from a first run.
Planner defaults to `openai`, finder defaults to `gemini`, and the three accepted values for each are `openai`, `gemini` and `ollama`. Both can be overridden on the command line, so a run can be entirely local, entirely remote, or mixed:
python main.py run "Open Google news" --platform=android --planner-model=openai --finder-model=geminipython main.py run "Open Reddit" --platform=android --planner-model=ollama --finder-model=gemini --image-quality=45The first run of all asks you to configure both roles through a `setup` command, where you pick the models and supply any API keys, and model-specific settings live in a `config/models.yaml` file whose keys you then export.
Four ways in: CLI, Gradio, Python API, REST
The same task can be started four ways, which is generous for an experimental project and useful when you are embedding it.
The CLI is the `click3` entry point declared in the project metadata, invoked as `./click3 run <task-prompt>`, and the package also installs as a wheel. The Gradio interface is the web surface. A Python API lets you import the framework directly, and the REST API exposes a single endpoint, `POST /execute`.
That endpoint takes a task prompt plus the platform, planner model, finder model and image quality, and its error contract is spelled out: a 400 with a detail string for an unsupported platform or model, and a 500 with a detail when execution fails, otherwise a 200 carrying a result object. The documented call is a curl to `http://localhost:8000/execute`, and the server behind it is started with `uvicorn api:app --reload`.
The clone command still points at yourusername
The installation instructions begin with a URL that was never filled in:
git clone https://github.com/yourusername/clickclickclick
cd clickclickclickIt is a template that survived into the published README, and anyone following the steps literally clones nothing. The rest of the sequence is more careful: create a virtual environment, activate it, then `pip install -r requirements.txt`.
Two other headings in the documentation are also empty or duplicated. The project structure heading has no content under it, the running-tasks section repeats the wheel install command instead of the run command, and the contributing notes give pre-commit its own three commands where a repository already ships a `.pre-commit-config.yaml`.
None of this stops the framework working, and it is worth knowing before you spend time assuming you misread something.
Image quality is the performance lever, not the model
The one tuning knob with a documented effect is `--image-quality`, a percentage from 1 to 100 that defaults to 100. Lower values shrink the screenshots the finder sees, which makes processing faster.
The Ollama guidance ties the two together: `qwen3.5:4b` is usable as a planner but weak as a finder, and the README suggests running it with `--image-quality=45` for better performance. So the recommendation for a small local model is to spend resolution on the model that needs it.
The three demos show what the thing is for, and they are unglamorous: drafting a Gmail message with a specific subject line, body length and tone, opening a mapping site and finding bus stops in a named town, and starting a timed chess game on a public board. Each is a task with a checkable outcome, which is the right way to evaluate this kind of framework.
mlx support exists only in requirements.txt
The newest release, v0.3.0 from 2024-12-26, is titled as adding Molmo support via mlx. The two packages that support would need, `mlx==0.21.1` and `mlx-vlm==0.1.4`, appear in `requirements.txt`.
They do not appear in the project metadata. The dependency list in the package file has no mlx entries at all, so an install from the built wheel would not get the model support that release announced. Anyone following the README installs from `requirements.txt` and never notices; anyone installing the package does notice.
The same file mixes dev tooling into runtime dependencies. `pytest`, `black` and `pre-commit` are all in the main list, alongside an optional `test` extra that pins pytest again at a lower bound, and two Google client libraries where one would usually be enough. The requirements file itself pins nothing but the two mlx packages, and the export script that generates it passes `--without-hashes`, so the pinned list is weaker than the one it was generated from.
Three tags in one December and a manifest nobody tagged
The release history is short and sits entirely in December 2024. v0.1.0-alpha is titled as the first working cut and is dated 2024-12-17 at 06:12, v0.2.0-alpha follows on the same day at 09:21, and v0.3.0 arrives on 2024-12-26. Two alpha tags three hours apart is the shape of a project finding its footing, and the version string in the package metadata has since moved to 0.3.1 without a tag.
Since then the repository has been pushed as recently as 2026-03-17, which is well past the newest release, and the README describes the code as highly experimental, expected to evolve in future commits, and to be used at your own risk.
So the honest summary of the state is: an active-enough experiment with a vendored commercial backer behind the instavm.io homepage, a December 2024 release line, and a Raspberry Pi readme sitting in the root that the main documentation never links to.
Editorial conclusion
clickclickclick fits a developer experimenting with agent-driven device automation who wants planner and finder to be different models, since that separation is the whole design and it is the cheapest way to trade cost against reliability. It does not fit production automation, and the README says as much in one sentence before the install steps. Before you build on it, note that the last commit is dated 2026-03-17 and the newest tag is v0.3.0 from December 2024, that the manifest version is 0.3.1 with no matching tag, and that the mlx dependencies the newest release added are declared only in requirements.txt rather than in the package metadata.
Frequently asked questions
What is clickclickclick?
A framework for autonomous Android and computer use driven by any LLM, local or remote. The README names three components, a planner, a finder and an executor, and the planner and finder are separately configurable between openai, gemini and ollama.
How do I install clickclickclick?
Clone the repository, create and activate a virtual environment, then run `pip install -r requirements.txt`. Model-specific settings go in config/models.yaml and the keys it names have to be exported. Python 3.11 or newer is required, and adb must be installed locally for Android work.
Which models does clickclickclick support?
Local models through Ollama, including Llama 3.2-vision and qwen3.5:4b, plus Gemini and GPT 4o. The README says Gemini 3.1 Flash-Lite used as both planner and finder gives the best current result, while a small local model works as planner but not reliably as finder.
How do I run clickclickclick as an API?
Start it with `uvicorn api:app --reload` and POST to `/execute` with a task prompt plus optional platform, planner model, finder model and image quality. Responses are 200 with a result object, 400 with a detail for an unsupported platform or model, and 500 with a detail when execution fails.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/instavm-clickclickclick)