Library / SDK
Mininglamp-AI/Mano-P avatar
Mininglamp-AI/Mano-P

Mano-P: Local GUI Agent for Apple Edge Devices

Mano-P: Open-source GUI-VLA agent for edge devices. #1 on OSWorld (specialized, 58.2%). Runs locally on Apple M4 Mac mini/MacBook — no data leaves your device.Mano-P 是一个开源 GUI-VLA 项目,支持在 Mac mini/MacBook 上或通过算力棒本地运行推理,实现纯视觉驱动的跨平台 GUI 自动化操作。数据完全本地处理,支持复杂多步骤任务规划与执行。

2,789 stars270 forksUnknownApache-2.0

At a glance

What is it?
Mano-P is an open-source GUI-VLA agent project from Mininglamp AI that runs inference locally on Apple M4 Mac mini or MacBook hardware, keeping all screenshots and task data on-device. Its Mano-CUA 1.1 model ranked first on the OSWorld benchmark at 58.2 percent, and a companion Cider SDK delivers W8A8 activation quantization that the MLX framework does not natively provide.
Who is it for?
Mano-P is the right project for developers and organizations who need cross-platform GUI automation without sending screenshots or task data to any cloud service. The local inference requirement means a Mac mini with M4 and 32 GB of RAM is the minimum hardware, or a compute stick connected via USB 4.0.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 98 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Private AI on the Edge: What Mano-P Is For

Most GUI automation agents rely on cloud-hosted vision models. Every screenshot sent for processing leaves the device. Mano-P is built for the opposite: a GUI-aware vision-language-action (VLA) model that runs inference entirely on Apple Silicon hardware. The README states that all screenshots and task data stay on-device and are not uploaded to external servers.

The name encodes the design intent. According to the README, Mano means hand in Spanish and P stands for Private. The project targets three groups. First, agent enthusiasts who want to build smarter computer-use workflows on top of Mano-CUA Skills without relying on external APIs. Second, developers with high security requirements who want to use a locally running GUI-VLA model as a foundation for custom Skills and Tools. Third, developers who need to fine-tune or quantize their own edge models using Mano-P's training approach, which is the third phase of the open-source release.

OSWorld Benchmark and the Mano-CUA 1.1 Results

The README reports that Mano-CUA 1.1 achieves a 58.2 percent success rate on OSWorld, ranking first among all specialized GUI agent models. The second-place result listed is opencua-72b at 45.0 percent. The README also reports a NavEval score of 41.7 on the WebRetriever Protocol I benchmark, placing it above Gemini 2.5 Pro Computer Use at 40.9 and Claude 4.5 Computer Use at 31.3.

These numbers come from the README and have not been independently verified here. The practical significance of the OSWorld gap is that 58.2 versus 45.0 is not a marginal improvement: it is 13.2 percentage points on a benchmark that represents real desktop task completion. The README attributes the result to the Mano-Action bidirectional self-reinforcement learning method and a three-stage training process: supervised fine-tuning, offline reinforcement learning, and online reinforcement learning.

Hardware Requirements and Deployment Methods

The README specifies two supported deployment methods. The first is direct deployment on a Mac mini or MacBook with an Apple M4 chip and at least 32 GB of RAM. The second uses a compute stick connected via a USB 4.0 port or higher, which makes local inference available on hardware that does not have a built-in M4 chip at sufficient memory capacity.

The current Mano-CUA-4B model achieves approximately 80 tokens per second decode speed on an Apple M5 Pro, according to the README. The companion Cider SDK adds W8A8 activation quantization primitives that MLX lacks natively. With Cider's W8A8 quantization, the README states that prefill speeds up by approximately 12.7 percent over the W8A16 baseline. The Cider SDK is not limited to Mano-P models; it works with any MLX-compatible model and delivers 1.4x to 2.2x prefill speedup over MLX W4A16 on Apple M5 Pro.

The README notes that deployment instructions for both methods will be released in the near future, and that additional deployment options are planned. As of the last push on 2026-06-25, the detailed setup guides were not yet published in the repository.

What the Three-Phase Open-Source Release Covers

Mano-P releases its code in three phases, according to the README. In the first phase, the Mano-CUA Skills are open-sourced. These are aimed at agent enthusiasts who want to build CUA task workflows that avoid the need for human intervention at bottlenecks.

In the second phase, the local-side models and SDK components of Mano-CUA are open-sourced. This targets developers who need the full GUI-VLA model running locally and want to build their own Skills and Tools on top of it. In the third phase, the training methodology and the pruning and quantization techniques used for Mano-P models are open-sourced, targeting developers who want to fine-tune their own on-device GUI-VLA models.

The top-level repository entries as of the last push are .gitignore, LICENSE, README.md, README_CN.md, and a pics/ directory. This is consistent with an early-phase release where code and model weights are not yet in the repository. Developers who need the model weights can find Mano-CUA-2.0-4B on Hugging Face under the `Mininglamp-2718` namespace and on ModelScope under `Mininglamp2718`.

Mano-AFK: Autonomous Software Construction with Mano-P

The README describes Mano-AFK as an application built on top of Mano-P. It drives a full product requirements document to code to deploy to test to fix loop using Mano-P as the local vision model for real-browser end-to-end testing. The loop takes a single natural-language prompt and produces a deployed, tested application with no human in the loop.

The specific loop is: requirement clarification, technical architecture design, code generation, local deployment, and multi-level testing including API interface testing, LLM-based page visual inspection, and end-to-end GUI automation testing driven by the VLA model. When a test fails, the system automatically locates the root cause, fixes the code, and deploys a new version. The README shows a video demonstration of this process for a fully automated application construction scenario.

This is a significant capability extension beyond basic GUI automation. It positions Mano-P as the perception and action layer in a larger autonomous software engineering pipeline, not only as a tool-use assistant for human-directed tasks.

Limitations and Where Mano-P Is Not the Right Choice

The hardware requirement is the first real constraint. A Mac mini with M4 and 32 GB of RAM is not a typical developer workstation, and not every organization has one available. The compute-stick path offers an alternative but the README does not yet document it in detail. Windows and Linux are not supported deployment targets for the local model.

The phased release model means the full project is not yet available. The training code, quantization techniques, and full SDK are still under controlled release. A team that wants to fine-tune Mano-P on their own data or adapt it to a custom application cannot do so yet. The repository structure as of 2026-06-25 reflects an early phase with minimal code in the repository itself.

The OSWorld benchmark measures task completion on desktop automation scenarios. It does not measure latency on every task type, error recovery behavior, or performance on tasks outside the benchmark distribution. The README provides no information about failure modes on tasks with unusual GUI layouts or deeply nested interfaces.

Mano-P Compared to Cloud-Based Computer Use Agents

Claude 4.5 Computer Use and Gemini 2.5 Pro Computer Use are the most directly comparable alternatives in terms of capability, and the README explicitly compares Mano-CUA 1.1 against both on the WebRetriever Protocol I benchmark. Claude Computer Use and Gemini Computer Use run models in the cloud: every screenshot is transmitted to Anthropic's or Google's infrastructure. Mano-P runs entirely on the local device.

The privacy trade-off is direct. Cloud-based computer use agents benefit from significantly larger models and continuous updates, but every interaction crosses a network boundary. Mano-P trades model scale for on-device privacy and the ability to operate without internet connectivity. For use cases in regulated industries, where screenshots may contain protected health information, financial records, or legal documents, the on-device execution is not just a preference but a compliance requirement.

Editorial conclusion

Mano-P is the right project for developers and organizations who need cross-platform GUI automation without sending screenshots or task data to any cloud service. The local inference requirement means a Mac mini with M4 and 32 GB of RAM is the minimum hardware, or a compute stick connected via USB 4.0. Teams that do not have this hardware, or that work on Windows or Linux exclusively, cannot run the edge model. The phased open-source release also means that the training methodology and quantization techniques, as well as full SDK components, are not yet available; the last push was on 2026-06-25, so these phases may still be in progress. Check the repository and the Feishu wiki documentation before committing to a production integration.

Frequently asked questions

What hardware does Mano-P require for local inference?

The README specifies direct deployment on a Mac mini or MacBook with an Apple M4 chip and at least 32 GB of RAM. A second path uses a compute stick connected via USB 4.0 or higher. Windows and Linux are not documented as supported deployment targets for the local model.

Where can I find the Mano-P model weights?

The README links to Mano-CUA-2.0-4B weights on Hugging Face under the Mininglamp-2718 namespace and on ModelScope under Mininglamp2718. The repository itself does not bundle the weights.

Does Mano-P send data to external servers?

No. The README explicitly states that all CUA operations are executed on the local Mac mini and will not be uploaded to external servers. No cloud API calls are required for inference once the model is deployed locally.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Mininglamp-AI/Mano-P on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mininglamp-ai-mano-p.svg)](https://hysenlabs.com/projects/mininglamp-ai-mano-p)