# KnowAct's GUI agent is packaged as nanobot-ai, installed without its desktop extra

> KnowAct from the Lychee Team at Harbin Institute of Technology Shenzhen ships GUIClaw, a personal GUI assistant that drives desktop, Android, iOS and HarmonyOS screens through a know, route, act and reflect loop. The README describes two installation paths, five backends and seven runtime directories in detail, and leaves three things worth reading twice: the distribution is named nanobot-ai, the default install omits the extra the default backend needs, and the permission policy it initialises for you is model guidance rather than enforcement.

**HITsz-TMG/KnowAct** — Recursive Self-Improvement Personal Assistant

- Repository: https://github.com/HITsz-TMG/KnowAct
- Website: https://shibosusu.github.io/KnowAct-GUIClaw/
- Stars: 486 · Forks: 41
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/hitsz-tmg-knowact

## The package is called nanobot-ai and the repository has no versioned release

The project has two names and three identities. The repository is KnowAct, the code lives in a subdirectory called GUIClaw, and the file that says so is explicit: the current Python distribution is named nanobot-ai and ships both packages. That name is the one a reader would need in order to install anything from an index, and it matches neither the repository nor the tool. The standalone path is careful to add that it uses only the guiclaw executable at runtime and that guiclaw does not import nanobot, which is a clean separation but leaves a reader looking for a package called guiclaw that does not exist under that name. The top level listing holds five entries, .github/, GUIClaw/, LICENSE, README.md and README_CN.md, so the repository root is nearly empty and both quick starts have to change directory first, into KnowAct/GUIClaw for the host path and into KnowAct for the standalone one. There is also no versioned release: the single release on the project is tagged Result and named GUIClaw-Result, dated 2026-07-12, and the README points at it as the place to read real device experimental logs and trajectories rather than as a package. The host path adds a third naming problem: three entry points are offered, nanobot webui, nanobot gateway and nanobot agent, of which only the first appears in a quick start, so a reader has no way to tell from the file whether the other two are alternative surfaces or unfinished. The project itself is MIT licensed and shows 0 open issues on the issue tracker, with 486 stars and 41 forks, and its topic list is empty. The README also carries a Chinese version of itself as a sibling file and credits the Lychee Team at Harbin Institute of Technology Shenzhen, which is the only place the project says who wrote it.

## The default install leaves out the extra the default backend needs

The standalone quick start shows two install lines and picks the first one:

```bash
git clone https://github.com/HITsz-TMG/KnowAct.git
cd KnowAct

# Android, HarmonyOS, or dry-run
uv tool install ./GUIClaw

# Use this instead for local desktop automation
# uv tool install './GUIClaw[desktop]'
```

The line the reader is told to use is annotated for Android, HarmonyOS or dry run, and the desktop variant is left commented out. That matters because the default backend is the desktop one. The commands table describes a bare guiclaw TASK as running with the default local backend and defines local as foreground desktop automation, and the defaults table repeats it: backend local for the standalone CLI, adb for the nanobot adapter. So the documented out of the box pairing is an install without the desktop extra pointed at the backend that needs it. The complete host path takes the opposite approach and pulls three extras at once, with uv sync --extra web --extra desktop --extra cjk, and its onboarding step is a separate wizard run afterwards. Requirements for the host are Python 3.11 or newer, uv, a multimodal model, and whatever the chosen backend needs. The two smoke tests that follow are the practical part of that path: an API key is exported into OPENAI_API_KEY, then one command runs with the dry run flag and a task that describes the current screen and finishes, and the next runs the same way over ADB with a task that opens Settings and enables Wi-Fi. A third form exists for callers that are not a human at a terminal, where the task arrives through a task flag alongside a json flag, for example asking to open Contacts and search for a name. The standalone tool is also described as callable from a terminal, a script, Hermes, OpenClaw or another host.

## Two config files, two naming conventions, and one setting with no row

Runtime configuration lives in two places depending on which path you took, and the two files do not share a naming style. The host reads a gui block inside ~/.nanobot/config.json, where the keys are camel case: maxSteps, enableSkillExecution, enablePromptSkillSelection, promptSkillTopK, promptShortcutOnly and promptSkillAppFilter, alongside a providers entry carrying apiKey and apiBase. The standalone CLI reads ~/.guiclaw/config.yaml, which is snake case throughout:

```yaml
provider:
  base_url: "https://api.example.com/v1"
  model: "your-vision-model"

max_steps: 15
stagnation_limit: 0
image_scale_ratio: 0.5
agent_profile: default
```

One key in that file has no row in the defaults table: stagnation_limit is set to 0 here and never mentioned again in the README, including in the pointer to the CLI and configuration reference that is supposed to hold complete defaults. The values that do appear are consistent across both columns, with 15 maximum steps and an image scale of 0.5, and the agent profile is called default when unset for the adapter. What a reader has to reconcile by hand is the casing, which is exactly the kind of thing that produces a silently ignored block. The two files also live in different homes, one under the nanobot workspace and one under a guiclaw directory of its own, and the host file carries a second requirement the file states in passing, namely that the provider must be configured for nanobot as well as referenced by the gui block, so a gui block naming a provider that does not exist is the first failure a new user meets, and nothing in the example shows the provider side of that pair being filled in.

## The example config enables what the defaults table calls disabled

The gui block in the host example sets enableSkillExecution to true and enablePromptSkillSelection to true, with promptSkillTopK at 5 and both the shortcut-only and app-filter flags off. The defaults table two sections later says something narrower. Its row headed prompt skill selection names a setting called enable_skill_execution, which is the skill execution switch rather than the prompt selection switch, and reports that setting as disabled by default in the standalone column and disabled by default again in the adapter column. A separate row for skill and memory extraction says those are separate YAML switches, also disabled by default on both sides. So the copy and paste example enables two capabilities the table describes as off, and the row that should describe the first of them appears to be about the second. The JSON also puts a placeholder model name, a placeholder API base and a placeholder API key in the same block, so the file is a template with defaults switched on rather than a working configuration, and the standalone YAML carries the same kind of placeholder pair under its provider heading.

## The permission policy it writes for you is guidance, not a control

When no POLICY entry exists, the first GUI task initialises one, and the defaults it picks are conservative: deny, cancel or defer any request unless the task explicitly authorises it. The entry lives in ~/.guiclaw/memory/policy.md as a structured block that a user can edit, and the same file is described as always-injected memory. What makes this section worth reading carefully is the paragraph that follows: the rules are model guidance, not guaranteed enforcement, and anyone who needs a mandatory restriction should use operating system controls, backend controls, host approval or a sandbox instead. The README's own callout agrees, warning that GUI automation can click, type, launch applications and change device state, and telling the reader to start with the dry run flag and to use a test device or account while validating. That is an unusually plain admission for a tool whose default backend is a live screen, and it is the reason the dry run mode sits in the commands table as a first class entry rather than as a debug flag. One thing that policy does not cover is the host path. When GUIClaw runs as a host it registers itself as a task type and the host routes screen work to it while keeping files, shell, web, MCP, memory and other tools inside the host runtime, so the deny, cancel and defer rules sit next to a set of host capabilities that the same README does not put under the same policy.

## Validated shortcuts get promoted into a Python module

Android shortcuts move through three stages, and the third one writes code. Static inference takes a source, which the commands table describes as a manifest or a manifest directory, and writes one cache file per package into ~/.guiclaw/shortcut_cache/. Runtime validation then launches the candidates on a device over ADB, optionally consulting a vision model first, which needs a DashScope key exported under a variable whose name begins with DASHSCOPE_A; those records land in ~/.guiclaw/shortcut_cache_validation/. The promote step is the one to notice: given a cache file and the validate flag plus promote, the command adds eligible validated shortcuts to the shared skill file, and that shared file is ~/.guiclaw/skill/skills.py, described as validated shortcuts followed by extracted skills. The commands table also exposes shortcuts SOURCE on its own, and separate help output for the task command and the shortcut command, which is a small sign that the shortcut pipeline is treated as a subsystem of its own. So a candidate inferred statically from a manifest, checked at runtime on a real device and judged by a vision model can end up appended to a Python module that the agent then imports, and the file never says what makes a candidate eligible for promotion or who reviews the result. The README also points at a CLI and configuration reference inside the repository for installation extras, every flag, complete defaults, backend setup, shortcut validation and subprocess integration, which is where the missing stagnation_limit row would belong, and at a separate adapter contract document for anyone embedding the tool rather than calling it.

## Self evolving memory sits behind switches that are off by default

The name promises recursive self improvement, and the machinery exists: an induced GUI memory bank at ~/.guiclaw/memory/gui_memory_bank.jsonl, a policy file beside it, and a run history directory at ~/.guiclaw/gui_runs/ that holds screenshots, compact trajectories and task results. Every run leaving a screen trace plus a trajectory plus a result is what makes replay and later extraction possible at all. The catch is in the defaults table, where the row for skill and memory extraction says separate YAML switches, disabled by default in the standalone column and disabled by default in the adapter column, matching the row above it for prompt skill selection. The architecture section describes the loop with four stages named know, route, act and reflect, and the extraction step is where the loop would close, since a run has to produce material that later runs can learn from. The architecture section is one sentence long, naming four stages and stopping there, so the recursive part of the description is carried by those two files rather than by anything the file explains about how memory is built, bounded or reviewed. As shipped, the learning half is off unless a reader turns it on, and the README does not say how large the memory bank gets or what it does when a screen layout changes underneath it. Seven directories and files make up that runtime state, all of them under a guiclaw home that the file says is kept outside the nanobot workspace: the configuration file, the run directory with its screenshots and trajectories, two shortcut directories for inferred candidates and for validation records, the skills module, the policy file and the memory bank.

## The benchmark sentence names its own base model twice

The performance claim sits in one long sentence near the top: GUIClaw with the open source Kimi-2.6 reaches 64.1 percent on the long horizon MobileWorld benchmark, described as outperforming all open agent frameworks and closed agents including Seed-2.0-Pro and GPT-5.5. The next sentence says the framework's knowledge memory and execution capabilities generalise across base models, and gives plus 8.5 percent on Kimi-2.6 and plus 16.2 percent on Qwen3.5-35B-A3B. Kimi-2.6 is therefore the model behind the headline number and also one of the models said to gain from it, and neither percentage is given a baseline, a metric definition or a comparison against the unmodified model. The repository description is shorter still, calling it a recursive self improvement personal assistant with no benchmark in it at all. The rest of the header points at a Chinese README, a paper on arXiv, a website with a demo anchor, a CLI reference and an adapter contract document, and the default branch main last moved on 2026-08-26. None of those links is referenced from the body of the file, so a reader who wants to check the 64.1 percent against the paper has to leave the README to do it, and the project ships no version of its own to compare against.

## Conclusion

This is a research group's GUI agent with the documentation of one, and the documentation is unusually honest about where it stops. Read it if you need screen automation across four platforms from one command line, and treat the two install paths as genuinely different products: the host path gives you a web UI and a tool registry, the standalone path gives you one executable that deliberately does not import the host. Before you run either, confirm that the extra set you installed matches the backend you intend to use, work out whether prompt skill selection and memory extraction are meant to be on, and put real enforcement in front of the permission policy, because the project says itself that the policy is guidance to a model rather than a control. The benchmark figure on the front page has no baseline stated and names its own base model twice, so treat it as a claim to check, not a result to plan around.

## FAQ

### What is HITsz-TMG/KnowAct?

KnowAct is a MIT licensed project from the Lychee Team at Harbin Institute of Technology Shenzhen. Its code lives in a GUIClaw subdirectory and provides a personal GUI assistant for desktop, Android, iOS and HarmonyOS, organised as a four stage loop of know, route, act and reflect.

### What is the KnowAct GUI tool called on a package index?

The Python distribution is named nanobot-ai and it ships both packages. The standalone path installs that distribution as a uv tool and then calls the guiclaw executable, which does not import nanobot at runtime.

### Which backends does the KnowAct guiclaw command support?

Five: local for foreground desktop automation, adb for an Android device or emulator, ios through WebDriverAgent, and hdc for a HarmonyOS device, plus a dry run mode that tests the model and agent loop without changing a real device. The standalone default is local and the nanobot adapter default is adb.

### Does KnowAct need permissions before it can change a device?

Not as an enforced control. When no POLICY entry exists the first GUI task writes a conservative policy of deny, cancel or defer unless the task authorises the request, and the README states plainly that those rules are model guidance rather than guaranteed enforcement, so mandatory restrictions need operating system, backend, host approval or sandbox controls.

### How does KnowAct use Android shortcuts?

Static inference writes one cache file per package, runtime validation launches the candidates on a device over ADB and can consult a vision model first, and the promote step adds eligible validated shortcuts to the shared skill file, which is a Python module at ~/.guiclaw/skill/skills.py.

### Does KnowAct have a versioned release?

No. The only release on the project is tagged Result and named GUIClaw-Result, dated 2026-07-12, and the README points at it for real device experimental logs and trajectories. The default branch main last moved on 2026-08-26.

## Sources

- [HITsz-TMG/KnowAct on GitHub](https://github.com/HITsz-TMG/KnowAct)
- [License: MIT](https://github.com/HITsz-TMG/KnowAct/blob/main/LICENSE)
- [Project website](https://shibosusu.github.io/KnowAct-GUIClaw/)
- [README](https://github.com/HITsz-TMG/KnowAct/blob/main/README.md)
- [Releases](https://github.com/HITsz-TMG/KnowAct/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hitsz-tmg-knowact
