# AI Manus: a self-hosted PlanAct agent that runs its tools in per-task Docker sandboxes

> Simpleyyt/ai-manus is an MIT-licensed Python agent system you deploy with Docker Compose, point at a tool-calling LLM, and drive from a web UI. The interesting part is the sandbox lifecycle, and the part to think about first is that the backend gets read access to the Docker socket.

**Simpleyyt/ai-manus** — AI Manus is a general-purpose AI Agent system that supports running various tools and operations in a sandbox environment.

- Repository: https://github.com/Simpleyyt/ai-manus
- Website: https://ai-manus.com
- Stars: 1,628 · Forks: 407
- Language: Python
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/simpleyyt-ai-manus

## The problem AI Manus solves, and who ends up running it

Most agent demos assume the agent can touch the machine it lives on: write files, open a browser, run shell commands. That is fine on a laptop and unacceptable on a server that also hosts something. AI Manus takes the opposite position. Every task gets its own sandbox, described in the README as an Ubuntu Docker environment that starts a Chrome browser plus API services for File and Shell tools. The agent works inside that container, not on your host.

The intended user is a developer or small team that wants a general-purpose agent with a web UI, session history and file handling, and is willing to run the supporting stack. The README's deployment section is explicit that the minimal path needs an LLM service and nothing else external, but the repository's own docker-compose.yml also brings up MongoDB and Redis for session state, so in practice you are running four or five containers: frontend, backend, sandbox (pulled, then stopped), mongodb and redis. The homepage points at a hosted demo at app.ai-manus.com if you want to look before installing.

This is not a library you import into an existing Python service. The backend is a service, the frontend is a service, and the agent loop lives behind a WebSocket. If what you actually want is a function you can call from your own code, the repository layout will feel like a lot of infrastructure for that.

## How the PlanAct loop, sandbox creation and WebSocket events fit together

The README lays out the request flow in five steps, and it is worth reading as a lifecycle rather than a feature list. The web client asks the server to create an agent. The server creates a sandbox through /var/run/docker.sock and returns a session ID. The sandbox boots, starting Chrome and the tool APIs. The client then sends user messages against that session ID, and the server forwards them to the PlanAct Agent.

Inside the agent, planning and execution are separated. The README states that the planner and executor submit structured results through native tool calls, and the Key Features section contrasts this with a JSON-in-prompt protocol, which it calls fragile. Concretely, the environment requirements name create_plan and complete_step as the structured output tools. That design choice is the reason the model requirement is narrow: a provider that does not support native tool or function calling cannot drive this loop at all. The .env.example lists built-in providers openai, deepseek, anthropic, ollama and orcarouter, and notes that OpenAI-compatible endpoints such as DeepSeek, OneAPI or vLLM work through openai plus API_BASE.

Tool calls reach the sandbox for Shell, Browser, File, Search and MCP. Events flow back over WebSocket, which is what makes live viewing possible. For the browser specifically, the README describes a chain: the sandbox's headless browser starts a VNC service through xvfb and x11vnc, websockify converts VNC to WebSocket, and the frontend's NoVNC component connects through the server's WebSocket forward. That is a real amount of machinery, and it is also why takeover of a running browser session is possible rather than aspirational.

Two configuration details shape behaviour more than they look. SANDBOX_TTL_MINUTES defaults to 30, so a long task can outlive its sandbox unless you raise it. And BROWSER_ENGINE accepts browser_use (the default, driving Chrome through the browser-use library's event bus) or playwright, a lighter engine over CDP that emits the same tool payload shape.

## Installing AI Manus with Docker Compose and running a first task

The README recommends Docker Compose and requires Docker 20.10+ with the Compose plugin. Start by cloning the repository and copying the environment template, because the backend reads all of its configuration from a .env file.

```bash
git clone https://github.com/simpleyyt/ai-manus.git
cd ai-manus
cp .env.example .env
```

Open .env and set at least API_KEY. The README's example also sets API_BASE and MODEL_NAME; the .env.example ships with MODEL_PROVIDER=openai and MODEL_NAME=deepseek-chat, so change the model name if you are not pointing at DeepSeek.

```ini
API_KEY=sk-xxxx
API_BASE=https://api.openai.com/v1
MODEL_NAME=gpt-4o
```

Now bring the stack up. The repository's docker-compose.yml defines frontend on port 5173, backend, a sandbox service whose command is /bin/sh -c "exit 0", plus mongodb and redis.

```bash
docker compose up -d
```

The README warns that seeing sandbox-1 exited with code 0 is normal: the sandbox service exists to pull the image locally, and the backend starts real sandboxes later through the Docker socket. When the stack is up, open http://localhost:5173 in a browser, log in, and start a session. A first task worth trying is the one the README demonstrates, asking the agent to write a Python example, because it exercises the File and Shell tools without depending on search or a live site.

If you are changing code rather than just running it, the development path is different. ./dev.sh up is equivalent to docker compose -f docker-compose-development.yml up, services run in reload mode, and the README lists the exposed ports: 5173 for the web frontend, 8000 for the server API, 5678 for debugpy, 8080 for the sandbox API, 5902 for VNC (mapped to 5900 inside the container) and 27017 for MongoDB. The README notes that debug mode starts only one sandbox globally, so it is not a way to test concurrent tasks.

```bash
./dev.sh up
./dev.sh down -v
./dev.sh b
```

The second and third commands matter when dependencies change in backend/pyproject.toml or frontend/package.json: down -v cleans up related resources, and b rebuilds images.

## Where AI Manus is the wrong tool

The clearest limitation is the Docker socket. The compose file mounts /var/run/docker.sock into the backend read-only, and the backend needs it to create sandboxes. Read-only reduces what a compromised backend can do, but a process that can talk to the Docker API is still a process with unusual power on the host. If your policy forbids socket mounts, this architecture does not have a documented alternative in the README; the roadmap mentions K8s and Docker Swarm multi-cluster deployment as future work, not as a shipped path.

Model compatibility is the second constraint. The environment requirements ask for native tool or function calling, and the README recommends DeepSeek and GPT models with reliable tool calling. A local model served through an OpenAI-compatible endpoint may work through the openai provider and API_BASE, but nothing in the README promises that a given model will produce usable create_plan and complete_step calls. Test your model before you plan around it.

Third, the sandbox is Linux and Docker. The roadmap lists mobile and Windows computer access for the sandbox as planned, which means it is not there yet. If your team develops on Windows hosts without a Linux Docker environment, the deployment story is harder than the README's three commands suggest.

Finally, consider whether you need an agent platform at all. If your task is a fixed sequence of API calls, a script with a retry loop is cheaper to operate than five containers, a session store and an LLM bill. AI Manus earns its complexity when the task genuinely requires browsing, file manipulation and shell work in an order you cannot predict.

## AI Manus compared with OpenManus-style agent frameworks

The obvious comparison is with OpenManus and similar Python agent frameworks, which share the goal of a general agent that can use tools but differ in what they ship. A typical framework is a library: you install it, write a script, choose tools, and run it in your own process. There is no web UI, no session database and no sandbox lifecycle, because the agent runs wherever your script runs.

AI Manus inverts that. The repository is organized as separate frontend, backend, sandbox and mockserver sub-projects, and the agent is reachable only through the server's API and WebSocket. You gain a UI, login and authentication, session history in MongoDB and Redis, background tasks, a sidebar Library that aggregates attachments and artifacts across sessions with type filters, search, favorites and preview, plus file upload and download. You lose the ability to embed the loop in an existing service without going through that API.

The sandbox is the other real difference. A library-style agent that runs shell commands runs them on your machine unless you build isolation yourself. AI Manus allocates a container per task, which is a stronger default and a heavier dependency. The mockserver sub-project is a small but telling detail: it exists so the stack can be developed and tested without spending on a real model, which is a sign the project is built to be run by its own maintainers, not just demonstrated.

If your need is a scripted agent inside an existing Python application, a library is the smaller commitment. If your need is a hosted agent your colleagues can log into and watch, the library leaves you building the UI, the session store and the isolation layer yourself.

## Maintenance, upgrade cost and the MIT licence

The repository is not archived, and the last push was on 2026-08-12, the same day as the v2.7.0 release. Releases have been reasonably close together: v2.5.0 on 2026-07-19, v2.6.0 on 2026-08-02, v2.7.0 on 2026-08-12. That is a project still being worked on, but it also means the surface moves. The compose file parameterizes image tags through IMAGE_REGISTRY and IMAGE_TAG, both defaulting to simpleyyt and latest, so a plain docker compose pull can move you across releases without a version pin.

The upgrade cost is concentrated in three places. Dependencies live in backend/pyproject.toml and frontend/package.json, and the README's development instructions say to run ./dev.sh down -v followed by ./dev.sh b when they change, which discards volumes. Configuration keys are read from .env, and .env.example carries a large commented block covering MongoDB, Redis, sandbox, browser engine and search engine settings; a release that adds a key will not break you, but one that renames MODEL_PROVIDER or LLM_PROVIDER behaviour will. Session data sits in the mongodb_data volume, named manus-mongodb-data in the compose file, so that is the thing to back up before a destructive upgrade.

AI Manus is MIT licensed. In practical terms that permits commercial use and modification, and the LICENSE file is at the repository root. This is not legal advice, and the licence covers the code in this repository; the container images and any model provider you connect to have their own terms, and the sandbox image pulls an Ubuntu base plus Chrome, whose redistribution terms are separate from the MIT grant.

## Conclusion

Adopt AI Manus if you want a browser- and terminal-capable agent you host yourself and you already run Docker, MongoDB and Redis. Do not adopt it if your only model is a chat endpoint without reliable native tool calling, since plans and step results are submitted through structured output tools such as create_plan and complete_step. Before anything else, read the compose file's mount of /var/run/docker.sock into the backend and decide whether that read-only socket is acceptable on the host you plan to use, then check whether SANDBOX_TTL_MINUTES=30 matches how long your tasks actually run.

## FAQ

### What is AI Manus?

It is a general-purpose AI Agent system that runs tools and operations inside a sandbox environment, written in Python and licensed under MIT. The README describes a PlanAct loop where a planner and executor submit structured results through native tool calls, with Terminal, Browser, File, Web Search and MCP tools available.

### Is AI Manus free or paid?

The source code is MIT licensed, so you can run it yourself without paying for the software. You still pay for whatever model provider you configure through API_KEY and API_BASE, and for the host running the Docker Compose stack.

### Is AI Manus a Chinese company?

The README does not describe the project as a company. AI Manus is a GitHub repository under Simpleyyt with an MIT licence, and the README states that the interface supports both Chinese and English.

### Why do people use AI Manus?

The README's stated draw is a self-hosted agent whose minimal deployment needs only an LLM service, with each task getting its own Docker sandbox and a web UI that supports live viewing and takeover of tools. Session history, background tasks and a cross-session Library are the other features it lists.

### How is AI Manus different from ChatGPT?

AI Manus is a deployable agent system rather than a chat product: it creates a sandbox per task, runs Shell, Browser and File tools inside it, and streams events back to a web UI over WebSocket. The README requires a model with native tool or function calling, since plans and step results are submitted through tools such as create_plan and complete_step.

## Sources

- [Official documentation](https://ai-manus.com)
- [Official README](https://github.com/Simpleyyt/ai-manus#readme)
- [Project repository](https://github.com/Simpleyyt/ai-manus)
- [Release notes](https://github.com/Simpleyyt/ai-manus/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/simpleyyt-ai-manus
