Muteki points eight agent CLIs at one CTF, from a coordinator that holds the host Docker socket
Project Muteki (無敵): autonomous multi-model CTF-solving AI agent swarm
At a glance
- What is it?
- Project Muteki is an AGPL-3.0 Python coordinator that schedules commercial coding agents at a single challenge and lets them share evidence. Its quickest setup is also its least isolated one, and the compose deployment hands the coordinator the host Docker daemon.
- Who is it for?
- Take it if you already keep agent CLIs signed in and want them aimed at a challenge you own, and prefer the container runtime over the four-command local path. Leave it alone if you need isolation, since the project states plainly that it does not isolate malicious challenges.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Local mode is the documented default, and it is the one without isolation
The setup that fits on one screen is four commands:
git clone https://github.com/FishCodeTech/muteki.git
cd muteki
./init.sh
./run.sh webIt wants uv, Python 3.13 or newer, Node.js/npm, and at least one supported Agent CLI installed and signed in. `init.sh` syncs the Python dependencies, the first Web start installs and builds the frontend, and the deck answers on http://127.0.0.1:3001 with the backend on port `8000`, both bound to loopback. `Ctrl+C` stops it.
The configuration walkthrough then tells you to pick Local for host CLIs, because Local mode reuses an Agent CLI login already present on the machine and it is the fastest route to a first run. That instruction sits directly against the project's own warning block, which says Muteki drives CLI Agents able to run commands, invoke security tools and access target services, that it does not isolate malicious challenges, and that a disposable VPS, VM or machine without sensitive data is where it should live, with shared hosts and production systems called out as places to avoid. The same block records that the author often runs it on his own computer because that is easier to set up, and that this does not remove the risk.
The compose file gives the coordinator the host Docker socket
The header comment of `docker-compose.yml` spells out the topology. `web-api` is the FastAPI coordinator, and it mounts the host docker socket to launch WORKER containers as SIBLINGS on the host daemon, explicitly not dind. An in-container supervisor dials back to `web-api:9100` over `muteki_net`. A second service, `ui`, is a Next command deck that proxies `/api` to `web-api`. Workers are not a compose service at all: `web-api` shells out to `docker run` for them per run.
Two things follow from that shape. Container creation authority sits with the process holding the socket, so whatever a run asks the coordinator to launch is launched against the host daemon rather than inside the compose project. And the volume entry that grants it that authority is unfinished: the line mounting the host socket stops partway through its container path, so where the mount lands inside `web-api` is not written down in the file. The network block at least names itself, `name: ${MUTEKI_WORKER_NETWORK:-muteki_net}`, commented as a fixed external-style name so workers join by name, overridable when several deployments share one host.
Four env variables keep the swarm from falling back to a local worker
Four settings are documented as the fixes for what the compose header calls the P2-v3 BLOCKERs. `MUTEKI_CONTROL_BIND=0.0.0.0` makes the receiver reachable by sibling workers. `MUTEKI_WORKER_NETWORK=muteki_net` with `MUTEKI_CONTROL_HOST=web-api` lets those workers resolve the receiver by its service alias. `${MUTEKI_HOST_DATA_ROOT}` is bind-mounted at the same path inside `web-api`, with `MUTEKI_SESSIONS_ROOT` under it, so the sibling mounts the HOST daemon resolves land on real host paths; the comment adds that `_mount_source()` is identity in this configuration. The fourth is `MUTEKI_IN_CONTAINER=1`, and its stated purpose is that the swarm hard-fails instead of falling back to a host-native local worker.
That last line is the one worth reading before deploying. A container deployment that cannot start its worker fails outright, by design. The reverse also holds for the local setup: nothing on the compose path quietly downgrades a run to a worker running directly on the host with your own Agent logins. The web image is built from `docker/web/Dockerfile` and tagged `muteki-web:latest`.
Port 8000 is published by default and the deck password is not set
`.env.example` gates the command deck behind `MUTEKI_WEB_PASSWORD`. Set, it locks the whole `/api` surface, described as runs, credential accounts, HITL, and the SSE and WS streams, and the browser keeps a signed session token with `MUTEKI_WEB_TOKEN_TTL=43200`, a twelve hour default, signed with `MUTEKI_WEB_AUTH_SECRET` when you supply one. Unset, it is documented as acceptable only for a loopback-only single-operator setup. It is marked REQUIRED the moment the backend binds anywhere but loopback, with `./run.sh web --host 0.0.0.0` and `docker compose` both named, and the backend refuses to start otherwise so the credential-account APIs are never exposed unauthenticated.
Compose lands on the REQUIRED side by default. It publishes `8000:8000` with a comment calling that operator and API access that should be removed to keep it compose-internal, and the header's own one-liner passes `MUTEKI_HOST_DATA_ROOT` and `MUTEKI_WEB_PASSWORD` together on the command line. Set the password before the first `docker compose up --build`. The same file notes that `.env` at the repo root is auto-loaded by the eval scripts and the web server, and that a variable already exported in your shell always wins, which is why an inline `MUTEKI_DEEPSEEK_API_KEY=... uv run ...` keeps working unchanged.
One vendor SDK sits in the dependency list of an eight-engine roster
The coordinator advertises eight Worker engines: Claude, Codex, Cursor, Pi, OMP, Kimi, Grok and OpenCode. The dependency list names one of them specifically. Alongside FastAPI, uvicorn with the standard extra, `sse-starlette`, pydantic, httpx, anyio, pyjwt and `mcp>=1.10,<2`, `pyproject.toml` carries `claude-agent-sdk>=0.1.0`, with MCP held below 2. The remainder is the CTF and forensics toolchain pulled into Python itself: gmpy2, sympy, pycryptodome, scapy, pyzbar, zxing-cpp, magika, ropgadget, capstone, r2pipe, lief, pillow, numpy, scipy and qrcode.
The same file requires Python 3.13 or newer, exposes a single console script, `muteki = "muteki.cli:main"`, and builds a hatchling wheel containing only the `muteki` package. The one endpoint written into the project file is `[tool.muteki] deepseek_base_url = "https://api.deepseek.com/v1"`, which is also the key the eval scripts demand. The only benchmark result naming its model is the hosted TSecBench run, where deepseek-flash placed 13th. Meanwhile the UI asks you to select and test a Planner endpoint and model of your choosing, and to configure the Titler if you need one.
Two hours of solving, 39 hours on the platform
The results section mixes two clocks. At RIFFHACK 2026 the swarm ran three hours without human takeover, solved every challenge, and placed eighth. On the iChunQiu Yunjing blackmaze range, which had seen no solves for three months, it took first blood in about two hours of actual solving; the same paragraph explains that the platform shows 39 hours because debugging and multi-Flag development were interleaved with that run. Both numbers hold under their own definition, and neither compares with the other, or with the NYU CTF Bench score of 200/200, which carries its own caveat that the figures reflect the tested setup and model versions rather than promising the same on every new challenge.
Multi-Flag shows up elsewhere as a product detail: the submission form asks whether a challenge carries a single Flag or several, and the blackmaze run was measured with more than one in play. The rest of the record is the Yunjing badge scenarios and Hack The Box challenges at Insane and Hard difficulty. The version history is short: v0.3.1 on 2026-08-20, v0.3.2 on 2026-08-22, and v0.4.0 on 2026-09-27, matching the version field in the project file, with the last push on 2026-09-30 and the repository not archived.
A container cannot inherit the CLI login you already have
Local mode reuses an Agent CLI login sitting on the host. A container cannot inherit that login automatically, so the container path asks you to bind an injectable account on the credentials page, and container mode needs a Worker image as well. The supported runtimes are a table of four rows. On macOS locally, `./ctf-tools/setup.sh` installs native CTF tools. With Docker on macOS you pull the full Worker image, which carries the CTF toolchain. On Ubuntu 24.04 locally, the installer takes `--with-ctf-tools`, or `./ctf-tools/setup-ubuntu.sh` does the same work. On Windows, the host carries an Ubuntu 24.04 VMware virtual machine and Muteki runs inside it. The tested platforms are macOS 26 and Ubuntu 24.04.
The macOS setup section is where the written record stops. It opens a dependency sentence about installing Homebrew, Node.js/npm and a third item whose name ends mid-word, so the rest of the macOS list is not written down anywhere in it. The credentials screen itself is specified closely enough to act on: check an existing CLI login or add a token, an API key or a custom Base URL, then run the real connection test instead of trusting a saved entry.
One task form is enabled while the Web pentest mode is switched off
The quick start opens by telling you which workspace you get. The home page shows the single-challenge CTF workspace by default. Conversation, competitions and custom extensions are described as still being tested and can be enabled as described below. The Web pentest mode is being reworked and is currently unavailable, even though the project subtitle reads Heterogeneous Multi-Model AI Agent Swarm and Autonomous Security Automation, and the introduction says the workbench will expand beyond CTF.
The repository is laid out for a wider surface than that one page: `apps/`, `cmd/`, `docker/`, `scripts/`, `skills/` and `ctf-tools/` at the root, a `.mcp.json.example` beside `docker-compose.yml` and `docker-compose.release.yml`, `README_CN.md` next to the English one, and `CHANGELOG.md` and `SECURITY.md` pointed at from the warning. What ships switched on is narrower: one task form taking a challenge name, the original description, a target URL, known facts and a Flag format, attachments by picker, paste or drop, a Web tools toggle, runtime and advanced options, and a `Dispatch swarm` button. Category is inferred when left empty, and since the first run prepares a workspace and Workers, the instruction is to watch the status rather than dispatch again.
Editorial conclusion
Take it if you already keep agent CLIs signed in and want them aimed at a challenge you own, and prefer the container runtime over the four-command local path. Leave it alone if you need isolation, since the project states plainly that it does not isolate malicious challenges. Before the first compose start, read SECURITY.md and set MUTEKI_WEB_PASSWORD, because the process holding that socket is the one your browser talks to.
Frequently asked questions
What does Project Muteki coordinate?
A Coordinator schedules eight Worker engines, Claude, Codex, Cursor, Pi, OMP, Kimi, Grok and OpenCode, toward one goal, manages their context, and lets them collaborate through shared evidence. The project is Python, licensed AGPL-3.0, and requires Python 3.13 or newer.
Does Project Muteki isolate the challenges it runs against?
No. The project states that it does not isolate malicious challenges, that its Workers can execute commands, invoke security tools and access target services, and that a dedicated disposable VPS, VM or machine without sensitive data is the recommended place to run it. It points to SECURITY.md and says to avoid shared hosts and production systems.
What does Muteki need before a real evaluation run?
A DeepSeek API key in .env as MUTEKI_DEEPSEEK_API_KEY. Without it the eval scripts abort and live tests skip. The Agent CLIs need their own logins, tokens, API keys or a custom Base URL entered on the credentials page, and a Planner endpoint plus model under Reasoning models.
Can Project Muteki run natively on Windows?
The tested platforms are macOS 26 and Ubuntu 24.04. For Windows the recommendation is to let Windows host an Ubuntu 24.04 VMware virtual machine and run the Ubuntu installer inside it, with Muteki and its Workers both living in the VM.
What happens if MUTEKI_WEB_PASSWORD is left unset in Project Muteki?
The whole /api surface, runs, credential accounts, HITL and the SSE and WS streams, stays ungated, which is documented as acceptable only for a loopback-only single-operator setup. The backend refuses to start on a non-loopback bind without it, and the signed session token default lifetime is 43200 seconds.
Should Muteki Workers run locally or in containers?
Local mode reuses an Agent CLI login already available on the host, which is why the quick start suggests it as the fast route. A container cannot inherit that login automatically, so you bind an injectable account on the credentials page and provide a Worker image that includes the CTF toolchain.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/fishcodetech-muteki)