Model or dataset
ASCIT31/Dark-Moon avatar
ASCIT31/Dark-Moon

DarkMoon proves findings with agent transcripts, and its compose file hands the stack the Docker socket

Autonomous AI pentesting engine across web, cloud, identity, CI/CD, IaC, databases, Active Directory, Kubernetes, IoT firmware and AI/LLM endpoints (OWASP LLM Top 10). Real exploits with proof for every finding. Privacy gateway: the LLM never sees your real IPs, hosts or creds; nothing leaves your perimeter.

993 stars166 forksPythonGPL-3.0

At a glance

What is it?
DarkMoon is a GPL licensed autonomous penetration testing engine that runs against infrastructure you own, dispatching around fifty specialist agents over an MCP interface. Two things in the repository deserve to be read together: the open source edition states plainly that exploitation is agent-asserted and that only the paid tier retests a fix, and the shipped compose file runs its container as root, on the host network, with the Docker socket mounted.
Who is it for?
This is a tool for a team that already owns the network it points at, and the GPL licence plus plain Markdown agent methods mean you can read exactly what each agent is instructed to do before you let it run. Two decisions belong to you rather than to the project.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Finds and proves, where the proof is the agent's own transcript

The positioning is explicit about scope. DarkMoon is autonomous penetration testing for infrastructure you own, you point it at an authorized target, and it runs the assessment on its own and documents every finding with the exact command and the raw output. The licence is GPLv3 and the project is self-hosted, and the unusual part is that every agent's methodology is plain Markdown, so the instructions can be read, diffed and forked before a run. The evidence claim then gets qualified in the same sentence. Each finding ships with its command and output, but exploitation is described as agent-asserted, and the machine-verified step is the Pro remediation retest, where the exploit is re-run to confirm a fix. In the open source edition, what you read is the agent's account of what it did, not an independent confirmation.

Root, host networking, and the Docker socket in the same service

The feature grid claims security by design on the grounds that the model never runs a command directly and every action flows through a controlled, logged interface. The compose file shipped at the root tells a different story about containment. The main service sets its user to root, uses host networking instead of a bridge network, adds two network capabilities, mounts the Docker socket from the host, mounts a kubeconfig directory read only, restarts unless stopped, and always pulls its image. That combination means a process inside the container can reach the host Docker daemon, which in practice is equivalent to root on the machine. The second service, which runs the agent loop, mounts the same Docker socket again. Whatever the privacy gateway does with your target addresses, the deployment itself holds the authority of the host it is scanning.

A host directory mounted over the tool source tree

The second service in the compose file is where the packaging gets unusual. Five host paths are bind mounted into it. Three of them point the same host directory at three different locations inside the container, covering the agent configuration, the local data share and an agents directory. Two more cover reports and sessions. The fifth mounts a host directory over a path deep inside the application itself, the workflows tool directory of the MCP server, and does it read write. That means the set of workflows the engine can run is whatever is in your working copy, not whatever shipped in the image, which is a deliberate extension point and also the reason an update to the image will not change your behaviour. The two services declare no ordering between them, so nothing waits for the database or network side to come up first. The compose file publishes a single port, 9222, which is the port conventionally used by browser debugging protocols.

One install script asks which model provider you use

Setup is three lines, and the third one is a script that configures your model provider interactively so you never edit the compose file by hand, then builds the whole stack:

bash
git clone https://github.com/ASCIT31/Dark-Moon.git
cd Dark-Moon
./install.sh

The script takes flags: the bare form skips the questionnaire when a provider file already exists, an init flag forces reconfiguration, and a help flag prints usage. Cloud providers are supported, naming Anthropic, OpenAI and OpenRouter among them, and so are two local runtimes, Ollama and llama.cpp, which is how the perimeter claim is meant to hold. A run is then started with a single target argument, and a log flag replays a session:

bash
./darkmoon.sh "TARGET: example.com"
./darkmoon.sh --log <session_id>

The stated prerequisites are Docker with Docker Compose and either a model API key or a local model. GPU setup, environment variables and the complete flag reference live in a separate documentation file rather than in the readme.

Fifty agents chosen by what recon found

Around fifty specialist agents sit behind one dispatching loop, and the project is explicit that the choice is signal-driven rather than fixed: DarkMoon decides which specialists to deploy from exactly what it detects on the target. The surfaces named span web applications, APIs, Active Directory, Kubernetes, three cloud providers, continuous integration pipelines, databases, IoT firmware and AI inference endpoints, chained end to end. A dedicated agent probes the AI endpoints against the OWASP LLM Top 10 using probes from an external adversarial testing harness. The offensive tooling itself is reached through MCP, with more than 140 tools orchestrated that way, and the privacy gateway sits in front of it: real addresses, hosts and credentials are tokenized into deterministic placeholders locally and rehydrated only at the moment a tool runs. Because the agent methods are Markdown, that file tree is the place to review what the harness is actually being asked to do.

libcurl is built from whatever the download page lists first

The image is built in a Go builder stage, and the environment is pinned tightly at the top: the frontend is non interactive, the toolchain is forced to the local one, modules are on, workspace mode is off, and the module proxy and checksum database both point at the public Go infrastructure. The apt layer installs the usual C toolchain plus a long list of development headers. Then one dependency is handled by hand, and this is the part worth noticing. Instead of taking a distribution package or a pinned tarball, the build fetches the libcurl download page, greps it for a file name matching a versioned archive pattern, takes the first match, configures that tarball with TLS and HTTP2 and file and FTP support, disables the static build, compiles it, installs it into an output prefix and then runs the resulting binary to print its version. The Go toolchain is pinned and libcurl is not, so the curl that ends up inside the image is whatever upstream published most recently when the build ran.

Six shell entry points, three compose files, two images

The repository surface is wider than the readme suggests. There are six shell entry points: the launcher, the interactive installer, a development installer, a setup script, and separate Python and Ruby setup scripts. There are three compose files, for the default stack, for development and for GPU, and two Dockerfiles, the main one and a second for the agent runtime image. Alongside them sit directories for configuration, for the MCP server, for bundled tools and for documentation, plus a contributing guide. That last detail, the language specific setup scripts, sits oddly against the repository's declared primary language being Python while the image is compiled from Go and the stack runs a separate agent container. Nothing in the readme explains when to reach for the Python or Ruby variants, and nothing explains which of the three compose files to pick, which is the practical gap for anyone evaluating it.

Editorial conclusion

This is a tool for a team that already owns the network it points at, and the GPL licence plus plain Markdown agent methods mean you can read exactly what each agent is instructed to do before you let it run. Two decisions belong to you rather than to the project. First, the open source edition reports what its own agent claimed, so treat a finding as a hypothesis until you have reproduced the command yourself, since machine verification arrives only with the paid tier. Second, the deployment gives the container root, host networking and the Docker socket, which is the same authority as the host it is scanning, so run it on a disposable machine or a dedicated virtual machine rather than on a workstation with production credentials in reach.

Frequently asked questions

Is DarkMoon's finding verification machine checked in the open source edition?

No. The project states that each vulnerability ships with the exact command and raw output, but that exploitation is agent-asserted, and that the machine-verified step is the Pro remediation retest where the exploit is re-run to confirm a fix.

What does the DarkMoon privacy gateway actually do?

It performs reversible local tokenization, turning real addresses, hosts, URLs and credentials into deterministic placeholders that the model reasons over. The real values stay on your perimeter and are rehydrated locally at the moment a tool runs, and a local model runtime such as Ollama or llama.cpp is supported so the assessment itself can stay offline.

What permissions does the DarkMoon docker compose stack need?

The main service runs as root, uses host networking, adds the raw socket and network administration capabilities, mounts the host Docker socket and a kubeconfig directory read only, and restarts unless stopped. The agent service mounts the Docker socket as well. That is host-level authority, so it belongs on a disposable machine.

How do I start a DarkMoon assessment?

Clone the repository, enter the directory and run the installer, which asks which model provider to configure and then builds the stack. After that a run is started with a single target argument in the form TARGET: followed by the host, and a log flag replays a stored session. Docker with Docker Compose and either a model API key or a local model are required.

Can I read what the DarkMoon agents are instructed to do?

Yes. The project states that every agent's methodology is plain Markdown that you can read, diff and fork, and the licence is GPLv3. The agent loop itself is dispatched from what reconnaissance detects, with about fifty specialists available and more than 140 tools reached over MCP.

Official sources

  1. ASCIT31/Dark-Moon on GitHub
  2. License: GPL-3.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ascit31-dark-moon.svg)](https://hysenlabs.com/projects/ascit31-dark-moon)