PentAGI: a self-hosted autonomous penetration testing agent stack in Go
PentAGI runs fully autonomous AI agents for complex penetration testing, self-hosted and configurable with OpenAI, Anthropic, Ollama, and other LLM providers.
At a glance
- What is it?
- PentAGI runs an LLM-driven pentest agent inside Docker sandboxes, with PostgreSQL plus pgvector storing every command and output. This review covers what it does, how the pieces fit, and where it stops being the right tool.
- Who is it for?
- Adopt PentAGI if you already run Docker hosts, hold API keys for one of the supported LLM providers, and want a self-hosted record of every command an agent executes against a target you are authorised to test. Do not adopt it if you need predefined attack campaigns, because the README states plainly that PentAGI is not a CALDERA-style Breach and Attack Simulation product and that agent-authored attack scripts are conceptual rather than implemented.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem PentAGI targets, and the people it is built for
A manual penetration test is a long chain of small decisions: enumerate, pick a tool, read the output, decide the next step. PentAGI's premise is that an LLM can drive that loop itself. The README describes the project as an autonomous and assistant-guided penetration testing platform for information security professionals and researchers, and the feature list is organised around that loop rather than around scanning: an agent that determines and executes steps, a memory system that keeps research results and successful approaches for later runs, and a delegation layer that splits work between specialised agents for research, development and infrastructure tasks.
The audience is narrow in a useful way. This is not a scanner you point at a URL from a CI job. It expects a Docker host, a PostgreSQL instance with the pgvector extension, and credentials for at least one supported model provider. The README lists more than twenty built-in security tools, naming nmap, metasploit and sqlmap, and it ships a browser component for pulling information from the web plus integrations with Tavily, Firecrawl, Traversaal, Perplexity, DuckDuckGo, Google Custom Search, Sploitus and Searxng. If your workflow is a single engineer running a handful of tools by hand and writing up findings afterwards, the operational surface here is larger than the task.
How the agent loop actually runs: sandboxes, pgvector and a supervisor
The architecture that the repository makes visible is a set of cooperating services rather than one binary. The root docker-compose.yml defines a pentagi service built from vxcontrol/pentagi, a pgvector service that pentagi depends on with a service_healthy condition, and named volumes for application data, SSL material, the Ollama model store and the PostgreSQL data directory. The pentagi container exposes 8443/tcp and publishes it on PENTAGI_LISTEN_IP and PENTAGI_LISTEN_PORT, defaulting to 127.0.0.1 and 8443. The README's quick start section is titled giving agents Docker without giving away the host, which tells you the design intent: tool execution happens in containers, not on the machine running the control plane.
State lives in PostgreSQL with pgvector. The README states that all commands and outputs are stored there, and that the same database backs the smart memory system for long-term storage of research results. Port publishing for per-flow sandboxes is controlled by DOCKER_PORTS_BASE, which the .env.example documents with a default base of 28000 and a range of two thousand ports per instance. The README also describes optional execution monitoring and intelligent task planning as a way to get better results from smaller models, and an optional Graphiti knowledge graph on Neo4j for semantic relationship tracking. Those are separate compose files, docker-compose-graphiti.yml, docker-compose-langfuse.yml and docker-compose-observability.yml, so the default stack stays smaller than the full feature list suggests.
Install and first run: Docker Compose, an LLM key, and the 8443 listener
The README points at Docker Compose as the deployment path and the repository ships both docker-compose.yml and .env.example. Start by copying the example environment file, because the compose file reads nearly every setting from the environment and defaults most of them to empty.
cp .env.example .envThen open .env and fill in one provider. The OpenAI variables are OPEN_AI_KEY and OPEN_AI_SERVER_URL, which the example file sets to https://api.openai.com/v1. Anthropic uses ANTHROPIC_API_KEY and ANTHROPIC_SERVER_URL, and Google uses GEMINI_API_KEY and GEMINI_SERVER_URL. Leave the others blank unless you intend to use them.
docker compose up -dWhen the stack is healthy, the UI is on the published port. With the defaults untouched that is https://127.0.0.1:8443, and because the container exposes 8443/tcp directly, changing PENTAGI_LISTEN_PORT in .env moves the host side without touching the container. The README's own section on what to do after login is where the workflow continues.
The README documents a second deployment concern that the .env.example spells out in comments. TENANT_ID is an optional namespace for instances that share one PostgreSQL, one worker node, one Neo4j and one Langfuse. It must match ^[a-z][a-z0-9_]{0,31}$ and an invalid value aborts startup. The file is explicit that TENANT_ID does not derive DATA_DIR, DOCKER_PORTS_BASE, the listen ports or INSTALLATION_ID, and that sharing DATA_DIR between instances overwrites flow data. Two instances on one host therefore need different DOCKER_PORTS_BASE values, since each instance owns a two-thousand-port window starting at its base.
Where PentAGI is the wrong tool
The README does something unusual and useful: it publishes a section called current capability boundaries. Read it before you plan a deployment. PentAGI is not a CALDERA-style Breach and Attack Simulation or adversary emulation product, and it does not ship predefined campaigns or attack plans. The README goes further and says that BAS-like agent-authored attack scripts should be treated as conceptual or future work rather than as an implemented feature. If your requirement is repeatable emulation of a named threat actor, this is the wrong project and the documentation says so.
Two smaller boundaries matter operationally. The flow report UI supports web view, copy to clipboard, Markdown download and PDF download, but JSON flow-report export is not documented as a supported output format. Teams that want to pipe findings into another system will be parsing Markdown or PDF rather than structured data. Second, the whole model layer is an external dependency. Provider flexibility exists through built-in providers and custom OpenAI-compatible endpoints, but the quality of the run still tracks the model behind it, and the README positions execution monitoring and task planning as the mechanism that makes smaller models viable. There is no offline mode in the default stack: the Ollama service is present in the compose volumes, yet the default configuration expects a provider key.
How PentAGI differs from a scripted pentest framework
The natural comparison is a scripted framework such as CALDERA, and the README makes the contrast itself. A scripted framework encodes an attack plan up front: you define the campaign, the steps and the expected outcome, and the tool replays it. PentAGI inverts that. The agent determines the steps, the memory system keeps what worked, and the delegation layer assigns research, development and infrastructure work to specialised agents. The output is a flow report rather than a campaign result, and the README's boundary section is careful to say that campaign-style emulation is not what this does today.
The trade-off is reproducibility. A scripted campaign produces the same steps on every run, which is what makes it usable as a control in a security programme. An agent that plans its own path does not, and the README's answer to that is monitoring and planning rather than a fixed plan. The second axis is isolation. PentAGI's README leads with sandboxed Docker execution and a section on giving agents Docker access without giving away the host, which is a different security posture from a framework that runs its actions from the control node. If your environment cannot grant the pentagi container access to the Docker socket, the tool execution model does not work as described.
Maintenance, licensing and what an upgrade costs you
The repository is not archived, and the last push was on 2026-05-29, which is the same timestamp as the v2.1.0 release. The release history shows v1.2.0 in February 2026, v2.0.0 in April 2026 and v2.1.0 in May 2026, so the project has been cutting major versions within a few months of each other. Treat that cadence as the upgrade cost: a v1 to v2 jump inside roughly three months means configuration and compose files are likely to move, and the .env.example is the file to diff first when you pull a new image tag.
The licence situation needs care rather than a summary. The repository's LICENSE file is MIT, and the project metadata lists MIT. At the same time the repository root contains an EULA.md, a NOTICE file and a licenses/ directory, and the Dockerfile generates a licence report for frontend dependencies into /licenses/frontend. The README also references LICENSE_KEY and INSTALLATION_ID environment variables described as being for communication with the PentAGI Cloud API. An MIT licence on the source and a separate end user licence agreement in the same tree is a combination worth reading in full before commercial deployment. This is not legal advice, and the only reliable answer comes from those files plus your own counsel.
On running cost, one concrete statement is supported: every agent step consumes tokens from a provider you pay for directly, and the default configuration expects an external key. The compose file also carries an Ollama volume, which suggests local inference is a supported path, and the README links a vLLM plus Qwen3.5-27B-FP8 guide under examples/guides for production local deployments. Neither the README nor the example file states hardware requirements for that path, so sizing is something you would have to establish yourself.
Editorial conclusion
Adopt PentAGI if you already run Docker hosts, hold API keys for one of the supported LLM providers, and want a self-hosted record of every command an agent executes against a target you are authorised to test. Do not adopt it if you need predefined attack campaigns, because the README states plainly that PentAGI is not a CALDERA-style Breach and Attack Simulation product and that agent-authored attack scripts are conceptual rather than implemented. Before pointing it at anything, verify three things: that your target is in scope, that your provider key and model are one the configuration actually accepts, and that you have set TENANT_ID and DOCKER_PORTS_BASE correctly if a second instance will share the same PostgreSQL, Neo4j or Langfuse.
Frequently asked questions
What is PentAGI?
PentAGI is a self-hosted penetration testing platform that runs autonomous AI agents inside sandboxed Docker environments. The README describes it as an autonomous and assistant-guided platform for security professionals and researchers, with built-in access to more than twenty security tools including nmap, metasploit and sqlmap.
How does PentAGI work?
An LLM-driven agent determines and executes penetration testing steps, delegating research, development and infrastructure work to specialised agents. Tool execution happens in Docker containers rather than on the host, and all commands and outputs are stored in PostgreSQL with the pgvector extension, which also backs the long-term memory system.
How do I install PentAGI?
The README points at Docker Compose. Copy .env.example to .env, fill in one provider such as OPEN_AI_KEY or ANTHROPIC_API_KEY, then run docker compose up -d. The pentagi container exposes 8443/tcp, published by default on 127.0.0.1:8443.
How do I set up PentAGI on one host with several instances?
Set TENANT_ID on each instance, matching ^[a-z][a-z0-9_]{0,31}$, so they can share one PostgreSQL, worker node, Neo4j and Langfuse. The .env.example warns that TENANT_ID does not derive DATA_DIR, DOCKER_PORTS_BASE, listen ports or INSTALLATION_ID, and that sharing DATA_DIR overwrites flow data, so give each instance its own DOCKER_PORTS_BASE inside its two-thousand-port window.
Is PentAGI free?
The repository's licence is MIT, but the root also contains an EULA.md, a NOTICE file and a licenses/ directory, and the README references LICENSE_KEY and INSTALLATION_ID for communication with the PentAGI Cloud API. Read those files before commercial use. Separately, agent runs consume tokens from whichever LLM provider you configure, which you pay for.
How do I use PentAGI after logging in?
The README has a dedicated section titled How to Use PentAGI After Login, and it is the place to start once the stack is up. The UI is served on the published port, 8443 by default, and the README also documents REST and GraphQL APIs with Bearer token authentication for programmatic access.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vxcontrol-pentagi)