Model or dataset
PurpleAILAB/Decepticon avatar
PurpleAILAB/Decepticon

Decepticon: an autonomous red team agent that writes its own rules of engagement

Autonomous Hacking Agent for Red Team

5,619 stars1,062 forksPythonApache-2.0

At a glance

What is it?
PurpleAILAB's Decepticon is a Python and LangGraph red team agent that generates RoE, ConOps and an OPPLAN before touching a target, then runs the kill chain from a terminal CLI. Here is what the repository actually documents, and where it stops.
Who is it for?
Adopt Decepticon if you already run scoped red team engagements and want the planning artefacts (RoE, ConOps, Deconfliction Plan, OPPLAN) generated as part of the run rather than bolted on afterwards. Do not adopt it if you want a single pip install that works without services: the SDK is a client and routes LLM calls and sandbox execution over HTTP to a LiteLLM proxy and a sandbox URL, so you need the Docker stack or your own equivalents.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Decepticon is aimed at: scanners that stop at the report

Most "AI plus hacking" projects in this space are wrappers around nmap. The README opens by mocking exactly that: a tool that "runs nmap and writes a report". Decepticon's claim is narrower and more specific. It positions itself as a professional autonomous red team agent that executes reconnaissance, exploitation, privilege escalation, lateral movement and C2 as a chain, adapting to whatever path opens up rather than walking a fixed checklist.

The audience is therefore not a developer who wants to scan their own staging box. It is a red team or a security consultancy running scoped engagements, where the deliverable includes the paperwork that authorises the activity. Decepticon generates an engagement package before a packet leaves the wire: RoE, ConOps, a Deconfliction Plan, and an OPPLAN mapped to MITRE ATT&CK. Every action is then supposed to run inside those rules. That framing is the actual differentiator, not the exploitation itself. Plenty of tools exploit; far fewer treat the authorisation boundary as a first-class artefact the agent must produce and then obey.

Whether that discipline holds in practice is not something the README demonstrates. It states the intent and links to an engagement workflow document. Treat the planning layer as the interesting design bet, and read docs/engagement-workflow.md before you trust it.

How the stack is put together: management plane, specialists, sandbox

The default start brings up what the README calls the core management plane: LiteLLM, PostgreSQL, Neo4j, Skillogy, LangGraph and a sandbox. The terminal CLI launches alongside it. That is the always-on layer. Specialist workloads are separate and come up on demand. The README names BloodHound CE, Sliver C2 and Ghidra MCP as examples, and says the orchestrator spawns them through calls such as ops_start("ad"). The web dashboard is also on-demand, started from inside the CLI with /web.

Two details in the repository files are worth reading as design decisions rather than boilerplate. First, docker-compose.yml defines a shared log-rotation anchor and applies it per service, capping long-running containers at three files of ten megabytes. The comment is blunt about why: Docker's default json-file driver ships no rotation, so logs grow until the host disk fills and the stack crashes. Someone has been burned by this. Second, the compose file parameterises container names through DECEPTICON_STACK_NAME, with a comment explaining that COMPOSE_PROJECT_NAME would collide with the project directory name and produce a doubled prefix. That is the kind of fix that only appears after real multi-stack use.

The management plane is also where the LLM traffic lands. LiteLLM runs on port 4000 by default, bound to 127.0.0.1, and reads its routing config from config/litellm.yaml. Model choice, credentials and usage accounting all pass through that service rather than through the agent process.

Installing Decepticon and running a first engagement

The prerequisites are Docker and Docker Compose v2. The README lists macOS on Apple Silicon and Intel, Linux on amd64 and arm64, and Windows on amd64 and arm64 either natively through PowerShell or through WSL2 with Ubuntu or Kali.

On macOS, Linux or WSL2, the installer is a shell script served from the project domain. It sets up the environment, after which the onboard wizard walks through provider, API key and model profile, and the bare command starts the core stack and drops you into the terminal CLI.

bash
curl -fsSL https://decepticon.red/install | bash
decepticon onboard
decepticon

On Windows the equivalent is a PowerShell one-liner, followed by the same two commands. The onboard step is not optional in practice: it is what writes your selected credentials into the environment file. The README notes that the installer copies .env.example to ~/.decepticon/.env, and that unselected keys keep their placeholder values and are ignored at runtime.

powershell
irm https://decepticon.red/install.ps1 | iex
decepticon onboard
decepticon

If you are building on top of the agents rather than driving the CLI, there is a separate path. The SDK is published on PyPI, and the README shows both the core install and an extra that pulls in the knowledge-graph attack-chain tools.

bash
pip install decepticon              # core SDK
pip install "decepticon[neo4j]"     # + the knowledge-graph attack-chain tools

The important caveat is stated in the README itself: decepticon is a client SDK. It ships agent factories, middleware, tools and skills, and routes LLM calls and sandbox execution to runtime services over HTTP through DECEPTICON_LLM__PROXY_URL and SANDBOX_URL. Running agents still requires those services. Point the URLs at the Docker stack or at your own equivalents; the pip install alone will not execute anything.

The benchmark numbers and what they do not tell you

The repository publishes results against the XBOW validation-benchmarks: 45 of 45 on easy, 50 of 51 on medium, 7 of 8 on hard, 102 of 104 across all levels. A per-challenge index, an attack-class matrix and LangSmith traces are linked from benchmark/results/README.md, and a comparison against other AI pentest agents lives in docs/benchmark-comparison.md.

Read those as a claim by the maintainers, not as an independent result. The benchmark repository is itself under the PurpleAILAB organisation, which does not make the numbers wrong but does mean the scoring pipeline and the agent share an owner. The hard tier is also small: eight challenges, seven passed. A single flipped result moves that number by more than twelve points, so the level-3 figure carries much less weight than the aggregate.

What the benchmark does not cover is the part Decepticon makes central. There is no published measure of whether the generated RoE and OPPLAN actually constrain the agent's behaviour, or how often the agent proposes an action outside the scope it wrote for itself. That is the claim most worth testing yourself, and the traces are the place to start.

Where Decepticon is the wrong tool

The install path is the first real limitation. This is not a tool you pip install and run. Even the SDK route requires a reachable LiteLLM proxy and a sandbox service, and the README says so plainly. If you want a library that runs in-process, this is a client with a service dependency, and that dependency is architectural rather than incidental.

The second limitation is operational weight. The default start brings up PostgreSQL, Neo4j, LiteLLM, LangGraph, a sandbox and a CLI. Specialist workloads are additional containers pulled on demand. That is a real footprint for a workstation, and it means the tool is a poor fit for a quick one-off check against a single host. A conventional scanner will finish before Decepticon's stack has finished starting.

The third is scope. The README states that every action runs inside the engagement rules the agent generates. That is a promise about behaviour, and the repository does not publish the enforcement mechanism or a test suite that verifies it. If your engagement letter does not authorise lateral movement or C2, an agent whose stated purpose is to pursue objectives "through whatever path opens up" is the wrong instrument, regardless of what the OPPLAN says. Use it where the authorisation is already broad and documented, and keep a human on the CLI.

How it differs from Strix, PentestGPT and the rest

The repository links its own comparison document, docs/benchmark-comparison.md, which covers Strix, PentestGPT, MAPTA, Cyber-AutoAgent and the commercial XBOW among others. The distinction the README draws is not about which model or which tool set. It is about the engagement layer and about interactivity.

On interactivity, the README makes a concrete point: real offensive tools are interactive, naming msfconsole, sliver-client and evil-winrm. Decepticon runs commands inside persistent sessions rather than one-shot invocations. That matters because a large part of real exploitation is reacting to what a shell prints back, which a stateless command runner cannot do.

On the engagement layer, the difference is that Decepticon generates RoE, ConOps, Deconfliction Plan and OPPLAN as part of the run, with MITRE ATT&CK mapping, before it starts. A scanner-shaped agent produces findings; Decepticon produces findings plus the authorisation trail. That is a workflow difference, not a capability difference, and it is the reason the tool is heavier to set up. If you do not need the trail, the extra machinery is cost without benefit.

Maintenance, licensing and the cost of keeping up

The repository is not archived. The last push was on 2026-08-30, which is recent enough that the project is being worked on, and the release history supports that: v1.1.38 on 2026-07-12, then v1.1.39 and v1.1.40 on 2026-07-27. The cadence is quick, and the version numbers are close together, which is typical of a project iterating on a container-based stack.

That cadence is also the upgrade cost. Because the default deployment is Docker Compose and the images are pulled by tag, with DECEPTICON_VERSION defaulting to stable, you are tracking a moving target unless you pin. The compose file parameterises the stack name, so running more than one version side by side is possible, but the README does not document a rollback procedure, and it does not describe a migration path for the PostgreSQL or Neo4j volumes between releases. Verify that before you run this against anything you cannot rebuild from scratch.

The licence is Apache-2.0, which permits commercial use and modification and includes a patent grant. The repository also carries a THIRD_PARTY_LICENSES.md and a TELEMETRY.md file, which suggests you should read both before deploying inside an organisation: the second tells you what the stack reports home. None of this is legal advice; check the terms against your own policy.

Editorial conclusion

Adopt Decepticon if you already run scoped red team engagements and want the planning artefacts (RoE, ConOps, Deconfliction Plan, OPPLAN) generated as part of the run rather than bolted on afterwards. Do not adopt it if you want a single pip install that works without services: the SDK is a client and routes LLM calls and sandbox execution over HTTP to a LiteLLM proxy and a sandbox URL, so you need the Docker stack or your own equivalents. Before committing, verify three things: that your provider key is one the onboard wizard accepts, that the specialist workloads you need (BloodHound CE, Sliver C2, Ghidra MCP) start on demand in your environment, and that your engagement authorises the kill-chain behaviour the OPPLAN describes.

Frequently asked questions

What is Decepticon and what does it do?

Decepticon is an autonomous red team agent from PurpleAILAB, written in Python and built on LangGraph. According to the README, it executes reconnaissance, exploitation, privilege escalation, lateral movement and C2 as a chain, and generates an engagement package with RoE, ConOps, Deconfliction Plan and OPPLAN before starting.

Does Decepticon have a hosted version, or do I have to self-host it?

Both. The README documents a Docker Compose install for macOS, Linux and Windows, and also links a hosted app at app.decepticon.red for users who do not want to run the stack themselves.

Can I install Decepticon with pip instead of Docker?

Yes, but only as a client SDK. The README states that pip install decepticon ships agent factories, middleware, tools and skills, and routes LLM calls and sandbox execution to runtime services over HTTP through DECEPTICON_LLM__PROXY_URL and SANDBOX_URL. Running agents still requires those services.

What are the prerequisites for installing Decepticon?

Docker and Docker Compose v2. The README lists support for macOS on Apple Silicon and Intel, Linux on amd64 and arm64, and Windows on amd64 and arm64, either natively through PowerShell or through WSL2 with Ubuntu or Kali.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. PurpleAILAB/Decepticon on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/purpleailab-decepticon.svg)](https://hysenlabs.com/projects/purpleailab-decepticon)