Guardian CLI: an AI pentest orchestrator that wraps 50 security tools behind a workflow DSL
Guardian is a production-ready AI-powered penetration testing automation CLI tool that leverages Google Gemini and LangChain to orchestrate intelligent, step-by-step penetration testing workflows while maintaining ethical hacking standards.
At a glance
- What is it?
- Guardian CLI (zakirkun/guardian-cli) is a Python 3.11+ command line tool that drives nmap, nuclei, sqlmap and dozens of other scanners through an AI planning layer. The interesting part is the DAG workflow format and the evidence trail, not the model choice.
- Who is it for?
- Adopt Guardian CLI if you already run authorised assessments and want a single command surface over nmap, nuclei and sqlmap with an inspectable evidence trail. Skip it if you need a mature, narrowly scoped scanner: the project describes itself as Beta in pyproject.toml and installs a large tool chain.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 95 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Guardian CLI targets: scanner output that nobody correlates
A normal web assessment ends with several tools writing separate files. nmap gives you ports, httpx gives you live hosts, nuclei gives you template hits, and a human stitches them together in a spreadsheet. Guardian CLI's pitch is that the stitching is the job, and that a language model can do the next-step selection better than a fixed script.
The audience is narrow. This is for a penetration tester or a small security team running authorised engagements who is comfortable with Python packaging, YAML, and the idea of an LLM reading tool output. It is not a managed scanning service and it does not replace a tester's judgement. The README states the legal position plainly: Guardian is designed for authorised security testing and educational purposes, and the operator is responsible for having written permission.
The project also carries a Beta classifier in pyproject.toml (Development Status :: 4 - Beta) even though the README calls it enterprise-grade. Those two statements sit in the same repository, and the classifier is the more useful signal.
How the planner, tool selector and analyst agents divide the work
The architecture described in the README is a set of specialised agents rather than one prompt. A Planner decides the sequence, a Tool Selector picks from the registered arsenal, an Analyst interprets output, and a Reporter assembles findings. On top of that sit debate roles: a Red Advocate, a Blue Advocate and a Judge, used for triage on ambiguous findings. The README claims the three-role debate reaches F1 at least five percentage points above a single-agent baseline, but no evaluation data is published in the repository, so treat that as a project claim rather than a measured result.
Two mechanisms are worth understanding before you trust any output. First, the RAG knowledge base: SQLite with FTS5 over CVE, CWE and MITRE ATT&CK feeds, which the README says exists to kill hallucinated CVE references. Second, prompt-injection defence: all tool output is wrapped in UNTRUSTED_TOOL_OUTPUT delimiters and ANSI escapes are stripped before it reaches a model. That second point matters more than it sounds. A scanner that prints attacker-controlled text into its output is a channel into your agent loop, and Guardian is one of the few tools in this space that names the problem.
There is also a cost control called judge model routing, described as a think_deeply swap-and-restore: a large model does the reasoning, a small model judges, and the README puts the saving at roughly 10x. The number is unverified, but the pattern is a real one.
The workflow DSL: depends_on, Jinja2 parameters and the DAG scheduler
This is the part of Guardian CLI that has the least to do with AI and the most to do with whether you can operate it. Workflows are YAML. Steps declare depends_on, and independent steps run in parallel up to max_parallel_tools. Parameters are Jinja2 templates resolved against prior step results, and the README gives the shape as parameters: {key: "{{ <id>.parsed.alive_hosts }}"}. Conditional steps use a when: clause gated on earlier output.
Parameter precedence is stated explicitly: workflow YAML beats the config block, which beats tool defaults. That ordering is the kind of detail that saves an afternoon, because it tells you where to look when a scan runs with the wrong wordlist.
Sessions are checkpointed atomically to session_<id>.json, which is what makes --resume possible. The README says resume picks up after the last completed step. It does not document what happens when a step is interrupted mid-execution, and that is the failure mode you will actually hit, because a long nmap or sqlmap run is exactly the step you want to kill and restart.
Analysis steps can name an agent directly with agent: debate | visual | analyst, so the workflow file is also where you decide how much money a finding is worth spending on.
Installing Guardian CLI and running a first workflow
The package is published as guardian-cli and requires Python 3.11 or higher. The console entry point is guardian, mapped in pyproject.toml to cli.main:app. The README's prerequisites section lists an AI provider API key as required, and .env.example shows the Google Gemini key as the documented default, with LangSmith tracing commented out.
Start by creating the environment file. Copy .env.example to .env and set the key; the rest of the file is optional and commented.
cp .env.example .envThe repository ships a pyproject.toml with a litellm extra and a dev extra containing pytest, black and ruff. Installing the package itself is not shown in the README's prerequisites section; the repository layout points to pyproject.toml as the build definition, and the console script guardian is declared there.
If you would rather not install the tool chain by hand, docker-compose.yml builds the image guardian-cli:latest and mounts reports, logs, config and workflows from the host. The service is named guardian in that file, so any command you run inside it goes through that service name.
cp .env.example .env
docker compose up -dThe README says 50 tools are registered but none are imported until needed, and that this keeps --help under 500ms. That is a claim about lazy loading, not a measurement you can take from the repository, but the design is visible in the tools/ package layout. Note that the Dockerfile installs a subset of the arsenal (the file header says 15 security tools) built from projectdiscovery images and an Alpine base with gcompat for Go binary compatibility.
The README does not give a canonical first scan command in the repository, so check QUICKSTART.md in the repository root for the intended entry point rather than guessing flags.
The confirmation gate, scope validation, and why safe mode is not a sandbox
Guardian CLI implements a confirmation gate: tools classified as Active+ (intrusive or destructive) require explicit user approval, and Safe Mode prevents destructive actions by default. Scope validation resolves DNS before deciding whether a target is in scope, and the README says this closes an SSRF-class bypass by blacklisting private RFC1918 ranges.
Read that as a guardrail, not a container. The tool still shells out to real scanners through asyncio subprocesses. If you approve an intrusive tool, it runs with your privileges on your network. The confirmation gate protects you from the agent deciding to be aggressive; it does not protect you from a misconfigured scope file.
The honest limitation is the model layer. An LLM choosing the next tool is a probabilistic decision, and Guardian's own mitigations (RAG grounding, debate triage, CVSS v3.1 recomputation that flags drift between a claimed score and the vector math) exist because the model gets things wrong. A tester who wants deterministic, reproducible, byte-identical runs across two invocations will find that property absent here by design. If your engagement requires a fixed, auditable scan plan, a plain nuclei or nmap invocation script is the better tool.
Guardian CLI versus a plain scanner wrapper or a manual tool chain
The realistic alternative is not another AI pentest agent. It is the thing most testers already have: a shell script or Makefile that runs nmap, httpx and nuclei in sequence and writes to a directory. That approach is deterministic, has no API cost, and produces output you can diff between runs. Guardian CLI trades all three for adaptive step selection and a structured evidence trail where every finding links to its source execution via execution_id.
That execution_id link is the strongest argument for the project. The README describes complete command history preserved with each finding, raw output snippets bound to findings, per-URL screenshots attached to web findings through a playwright_screenshot tool, and a per-agent ledger of token usage, cost and thinking chain. If you have ever tried to explain to a client why a finding is real three weeks after the scan, that traceability is the feature you are buying.
On output, Guardian exports SARIF v2.1.0 with security-severity and dedup fingerprints derived from execution_id, plus direct DefectDojo REST upload and Slack webhook posts with severity colour coding, all triggered from a single report invocation.
A narrower comparison is a single-purpose scanner with an API. Those give you one category of finding with no orchestration layer and no model cost. Guardian is the wrong tool when the engagement is one scanner's job.
Licence, maintenance and the real upgrade cost
The repository's LICENSE file is present, and pyproject.toml declares license = {text = "MIT"} with the OSI MIT classifier. GitHub reports the licence as NOASSERTION because the classifier and the file have not been matched automatically, so read LICENSE yourself before you redistribute anything. MIT is permissive, but it carries no warranty, which matters for a tool that runs intrusive scanners on your behalf.
The last push to the default branch was on 2026-06-27. The most recent release listed is v2.1, tagged on 2026-02-27, following v2 on 2026-02-03. Version numbers in the repository are inconsistent: pyproject.toml still declares version 0.1.0 while the release tags say v2.1, and the Dockerfile labels the image version 1.0.0. Do not rely on the package metadata to tell you which build you are running.
The upgrade cost is dominated by the tool chain, not the Python package. The Dockerfile pulls projectdiscovery/httpx, projectdiscovery/subfinder, projectdiscovery/nuclei, ghcr.io/oj/gobuster, owaspamass/amass and zricethezav/gitleaks as build stages and compiles ffuf from source with go install. Those upstream images move independently of Guardian, so an image rebuild can change scanner behaviour without a Guardian release. Pin the image digest if you need reproducibility. Note also that the README lists 50 tools across categories including Active Directory, mobile Android and LLM red-team, while the Dockerfile header describes 15 tools, so a container run will not have the full arsenal.
Editorial conclusion
Adopt Guardian CLI if you already run authorised assessments and want a single command surface over nmap, nuclei and sqlmap with an inspectable evidence trail. Skip it if you need a mature, narrowly scoped scanner: the project describes itself as Beta in pyproject.toml and installs a large tool chain. Before trusting a run, verify that the confirmation gate actually blocks an Active+ tool on your machine, check whether --resume works after you kill a session mid-step, and confirm which of the 50 tools your environment really has, since the Dockerfile installs a subset.
Frequently asked questions
How do I install Guardian CLI?
Guardian CLI requires Python 3.11 or higher and is packaged as guardian-cli, with the console entry point guardian defined in pyproject.toml. Copy .env.example to .env and set GOOGLE_API_KEY; the repository also ships a docker-compose.yml that builds the image guardian-cli:latest.
Does Guardian CLI run locally or does it need a cloud service?
The README lists Ollama (local) and OpenAI-compatible endpoints such as vLLM and LM Studio among the supported providers, so a local model is possible. The default documented key in .env.example is a Google Gemini key, which sends tool output to a remote provider.
What is CLI in cyber security?
In this context it means a command line tool that automates security work, which is what Guardian CLI is: a Python application invoked as guardian that orchestrates scanners rather than a graphical dashboard. Guardian CLI specifically wraps tools such as nmap, nuclei and sqlmap behind a YAML workflow format.
What is the difference between a CLI and an API in Guardian CLI?
Guardian CLI is the CLI side: you invoke guardian and it runs workflows locally. The API surface it does use is outbound, calling AI providers and exporting results to DefectDojo over REST and to Slack via webhook.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zakirkun-guardian-cli)