Model or dataset
Vasco0x4/AIDA avatar
Vasco0x4/AIDA

AIDA isolates its agent in a container, and --localhost takes that away with one flag

Turn any LLM into an autonomous pentester. You define the scope, the agent does the work, you review the findings.

492 stars82 forksJavaScriptAGPL-3.0

At a glance

What is it?
An autonomous pentesting agent that claims four published CVEs, where the defensible part is the Docker socket proxy that denies run, create and volumes, the least defensible is a flag that moves command execution onto the host, the database defaults to aida over aida, and a public deploy flag sits next to an instruction not to expose the dashboard.
Who is it for?
AIDA is a serious piece of work for the case it is built for, which is an assessment you are authorised to run, and the parts worth studying are the containment ones: the Docker socket is reached through a proxy that permits inspect, exec and start and refuses run, create and volume operations, the database port is bound to the loopback interface with a note that the mapping exists only for host-side tools, and findings are auto-scored so a retest is comparable.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 17 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Only against systems you have written permission to test

AIDA runs full security assessments end to end against web applications, APIs and infrastructure, and the entry point is three commands:

bash
git clone https://github.com/Vasco0x4/AIDA.git
cd AIDA
./start.sh

That brings up a dashboard on port 31337, and the project's own framing is that you define the scope and review the findings. The word pentester in the title is literal: the container it ships carries nmap, sqlmap, ffuf and nuclei, the agent installs a tool itself when one is missing, and it generates and executes its own Python rather than only invoking what is on the image. The scope is expressed as a natural-language instruction to the agent naming what to include and what to exclude, and the tool is designed to hold that boundary. None of that makes it something to point at a third-party system. Use it on infrastructure you own or on a scope you have written agreement for, get that agreement in writing before the first scan, and treat every subdomain and endpoint you did not name as out of scope even if the agent finds it. The dashboard warning in the same file is worth repeating: run it locally or on your LAN and do not expose the dashboard to the public.

--localhost removes the container that makes the tool defensible

The launcher has four documented modes and one of them undoes the central safety property. `start.sh` with no arguments is the normal path. `--dev` gives hot reload for contributors, `--lan` shares across the local network over HTTPS with a self-signed certificate, and `--domain x.com` performs a public deploy with Let's Encrypt. The fourth is `--localhost`, described in the flag list as running commands directly on the host with no pentest container, and it points at a documentation section for the details. Consider what that means for the rest of the page. The agent's `execute()` tool runs any command in the pentesting container, `python_exec` runs Python in it, and the environment is a two gigabyte image that starts automatically. With `--localhost` those tools have the privileges of whoever ran the script, on whatever machine is hosting the AI client, and the container boundary that made an autonomous agent acceptable to point at a network is simply not there. That is a decision to make deliberately, in writing, per engagement, not a convenience flag.

The Docker socket is reached through a proxy that denies run, create and volumes

This is the design decision that makes the rest defensible, and it is stated in the compose file rather than in a security page. Letting an agent run commands in a container normally means handing it the Docker socket, which is root-equivalent on the host and can start anything. The backend does not get the raw socket. A comment in the compose file says the docker CLI is routed through a socket proxy instead, and that the proxy allows only inspect, exec and start, with no run, no create and no volume operations. So the agent can start the assessment container and run commands inside it, and it cannot start a different container, cannot mount a host path, and therefore cannot climb out. The v1.1.0 notes list that proxy alongside path traversal prevention and a localhost-only database as the release's security hardening. If you deploy this yourself, that proxy is the component to verify first, because it is the one that bounds the blast radius of a misbehaving or manipulated model.

--domain puts the dashboard on the internet with a real certificate

The public deploy flag and the warning against public exposure are in the same document, and the tension is worth naming rather than resolving. `start.sh --domain x.com` issues a Let's Encrypt certificate and serves the dashboard, which is a genuinely convenient thing to be able to do from a phone on a client site. Against it, the release notes say in bold to run locally or on your LAN and not to expose the dashboard to the public. What stands between a public deployment and an unauthorised reader is the authentication added in v1.1.0: JSON Web Tokens, admin and user roles, and a first-run setup wizard. That is a reasonable control, and it is a single control. The compose file also sets a CORS origin list that the LAN mode fills with the LAN address plus localhost, so the cross-origin surface changes shape depending on how it was started, and a first-run wizard is exactly the kind of thing that gets skipped on a second install.

The database defaults to aida over aida, bound to loopback

The compose file spells out its own defaults, which is rare and worth reading closely. PostgreSQL 16 on Alpine takes its user, password and database name from variables with fallbacks: the user defaults to `aida`, the password defaults to `aida`, and the database defaults to `aida_assessments`. An assessment database holds findings, credentials and raw command output, so a default password of the product name is the wrong place to be casual. The port mapping is the part that mitigates it, and the comment above it is explicit: the mapping is bound to `127.0.0.1`, the database must never be reachable from the LAN, the backend reaches postgres over the internal Docker network, and the published port exists only so host-side tools like psql, pgAdmin or DBeaver can connect, with a note to delete the line to lock it down further. That is the right default posture, undermined only by the credential fallback and by anyone changing the bind address without reading the comment.

Stored credentials get spliced into requests, and findings can leave by message

Two features move data across a trust boundary, and both are on by default in the sense that they are part of the tool. The first is credential handling: `credentials_add` stores credentials that are then auto-injected into outgoing requests through `{{PLACEHOLDER}}` substitution, and `http_request` performs that substitution for you. The convenience is real, since it is how a session cookie reaches twenty endpoints without being retyped, and the cost is that the model composing a request does not have to know the secret to include it. A request the agent is induced to send to the wrong host carries the credential with it, so the set of hosts you have stored credentials for is a security decision, not a convenience setting. The second is notifications: v1.1.0 adds Telegram, Slack and email with an optional PDF attachment, and a finding is documented with the commands used, the raw output and the surrounding context. Assessment output leaving the host through a chat channel is worth deciding on per engagement.

--cli auto detects Qwen, which the support table does not list

Model support is presented as model-agnostic, with any tool-calling model said to work, and the table lists Claude Code, the OpenAI Codex CLI and Kimi Code CLI as automatic through the Python launcher, an external OpenAI-compatible endpoint through a base URL flag, and Claude Desktop, ChatGPT Desktop and Gemini CLI through MCP configuration. The auto-detect line names one more. `python3 aida.py --assessment "target-corp" --cli auto` is documented as auto-detecting Claude, Codex, Kimi, or Qwen, and Qwen has no row in the support table. So the discovery path covers a client the compatibility list does not describe, which is the kind of gap that matters here because the quality of the engagement is attributed to the model. The dashboard listens on port 31337, the frontend development server on 5173, and the persistent notebook is what lets an engagement stop and resume days later for a retest or a handoff.

The backend is a mutable latest tag, and the CLI needs three packages

The backend service in the compose file is defined twice at once, with an image of `vasco0x4/aida-backend:latest` and a build context pointing at `./backend` with its own Dockerfile. That means the artefact running your assessments comes from a mutable tag in a personal registry namespace rather than a pinned digest, and locally it can be rebuilt from a context you have read. For a tool whose whole job is executing commands, the backend is the component you would least like swapped underneath you between runs. The other dependency detail is smaller but has cost people time: `requirements.txt` holds click, httpx and rich, and its own comment says these are the AIDA CLI dependencies for the local launcher and are separate from the backend dependencies. The backend's own requirements are not in that file, so a reader looking for them will not find them there.

Editorial conclusion

AIDA is a serious piece of work for the case it is built for, which is an assessment you are authorised to run, and the parts worth studying are the containment ones: the Docker socket is reached through a proxy that permits inspect, exec and start and refuses run, create and volume operations, the database port is bound to the loopback interface with a note that the mapping exists only for host-side tools, and findings are auto-scored so a retest is comparable. Two flags decide how much of that you get. `--localhost` runs agent commands on the host with no pentest container, which is the difference between a sandboxed assessment and an autonomous agent with an unrestricted shell on your machine, so treat it as a decision rather than a convenience. And `--domain` puts the dashboard on the public internet with a real certificate, against the project's own instruction not to expose it, so the authentication added in v1.1.0 becomes the only barrier. Read LICENSE and the notification settings before an assessment runs, because findings travel. The newest release is v1.1.0 from 15 April 2026 and the last commit on the default branch main is dated 15 September 2026.

Frequently asked questions

What does AIDA need before it will run?

Docker Desktop and an AI client such as Claude, Gemini or GPT. You clone the repository, change into it and run ./start.sh, and the dashboard comes up on http://localhost:31337. The assessment is then launched with python3 aida.py and a name, with --cli to pick a client or --cli auto to detect one.

What is the difference between the default mode and --localhost in AIDA?

The default runs agent commands inside the built-in aida-pentest container. --localhost runs commands directly on the host with no pentest container, which removes the isolation boundary that keeps an autonomous agent's command execution off the machine running it.

How does AIDA keep the agent from escaping its container?

The docker CLI is routed through a socket proxy rather than the raw daemon socket, and the proxy allows only inspect, exec and start, with no run, create or volume operations. The v1.1.0 notes list that proxy, path traversal prevention and a localhost-only database as the release's security hardening.

How are AIDA findings scored and exported?

The add_card tool logs a finding and auto-scores it with CVSS 4.0, and every confirmed vulnerability is documented with the commands used, the raw output and full context. v1.1.0 adds one-click PDF export per assessment, an attack timeline, a cross-assessment findings view, and notifications through Telegram, Slack or email with an optional PDF attachment.

Official sources

  1. Issues
  2. License: AGPL-3.0
  3. README
  4. Releases
  5. Vasco0x4/AIDA on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vasco0x4-aida.svg)](https://hysenlabs.com/projects/vasco0x4-aida)