AIRecon: a local LLM pentest agent with a Kali sandbox and a Caido bridge
AIRecon is an autonomous cybersecurity agent that combines a self-hosted Large Language Model (Ollama) with a Kali Linux Docker sandbox and a Textual TUI. It is designed to automate security assessments, penetration testing, and bug bounty reconnaissance — without any API keys or cloud dependency.
At a glance
- What is it?
- AIRecon wires an Ollama model to a Kali Linux Docker sandbox, a Textual TUI and a four-phase recon pipeline, with no API keys and no cloud calls. The catch is that the local model has to be good enough to call tools reliably.
- Who is it for?
- Adopt AIRecon if you already have a GPU that can hold a 32B-class model and you want target data to stay on your own machine during recon. Skip it if your hardware tops out below 8B parameters or if you need a stable, released tool: the newest release is 0.1.7-beta from 2026-04-03.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The bill that AIRecon is trying to avoid
Autonomous reconnaissance is a loop, not a request. The agent observes, picks a tool, reads output, and repeats. The README states that recursive workflows of this kind can require thousands of LLM calls per session, and that is where hosted models stop making financial sense. AIRecon's answer is to remove the meter entirely: the model runs under Ollama on your own hardware, and the README describes operation as completely offline with no API keys required.
The audience is narrow and specific. It is a penetration tester or bug bounty hunter who already owns a GPU, is comfortable with Docker, and cares that target intelligence and tool output never leave the machine. It is not aimed at someone who wants a hosted agent they can start from a laptop with no local compute. The README's own comparison table makes that trade explicit: no API keys, no target data sent to cloud, works offline, native Caido integration, session resume, and a local knowledge base.
RECON, ANALYSIS, EXPLOIT, REPORT, and the checkpoints between them
The pipeline is four named phases: RECON, ANALYSIS, EXPLOIT, REPORT. According to the README, each phase carries its own objectives, recommended tools, and automatic transition criteria, and phase enforcement is soft, meaning the agent is guided but never blocked. That is a deliberate design choice with a real consequence: a model that decides to skip ahead is allowed to. Nothing in the README suggests a phase gate that refuses to advance.
Checkpoints fire on an iteration schedule rather than a phase schedule. Every 5 iterations the agent evaluates its phase, every 10 it self-evaluates, and every 15 it compresses context. That cadence matters for anyone trying to predict cost and latency, because context compression is the mechanism that keeps a long session inside a local model's window.
The learning story is worth reading carefully because the README is blunt about it: AIRecon does not fine-tune the LLM. What it calls learning is local telemetry. Sessions, findings, patterns, target intel, tool usage, model performance, skill usage and attack-chain discoveries go into a SQLite database at ~/.airecon/memory/airecon.db. Adaptive state lives at ~/.airecon/learning/global_learning.json, holding tool performance stats, strategy patterns, an observation log and distilled insights. Per-target memory files under ~/.airecon/memory/by_target/ record endpoints, vulns, WAF bypasses, sensitive params and auth endpoints.
That telemetry feeds back into behavior. On session start, memory context is injected. Every 8 iterations, learned patterns and similar findings can be re-injected based on detected technology. Adaptive tool ranking reorders tools by historical success and failure. Payload memory, when enabled, skips payloads that repeatedly failed for the same target and parameter. This is retrieval and ranking, not model improvement, and the README says so.
Installing AIRecon and running a first scoped session
The package is published as airecon, with a poetry script entry point at airecon.__main__:main, and the README badge states Python 3.12+. The repository ships both a pyproject.toml and a requirements.txt, so either path is available. The pyproject declares the runtime dependencies directly, including textual, httpx, fastapi, uvicorn, pydantic, playwright and cvss.
Start by pulling a model that supports native tool calling. The README's table gives the exact pull commands and the VRAM each one needs.
ollama pull qwen3.5:35bThe README calls Qwen3.5 35B the recommended option for most users at 20 GB of VRAM, with Qwen3.5 9B at 6 GB listed as the minimum viable choice and expected to produce frequent errors. The 122B entry needs 48+ GB.
If your local VRAM is below the minimum, the repository includes a Colab notebook at scripts/airecon_colab.ipynb that runs Ollama on a free T4 GPU and exposes it through a cloudflared tunnel. The README says to select Runtime, then Change runtime type, then T4 GPU, run all cells top to bottom, and copy the snippet printed in Cell 6 into ~/.airecon/config.yaml:
ollama_url: "https://xxxx.trycloudflareThat is the whole point of the config key: AIRecon's client speaks the OpenAI /v1/chat/completions wire format over httpx, and the pyproject notes that the openai SDK is intentionally not a dependency, so any OpenAI-compatible gateway can sit behind ollama_url. The same file is where you point AIRecon at a local vLLM or LiteLLM instance instead of Ollama.
Once the model answers, launch the TUI and give it a scoped target. The README does not print a full end-to-end command transcript, so treat the first session as a capability check rather than a production run: watch whether the agent actually invokes tools such as http_observe and execute, or whether it narrates results it never obtained. That single observation tells you more about your model choice than any benchmark table.
The model is the failure mode
AIRecon's most important limitation is stated in its own README with unusual directness. Tool calling support is required, and models without native function calling make AIRecon completely non-functional, because the agent cannot execute http_observe, execute, or browser actions. This is not a degradation. It is a dead tool.
The size guidance is equally specific. Models at or above 32B are described as reliable for full recon pipelines with good tool calling accuracy. Models between 8B and 14B are usable for simple tasks with 20 to 40 percent tool call errors and hallucinations expected. Models below 8B are called technically workable but unreliable, with frequent hallucinated tool output, invented CVEs, and skipped scope rules. The README names DeepSeek R1 as producing incomplete function calls.
Read those numbers as a hardware requirement, not a suggestion. A 20 GB VRAM floor for the recommended model puts AIRecon out of reach for a large share of laptops, and the Colab route trades local privacy for a third-party GPU, which undercuts part of the reason to choose this tool. Capabilities are auto-detected through ollama show metadata, so a mismatched model will be identified rather than silently misused, but detection does not fix it.
There is a second, quieter limitation: the whole project is beta. The newest release listed is 0.1.7-beta from 2026-04-03, and the pyproject version string is 1.7.1-beta. The README does not document rollback, and it does not describe what happens to an existing ~/.airecon/memory/airecon.db when the schema changes between versions.
Where AIRecon sits next to other agent stacks
The obvious alternative is a cloud agent built on GPT-4, Claude or Gemini, which the README names directly in its cost argument. The difference is not quality, it is where the data goes and who pays per call. A cloud agent needs no local GPU, starts faster, and generally calls tools more reliably at the top end of the model range. AIRecon needs a GPU, needs Docker, and inherits whatever tool-calling discipline your local model has. In exchange, target data stays on disk and a long recursive session costs electricity rather than tokens.
The second comparison is to running a Kali container by hand with your own scripts and notes. That is more work per engagement, but it has no model dependency at all and no chance of a hallucinated finding reaching a report. AIRecon's value over that baseline is the RECON to REPORT structure, the memory that survives between sessions, and the Caido integration, which the README describes as five built-in tools covering list, replay, automate with §FUZZ§, findings, and scope. If you do not use Caido, that specific advantage disappears.
The extended knowledge bases sit outside the main repository. airecon-skills is described as a community library with 57 additional CLI-based playbooks for CTF, bug bounty and pentesting, and airecon-dataset indexes roughly 1.09 million security records into local SQLite FTS5 so the LLM can call dataset_search before attempting unfamiliar techniques. Both are optional, and both are separate projects with their own maintenance.
Licence, upgrade cost, and what the beta label means in practice
AIRecon is MIT licensed, per both the LICENSE file and the pyproject license field. That is permissive: you can use it commercially, modify it, and redistribute it, provided the copyright notice and permission notice are kept. The licence covers the AIRecon code. It does not cover the models you pull through Ollama, the Kali container image, Caido, or the optional dataset and skills repositories, each of which carries its own terms. That is a factual boundary, not legal advice, and anyone shipping AIRecon inside a commercial service should read the terms of those components separately.
Upgrade cost is where the beta status bites. Three releases are listed in quick succession: v0.1.5-beta on 2026-03-05, v0.1.6-beta on 2026-03-17, and 0.1.7-beta on 2026-04-03. The last push to the repository was on 2026-06-20. That is a project moving in small increments rather than a frozen artifact, and the README does not describe a migration path for the on-disk state under ~/.airecon/. If you keep long-lived target memory, back up ~/.airecon/memory/airecon.db before upgrading, because nothing in the documentation promises the schema is stable across beta versions.
There is also a versioning inconsistency worth noticing: the release tags read as 0.1.x while the pyproject version reads 1.7.1-beta. Anyone pinning a version from a package index rather than a git tag should confirm which number they are actually installing.
Editorial conclusion
Adopt AIRecon if you already have a GPU that can hold a 32B-class model and you want target data to stay on your own machine during recon. Skip it if your hardware tops out below 8B parameters or if you need a stable, released tool: the newest release is 0.1.7-beta from 2026-04-03. Before trusting it on a real engagement, run one scoped target end to end, check that tool calls actually execute rather than being hallucinated, and confirm the model you pulled supports native function calling.
Frequently asked questions
Does AIRecon require API keys or a cloud connection?
No. The README states that AIRecon runs on a self-hosted Ollama model and operates completely offline with no API keys required, and that target data is not sent to the cloud.
What model does AIRecon need to work properly?
It requires a model with native tool calling and extended thinking blocks. The README recommends Qwen3.5 35B for most users at 20 GB of VRAM, calls 9B the minimum viable option, and warns that models without tool calling make AIRecon completely non-functional.
Can AIRecon run without a local GPU?
Yes, through the Colab notebook at scripts/airecon_colab.ipynb, which runs Ollama on a free T4 GPU and exposes it over a cloudflared tunnel. You then set ollama_url in ~/.airecon/config.yaml to the tunnel address.
Does AIRecon fine-tune the model on your engagements?
No. The README says AIRecon does not fine-tune the LLM, and that its learning is local telemetry stored in ~/.airecon/memory/airecon.db and ~/.airecon/learning/global_learning.json that guides tool ranking and avoids repeated failed paths.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pikpikcu-airecon)