# LuaN1aoAgent v2: a Planner-Executor-Observer pentest agent with an evidence graph

> LuaN1aoAgent v2 is a TypeScript rewrite of a Python pentest agent, built on the Pi SDK. It splits work across a Planner, an Executor and an Observer, and requires every confirmed vulnerability to trace back to stored evidence.

**SanMuzZzZz/LuaN1aoAgent** — LuaN1aoAgent is a fully autonomous AI-driven penetration testing agent powered by graph-based cognitive reasoning.

- Repository: https://github.com/SanMuzZzZz/LuaN1aoAgent
- Stars: 1,323 · Forks: 189
- Language: TypeScript
- License: AGPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/sanmuzzzzz-luan1aoagent

## What LuaN1aoAgent v2 solves, and for whom

Most LLM pentest agents keep their reasoning inside a conversation. The model observes something, concludes something, and the conclusion lives in a message history that nobody can audit afterwards. LuaN1aoAgent v2 is built against that pattern. The README states the design principle directly: every important conclusion must remain traceable to persisted events, artifacts, and graph evidence.

The audience is narrow and explicit. The repository describes the project as being for autonomous, authorized security research. If you are running an engagement where you can point a tool at a target and defend that decision in writing, this is aimed at you. If you want a scanner that produces a report from a single command, the graph machinery here is overhead you will pay for and not use.

The v2 line is not an in-place refactor. The README calls out that it is a new implementation with different configuration, persistence, Agent lifecycle, and observability contracts from the Python v1 runtime, and it declines to carry v1 benchmark numbers forward. That is an unusually blunt statement for a project README, and it sets the expectation correctly: existing v1 deployments do not migrate by changing a version pin.

## How the Planner, Executor and Observer split work

Three roles hold the runtime. The Planner reads compact task, reasoning and operation graph views, and creates or patches goal-level tasks rather than prescribing low-level actions. It owns dependencies, priority, independent-task concurrency, scope and task budgets, and it submits decisions through a terminating tool named planner_submit. The README notes that the Planner reconciles ready tasks against available capacity after graph changes and task handoffs, without waiting for an entire parallel wave to finish.

The Executor receives a bounded TaskEnvelope and picks its own tool strategy. It records public intent, tool input, tool output, usage, errors and final results, and it preserves large outputs as immutable artifacts instead of inflating Agent context. One detail matters for anyone reasoning about cost and state: the Executor reuses the same persisted Pi session lineage and workspace across epochs of one Task, while different Tasks stay isolated. It submits through task_result_submit.

The Observer runs in two modes. Supervisor is on the hot path and inspects recent Executor actions to decide whether to continue, checkpoint, stop, or return control to the Planner. Projector runs asynchronously and converts normalized observations into graph deltas. Each Observer invocation uses a fresh Pi session without sharing hidden model history, and the two modes submit through control_submit and graph_delta_submit respectively.

The graph itself is the interesting part. Evidence nodes connect to Hypothesis nodes through supports or contradicts edges. A Hypothesis confirms a Vulnerability, and a Vulnerability is exploited by an Exploit. Confirmed Vulnerability nodes and successful Exploit nodes cannot be written without evidence references, so the graph refuses an unsupported claim at write time rather than flagging it in a later review. Hypotheses stay distinct from confirmed findings by construction.

## Installing LuaN1aoAgent v2 and running a first task

The repository ships install.sh at the top level, and package.json exposes the build and start scripts. Node.js 25 or newer is required according to the README badge, and the runtime is the Pi SDK package @earendil-works/pi-coding-agent.

Start by copying the environment template. The file itself says to copy it to .env and fill in real values, and notes that .env is gitignored.

```bash
cp .env.example .env
```

The required keys are LLM_API_KEY, LLM_API_BASE_URL, LLM_DEFAULT_MODEL and LLM_API_TYPE. The example file normalizes a trailing /chat/completions or /responses away, so a full endpoint URL is accepted.

```bash
LLM_API_KEY=your-api-key
LLM_API_BASE_URL=https://api.openai.com/v1
LLM_DEFAULT_MODEL=your-model-id
LLM_API_TYPE=openai-completions
```

One warning in that file deserves repeating because it causes silent failures rather than errors. The default max_completion_tokens is 32768, and the comment states you should keep it well above 8192 because reasoning models burn reasoning tokens inside the completion budget, and a small cap truncates structured submissions. If the Planner or Executor stops returning graph operations, that cap is the first thing to check.

Thinking behaviour is configured separately. LLM_THINKING accepts off, minimal, low, medium or high, and LLM_THINKING_FORMAT selects the wire format. The example file states that the zai format sends thinking:{type:...} and is verified on glm-5.2, with thinking off producing zero reasoning tokens. Other vendors need deepseek, qwen or openai to match.

Build and run:

```bash
npm install
npm run build
npm start
```

npm run build runs build:server and build:web in sequence. npm start runs node dist/src/cli.js. The web interface is a separate entry point, npm run web, which starts dist/src/web-server.js, and npm run web:dev serves the Vite dev build. The build:executor-image, build:network-image and build:traffic-proxy scripts produce the executor container, the network container and the Go traffic proxy binary respectively, and each needs Docker or a Go toolchain present.

## Where the design costs you: provider coupling and unverified benchmarks

The per-role configuration is flexible on paper and demanding in practice. Each role falls back to the shared LLM_* values, but you can give the Executor its own endpoint and key through LLM_EXECUTOR_BASE_URL and LLM_EXECUTOR_API_KEY. That is useful when you want a cheap model doing tool loops and an expensive one planning. It also means the quality of the whole run depends on a provider combination nobody has published results for. The example file names glm-5.2 and deepseek-v4-pro as sample model IDs, and states that the zai thinking format is verified on glm-5.2. Everything else is left to you to validate.

Admission control is a shared FIFO limit. LLM_PROVIDER_MAX_CONCURRENT defaults to 3 and applies to roles using the same provider. Raise it and you may hit rate limits; lower it and the independent-task concurrency the Planner manages has nothing to run against.

The benchmark position is the sharpest limitation. The README states plainly that benchmark results reported by v1 are not automatically attributed to v2, and that v2 results will be published only after reproducible reruns on a frozen release. As of the v2.0.0 release on 2026-07-20, no v2 benchmark numbers are presented. Anyone comparing this against another agent is comparing architecture descriptions, not measured outcomes.

The last push to the repository was on 2026-08-24. The project is not archived, and there is a gap between the release and the most recent commit that the README does not explain.

## How it differs from other autonomous pentest agents

The related searches around this project point at a crowded field: GHOSTCREW, RedAmon, Penligent, Shannon, AutoAgent. The README does not compare LuaN1aoAgent against any of them, so the honest difference to describe is architectural rather than competitive.

The common shape in this category is a single agent loop with a shared message history, sometimes with sub-agents that inherit that history. LuaN1aoAgent v2 inverts it. The Planner never sees raw tool output; it reads compact graph views. The Executor never decides scope; it gets a bounded envelope. The Observer runs with a fresh Pi session each time and does not share hidden model history with the roles it inspects. Large tool outputs become immutable artifacts rather than context.

That buys auditability and costs tokens. Three roles plus two Observer modes means more model calls per unit of work than a single-loop agent doing the same reconnaissance. The return is that a confirmed finding carries a reference to the events that support it, and a hypothesis cannot be promoted to a confirmed vulnerability without one. If your workflow already ends in a human writing the report from raw logs, you are paying for a property you will not use.

## Licence and the cost of tracking a v2 runtime

The project is licensed AGPL-3.0, and package.json declares AGPL-3.0-only. That is the network-copyleft variant. If you modify LuaN1aoAgent and let users interact with it over a network, the licence's source-availability obligation is the part to read carefully with your own counsel. Running it unmodified against your own authorized targets is a different situation from shipping a modified version as a hosted service. Nothing here is legal advice.

The upgrade cost is the more practical concern. Because v2 changed configuration, persistence and Agent lifecycle contracts relative to v1, and because v1 is published as a separate Legacy release, there is no migration path described in the README. Scripts that read v1 state files or drive v1 configuration will not carry over. The .env.example keys are v2 keys, and the per-role override naming follows a ROLE in PLANNER/EXECUTOR/SUPERVISOR/PROJECTOR pattern that a v1 deployment would not have used.

Maintenance burden sits mostly in the provider layer. The thinking-format setting is vendor-specific and the example file only claims verification for one format on one model family. A provider changing its reasoning parameter is the kind of change that breaks structured submissions quietly, which is why the token-cap warning in the same file is worth keeping in mind alongside it.

## Conclusion

Adopt LuaN1aoAgent v2 if you run authorized engagements and want agent conclusions tied to persisted events and artifacts rather than to a chat transcript. Do not adopt it as a drop-in upgrade from the Python v1 runtime: the README states v2 has different configuration, persistence, Agent lifecycle and observability contracts, and that v1 benchmark results are not attributed to v2. Before pointing it at anything, check that your provider supports the thinking wire format you configure through LLM_THINKING_FORMAT, and confirm that the Node.js 25+ requirement matches your host.

## FAQ

### What is LuaN1aoAgent and who is it for?

It is an autonomous security agent for authorized penetration testing and security research, built in TypeScript on the Pi SDK. It is aimed at teams running authorized engagements who need every conclusion traceable to persisted events, artifacts and graph evidence.

### How do I install LuaN1aoAgent v2?

The repository provides install.sh at the top level, and package.json defines the build pipeline. Copy .env.example to .env, fill in LLM_API_KEY, LLM_API_BASE_URL, LLM_DEFAULT_MODEL and LLM_API_TYPE, then run npm install, npm run build and npm start. Node.js 25 or newer is required.

### Is LuaN1aoAgent v2 a drop-in upgrade from v1?

No. The README states v2 is not an in-place refactor of the Python v1 runtime and has different configuration, persistence, Agent lifecycle and observability contracts. It also states that v1 benchmark results are not automatically attributed to v2. v1 is published separately as a Legacy release.

### Which LLM providers does LuaN1aoAgent v2 work with?

The configuration accepts any OpenAI-compatible endpoint through LLM_API_BASE_URL, with LLM_API_TYPE set to openai-completions or openai-responses. LLM_THINKING_FORMAT selects the thinking wire format, and the example file states the zai format is verified on glm-5.2, with deepseek, qwen and openai available for other vendors.

## Sources

- [Issues](https://github.com/SanMuzZzZz/LuaN1aoAgent/issues)
- [License: AGPL-3.0](https://github.com/SanMuzZzZz/LuaN1aoAgent/blob/main/LICENSE)
- [README](https://github.com/SanMuzZzZz/LuaN1aoAgent/blob/main/README.md)
- [Releases](https://github.com/SanMuzZzZz/LuaN1aoAgent/releases)
- [SanMuzZzZz/LuaN1aoAgent on GitHub](https://github.com/SanMuzZzZz/LuaN1aoAgent)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sanmuzzzzz-luan1aoagent
