# cain-agent: scope.yaml generated from one target, then mounted on every tool call

> An Apache licensed Python tool that presents itself as an AI penetration testing engineer for authorized assessments, built on the Claude Agent SDK with per-stage model routing. Its safety case rests on a PreToolUse hook rather than on the model behaving, which is the part to scrutinize.

**cdxiaodong/cain-agent** — Real-world AI penetration testing engineer for authorized assessments — built-in cloud module covering AWS/Azure/GCP + Aliyun/Tencent/Huawei clouds. Built on Claude Agent SDK

- Repository: https://github.com/cdxiaodong/cain-agent
- Stars: 961 · Forks: 211
- Language: Python
- License: Apache-2.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/cdxiaodong-cain-agent

## scope.yaml is generated from one target, then enforced per tool call

The stated design principle is that deterministic engineering constrains agent freedom: stage transitions, scope enforcement and dangerous-operation circuit breakers are hard constraints, while path selection and evidence analysis are left to the agent. The scope itself is derived, not typed in by hand. A bootstrap stage turns the target into scope.yaml, and every tool call afterwards is judged against that file.

Enforcement is a PreToolUse hook that blocks any tool call whose target falls outside scope, and the feature list pairs that with read-only by default and credential redaction. The run invocation is short:

```bash
cain-agent run \
  --target https://app.example.com \
  --total-budget 1800
```

The documented flags are --target as required, --workspace for the state directory with a default of ./workspace, --total-budget in wall-clock seconds, and --idle-timeout per step. Those two timeouts are the entire documented stopping mechanism; there is no interactive abort described anywhere in the flags.

One detail deserves attention. The architecture diagram draws the CLI line as cain-agent run --target <t> [--dry-run], so a dry run exists in the box art, but --dry-run does not appear in the flag list. The written flags cover four options and none of them is dry run.

The diagram itself is short. It shows the CLI, the Scope Bootstrap box that maps target to scope.yaml, and then the fence closes, so the layers below the guard are described in prose rather than drawn.

## Per-stage routing cannot downgrade the scope guard

Model routing is per stage, and the documentation is explicit that this does not weaken the guard. Reconnaissance is mostly repetitive enumeration while the test stage needs a high-capability model, so the two discovery stages can each get their own engine and model. The report stage already had its own channel through --pi-validation-provider and --pi-validation-model.

That gives three separately routable stages: recon, test and report. The fallback rules are spelled out: --recon-backend and --test-backend override the engine and fall back to --backend, while --recon-provider, --recon-model, --test-provider and --test-model fall back to --pi-provider and --pi-model and are ignored by the claude backend. The validation channel falls back the same way.

Two guarantees are stated in bold, and they are the reason to read this section at all. Zero default change: with no stage flags the two stages share the same discovery session, behaving exactly as configuring --backend alone would. No scope downgrade: whatever the mix, every execution channel mounts the same ScopeGuardHook, so hard scope enforcement and the finder is not validator dual-session semantics are preserved.

That last guarantee is the load-bearing one. An agent that proposes findings and an agent that accepts them are separate sessions, and swapping engines per stage is explicitly not a way to get an unvalidated channel.

## The pi backend adds a Node toolchain to a Python tool

The default backend is claude, driven by the Claude Agent SDK that is also the core dependency. A second backend named pi is opt-in through --backend pi, and it is not free: it requires Node.js 20 or newer plus a one-time bridge install inside the repository's own toolchain directory.

```bash
npm ci --prefix toolchain/pi
export ANTHROPIC_API_KEY="your-api-key"
cain-agent run --target https://app.example.com --backend pi
```

So a Python 3.11 project now expects a Node runtime and a built toolchain directory before the second engine works at all. Provider selection on that channel runs through --pi-provider and --pi-model, with the API key passed in the environment variable each provider expects: OPENAI_API_KEY, GEMINI_API_KEY, DEEPSEEK_API_KEY or OPENROUTER_API_KEY are the named examples.

There is also a gateway path for anything speaking the Anthropic Messages API. Setting PI_BASE_URL plus ANTHROPIC_AUTH_TOKEN as the bearer credential points the pi channel at a third-party endpoint instead of a vendor. That is a meaningful configuration choice: the bearer token for your gateway goes into the same environment as the scope guard, so a shared or logged environment exposes both.

Further provider and model detail is deferred to the pi bridge guide under toolchain/pi rather than kept in the main README.

## Six clouds are named, and the cloud extra declares two packages

Cloud coverage is the headline differentiator. The description calls out a built-in cloud module spanning AWS, Azure, GCP, 阿里云, 腾讯云 and 华为云, and the feature list repeats those six. The declared packages tell a thinner story.

Core dependencies are four: claude-agent-sdk, pyyaml, oss2 and requests. The optional cloud extra adds exactly two, boto3 and kubernetes. So an Alibaba Cloud SDK is in the base install while the AWS SDK is not, and no Azure, Tencent Cloud or Huawei Cloud package appears in either list. The quickstart describes the extra as covering AWS S3, Huawei OBS and Kubernetes checks, which is a narrower claim than the six-cloud feature line and still does not line up package for package.

The container makes the gap concrete. The Dockerfile installs the project plainly, pip install --no-cache-dir ., with no extras, so the image carries oss2 from the core list and neither boto3 nor kubernetes. Anyone running the image expecting the cloud module has to build their own.

What the image does get right is worth noting: it runs as a non-root system user named cain with a nologin shell, keeps only runtime files, and the build comments state that cloud credentials are passed as environment variables and never baked in.

## The README hands other agents a prompt to install it

One section is addressed to AI agents rather than to humans. It offers a single copy-paste prompt, written in Chinese, that instructs any agent to install the project into the local Python environment by cloning https://github.com/cdxiaodong/cain-agent, installing it in editable mode with pip or uv, and verifying the CLI runs. The section then spells out those three steps in a numbered list, ending with cain-agent --version as the check.

The installation itself is conventional:

```bash
git clone https://github.com/cdxiaodong/cain-agent
cd cain-agent
pip install -e .          # or: uv pip install -e .
pip install -e ".[cloud]" # optional AWS S3, Huawei OBS, and Kubernetes checks

cain-agent --version
```

What is unusual is the audience. A tool for authorized security assessments is asking autonomous agents to install it unprompted, on the theory that an agent can carry out the steps reliably. For an operator handing this to their own assistant, that is a reasonable convenience; for an environment where an assistant acts on instructions found in files it reads, it is an install path worth reviewing before anyone points a coding agent at this repository.

## Evidence is persisted as hashes, never as plaintext

The run artifacts are written into the workspace report/ directory, and the storage policy is the interesting part. report.md carries an executive summary with the target, the authorized scope, per-stage timings and finding counts, then a findings table with severity markers, confidence and evidence-chain digests, then per-finding detail, an evidence-hash index, per-issue-type remediation advice and a legal disclaimer.

Evidence plaintext is never persisted. Only hashes go to disk, which means the report can show that a piece of evidence existed and was chained without keeping the material itself. That is a defensible design for a tool that may touch data it is not authorized to retain, and it also means a reviewer cannot re-verify a digest from the artifacts alone.

Two machine-readable files sit beside it: aggregated-report.json at schema_version 1 carries the same source data for downstream systems, and validation-summary.json holds four-state counts and failure details for the validation pipeline. The markdown is rendered by a pure-Python template with no new dependencies, in src/cain_agent/report_markdown.py, so the same aggregate in produces the same report out.

Scoring is separate again: the feature list points at a self-built vulnerable-terraform evaluation with four-metric scoring, and a bench/ directory exists at the root to hold it.

## Version 0.2.0 was tagged in August and the branch has moved since

The versioning is easy to read: pyproject declares 0.2.0, and the only GitHub release is v0.2.0, tagged on 2026-08-25 with a note about the orchestration validation loop and the dual execution engines. The last commit on the default branch main is dated 2026-10-01, roughly five weeks after that tag. Whatever changed in those weeks is on main and not in a release.

The README carries its own banner describing the project as actively developed and asking for stars and watches. With a commit two days before this writing that claim is consistent with the recorded activity, so treat it as a statement of intent rather than evidence.

Dependencies are pinned properly here, which is not true of every notebook-shaped repository: a uv.lock file sits at the root, and pyproject requires Python 3.11 or newer with setuptools 68 as the build backend. Tooling is ruff at line length 110 selecting E, F, I, UP, B and SIM, pyright in basic mode covering src and tests, and pytest pointed at tests.

The tree also carries CHANGELOG.md, ROADMAP.md and a Chinese README at README.zh-CN.md, plus directories named exploits, skills, tasks, templates, tools, tests and toolchain. The pyright configuration explains one subtlety in a comment: bench is kept out of the include list even though tests import modules from it through sys.path.insert, so extraPaths carries bench to keep static analysis aligned with runtime.

## Conclusion

cain-agent fits an authorized engagement where a client has agreed in writing to a target list, an operator can supply a wall-clock budget, and the deliverable is an evidence chain rather than a flag. It does not fit anyone who wants an autonomous agent pointed at a range of hosts, because the entire safety argument reduces to one hook and one scope file. Verify four things before a run. Read scope.yaml yourself after bootstrap, since every later tool call is judged against it and nothing else. Confirm which engine each stage uses, because the pi backend adds a Node 20 toolchain and a gateway token path alongside the default claude backend, and the per-stage flags fall back to the global ones in ways that are easy to misread. Set --total-budget and --idle-timeout on the first run, since those timeouts are the only documented stop. And check the cloud dependency story, because the six clouds named in the feature list are not matched by the declared packages, and the Docker image installs the project without the cloud extra. Version 0.2.0 was tagged on 2026-08-25 and the branch has moved since, with the last commit on 2026-10-01.

## FAQ

### What is cain-agent?

An Apache 2.0 licensed Python project that describes itself as an AI penetration testing engineer for authorized security assessments, built on the Claude Agent SDK. It walks a deterministic attack pipeline and treats scope as an engineering constraint rather than something the model is asked to respect.

### How do I install cain-agent?

Clone the repository, then run pip install -e . or uv pip install -e ., with pip install -e ".[cloud]" for the optional cloud checks, and confirm with cain-agent --version. Python 3.11 or newer is required, and the Dockerfile builds a python:3.11-slim image whose entrypoint is the CLI running as a non-root user.

### Which clouds does the cain-agent cloud module cover?

The feature list names AWS, Azure, GCP, 阿里云, 腾讯云 and 华为云. The declared packages are thinner: the cloud extra adds boto3 and kubernetes, oss2 sits in the core dependencies, and no Azure or Huawei Cloud SDK appears in either list.

### How does cain-agent keep an agent inside its scope?

A bootstrap stage turns the target into scope.yaml, and a PreToolUse hook blocks any tool call whose target falls outside scope. Runs are read-only by default, credential redaction is on, and every execution channel, including the per-stage pi channels, mounts the same ScopeGuardHook.

### Does cain-agent have a stop switch or a dry run?

The documented bounds are --total-budget in wall-clock seconds, shown as 1800 in the documented example, and --idle-timeout per step, plus --workspace for the state directory defaulting to ./workspace. A [--dry-run] placeholder appears in the architecture diagram but is absent from the flag list, and no interactive abort is described.

## Sources

- [cdxiaodong/cain-agent on GitHub](https://github.com/cdxiaodong/cain-agent)
- [Issues](https://github.com/cdxiaodong/cain-agent/issues)
- [License: Apache-2.0](https://github.com/cdxiaodong/cain-agent/blob/main/LICENSE)
- [README](https://github.com/cdxiaodong/cain-agent/blob/main/README.md)
- [Releases](https://github.com/cdxiaodong/cain-agent/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/cdxiaodong-cain-agent
