# Strix: an autonomous AI pentesting agent you run from the CLI

> Strix is an Apache-2.0 Python tool that drives LLM-backed agents against a target directory or application to find and validate vulnerabilities. It installs with a shell script, needs Docker and an LLM API key, and its own metadata labels it alpha.

**usestrix/strix** — Autonomous AI penetration-testing agents that run your code, find vulnerabilities, and verify them with real proofs of concept, plugging into CI/CD pipelines.

- Repository: https://github.com/usestrix/strix
- Website: https://strix.ai
- Stars: 65,555 · Forks: 7,193
- Language: Python
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/usestrix-strix

## What Strix actually automates

The README frames Strix as autonomous AI penetration testing agents that "run your code dynamically, find vulnerabilities, and validate them through actual proofs-of-concept". The stated audience is developers and security teams who want faster testing than manual pentesting and fewer false positives than static analysis. That framing matters, because the two alternatives it positions against have different costs. Manual pentesting is slow and expensive but produces human-verified findings. Static analysers are fast and cheap but report patterns, not exploits.

Strix tries to occupy the middle: an agent that executes the target, attempts exploitation, and only reports what it could demonstrate. The README lists use cases including application security testing, rapid penetration testing, bug bounty automation, and CI/CD integration. The last of those is the most operationally interesting, since it implies the tool is expected to run unattended on every pull request rather than as a one-off engagement.

## Multi-agent orchestration and the Docker sandbox

The mechanism described in the README is a graph of agents. Specialised agents handle reconnaissance, exploitation, and post-exploitation, and the README says they collaborate and scale. The agents are given an offensive toolkit rather than a single prompt: an HTTP interception proxy built on Caido, an automated browser for XSS and CSRF flows, a shell for exploit development, a Python sandbox for writing proofs-of-concept, and reconnaissance tooling for attack surface mapping. Findings are structured with CVSS scoring and OWASP classification.

The dependency list in pyproject.toml confirms the shape of this. The package depends on openai-agents with the litellm extra, the docker SDK, caido-sdk-client, cvss, and reportlab for PDF generation. So the runtime is a Python agent loop, the model access is routed through LiteLLM, and the target work happens inside a Docker sandbox. The README notes that the first run automatically pulls the sandbox Docker image, which is the point where a locked-down network will fail.

## Installing Strix and running a first scan

The README gives a three-step quick start. The prerequisites are a running Docker daemon and an LLM API key from a supported provider; the docs site lists OpenAI, Anthropic, and Google among them. The install script is served from the project's own domain, so you are piping a remote script into bash, which is worth noting for anyone who reviews installers before running them.

```bash
curl -sSL https://strix.ai/install | bash
```

Next, point Strix at a model and a key through environment variables. The README uses these exact names, and the model string follows a provider/model format.

```bash
export STRIX_LLM="openai/gpt-5.4"
export LLM_API_KEY="your-api-key"
```

Then run the assessment against a directory. The README's example targets a local app folder rather than a URL.

```bash
strix --target ./app-directory
```

According to the README, the first run pulls the sandbox image, and results land in strix_runs/<run-name>. That directory is what you read afterwards: the findings, and in the platform version, reproduction steps. There is also a second entry point for coding agents. Running the npx command below installs nine skills, including penetration-testing-with-strix and fix-security-vulnerabilities-with-strix, for SKILL.md-compatible agents such as Claude Code, Cursor, and Codex.

```bash
npx skills add usestrix/strix
```

For contributors, the Makefile provides uv-based targets: make install runs uv sync --no-dev, make dev-install runs uv sync, and make setup-dev adds pre-commit hooks. The package requires Python 3.12 or newer.

## Where the alpha label and the LLM dependency bite

Two constraints stand out. The first is in pyproject.toml: the classifier reads "Development Status :: 3 - Alpha". The project's own metadata does not claim production stability, even though the README markets CI/CD integration. Those two statements are not contradictory, but they set expectations. A tool that blocks pull requests on findings should be judged on its false positive rate, and an alpha project is exactly where that rate is still moving.

The second constraint is the LLM. Every scan sends code, traffic, or both through a third-party model provider, and the API key is a hard prerequisite. For regulated codebases, or for anyone whose threat model includes the model vendor, that is a non-starter regardless of how good the findings are. Cost is also per-token and unbounded by the tool itself. The README does not document a spending cap, a token budget, or a maximum scan duration, so a large target with a chatty agent loop is a budget risk you have to manage outside Strix.

The dependency pins show the team is aware of packaging fragility. The cryptography pin carries a comment explaining that 49.x drops the universal2 macOS wheel and breaks the Intel macOS release build. That is a sign of active release engineering, and also a sign that the supported platform matrix is narrow enough to need hand-tuning.

## Strix against a conventional DAST scanner

The obvious alternative is a classic dynamic scanner such as OWASP ZAP or Burp Suite's scanner. The difference is in what drives the test. A conventional scanner walks a fixed catalogue of checks: it sends payloads, matches responses against signatures, and reports anything that looks anomalous. Coverage is predictable and the run is deterministic and cheap. It will not chain a leaked token into a privilege escalation, because no rule describes that chain.

Strix inverts this. The agent decides what to try next based on what it has already seen, which is how it can produce a working proof-of-concept for a logic flaw. The trade-off is the opposite of a scanner's: coverage is not guaranteed, the same target can produce different results across runs, and each run costs model tokens. If your requirement is a repeatable regression gate with a known check list, a conventional scanner is the better instrument. If your requirement is finding a chained exploit that no signature describes, the agentic approach is the one that can plausibly get there.

## Licence, releases and the cost of keeping up

Strix is Apache-2.0, which permits commercial use, modification, and redistribution provided the licence and notices are preserved. That is a permissive choice, and it means you can embed the CLI in an internal pipeline without a licensing conversation. It also means the project carries no copyleft obligation back to you. This is a description of the licence text, not legal advice; have counsel review anything you ship.

The release cadence visible in the repository is fast. Three releases landed within four days in August 2026, and pyproject.toml declares version 1.6.2 while the most recent tagged release listed is v1.5.3, so the packaging file runs ahead of the tags. The last push to the default branch was on 2026-08-10. Fast iteration on an alpha tool means upgrade cost is real: the openai-agents dependency is pinned below 0.20, LiteLLM moves quickly, and any of those bumps can change agent behaviour rather than just fixing a bug. Pin the version you validate against, and re-run your own target after each bump rather than trusting the changelog.

## Conclusion

Strix fits teams that already run containerised workloads and want an LLM-driven pass over an application before a release, with findings written to strix_runs/<run-name>. It does not fit anyone who cannot send source code or traffic to a third-party LLM provider, or who needs a stable, versioned tool rather than an alpha. Before adopting it, confirm the Docker sandbox image pulls cleanly in your network, check which LLM providers docs.strix.ai lists for your region, and read the strix_runs output on a throwaway target to see how it reports proof-of-concept evidence.

## FAQ

### How to install Strix?

The README's quick start runs a shell installer served from strix.ai, then exports STRIX_LLM and LLM_API_KEY before the first scan. Docker must be running, and the first run pulls the sandbox image automatically.

### How to use Strix?

After installation, the README shows a single command, strix --target ./app-directory, which starts an assessment against that path. Results are saved under strix_runs/<run-name> for you to read afterwards.

### What is Strix?

In this repository, Strix is an open-source AI penetration testing tool: autonomous agents that run a target dynamically, find vulnerabilities, and validate them with proofs-of-concept. It is written in Python and distributed under Apache-2.0.

### What does Strix mean?

The repository does not explain the origin of the name, and its documentation covers the tool rather than the word. Anything about the Latin or Greek senses of strix is outside what this project's material states.

## Sources

- [Official documentation](https://strix.ai)
- [Official README](https://github.com/usestrix/strix#readme)
- [Project repository](https://github.com/usestrix/strix)
- [Release notes](https://github.com/usestrix/strix/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/usestrix-strix
