# Prompt Fuzzer (ps-fuzz): Testing and Hardening Your System Prompt

> Prompt Fuzzer is an MIT-licensed Python tool that attacks a GenAI application's system prompt with 16 LLM-driven attack categories and scores the result. It is useful before you ship a prompt, and it is not a runtime guardrail.

**prompt-security/ps-fuzz** — Make your GenAI Apps Safe & Secure :rocket: Test & harden your system prompt

- Repository: https://github.com/prompt-security/ps-fuzz
- Website: https://www.prompt.security/fuzzer
- Stars: 715 · Forks: 104
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/prompt-security-ps-fuzz

## What Prompt Fuzzer actually assesses

The project's own description is narrow: it assesses the security of a GenAI application's system prompt against dynamic LLM-based attacks, then gives you a security evaluation based on how those attack simulations ended. The unit under test is the prompt, not the application. That distinction matters more than it first appears. A model can be perfectly aligned while your system prompt leaks its instructions, and a system prompt can be tightly written while the surrounding application happily forwards user text into a tool call.

Prompt Fuzzer sits on the first of those problems. You paste in a system prompt, the tool generates attacks tailored to your application's configuration and domain, and you get back a picture of which attack families got through. The README also describes a Playground chat interface for iterating on the prompt, so the intended loop is attack, edit, re-attack. The audience is developers and QA engineers who write prompts and want evidence that the current revision holds up, rather than a one-time audit report.

The README carries a warning worth reading twice: using the Fuzzer consumes tokens. Every attack is a real model call against a real provider, and the multi-threaded mode multiplies that. This is a paid activity dressed as a test suite.

## The attack model: LLM-generated probes against your endpoint

There are two model roles in a run. The attack provider and attack model generate the adversarial prompts. The target provider and target model receive them and produce the responses that get judged. Both are configurable from the command line, which means you can run a cheaper or open model as the attacker while pointing the target at the deployment you actually care about. That split is the core architectural decision in the tool, and it is a sensible one.

The README lists 16 supported attack categories, grouped under jailbreak, prompt injection, RAG and vector database attacks, and system prompt extraction. RAG tests are separated out in the options list: --embedding-provider and --embedding-model are described as required for RAG tests, with separate base URL flags for Ollama and OpenAI embeddings. That tells you the RAG path needs its own configuration and will not run from the basic setup.

Provider coverage is broad. The README documents 16 providers through environment variables, including ANTHROPIC_API_KEY, AZURE OPENAI_API_KEY, COHERE_API_KEY, GOOGLE_API_KEY, QIANFAN_AK and QIANFAN_SK, and YC_API_KEY. For self-hosted work there are --ollama-base-url and --openai-base-url flags, so an OpenAI-compatible endpoint is a supported target. Attack data ships inside the package: pyproject.toml declares ps_fuzz = ["attack_data/*"] as package data, so the probe corpus travels with the wheel.

## Installing Prompt Fuzzer and running a first batch test

The README requires Python 3.10 or newer, and pyproject.toml sets requires-python = ">=3.10". Install from PyPI under the distribution name prompt-security-fuzzer, which is not the same as the repository name. The console script it exposes is prompt-security-fuzzer.

```bash
pip install prompt-security-fuzzer
```

Before launching, the tool needs a provider key. The README says the default is OPENAI_API_KEY, and that you can either export it or place it in a .env file in the current directory. python-dotenv is a declared dependency, which is consistent with the .env route.

```bash
export OPENAI_API_KEY=sk-123XXXXXXXXXXXX
prompt-security-fuzzer
```

Launched with no arguments, the tool is interactive: it asks for your system prompt and then starts testing. For unattended runs the README documents the -batch or -b flag, which bypasses the interactive steps. Combined with the provider and model flags, that is the shape a CI job would take.

```bash
prompt-security-fuzzer -b \
  --attack-provider openai \
  --attack-model gpt-4o \
  --target-provider openai \
  --target-model gpt-4o \
  -n 10 -t 4
```

The -n flag sets the number of different attack prompts and -t sets worker threads. Two other flags are worth knowing before you run anything: --list-providers and --list-attacks print the available options and exit, which is the fastest way to confirm that the provider name you intend to pass is one the tool recognises. The README does not spell out the exact provider identifier strings, so check with --list-providers rather than guessing from the environment variable names.

What you should see is a scored evaluation of the attack outcomes, plus a Playground chat interface you can use to iterate on the prompt. The README points to system_prompt.examples in the repository for prompts of varying strength, which is the right place to look if you want a baseline to compare your own prompt against.

## Where Prompt Fuzzer stops being the right tool

The most important limitation is scope. Prompt Fuzzer evaluates a system prompt. It does not sit in the request path, inspect live traffic, or block anything. If your actual problem is a user jailbreaking your production assistant at 3am, this tool will not help you at that moment; it helps you find the weakness beforehand and rewrite the prompt. Teams sometimes install a fuzzer expecting a firewall and are disappointed by a report.

The second limitation is cost and variance. Every run spends tokens on both the attack model and the target model, and the README flags this explicitly. Attack generation is itself a model call, so results are not perfectly deterministic across runs even with the same prompt. The -a or --attack-temperature flag exists precisely because that variance is real. A single green run is weaker evidence than it looks.

The third is configuration surface. RAG tests need an embedding provider and embedding model, and the README notes those are required for RAG tests. If you do not have an embedding endpoint available, that whole attack category is out of reach. The README's Docker section is marked coming soon, so container-based installation is not documented yet; the pip path and the release wheel are what the README actually describes. There is no documented rollback or prompt-versioning mechanism, so tracking which prompt revision produced which result is on you.

## How it differs from promptfoo and static prompt review

promptfoo is the closest well-known alternative and appears in the searches around this project. The difference in approach is about where the attack corpus comes from. Prompt Fuzzer generates its probes with an LLM at run time and tailors them to your application's configuration and domain, which is what the README means by dynamic attacks. promptfoo's model is centred on declarative test configuration: you write the cases and assertions, and the tool executes them. That gives you reproducibility and a clear diff when a test flips, at the cost of only testing what you thought to write down.

Neither is strictly better. A generated corpus finds things you would not have imagined; a declarative suite tells you precisely which named case regressed. If you already run promptfoo, adding Prompt Fuzzer is not redundant, because the two produce different kinds of evidence. If you have neither and you only want a fixed regression suite in CI, promptfoo's determinism is the easier property to build on.

The other alternative is plain manual review, reading the prompt and reasoning about it. That is free and catches obvious instruction-leakage phrasing. It will not tell you whether a Baichuan, Cohere or Yandex model routes around your instructions differently, which is the kind of thing a multi-provider fuzzer is built to surface.

## Licence, maintenance and the cost of keeping up

The licence is MIT, declared in both the README badge and pyproject.toml, and the LICENSE file sits at the repository root. MIT is permissive: you can use, modify and redistribute the tool, including inside commercial work, provided the copyright notice and licence text are preserved. That is a description of the licence terms, not legal advice; if the tool ends up embedded in a product you ship, have counsel read the actual LICENSE file rather than a summary of it.

The repository is not archived, and the last push was on 2026-08-19, with release 2.1.1 tagged on 2026-08-18. The dependency set is the real maintenance surface. pyproject.toml pins or bounds a large stack: langchain, langchain-community, langchain-core, langchain-openai, langchain-ollama, chromadb, langchain-chroma, openai, fastparquet, pandas, tiktoken and others. LangChain's release cadence means an unpinned environment will drift, and the langchain-community bound at <0.5.0 in particular is the kind of range that eventually forces an upgrade. Installing into a virtual environment is the practical mitigation.

Upgrading the tool itself is a pip install away, but upgrading it also upgrades that stack. Read CHANGELOG.md before bumping, since the jump from 2.0.0 to 2.1.x involved a version scheme change as well as a release.

## Conclusion

Adopt Prompt Fuzzer if you own a system prompt and want a repeatable pre-release check against jailbreaks, prompt injection, system prompt extraction and RAG poisoning. Skip it if you need inline request filtering, since the README describes an assessment tool with a Playground, not a proxy. Verify first that Python 3.10 or newer is available, that you have budget for the token consumption the README warns about, and that your target endpoint accepts the --target-provider and --target-model you intend to point it at.

## FAQ

### What is Prompt Fuzzer (ps-fuzz) and who is it for?

It is an MIT-licensed Python tool that assesses a GenAI application's system prompt against dynamic LLM-based attacks and returns a security evaluation. The README targets developers and QA engineers who write system prompts and want to harden them.

### How do I install Prompt Fuzzer?

Install the PyPI distribution named prompt-security-fuzzer with pip; it requires Python 3.10 or newer. The README also points to the PyPI package page and to wheel files attached to GitHub releases.

### Does using Prompt Fuzzer cost money?

Yes. The README carries an explicit warning that using the Fuzzer leads to the consumption of tokens, because both the attack model and the target model are called during a run.

### Which LLM providers does Prompt Fuzzer support?

The README documents 16 providers configured through environment variables, including OpenAI, Anthropic, Azure OpenAI, Cohere, Google, Baichuan, Qianfan and YandexGPT, plus Ollama and OpenAI-compatible base URLs for self-hosted endpoints.

### What attack categories does Prompt Fuzzer cover?

The README groups the 16 supported attacks under jailbreak, prompt injection, RAG and vector database attacks, and system prompt extraction. RAG tests additionally require an embedding provider and embedding model.

### Can Prompt Fuzzer run without the interactive prompts?

Yes. The README documents the -batch or -b flag, which runs the fuzzer in unattended mode and bypasses the interactive steps, alongside flags for the attack and target providers and models.

## Sources

- [License: MIT](https://github.com/prompt-security/ps-fuzz/blob/main/LICENSE)
- [Project website](https://www.prompt.security/fuzzer)
- [prompt-security/ps-fuzz on GitHub](https://github.com/prompt-security/ps-fuzz)
- [README](https://github.com/prompt-security/ps-fuzz/blob/main/README.md)
- [Releases](https://github.com/prompt-security/ps-fuzz/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/prompt-security-ps-fuzz
