# PromptInject is a 2022 research framework that still points at GPT-3

> PromptInject assembles prompts modularly to measure how a language model reacts to adversarial input, and it does so for the two attack classes named in its paper: goal hijacking and prompt leaking. Read it as a research artefact, because the framework, its dependency floor and its target model all date from 2022.

**agencyenterprise/PromptInject** — PromptInject is a framework that assembles prompts in a modular fashion to provide a quantitative analysis of the robustness of LLMs to adversarial prompt attacks. 🏆 Best Paper Awards @ NeurIPS ML Safety Workshop 2022

- Repository: https://github.com/agencyenterprise/PromptInject
- Stars: 526 · Forks: 57
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/agencyenterprise-promptinject

## It is a prompt composer, not an attack library

PromptInject does not ship a catalogue of exploits. What it provides is a way of assembling prompts modularly so that a model's resistance to adversarial input can be measured quantitatively rather than argued about. The paper behind it calls the approach a prosaic alignment framework for mask-based iterative adversarial prompt composition, which is a dense way of saying that a template prompt with a placeholder is filled with candidate inputs and the model's answers are scored.

That framing is the point. A tool that tries a handful of handcrafted attacks and reports whether they worked tells you very little, because a single success or failure is indistinguishable from sampling noise. A composer plus a scoring step produces numbers you can compare across prompts and across runs. The repository layout matches: a promptinject package, a tests directory, notebooks, and nothing resembling an attack payload collection.

## Two attack classes that differ only in what the attacker wants printed

The paper separates adversarial input into goal hijacking and prompt leaking, and the distinction is entirely about the attacker's objective. In goal hijacking the substituted input redirects the model to print a specific target string, and the paper is careful to add that this target string may itself contain malicious instructions, so the attack can be used as a delivery mechanism rather than only as a nuisance. In prompt leaking the objective is to make the model print the application prompt itself.

Figure 1 of the paper lays this out as a diagram: the application prompt is a box with a user input placeholder, ordinary input arrives in one style, and each attack arrives in another, with the resulting model output beside it. Both attacks share the same mechanism of substituting for the placeholder. Only the target differs, which is why one framework can study both without carrying two sets of machinery.

## The threat model is an ordinary user with bad intentions, not a researcher

The framing of the original paper deserves attention because it sets expectations for what the framework is measuring. It argues that studies of vulnerabilities emerging from malicious user interaction were scarce, and that even low-aptitude but sufficiently ill-intentioned agents can exploit the stochastic nature of a deployed model to create long-tail risks. The target named in the abstract is GPT-3, described as the most widely deployed language model in production at the time.

So this is a claim about a population of attackers who are not trying to break the system and mostly will not. That matters for how you use the framework today. A quantitative score derived from handcrafted inputs measures a model's resistance to that population, not to an adversary who has read your system prompt and has time, and the two are very different numbers.

## The only usage documentation is a notebook

The usage section consists of one sentence pointing at notebooks/Example.ipynb. There is no API reference, no command-line interface, no documentation directory, and no changelog in the top-level entries. What the repository contains is the promptinject package, the tests, the notebook, an images directory for the paper's figures, a poetry lock file, the manifest, and the usual governance files.

That shape is normal for a paper artefact and it sets the expectation correctly: you are expected to read a notebook to learn the interface. It also means there is no upgrade path to follow. Nothing in the repository describes what changed between versions, because the manifest has sat at 0.1.0 and there are no tagged releases to compare.

## Installed from git, against a pre-1.0 OpenAI client

There is no package on PyPI to depend on. Installation is a single git reference:

```bash
pip install git+https://github.com/agencyenterprise/PromptInject
```

The dependency floor explains why that matters. The manifest requires Python ^3.9 and pins openai to ^0.25.0, which in Poetry means a range below 0.26.0, on a library that has since moved well past its 1.0 line. Alongside it sit rapidfuzz ^2.13.2 for fast approximate string matching, tqdm ^4.64.1 and pandas ^1.5.1. Because the caret range on a 0.x version only advances the minor component, the install resolves to a single narrow band of that old client rather than tracking later work.

## Dev tooling is declared in Poetry's older section and built for notebooks

The development dependencies sit under [tool.poetry.dev-dependencies] rather than a Poetry dependency group, which is the older spelling of that configuration. The tools listed are black with its jupyter extra at ^22.10.0, jupyterlab ^3.5.0, jupyterlab-code-formatter ^1.5.3 and isort ^5.10.1.

Every one of those is a notebook or formatting tool, and none is a testing, type-checking or coverage tool. The test suite runs without a dedicated test dependency declared at all. This is consistent with the rest of the repository: a small framework, verified by a notebook, built for a paper demonstration rather than maintained as a library with a compatibility surface.

## The version that matters is the paper's, published in 2022

PromptInject exists to support Ignore Previous Prompt: Attack Techniques For Language Models, by Fabio Perez and Ian Ribeiro, published on arXiv in 2022, and the repository description records a best paper award at the NeurIPS Machine Learning Safety Workshop that same year. The framework's own version is 0.1.0. The last push was on 2026-04-27, and the repository has no GitHub releases at all.

So the currency question is settled by the subject matter rather than by the commit history. The model analysed in the abstract is GPT-3, and the client library is from the same era. The ideas transfer to current models; the measurements do not transfer without redoing them against whatever API you actually intend to ship.

## Conclusion

PromptInject earns a place on a reading list rather than in a dependency list. Its two contributions are still the useful ones: a vocabulary that separates changing what a model does from making it disclose what it was told, and an insistence on measuring that quantitatively rather than by anecdote. What the repository cannot give you is currency. The client library is pinned to a pre-1.0 OpenAI SDK, the manifest version is 0.1.0, there are no published releases to pin against, and the model under discussion is GPT-3. If you want to reproduce the original measurements, take the paper and pin your own environment deliberately, because the framework will not do it for you.

## FAQ

### What is PromptInject?

A Python framework that assembles prompts in a modular fashion to provide a quantitative analysis of how far a language model resists adversarial prompt attacks. It accompanies the paper Ignore Previous Prompt: Attack Techniques For Language Models.

### Which two attacks does PromptInject study?

Goal hijacking, where the substituted input redirects the model to print a specific target string that may itself contain malicious instructions, and prompt leaking, where the objective is to make the model print the application prompt.

### How do I install PromptInject?

From the git repository, with pip install git+https://github.com/agencyenterprise/PromptInject. It is not on PyPI, the repository has no GitHub releases, and the manifest version is 0.1.0.

### What does PromptInject need to run?

Python 3.9 or newer, with openai ^0.25.0, rapidfuzz, tqdm and pandas declared as dependencies. The development tools are black, jupyterlab, jupyterlab-code-formatter and isort.

### Is PromptInject still being worked on?

There is no published release and the last push was on 2026-04-27. The project supports a 2022 paper and its framework targets GPT-3, so it is best read as a research artefact rather than a current library.

## Sources

- [agencyenterprise/PromptInject on GitHub](https://github.com/agencyenterprise/PromptInject)
- [Issues](https://github.com/agencyenterprise/PromptInject/issues)
- [License: MIT](https://github.com/agencyenterprise/PromptInject/blob/main/LICENSE)
- [README](https://github.com/agencyenterprise/PromptInject/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/agencyenterprise-promptinject
