# AgentDojo's headline detector needs a second install and its example pins a 2024 model

> A benchmark for evaluating prompt injection attacks and defenses on LLM agents, from ETH Zurich and Invariant Labs, installed with one pip command and run through a script with four independent axes. The prompt injection detector sits behind an optional extra, the dependencies carry lower bounds and no ceilings, and the package still warns that its API may change.

**ethz-spylab/agentdojo** — A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents.

- Repository: https://github.com/ethz-spylab/agentdojo
- Website: https://agentdojo.spylab.ai/
- Stars: 891 · Forks: 234
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ethz-spylab-agentdojo

## A 0.1.x package whose own page warns the API may change

The quickstart is one command, and the next thing on the page is a warning:

```bash
pip install agentdojo
```

The note attached to it says the API of the package is still under development and might change in the future. That is consistent with the version scheme. The package declares 0.1.35 and is classified as Development Status 4 - Beta, with classifiers naming Python 3.10, 3.11 and 3.12 against a requires-python floor of 3.10.

The release history is where the warning earns its keep. The three most recent tags are v0.1.33 on 2025-05-14, v0.1.34 on 2025-06-02 and v0.1.35 on 2025-10-27. The version in the packaging metadata matches the newest tag exactly, so the numbers are consistent, but the last push to the repository is dated 2026-06-02, roughly eight months after that tag. The tree you would clone today is not the tree that produced v0.1.35.

So there are two kinds of stability to keep apart. The version number and the tag agree, which means nothing has been force-renumbered. What is not pinned is the relationship between the branch and the release, and a reader installing from the index gets the October 2025 build while a reader cloning gets eight months of later work under the same-looking version.

## The GPU classifier sits on a package that only makes API calls

One line in the project classifiers reads `Environment :: GPU`. It is the only environment declaration in the metadata.

Now look at what the package actually depends on for its work: the OpenAI SDK, the Anthropic SDK, the Cohere SDK, the Google generative AI SDK, and langchain. Every one of those reaches a hosted model over a network call. Nothing in the dependency list computes locally, and there is no numerical, machine learning or scientific package among the fourteen entries. The library runs attacks and defenses by prompting models you pay someone else to run.

So the GPU declaration describes neither a hardware requirement nor an acceleration path. It is a metadata field that has drifted from the package's nature, and it is the kind of line that shows up in package indexes and deployment filters, where it can mislead someone deciding whether a host needs a GPU before reading the dependency list.

The other classifiers are more accurate about intent: Topic Scientific/Engineering :: Artificial Intelligence, Topic Security, and Intended Audience for both Developers and Science/Research. Security is classified as a topic rather than as a use case, which is the neutral way to describe a benchmark whose entire purpose is attacking agents.

## The prompt injection detector is not in the default install

The name of the project contains the attack class it studies, and the most on-topic component is gated behind an extra. The page states that if you want to use the prompt injection detector you need to install the `transformers` extra, and gives the command:

```bash
pip install "agentdojo[transformers]"
```

The default `pip install agentdojo` does not bring it. That matters for anyone trying to reproduce the defensive side of the benchmark, because a detector that runs as a local classifier and one that calls a model are different experiments, and the page does not say which of the two the published results used.

The install section is otherwise two commands long, so this is the only capability split the front page mentions. There is no indication of what the extra costs to install, since `transformers` is a large dependency for what is most likely a single classification model, and no note on whether the detector needs a GPU of its own to run locally.

Read alongside the GPU classifier, the picture is coherent: the benchmark harness itself is API-driven and light, and anything that runs a model locally is optional.

## Fourteen dependencies with lower bounds and no ceilings

The dependency list is fourteen entries long and every one of them is a lower bound with no upper cap: `openai>=1.59.7`, `anthropic>=0.47.0`, `cohere>=5.3.4`, `google-genai>=1.15.0`, `langchain>=0.1.17`, `pydantic[email]>=2.7.1`, `docstring-parser>=0.15`, `tenacity>=8.2.3`, `typing-extensions>=4.11.0`, `pyyaml>=6.0.1`, `click>=8.1.7`, `deepdiff>=8.6.1`, `rich>=13.7.1` and `python-dotenv>=1.0.1`.

Four of those are model provider SDKs and a fifth is an agent framework. That is the dependency profile of a benchmark whose subject is agent behavior: it must be able to drive whatever agent is under test, so it reaches for the major frameworks rather than pinning one. The cost is that a major release from any provider, or from langchain, can change the tool-calling behavior the benchmark depends on without any version of AgentDojo changing.

The rest of the list is more telling about the design. `docstring-parser` suggests tool definitions are carried in docstrings rather than in hand-built schemas, `click` provides the command line, `tenacity` handles retries against rate-limited APIs, `rich` renders terminal output, and `deepdiff` compares structures, which fits a harness whose job is diffing an agent's behavior with and without a defense. `python-dotenv` is where the provider keys come from.

The repository also carries a `uv.lock`, so a development checkout resolves to exact versions even though the published metadata does not.

## Four independent axes, and the documented run is pinned to a 2024 snapshot

The benchmark script takes four things that can each vary, and the documented example sets all four explicitly:

```bash
python -m agentdojo.scripts.benchmark -s workspace -ut user_task_0 \
    -ut user_task_1 --model gpt-4o-2024-05-13 \
    --defense tool_filter --attack tool_knowledge
```

A suite (`-s`), one or more user tasks (`-ut`, repeatable), a model, a defense and an attack. The suite is the `workspace` suite, the defense is the tool filter and the attack is the one with tool knowledge. Because the axes are independent, one command fixes a single cell of a matrix, and the page's own guidance for usage is the script's `--help` flag rather than a written reference.

The model is the detail worth pausing on. `gpt-4o-2024-05-13` is a dated snapshot, not an alias, so the example is reproducible in a way that `gpt-4o` would not be. That is the right choice for a benchmark, and it also means anyone re-running the published numbers needs that exact snapshot to still exist.

The second command drops the task selection and runs all suites and tasks:

```bash
python -m agentdojo.scripts.benchmark --model gpt-4o-2024-05-13 \
    --defense tool_filter --attack tool_knowledge
```

So the full run has no cap on how many tasks execute. Since every task is an agent trajectory against a hosted model, the cost of that second command scales with suites times tasks times turns, and none of those counts is stated on the page.

## Six authors on the paper, four in the package metadata

The paper has six authors: Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner, Marc Fischer and Florian Tramèr, split between ETH Zurich and Invariant Labs, with Balunović, Beurer-Kellner and Fischer carrying both affiliations.

The packaging metadata names four of them as authors, Debenedetti, Zhang, Balunovic and Beurer-Kellner, with Debenedetti as the sole maintainer. Marc Fischer and Florian Tramèr are absent from that list.

Neither version is wrong. Package metadata conventionally lists contacts rather than every contributor, and a benchmark package maintained by the first author is normal. But the two lists disagree, which matters in practice: automated citation tools and dependency metadata harvest the packaging authors, so a paper built on this package would credit four people where the paper credits six.

The citation situation is otherwise well handled, which is more than most repositories in this class manage. A `CITATION.bib` sits in the root, the page asks users to cite the work, and the venue is given as the Thirty-eighth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, with the OpenReview forum identifier included and an arXiv link alongside. This one went through a peer-reviewed track rather than living as a preprint.

One affiliation detail is worth naming because it touches where results live. Results are published on the project's own results page and are also listed in the Invariant Benchmark Registry, which is operated by Invariant Labs, the same organization two of the authors are affiliated with.

## Benchmark runs are committed to the repository, and the front page does not mention it

The repository root holds `runs/`, alongside `notebooks/`, `tests/`, `util_scripts/`, `examples/`, `docs/` and `src/`. A committed runs directory means benchmark output lives in version control, not only on the results site, which is the more useful artifact for anyone checking a number.

The examples directory is worth a separate look. Alongside a readme and two files named `attack.py` and `pipeline.py`, there is `functions_runtime.py` and a `counter_benchmark/` directory. A counter-benchmark is an attack of the benchmark itself, which is the only part of this repository that lets you test whether the evaluation is measuring what it claims to. Nothing on the front page points at any of it; the page links the paper, the results page and the documentation instead.

The development tooling around it is conventional and modern. There is a `uv.lock` for reproducible resolution, a `.python-version` file pinning the interpreter, a `.devcontainer/` definition for a containerised environment, a `.pre-commit-config.yaml`, and `mkdocs.yml`, which is what publishes the documentation site at agentdojo.spylab.ai. There is also a `CLAUDE.md` at the root, an instruction file for coding agents, which tells you something about how the project expects to be worked on.

So the repository ships far more than the 214-word front page advertises: committed run data, a counter-benchmark, and a pinned dev environment.

## Conclusion

AgentDojo earns a place if you need a published, citable benchmark for prompt injection rather than a suite you build yourself, since the results are on a public site and the paper went through a peer-reviewed track. Three things to check before you rely on a number. Reproducibility rests on the model snapshot you pass, and the documented example uses a 2024-05-13 snapshot, so pin deliberately rather than accepting whatever alias a provider currently resolves. The dependencies carry lower bounds and no upper bounds across four model provider SDKs plus langchain, which means a provider release can change what the benchmark is measuring. And the prompt injection detector, which sounds central to the name, is not in the default install: it needs the transformers extra. For anyone planning to extend the suite, the committed `runs/` directory and the counter-benchmark examples are the most useful part of the repository, and they are not mentioned on the front page.

## FAQ

### what is agentdojo

AgentDojo is a Python package for evaluating prompt injection attacks and defenses for LLM agents, built at ETH Zurich and Invariant Labs and installed with `pip install agentdojo`. Its paper is AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents, available at arxiv 2406.13352.

### how to use agentdojo

Install with `pip install agentdojo`, then run `python -m agentdojo.scripts.benchmark` with a suite, one or more user tasks, a model, a defense and an attack. The script's own usage comes from its `--help` flag, and the prompt injection detector needs the separate `agentdojo[transformers]` extra.

### Which model providers does agentdojo depend on?

Four provider SDKs are hard dependencies, openai, anthropic, cohere and google-genai, alongside langchain. None carries an upper version bound, so a major release from a provider or from langchain can change the agent behavior the benchmark depends on without AgentDojo itself changing version.

### Does the agentdojo example pin a specific model version?

Yes. The documented run passes `gpt-4o-2024-05-13`, a dated snapshot rather than an alias, together with the tool_filter defense and the tool_knowledge attack. The package version is 0.1.35 and its most recent tag is dated 2025-10-27.

### Where are agentdojo results published?

On the documentation site's dedicated results page, and also in the Invariant Benchmark Registry at explorer.invariantlabs.ai. Two of the paper's six authors are affiliated with Invariant Labs, the organization that operates that registry.

## Sources

- [ethz-spylab/agentdojo on GitHub](https://github.com/ethz-spylab/agentdojo)
- [License: MIT](https://github.com/ethz-spylab/agentdojo/blob/main/LICENSE)
- [Project website](https://agentdojo.spylab.ai/)
- [README](https://github.com/ethz-spylab/agentdojo/blob/main/README.md)
- [Releases](https://github.com/ethz-spylab/agentdojo/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ethz-spylab-agentdojo
