Model or dataset
sunblaze-ucb/exploitgym avatar
sunblaze-ucb/exploitgym

ExploitGym: a benchmark for AI agents that write real exploits

ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.

1,120 stars145 forksCApache-2.0

At a glance

What is it?
ExploitGym packages 869 real-world vulnerabilities into a runnable benchmark that scores whether an AI agent can turn a bug into a working exploit. Here is what the repository actually ships, how the controller and firewall fit together, and where the setup will fight you.
Who is it for?
Adopt ExploitGym if you already run containerised Linux and want a scored, reproducible target for exploit-writing agents rather than a chat transcript. Skip it if you need a quick local demo or you cannot give the runner Docker, GDB and outbound network isolation.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What ExploitGym measures that a chat transcript does not

Most agent evaluations end with a model saying it found a bug. ExploitGym ends with a flag. The README describes the project as "a large-scale, realistic benchmark built from real-world vulnerabilities across userspace programs, Google's V8 engine, and the Linux kernel, designed to evaluate AI agents' ability to develop exploits." The unit of work is a task drawn from a real vulnerability, and the agent has to produce an exploit that the harness can verify.

That framing sets the audience. It is not for people who want to learn exploit development, and it is not a scanner you point at your own code. It is for researchers and evaluation teams who need a repeatable way to compare agents, and for the people building those agents who want a signal that survives contact with a hardened target. The three target families matter because they fail differently: a userspace crash gives the agent a short path to control flow, while the kernel and V8 tasks demand a different class of reasoning and a much longer chain.

The released benchmark is v1.0 with 869 instances, per the README, and the canonical task list for that release is data/task_ids/v1.txt. The repository also ships a sample list at data/task_ids/sample.txt, which is what the quick start uses. Evaluating against the sample and reporting the result as a v1.0 score is the most obvious way to produce a number that means nothing.

The controller, the firewall and the LLM proxy

ExploitGym is not a single process. The quick start starts three services: a controller, a firewall, and an LLM proxy. pre_run.py runs readiness checks and starts all three, auto-detecting any that are already running, which means you can also start them by hand and pre_run.py will leave them alone. The hand-start path is documented in docs/eval.md.

The controller is the piece with the interesting design. The README states that no secret is baked into the repository: the controller generates its token salt, flag seed, and API key on startup, or takes them from environment variables. Both ends must use the same values. That is why the quick start prints four variables and asks you to export them before running the agent: CYBERGYM_ADMIN_KEY, CYBERGYM_SERVER_SALT, CYBERGYM_SERVER_FLAG_SEED and CYBERGYM_SERVER_API_KEY. The runner needs them to mint task tokens and derive expected flags.

The flag seed is the part worth pausing on. If the expected flag is derived from a seed rather than stored, then a fresh run with a fresh seed produces a different expected flag for the same task. That is good hygiene for a public benchmark, and it also means you cannot cache a flag between runs. The firewall is the other half of the isolation story: the README points to docs/firewall.md for outbound network isolation for agent containers, and the quick start pulls ubuntu/squid:latest as the Squid image. An agent that can reach the internet can fetch a public exploit for a known CVE, so outbound isolation is not decoration here, it is what keeps the task about the agent rather than about its search engine.

Installing ExploitGym and running your first task

The README gives a seven-step quick start, and it is worth following in order the first time. Python dependencies come from uv, and the proxy extra pulls in the LLM proxy stack.

bash
uv sync --extra proxy

The next step builds runtime artifacts (gdb, socat, nc, node plus the agent CLIs) and extracts task data. This is the slow one, and it is where system dependencies bite.

bash
bash scripts/setup/setup_data.sh
bash scripts/setup/validate.sh

validate.sh is the install check. If it fails, docs/setup.md is where the system dependencies, GDB, static node and agent CLIs are described. After that, pull the Squid image and the Docker images for the tasks you intend to run. The pull script takes a task list file, and the repository ships sample.txt for exactly this purpose.

bash
docker pull ubuntu/squid:latest
uv run scripts/setup/pull_images.py data/task_ids/sample.txt

Now export your model provider keys and start the three services. pre_run.py runs the readiness checks and starts the controller, firewall and LLM proxy.

bash
export OPENAI_API_KEY=...
export ANTHROPIC_API_KEY=...
uv run scripts/setup/pre_run.py data/task_ids/sample.txt

pre_run.py prints the environment variables the runner needs. Export them, then check the runner's interface before launching anything.

bash
export CYBERGYM_ADMIN_KEY=...
export CYBERGYM_SERVER_SALT=...
export CYBERGYM_SERVER_FLAG_SEED=...
export CYBERGYM_SERVER_API_KEY=...
uv run examples/run_agent.py --help

The README does not document what --help prints, so treat that command as the point where you read the runner's own usage rather than a step with a known output. The package also installs three console scripts for rendering agent streams: claude-stream-render, codex-stream-render and gemini-stream-render, each mapped to a module under cybergym.evaluation.agents. Those exist because a long exploit-writing run produces a lot of output, and reading it as raw JSON is unpleasant.

Where ExploitGym will waste your afternoon

The dependency chain is the first real cost. You need Docker for the target images, GDB and socat and nc for the runtime artifacts, and a static node build for the agent CLIs. That is a lot of surface area before a single task runs, and setup_data.sh is the script that has to get all of it right. The README pushes the detail into docs/setup.md rather than inlining it, which is honest but means the quick start is not self-contained on a bare machine.

The secret handling is the second trap. Because the controller generates its salt, flag seed and API key at startup, restarting the controller without exporting the same values invalidates the tokens the runner already holds. The README is explicit that both ends must use the same values, and it links to the Controller secrets section of docs/eval.md. If you restart services casually and see authentication failures, that mismatch is the first thing to check, not the agent's behaviour.

Third, the LLM proxy is a hard dependency for the documented path. examples/run_agent.py is the runner the README points at, and the proxy is what it talks to. If your model provider is not one the proxy supports, the quick start does not describe a fallback. The README does not document rollback or a partial teardown either, so cleaning up a half-finished run means understanding the three services yourself.

Finally, the scope is deliberately narrow. ExploitGym evaluates exploit development against known vulnerabilities in a controlled harness. It is the wrong tool for finding unknown bugs in your own codebase, for measuring whether a model will refuse a request, or for anything where you need a result in minutes rather than after a full setup and image pull.

ExploitGym against a general-purpose agent benchmark

The closest comparison is a general software-engineering agent benchmark, the kind built from GitHub issues and unit tests. The difference is not difficulty, it is the verification signal. Those benchmarks check whether a patch makes a test suite pass, so the agent can iterate against the same signal the grader uses. ExploitGym's targets are real vulnerabilities with system defenses in play, and docs/defenses.md covers disabling them (ASLR, for example). An agent that produces a crash has not finished; it has to turn that into something the harness accepts as an exploit.

That shifts where the difficulty lives. In a test-based benchmark, the hard part is understanding the code. Here, the hard part is the chain: getting a primitive, stabilising it across runs, and doing it while the firewall keeps the agent from reading a public writeup. The task families reinforce this. A V8 task and a kernel task share almost no technique, so a single aggregate score hides which kind of work an agent is actually good at.

A second comparison is a capture-the-flag benchmark. CTF challenges are usually hand-built, self-contained puzzles with a designed solution path. ExploitGym's tasks derive from real vulnerabilities, and the README notes that the bundled task data under data/tasks/ comes from external upstreams and retains their licences. Real provenance is a strength for realism and a complication for redistribution, which is exactly why DATA_LICENSE.md exists next to LICENSE.

Maintenance, licence and what a run costs you

The last push to the repository was on 2026-08-06, and the repository is not archived. The README states that the released benchmark is actively maintained and that the current release is v1.0 with 869 instances, with CHANGELOG.md holding the version history. There are no retrieved releases, so version tracking happens through CHANGELOG.md and the task list files rather than through tagged artifacts.

The licence split is the part to read carefully. The source code is Apache-2.0, and pyproject.toml declares license = "Apache-2.0" with license-files = ["LICENSE"]. The task data is not covered by that: the README says the bundled task data under data/tasks/ derives from external upstreams and retains their respective licences, pointing at DATA_LICENSE.md. If you plan to redistribute the tasks, or to publish a leaderboard built on them, that file is the one that governs, not LICENSE. This is a description of what the repository states, not legal advice.

Upgrade cost is mostly re-pulling images. Moving from one release to the next means a new canonical task list, which the README names as data/task_ids/v1.txt for the current release, and a corresponding pull of target images through scripts/setup/pull_images.py. Because the controller regenerates its secrets, an upgrade also means re-exporting the four environment variables for any long-running setup. There is no released migration tooling described in the README.

Editorial conclusion

Adopt ExploitGym if you already run containerised Linux and want a scored, reproducible target for exploit-writing agents rather than a chat transcript. Skip it if you need a quick local demo or you cannot give the runner Docker, GDB and outbound network isolation. Before trusting any number, confirm which task list you evaluated against, since the README names data/task_ids/v1.txt as the canonical list for the v1.0 release and data/tasks/ carries upstream licences that differ from Apache-2.0.

Frequently asked questions

What is the ExploitGym benchmark?

ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities across userspace programs, Google's V8 engine and the Linux kernel, designed to evaluate AI agents' ability to develop exploits. The current release is v1.0 with 869 instances.

What is ExploitGym?

It is a benchmark repository from sunblaze-ucb, licensed Apache-2.0 for the source code, that scores AI agents on turning real vulnerabilities into exploits. Tasks are run against Docker target images through a controller, a Squid firewall and an LLM proxy.

How do I install ExploitGym and run an agent?

The README's quick start runs uv sync --extra proxy, then scripts/setup/setup_data.sh and scripts/setup/validate.sh, pulls the Squid and target images, starts the services with scripts/setup/pre_run.py, exports the four printed environment variables, and launches examples/run_agent.py.

Does ExploitGym need API keys for a model provider?

Yes. The quick start exports OPENAI_API_KEY and ANTHROPIC_API_KEY before starting the LLM proxy, and the runner also needs CYBERGYM_ADMIN_KEY, CYBERGYM_SERVER_SALT, CYBERGYM_SERVER_FLAG_SEED and CYBERGYM_SERVER_API_KEY, which pre_run.py prints.

Can I evaluate against just a few ExploitGym tasks?

Yes. The image pull script and pre_run.py both accept a task list file, and the repository ships data/task_ids/sample.txt for that purpose. The canonical list for the v1.0 release is data/task_ids/v1.txt.

Is the ExploitGym task data covered by the Apache-2.0 licence?

No. The README states that the source code is Apache-2.0, while the bundled task data under data/tasks/ derives from external upstreams and retains their respective licences, documented in DATA_LICENSE.md.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. sunblaze-ucb/exploitgym on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sunblaze-ucb-exploitgym.svg)](https://hysenlabs.com/projects/sunblaze-ucb-exploitgym)