# CyberGym: a benchmark that scores AI agents on real vulnerability analysis

> CyberGym is an Apache-2.0 Python and Docker framework from sunblaze-ucb that turns known CVEs into graded proof-of-concept tasks. The interesting part is not the scoreboard, it is the plumbing: a local server, a firewall, and a 240GB dataset you have to host yourself.

**sunblaze-ucb/cybergym** — CyberGym is a large-scale, high-quality cybersecurity evaluation framework designed to rigorously assess the capabilities of AI agents on real-world vulnerability analysis tasks.

- Repository: https://github.com/sunblaze-ucb/cybergym
- Website: https://cybergym.io
- Stars: 923 · Forks: 113
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/sunblaze-ucb-cybergym

## The gap CyberGym is trying to fill in agent evaluation

Most agent benchmarks ask a model to answer a question. CyberGym asks it to produce a file that crashes a program. The repository describes the project as a large-scale, high-quality cybersecurity evaluation framework designed to rigorously assess the capabilities of AI agents on real-world vulnerability analysis tasks. The unit of work is a task such as arvo:10400, drawn from the ARVO and OSS-Fuzz identifiers visible in the task list.

That choice sets the audience. If you are building or selecting an agent that reads a codebase, reasons about a memory-safety bug, and writes a proof of concept, CyberGym gives you a scoring harness. If you are evaluating a chat assistant that explains CVEs in prose, the harness measures something you do not care about, and the setup cost will not be worth it. The framework is also not a scanner and not a patch generator. Nothing in the README suggests it fixes anything; it grades submissions against pre-patch and post-patch behaviour, which is the distinction the FAQ file is said to cover.

## How a task moves from dataset to server to scored PoC

The data flow has three pieces. A dataset clone supplies the task material. A local server accepts proof-of-concept files and runs them. A task generator produces a working directory that the agent consumes.

The generator writes four items into an output directory: description.txt, README.md, repo-vul.tar.gz, and submit.sh. The agent reads the description and the vulnerable source, writes a PoC, and calls submit.sh. The server executes that PoC inside the prepared environment and returns a JSON record containing task_id, exit_code, output, and poc_id. Later, scripts/verify_agent_result.py reads the SQLite database at the path you passed as --pocdb_path and reports fields including poc_hash, poc_length, vul_exit_code, and fix_exit_code.

Those last two fields are the design decision that matters. A submission is interesting when it behaves differently against the vulnerable build and the fixed build, which is why the framework needs both versions of every target. That is also why the full server data is described as roughly 10TB: the compilation environments are the expensive part, not the PoCs. The binary-only mode exists precisely to avoid that cost.

## Installing CyberGym and generating your first task

The README requires a Python and Docker environment, and pyproject.toml sets requires-python to >=3.12. Install the server and task-generation extras first.

```bash
pip3 install -e '.[dev,server]'
```

Then fetch the benchmark data. The README puts the clone at roughly 240GB, and it uses Git LFS, so install that before cloning.

```bash
git lfs install
git clone https://huggingface.co/datasets/sunblaze-ucb/cybergym cybergym_data
```

If you only need static analysis and not the dynamic compilation environment, the README offers a binary-only path that it says takes about 130GB. It downloads runners with a script and then fetches and extracts an archive.

```bash
python scripts/server_data/download_binary_only_runners.py
wget https://huggingface.co/datasets/sunblaze-ucb/cybergym-server-binary/resolve/main/cybergym-server-data.7z
7z x cybergym-server-data.7z
```

Before starting the server, pick the bind address. The README is explicit that the right address is the gateway of the Docker network your agent containers run on, and it gives a command to query it rather than hardcode it.

```bash
HOST=$(docker network inspect cybergym-internal -f '{{(index .IPAM.Config 0).Gateway}}')
echo $HOST
```

Start the PoC submission server on port 8666, pointing it at a mask map, a log directory, and a SQLite database inside that directory.

```bash
PORT=8666
POC_SAVE_DIR=./server_poc
python3 -m cybergym.server \
    --host $HOST --port $PORT --mask_map_path mask_map.json \
    --log_dir $POC_SAVE_DIR --db_path $POC_SAVE_DIR/poc.db
```

Now generate a task and submit a placeholder file to confirm the loop works. The README's example uses arvo:10400 and difficulty level1.

```bash
SERVER_IP=$HOST
SERVER_PORT=8666
TASK_ID='arvo:10400'
OUT_DIR=./cybergym_tmp
CYBERGYM_DATA_DIR=./cybergym_data/data
python3 -m cybergym.task.gen_task \
    --task-id $TASK_ID --out-dir $OUT_DIR --data-dir $CYBERGYM_DATA_DIR \
    --server "http://$SERVER_IP:$SERVER_PORT" --mask-map mask_map.json \
    --difficulty level1
echo -en "\x00\x01\x02\x03" > $OUT_DIR/poc
bash $OUT_DIR/submit.sh $OUT_DIR/poc
```

The expected return is a JSON object with the task id, an exit code, the target's stdout, and a poc_id. The output in the README shows the fuzzer running one input and reporting that it executed the file in 3 ms. A four-byte file that does not crash anything is a successful round trip, not a successful exploit.

## The firewall, the API key, and the warning you should read twice

The README carries a warning in block capitals: deploy everything locally and do not expose any part of CyberGym to the public internet. It then explains why, noting that 0.0.0.0 binds every interface and that on a machine with a public IP this exposes a partly unauthenticated service.

The framework ships a firewall for this. python3 -m cybergym.firewall start and status both print the Docker gateway as host_gateway, and FirewallProxyManager.start() adds that gateway to NO_PROXY and to the proxy's IP allowlist so agent traffic to the server bypasses Squid. The README also warns that the two gateways are not interchangeable: cybergym-internal is created with internal=True, so containers on it have no route to the default bridge, and a server bound to 172.17.0.1 is unreachable from firewalled agents.

There is an API key, and the README is blunt about it. The value shown in the documentation, cybergym-030a0cd7-5908-4862-8ab9-91f2bfc7b56d, is described as the public default placeholder key, to be overridden via CYBERGYM_API_KEY on both the server and the client. The README states directly that it is not a substitute for keeping the server private. Treat that as the honest reading: the key is a speed bump, not an authentication boundary.

## Where CyberGym gets expensive or simply does not fit

Storage is the first wall. The dataset clone is roughly 240GB, the binary-only server data is about 130GB, and the full server data is described as roughly 10TB. A subset download script exists for ten named tasks, five of which the README says an agent can solve and five of which it says are not easy. That subset is the honest starting point for anyone who wants to see whether the harness runs before buying disks.

The second wall is the network topology. Because the server must be reachable from agent containers and the host but not from the outside, you are configuring Docker networks and a Squid proxy before you evaluate anything. Teams used to a hosted API will find this heavy. Teams used to fuzzing infrastructure will recognise it.

The third limitation is scope. Every task in the visible material is a crash-reproduction task with a pre-patch and post-patch build. A denial-of-service bug, a logic flaw, or a misconfiguration has no exit-code signal here. If your agents work on those classes, CyberGym is the wrong harness, and no amount of tuning the mask map will change that.

Finally, the packaging is thin in one visible place: pyproject.toml lists no base dependencies at all, with everything behind the dev and server extras. Installing the bare package gets you the import path and little else.

## What CyberGym replaces, and what it does not

The obvious alternative is the family of static vulnerability benchmarks built from curated code snippets with labelled answers. Those are cheap to run and easy to grade, and they measure whether a model can recognise a pattern. CyberGym measures whether it can produce an input that changes a program's behaviour, which is a different and harder claim.

The second alternative is running your own fuzzing or capture-the-flag setup. That gives you full control and no dataset download, but you own the scoring logic, the pre-patch and post-patch comparison, and the reproducibility story. CyberGym's contribution is that comparison being standardised, with a submission protocol and a verification script that reads a SQLite database rather than trusting a transcript.

A third point of comparison is the hosted leaderboard. The repository includes SUBMISSION.md, described as guidelines for leaderboard submissions covering cost reporting and writeup requirements. If you only want a number next to a model name, that document is the relevant one, and you may never need to install anything. The framework is for people who want the number to be reproducible on their own hardware.

## Licence, maintenance, and what a version bump costs you

CyberGym is licensed under Apache-2.0, and pyproject.toml declares version 0.2.0. Apache-2.0 is a permissive licence with a patent grant and a requirement to preserve notices; it does not impose copyleft on your own code. That matters here because the repository bundles or downloads third-party material: the dataset on Hugging Face, the server binary archive, and Docker images containing compilation environments for ARVO and OSS-Fuzz targets. Those components carry their own terms, and the repository does not restate them. Anyone redistributing a packaged environment should read the upstream licences rather than assume Apache-2.0 covers everything in the download.

The last push to the default branch was on 2026-08-28. There are no retrieved releases, so the version in pyproject.toml is the only version signal available, and upgrades will most likely mean pulling main rather than installing a tag. The practical cost of that is data drift: the task generator, the mask map, and the server data are separate artefacts, and the README does not document a compatibility matrix between them. If you pin a dataset and later move the code forward, re-run the ten-task subset before trusting a full run. That is a concrete check, not a general caution.

## Conclusion

Adopt CyberGym if you already have a lab machine with Docker, a few hundred gigabytes of disk, and a team that wants to compare agents on reproduction tasks rather than on multiple-choice questions. Do not adopt it if you want a hosted leaderboard you can query from a laptop, or if you cannot keep the PoC submission server off the public internet, since the README warns that binding 0.0.0.0 exposes a partly unauthenticated service. Before you commit, verify three things: that your Python is 3.12 or newer, that you can pull the dataset from the Hugging Face repository sunblaze-ucb/cybergym, and that you can reach the gateway of the cybergym-internal Docker network, because the README states the two gateways are not interchangeable and a server bound to the default bridge is unreachable from firewalled agents.

## FAQ

### How does CyberGym work?

It generates a task directory containing description.txt, README.md, repo-vul.tar.gz, and submit.sh, then runs a local server that executes submitted PoC files against prepared vulnerable and fixed builds and returns a JSON result with task_id, exit_code, output, and poc_id.

### What is the CyberGym benchmark?

CyberGym is a cybersecurity evaluation framework that assesses AI agents on real-world vulnerability analysis tasks, distributed as an Apache-2.0 Python project with a dataset hosted on Hugging Face under sunblaze-ucb/cybergym.

### What is CyberGym?

The repository describes it as a large-scale, high-quality cybersecurity evaluation framework designed to rigorously assess the capabilities of AI agents on real-world vulnerability analysis tasks, with a website at cybergym.io and an arXiv paper numbered 2506.02548.

## Sources

- [Issues](https://github.com/sunblaze-ucb/cybergym/issues)
- [License: Apache-2.0](https://github.com/sunblaze-ucb/cybergym/blob/main/LICENSE)
- [Project website](https://cybergym.io)
- [README](https://github.com/sunblaze-ucb/cybergym/blob/main/README.md)
- [sunblaze-ucb/cybergym on GitHub](https://github.com/sunblaze-ucb/cybergym)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sunblaze-ucb-cybergym
