Model or dataset
paradigmxyz/evmbench avatar
paradigmxyz/evmbench

evmbench: Paradigm and OpenAI's Agent Harness for Smart Contract Audits

Collab with OpenAI. A benchmark and harness for finding and exploiting smart contract bugs

457 stars74 forksTypeScriptApache-2.0

At a glance

What is it?
evmbench is a self-hosted web application and evaluation harness, developed in collaboration between Paradigm and OpenAI, that runs a Codex-based agent against uploaded smart contract source code and delivers a structured vulnerability report. It pairs a Next.js frontend with a FastAPI backend, PostgreSQL, RabbitMQ, and a containerized worker that executes the audit agent.
Who is it for?
evmbench fits security researchers and teams who want to evaluate or benchmark how well LLM agents find and exploit smart contract vulnerabilities, either by running the UI workflow or by contributing to the companion detect evaluation. The three-hour default worker timeout and the Docker or Kubernetes infrastructure overhead make it better suited to targeted audits than to automated per-commit scanning.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What evmbench Does and Who It Is For

Smart contract auditing is a manual, expensive process. Automated tools that flag potential vulnerabilities help auditors triage faster, but most existing approaches rely on pattern matching or symbolic execution rather than agent-level reasoning about contract logic. evmbench tests a different approach: it runs an LLM-based agent (Codex in detect-only mode) against uploaded contract source code, instructs the agent to find and demonstrate vulnerabilities, and returns a structured report with the results.

The repository serves two purposes simultaneously. First, it is a companion interface to the detect evaluation in OpenAI's frontier-evals repository, which is included as a pinned submodule at frontier-evals/. Second, it is a self-contained application that security teams can deploy and use directly. Engineers upload a zip archive of contract files through a Next.js UI, choose a model and reasoning level, provide an OpenAI API key, and receive a rendered vulnerability report with file navigation and annotations.

The audience is security researchers benchmarking LLM agent capability on smart contract auditing, and teams that want to incorporate LLM-based pre-audit screening into their workflow without building the orchestration themselves.

Architecture: Six Services and the Worker Flow

The system is built from six distinct services coordinated through a job queue. A request starts at the Next.js frontend when a user submits a contract archive. The frontend POSTs to /v1/jobs/start on the FastAPI backend (port 1337), which creates a job record in PostgreSQL, stores a secret bundle in the Secrets Service (port 8081), and publishes a message to RabbitMQ. The Instancer consumes that message and starts a worker container either in local Docker or on a Kubernetes backend.

The backend API at port 1337 handles job submission, status queries (GET /v1/jobs/{id}), job history (GET /v1/jobs/history), and authentication. A dedicated Secrets Service at port 8081 stores and serves per-job bundles containing the uploaded zip and credential material. A Results Service at port 8083 receives worker output, validates that the JSON contains the expected vulnerabilities array structure, and persists it to the database. An optional oai_proxy service at port 8084 handles proxy-token mode.

The worker is where the audit happens. It fetches its bundle from the Secrets Service, unpacks the zip to an audit/ directory, then runs Codex in detect-only mode using the prompt from backend/worker_runner/detect.md, which is copied to $HOME/AGENTS.md inside the container. The model map in backend/worker_runner/model_map.json translates UI model keys to actual Codex model IDs. The agent writes its findings to submission/audit.md, which the worker uploads to the Results Service.

Setting Up evmbench Locally

Prerequisites are Docker and Bun. The worker and base images must be built before starting the stack:

bash
cd backend
docker build -t evmbench/base:latest -f docker/base/Dockerfile .
docker build -t evmbench/worker:latest -f docker/worker/Dockerfile .

Then start the backend stack:

bash
cp .env.example .env
docker compose up -d --build

The placeholder secrets in .env.example are sufficient for local development. The README notes that internet-exposed deployments require replacing them with strong values. Then start the frontend development server:

bash
cd frontend
bun install
bun dev

The frontend is available at http://127.0.0.1:3000 and the backend configuration endpoint at http://127.0.0.1:1337/v1/integration/frontend. The Makefile provides additional development tasks: backend-lint runs ruff over the Python backend, backend-typecheck runs py_compile on the hot paths (api/app.py, docker/worker/init.py, and the instancer backends), and docker-build-images builds all four images in sequence.

Security Model and OpenAI Credential Handling

evmbench runs an LLM-driven agent against code uploaded by users, which means the worker runtime handles untrusted input. The README is explicit: treat the worker filesystem, logs, and outputs as an untrusted environment. The SECURITY.md file contains the full trust model and operational guidance.

For OpenAI credentials, two modes are available. In direct BYOK (bring-your-own-key) mode, which is the default, the worker container receives a plaintext OPENAI_API_KEY or CODEX_API_KEY. In proxy-token mode, the worker receives only an opaque token and routes all OpenAI requests through the bundled oai_proxy service at port 8084, keeping the plaintext API key outside the worker container. Enabling proxy-token mode requires setting two environment variables:

bash
cd backend
cp .env.example .env
# set BACKEND_OAI_KEY_MODE=proxy and OAI_PROXY_AES_KEY=...
docker compose --profile proxy up -d --build

The maximum time a worker is allowed to run is controlled by EVM_BENCH_CODEX_TIMEOUT_SECONDS, which defaults to 10800 seconds (three hours). Long or complex contract audits may require raising this limit.

What the Agent Produces and Its Limits

The agent's job in detect-only mode is to identify vulnerabilities in the uploaded contract source and write a structured report. The output must contain parseable JSON with a vulnerabilities array. The Results Service validates this structure before persisting anything to the database. Beyond that constraint, the agent's behavior depends on the Codex model version selected, the reasoning level, and the complexity of the uploaded contracts.

The detect-only label matters: the agent is instructed to find and report vulnerabilities, not to deploy or execute any contract. However, the audit prompt in backend/worker_runner/detect.md does ask the agent to demonstrate exploitability, which is standard in security research (a vulnerability that cannot be exploited ranks lower priority than one that can).

There are practical constraints. The three-hour default timeout is generous but not unlimited; a very large codebase or a high reasoning level may exhaust it. The system requires an OpenAI API key, so the cost of each audit scales with the number of tokens the agent uses at the selected reasoning level. The Results Service validates structure, not accuracy; it does not verify whether the vulnerability descriptions are correct or whether the exploits the agent proposes would actually work. The backend directory also includes a prunner component for optional cleanup of stale worker containers, and the deploy/ directory covers production deployment considerations beyond the local setup.

Comparison to Slither and the Maintenance Record

The most widely used open source alternative in the smart contract security space is Slither, a static analysis framework for Solidity developed by Trail of Bits. Slither applies rule-based detectors that run in seconds without any external API calls. evmbench runs an LLM agent that reasons about contract behavior, which can catch higher-level logic flaws and multi-step attack paths that pattern-matching detectors miss, but at a cost of minutes to hours per audit and a variable per-run API bill. The two tools address different parts of the audit workflow: Slither suits continuous integration checks for known vulnerability classes; evmbench suits deeper, targeted analysis where an agent-level read of contract logic is worth the time and cost.

The last push to the repository was on 2026-09-18. The repository is not archived. The frontier-evals submodule points to the pinned evaluation code on OpenAI's repository. The licence is Apache-2.0, which permits commercial use and modification without requiring derivative works to be open-sourced, provided the licence and NOTICE file are preserved.

Editorial conclusion

evmbench fits security researchers and teams who want to evaluate or benchmark how well LLM agents find and exploit smart contract vulnerabilities, either by running the UI workflow or by contributing to the companion detect evaluation. The three-hour default worker timeout and the Docker or Kubernetes infrastructure overhead make it better suited to targeted audits than to automated per-commit scanning. Treat any deployment that accepts uploads from untrusted users as an untrusted runtime per the guidance in SECURITY.md. The last push was on 2026-09-18; the Apache-2.0 licence permits commercial use and does not require modifications to be open-sourced.

Frequently asked questions

What is evm bench?

evmbench is a benchmark and agent harness developed by Paradigm in collaboration with OpenAI for evaluating how well LLM agents can find and exploit smart contract vulnerabilities. It provides a web UI for uploading contract source, selecting a Codex model, and receiving a structured vulnerability report.

Does evmbench require a Kubernetes cluster to run?

No. The default setup uses Docker Compose with the Instancer starting worker containers locally. Kubernetes is listed as an optional backend for the Instancer and is documented in the deploy/ directory, but local Docker is sufficient for development and small-scale use.

What does the proxy-token mode in evmbench protect against?

Proxy-token mode keeps the plaintext OpenAI API key outside the worker container. Instead of receiving the actual key, the worker gets an opaque token and routes all OpenAI API calls through the bundled oai_proxy service. This limits the exposure of the API key if a worker container is compromised by untrusted contract code.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. paradigmxyz/evmbench on GitHub
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/paradigmxyz-evmbench.svg)](https://hysenlabs.com/projects/paradigmxyz-evmbench)