Model or dataset
google/oss-fuzz-gen avatar
google/oss-fuzz-gen

OSS-Fuzz-Gen: LLM-Powered Fuzz Target Generation and Evaluation

LLM powered fuzzing via OSS-Fuzz.

1,442 stars224 forksPythonApache-2.0

At a glance

What is it?
OSS-Fuzz-Gen is an Apache-2.0 Python framework from Google that uses large language models to generate fuzz targets for real-world open-source projects and benchmarks them via OSS-Fuzz. It evaluates generated targets on four metrics and has reported 30 bugs across a range of C and C++ projects.
Who is it for?
Security researchers and engineers who work on open-source C or C++ projects and want to explore LLM-assisted fuzz target generation will find OSS-Fuzz-Gen a useful research baseline. The framework generated valid targets for 160 C/C++ projects in one experiment, with a maximum 29% coverage increase.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OSS-Fuzz-Gen does and who it targets

OSS-Fuzz-Gen addresses a specific bottleneck in large-scale fuzz testing: writing the fuzz targets themselves. A fuzz target is a small program that feeds generated inputs to a function under test; writing one requires understanding the function's expected input format and call sequence. The README describes the framework as generating fuzz targets for real-world C, C++, Java, and Python projects with various large language models and then benchmarking them via the OSS-Fuzz platform.

The intended users are security researchers, open-source maintainers enrolled in OSS-Fuzz, and teams that want to experiment with automated harness generation at scale. The framework does not produce ready-to-ship security reports; the README notes that generated targets are evaluated and findings must be triaged before disclosure.

Supported language and model scope

The README lists four target programming languages for fuzz target generation: C, C++, Java, and Python. Evaluation runs against OSS-Fuzz, which manages the build infrastructure and fuzzing runners for enrolled projects.

For LLM backends, the README lists: Vertex AI code-bison, Vertex AI code-bison-32k, Gemini Pro, Gemini Ultra, Gemini Experimental, Gemini 1.5, OpenAI GPT-3.5-turbo, GPT-4, GPT-4o, GPT-4o-mini, GPT-4-turbo, and Azure-hosted variants of GPT-3.5-turbo, GPT-4, and GPT-4o. All listed backends are cloud-hosted; the framework does not document support for local models.

Each backend is accessed through the llm_toolkit.models module. The pipeline is structured in stages (writing, execution, analysis) implemented in the stage/ directory, which allows swapping backends or modifying individual stages without rewriting the full experiment runner.

Setting up the environment and running an experiment

The README directs users to the USAGE.md file for the full setup guide. The Dockerfile shows the base environment the framework expects: Debian 12, Python with a virtual environment, gcloud CLI for accessing Vertex AI, and Docker for running OSS-Fuzz containers:

dockerfile
FROM debian:12
RUN apt-get update && \
    apt-get install --no-install-suggests --no-install-recommends --yes \
    python3-venv \
    gcc \
    libpython3-dev \
    git \
    apt-transport-https \
    ca-certificates \
    gnupg \
    curl \
    wget2 \
    clang-format && \
    python3 -m venv /venv

The repository contains two top-level experiment scripts: run_one_experiment.py for a single target and run_all_experiments.py for bulk runs. A pipeline.py entry point ties the stages together. Individual agents can also be executed or evaluated in isolation using the agent execution framework documented in agent_tests/readme.md, which the README describes as useful for developing and testing individual components without running a full experiment.

The pyproject.toml declares the package as OSS-Fuzz-gen version 0.1.0, requiring Python 3.10 or higher. The setup uses setuptools with setuptools-scm for versioning.

Evaluation metrics and what the experiment data shows

Generated fuzz targets are measured against four metrics: compilability (whether the target builds), runtime crashes (whether it crashes on standard inputs), runtime coverage (how much of the target function the fuzzer reaches), and runtime line coverage diff against existing human-written OSS-Fuzz targets for the same project.

The fourth metric is the most informative for assessing whether a generated harness adds value beyond what already exists. The README reports a sample experiment from January 2024 covering 1,300 or more benchmarks across 297 open-source projects; it found that the framework managed to generate valid fuzz targets with nonzero coverage increase for 160 C/C++ projects, with a maximum line coverage increase of 29% over existing human-written targets.

The README also lists 30 bugs reported across real projects through this work, including CVE-2024-9143 in OpenSSL and OOB read or write findings in cJSON, libplist, hunspell, zstd, gdbm, gpac, sqlite3, htslib, libical, and others. Each row in the bug table identifies the project, bug type, model used, prompt builder variant, and target oracle strategy.

These numbers describe a research snapshot, not a continuously updated benchmark. Results will vary by model, project, and configuration.

Benchmark sets and the research graph for targeted hunts

The benchmark-sets/ directory contains the project and function targets used in experiments. The README references an all set of over 1,300 benchmarks from 297 open-source projects. These benchmark definitions are what the experiment scripts iterate over when running a full evaluation.

For one-off or targeted experiments, the framework supports generating custom research graph topologies at runtime rather than following the fixed workflow. This allows a researcher to focus the LLM on a specific subsystem or class of vulnerability without modifying the benchmark set definitions.

The agents can also be executed individually through the agent execution framework. This is useful when a researcher wants to test a single phase, such as the prompt builder or the crash triager, in isolation before committing to a full experiment run.

Note that experiment reports are not public, as the README states they may contain undisclosed vulnerabilities. Any findings produced by the framework should be handled through the appropriate responsible disclosure process before being shared.

Maintenance status and the wrong-tool cases

The last push to the repository was on 2026-03-17. There are no GitHub releases; the version in pyproject.toml is 0.1.0.

For projects that are not enrolled in OSS-Fuzz, the evaluation pipeline has no target to run against. OSS-Fuzz handles the build and execution infrastructure; OSS-Fuzz-Gen sits on top of it and does not provide its own fuzzing engine or corpus management. Teams working on projects outside the OSS-Fuzz ecosystem would need to adapt the framework substantially.

For teams that want deterministic, rule-based fuzz target coverage without LLM dependencies, manually written libFuzzer harnesses are the common alternative. libFuzzer runs locally without cloud API access, produces reproducible results on the same input corpus, and does not require a verified LLM backend. The trade-off is that writing good harnesses requires developer time and knowledge of the target API.

OSS-Fuzz-Gen is appropriate for researchers exploring LLM-based test generation and for maintainers of OSS-Fuzz-enrolled projects who want to experiment with extending their existing fuzz coverage. It is not appropriate as a drop-in replacement for human harness authorship in a production security program.

Editorial conclusion

Security researchers and engineers who work on open-source C or C++ projects and want to explore LLM-assisted fuzz target generation will find OSS-Fuzz-Gen a useful research baseline. The framework generated valid targets for 160 C/C++ projects in one experiment, with a maximum 29% coverage increase. Teams that want deterministic, rule-based analysis rather than probabilistic LLM output should look at libFuzzer with manually written harnesses, which requires no cloud model API and produces reproducible results. The last push to the repository was on 2026-03-17. Before starting, confirm that cloud API credentials are in place, that the target project is enrolled in OSS-Fuzz, and that the team has a process for manually triaging generated findings before disclosing them.

Frequently asked questions

What programming languages does OSS-Fuzz-Gen support for fuzz target generation?

The README lists C, C++, Java, and Python as the supported target languages. Generated targets are evaluated via the OSS-Fuzz platform, which manages build infrastructure for enrolled open-source projects.

Are vulnerability reports from OSS-Fuzz-Gen experiments publicly available?

The README states that experiment reports are not public because they may contain undisclosed vulnerabilities. Findings should be handled through responsible disclosure before being shared externally.

How does OSS-Fuzz-Gen measure the quality of a generated fuzz target?

Generated targets are evaluated on four metrics: compilability, runtime crashes, runtime coverage, and runtime line coverage diff against existing human-written targets already in OSS-Fuzz for the same project. The fourth metric shows whether the generated target adds coverage beyond what is already present.

Official sources

  1. google/oss-fuzz-gen on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-oss-fuzz-gen.svg)](https://hysenlabs.com/projects/google-oss-fuzz-gen)