Model or dataset
chris-koch-penn/gpt3_security_vulnerability_scanner avatar
chris-koch-penn/gpt3_security_vulnerability_scanner

GPT-3 Security Vulnerability Scanner: An LLM Cybersecurity Benchmark Experiment

GPT-3 found hundreds of security vulnerabilities in this repo - (this was the first real LLM cybersecurity eval!)

603 stars103 forksPHPLicense varies

At a glance

What is it?
This repository documents an early experiment in which GPT-3 (text-davinci-003) was used to scan 129 vulnerable code files across 35 vulnerability categories, finding 213 issues while a commercial scanner found 99, with only 4 false positives in a sample of 60 GPT-3 results.
Who is it for?
This repository is the right reference for researchers studying how LLMs perform on static vulnerability detection, security educators who want a corpus of labeled vulnerable code examples, and anyone comparing early GPT-3 results to more recent model evaluations. It is not a production-ready scanner: it has no install guide, no active maintenance as a tool, and the underlying GPT-3 variant (text-davinci-003) has since been superseded by newer models.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 113 days ago.
What is it written in?
Mainly PHP, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the Repository Contains and Why It Was Created

The repository was set up as a target for GPT-3 to analyze. Each top-level folder is named after a vulnerability type: Buffer Overflow, SQL Injection, XSS, Command Injection, Path Traversal, SSRF, XXE, and 29 more. Every folder holds PHP, C, C#, Python, or JavaScript files that intentionally contain one or more vulnerabilities of the named type. The README in each folder records GPT-3's analysis output for the files in that folder.

The motivation described in the README was to test whether GPT-3 could perform static analysis comparable to a commercial scanner. GPT-3 found 213 vulnerabilities across the 129 vulnerable files. A commercial tool from a cybersecurity company found 99 issues on the same codebase. The README states that after manually reviewing a random sample of 60 of GPT-3's 213 findings, only 4 were false positives, a rate of about 6.7%. Both tools had many false negatives.

The README explicitly labels this experiment as the first real LLM cybersecurity evaluation, a claim that was credible at the time given when the GPT-3 variant used (text-davinci-003) was state-of-the-art. The finding that an LLM could outperform a commercial scanner on a labeled dataset was notable.

The 4,000-Token Window Constraint and File-by-File Approach

The text-davinci-003 variant of GPT-3 had a context window of 4,000 tokens, approximately 3,000 English words. This was not enough to scan an entire repository at once. The experiment scanned each file independently, sending them to the model one at a time.

The README notes the consequence: GPT-3 could not reason about vulnerabilities that span multiple files, such as a deserialization sink in one file consuming data from an untrusted source in another. For such cross-file vulnerabilities, the model would need to see both files simultaneously to trace the data flow.

In practice, this was less limiting than it might appear for this particular codebase. The vulnerable files use common libraries like express.js, Flask, and the Python standard library. GPT-3 appeared to have learned enough about these libraries during training to infer what imported functions do, allowing it to identify vulnerabilities that depended on understanding library behavior without inspecting the library source itself. The README frames this as analogous to how some commercial scanners handle imports in static analysis.

Examining the Vulnerable Code Examples

The repository's value as a benchmark comes from the labeled code examples. The README gives three examples of increasing complexity.

The simplest is a C program with a format string vulnerability:

c
#include <stdio.h>

int main(int argc, char **argv) {
    printf(argv[1]);

    return 0;
}

GPT-3 identified both the format string vulnerability (passing user input directly as the printf format string) and an unvalidated input concern. The README confirms both findings as correct.

A C# log injection example shown in the README is less trivial: an ASP.NET controller that passes a URL path parameter directly to a logger without sanitization. GPT-3 correctly identified the log injection vulnerability in this case.

The third example is a C image-processing function with multiple out-of-bounds reads and writes involving memcpy without bounds checking. GPT-3 found four of five vulnerabilities correctly; the one false positive was flagging the fopen return value as unchecked when it was in fact checked in the immediately following line. This kind of local reasoning failure is characteristic of the limited context reasoning that 4,000-token windows impose.

Vulnerability Categories and Repository Organization

The repository spans 35 vulnerability categories. The top-level entries include Buffer Overflow, Code Execution, Code Injection, Command Injection, Connection String Injection, Denial Of Service, File Inclusion, Format String Attacks, IDOR, Insecure File Uploads, Integer Overflow, LDAP Injection, Log Forging, NoSQL Injection, Open Redirect, Out of Bounds, PHP Object Injection, Path Traversal, PostMessage Security, Prototype Pollution, ReDoS, Resource Injection, SQL Injection, SSRF, Sensitive Data Exposure, Server Side Template Injection, Symlink Attack, Unsafe Deserialization, Use After Free, XPATH Injection, XSS, XXE, and Zip Traversal.

The README in each folder contains GPT-3's analysis output for the files in that folder. The files themselves are code samples, some trivial and some realistic snippets a developer might encounter in a production codebase. The README acknowledges that these are snippets and therefore lack the context of a larger codebase.

The repository also includes count_files.py and summarize_results.py at the top level, which are helper scripts for counting files and aggregating results across folders.

Limitations as a Scanner and as a Benchmark

As a production scanner, this repository has significant gaps. There is no command-line interface for running GPT-3 against a new codebase. The README describes the methodology but does not provide runnable scripts for repeating the experiment with current models. The summarize_results.py and count_files.py scripts process existing results but do not drive the LLM scanning process itself.

The text-davinci-003 model used in the experiment is no longer the state-of-the-art. Newer models with larger context windows can process entire files together or even multiple files at once, which would affect the results on cross-file vulnerabilities. Repeating the experiment with a current model would require rewriting the scanning code from scratch.

As a benchmark dataset, the repository has a known limitation: the files are labeled by vulnerability category, but the labeling was done by the repository creator rather than by a formal vulnerability database. The comparison against one unnamed commercial tool provides a single data point. The README does not specify which tool was used or what version it was.

Semgrep, an open-source static analysis tool, takes a different approach: it uses rule patterns written by security engineers rather than LLM prompting. Semgrep can be run locally with no API cost and integrated into CI pipelines. Its false positive rate depends on rule quality rather than model training data. For production use, Semgrep or a rules-based scanner provides more predictable behavior than an LLM queried per-file.

Using the Repository as a Learning Resource

The most direct use of this repository is as a catalog of vulnerability examples with LLM-generated analysis attached. A security educator can use the labeled folders to teach students what each vulnerability type looks like in code, then discuss why GPT-3's analysis was or was not accurate for each case.

For researchers benchmarking newer LLMs on vulnerability detection, the repository provides 129 files across 35 categories with a baseline result (GPT-3 finding 213 of an unknown total, with 4 false positives in a 60-item sample). Running a newer model against the same files and comparing precision and recall would produce a concrete, reproducible comparison.

The repository has no GitHub releases and is licensed under an unspecified licence (the repository metadata does not record one). The last push was on 2026-06-09. There is no active maintenance as a scanner tool, though the code samples themselves do not require updates to remain valid as vulnerability examples.

Editorial conclusion

This repository is the right reference for researchers studying how LLMs perform on static vulnerability detection, security educators who want a corpus of labeled vulnerable code examples, and anyone comparing early GPT-3 results to more recent model evaluations. It is not a production-ready scanner: it has no install guide, no active maintenance as a tool, and the underlying GPT-3 variant (text-davinci-003) has since been superseded by newer models. The repository's 35 vulnerability categories with human-labeled examples remain useful as a benchmark dataset independent of any particular model. The last push was on 2026-06-09.

Frequently asked questions

Can the gpt3_security_vulnerability_scanner be used as a standalone scanner for new codebases?

No. The repository documents a methodology and contains vulnerable code samples with GPT-3's output, but it does not include a CLI or scripts for running GPT-3 against a new codebase. Repeating the experiment requires writing code to call the OpenAI API file-by-file, following the approach described in the README.

How accurate was GPT-3 in this security vulnerability experiment?

After manually reviewing a random sample of 60 of GPT-3's 213 findings, the README reports only 4 false positives, a rate of about 6.7% in that sample. Both GPT-3 and the commercial comparison tool had many false negatives, meaning neither found all the vulnerabilities present.

What vulnerability types does the gpt3_security_vulnerability_scanner repository cover?

The repository contains 35 vulnerability categories including Buffer Overflow, SQL Injection, XSS, Command Injection, Path Traversal, SSRF, XXE, IDOR, Prototype Pollution, Use After Free, and others. Each category has a folder with labeled vulnerable code files and GPT-3's analysis in the folder's README.

Official sources

  1. chris-koch-penn/gpt3_security_vulnerability_scanner on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/chris-koch-penn-gpt3-security-vulnerability-scanner.svg)](https://hysenlabs.com/projects/chris-koch-penn-gpt3-security-vulnerability-scanner)