# github/advisory-database: the OSV-format advisory corpus behind GitHub's CVE and GHSA data

> GitHub's advisory database is a CC-BY-4.0 corpus of CVE and GHSA records stored as individual OSV files, curated by GitHub's security advisory team. It is a data repository, not a scanner, and its supported-ecosystem list decides what contributions get merged.

**github/advisory-database** — Security vulnerability database inclusive of CVEs and GitHub originated security advisories from the world of open source software.

- Repository: https://github.com/github/advisory-database
- Stars: 2,479 · Forks: 785
- Language: Unknown
- License: CC-BY-4.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/github-advisory-database

## What github/advisory-database actually is, and who it is for

This repository is a corpus, not a service. The README describes it as "a database of CVEs and GitHub-originated security advisories affecting the open source world", free and open source, and explicitly "a tool for and by the community". Every advisory GitHub acknowledges is stored as an individual file, formatted in the Open Source Vulnerability (OSV) schema. The top-level layout reflects that: an advisories/ directory holds the records, with CONTRIBUTING.md, SECURITY.md, LICENSE.md and a .github/ directory alongside it.

The audience is narrower than the name suggests. You are the right reader if you build dependency tooling, mirror advisory data into your own pipeline, or want to audit how a specific package's vulnerability record was written and sourced. You are the wrong reader if you want something that watches your lockfile and tells you what to patch today. Nothing in the README describes a scanner, a CLI or a hosted query endpoint. The repository's job is to hold records that other tools read.

The goals section is unusually plain about the intent: provide a free and open-source repository of advisories, let the community crowd-source knowledge about them, and surface vulnerabilities "in an industry-accepted formatting standard for machine interoperability". That third goal is the one that shapes everything else. OSV is the interchange contract, and the repository is the store.

## How OSV files, GHSA IDs and database_specific fit together

The mechanism is file-per-advisory. Each record lives in the repository as its own OSV document, which means a change to one advisory is a diff against one file and a pull request against one path. That is a deliberate trade: it makes review and history legible, and it makes bulk queries something you do by reading many files rather than by hitting an index.

Identity is handled by the GHSA ID. The README states that a GHSA ID is assigned when an advisory is created on GitHub or added from any supported source, and that the syntax is GHSA-xxxx-xxxx-xxxx, where each x is a letter or digit drawn from the set 23456789cfghjmpqrvwx. Outside the GHSA prefix the characters are randomly assigned and all letters are lowercase. The README gives this regular expression for validating a GHSA ID:

```
/GHSA(-[23456789cfghjmpqrvwx]{4}){3}/
```

Note what that character set excludes: 0, 1, a, b, e, i, k, l, n, o, s, t, u, y, z. An ID containing any of those is not a valid GHSA ID, which is a useful check when you ingest records from mixed sources.

The other structural piece is database_specific. OSV allows a database_specific JSON object at the level of an affected package, its affected ranges, and the vulnerability as a whole. The README quotes the spec's position that these fields hold extra information "as defined by the database from which the record was obtained" and that their meaning and format is "entirely defined by the database [of record]". GitHub uses a number of these values for its own purposes. The practical consequence: if you parse the standard OSV fields you get portable data, and if you depend on anything inside database_specific you are depending on GitHub's private conventions, which the OSV spec does not constrain.

## Where the records come from before GitHub curates them

The database is an aggregator with a stated source list, not a primary reporter. The README names security advisories reported on GitHub, the National Vulnerability Database, the npm Security Advisories Database, the FriendsOfPHP Database, the Go Vulnerability Database, the Python Packaging Advisory Database, the Ruby Advisory Database, the RustSec Advisory Database, and community contributions to this repository. It also invites suggestions for further sources via an issue.

That upstream mix explains a quality pattern you should expect. An advisory imported from NVD and an advisory written directly on GitHub do not arrive with the same level of detail, and the repository carries both. The README's guidance on references reflects this: references are meant to be supplemental and relevant, primary sources from the CVE or GHSA authors are preferred, code or documentation is added when applicable, and secondary write-ups are generally avoided unless the upstream source provided them. On fix commits the README is direct: advisories are about specific build artifacts rather than the project in general, a fix-commit reference helps downstream readers judge impact, and duplicate fix-commit contributions are rejected because irrelevant duplicates burden the reader.

So the curation layer is doing editorial work, not just import. A pull request is reviewed and merged or closed by GitHub's internal security advisory curation team, and if the advisory originated from a GitHub repository, the original publisher is mentioned for optional commentary. That mention step is worth knowing about if you file a correction: your change may sit while someone else is invited to weigh in.

## Getting the data: clone the repository and read one advisory

There is no package to install. The README gives no install command, no CLI and no server; the data is the repository. Getting a working copy means cloning it, and the checkout is large because every advisory is its own file under the advisories/ directory listed in the repository layout.

After cloning, you should see the advisories/ directory at the top level, alongside CONTRIBUTING.md, SECURITY.md, LICENSE.md and .github/. The README documents no aggregate index, manifest or query interface, so the tree itself is the entry point.

To see the record shape, open a single file. The README states each advisory is an OSV document, so the top-level keys are the OSV ones and the GitHub-specific extras sit under database_specific. Reading one file as JSON is the quickest way to check what a record actually carries before you write a parser, and the README's own validation example for GHSA IDs is a regular expression you can apply to the records you pull:

```
/GHSA(-[23456789cfghjmpqrvwx]{4}){3}/
```

That pattern matches the GHSA-xxxx-xxxx-xxxx form the README specifies, where each character comes from the set 23456789cfghjmpqrvwx. Run it against the identifiers in the files you ingest, and anything that fails is either not a GHSA ID or malformed. If you plan to process the whole set, expect to walk the advisories/ tree and parse each file independently.

## The ecosystem list is the real boundary, and it is enforced

This is the limitation that matters most, and the README states it without hedging: "we cannot accept community contributions to advisories outside of our supported ecosystems". The reason given is that the curation team reviews each contribution thoroughly and needs to be able to assess each change. The supported list is Composer, Erlang, GitHub Actions, Go, Maven, npm, NuGet, pip, Pub, RubyGems, Rust and Swift, each tied to a package registry, with Swift namespaced by DNS.

The README explains the logic: ecosystems are "the namespace used by a package registry", focused on packages that tend to be dependencies in software development. That definition quietly excludes a lot. Operating system packages, language runtimes distributed outside a registry, firmware, and vendored code have no obvious home here. If your dependency graph is mostly outside the twelve, this database is a partial answer at best, and you cannot fix the gap by submitting a pull request. The README's own route for that case is to open an issue suggesting a new ecosystem and let it be discussed.

A second boundary is subtler. Because the curation team must be able to assess each change, corrections to imported records can be slow or declined when the upstream source is the authority. The README's stance on duplicate fix commits shows the same instinct: the reader's burden is treated as a real cost, and additions that do not reduce it get closed.

## GitHub Advisory Database vs NVD: aggregator against upstream authority

The comparison people reach for is NVD, and the relationship is not symmetric. NVD is one of the sources the README lists for this database, so the GitHub Advisory Database is partly downstream of it. Choosing between them is choosing between a normalized aggregator and the upstream record.

The difference in approach is the format and the scope. NVD publishes CVE records in its own schema; this repository republishes CVE-derived and GitHub-originated advisories as OSV files, which is the format the README calls an industry-accepted standard for machine interoperability. If your tooling already speaks OSV, the mapping work is done for you here. If you need the authoritative CVE record with its NVD-specific scoring and metadata, this repository is a second stop, not a replacement.

The second difference is coverage. This database adds GitHub-originated advisories and imports from npm, FriendsOfPHP, Go, Python, Ruby and Rust advisory databases, so it can carry entries that exist as GHSA records without a matching CVE. That is a genuine advantage for package-level dependency work. The cost is that you inherit GitHub's curation decisions and its database_specific conventions on top of whatever the upstream source said. For a security review where provenance matters, read the references field and follow it back rather than treating the OSV file as the last word.

## Licence, reuse and what a fork obliges you to do

The repository is licensed under CC-BY 4.0, and the README points to GitHub's terms for additional products and features for the full terms rather than restating them.

CC-BY-4.0 is an attribution licence, and that has practical consequences for anyone mirroring the data. If you redistribute the corpus, or a derived dataset built from it, the licence's attribution condition travels with it. This is not the same as a permissive software licence, and it is not the same as public domain. The README does not spell out a required attribution string, a notice file format, or how attribution should be handled when records are merged with data from other sources. Those specifics are in the linked terms, and the honest position is that this article cannot summarize them for you. If you are building a product on top of the corpus, read the linked terms and, where the stakes justify it, take advice; nothing here is legal advice.

One more reuse consideration is structural rather than legal. Because each advisory is a separate file, a mirror is a large tree of small documents. Keeping it current means tracking upstream changes across that tree, and the README documents no release process, no versioned snapshot and no changelog feed to subscribe to. There is also no releases data for this repository, so there is no tagged artifact to pin against. Plan for a refresh that re-reads the tree rather than one that downloads a bundle.

## Contributing a correction without having it closed

The README gives two routes. The first is the web form: from any advisory on github.com/advisories, click "Suggest improvements for this vulnerability", edit the form, and submit to open a pull request. The second is a pull request directly against a file in this repository, following CONTRIBUTING.md. Both land in the same review queue, where GitHub's internal security advisory curation team merges or closes them, and where the original publisher of a GitHub-originated advisory may be mentioned for optional commentary.

Before you write, check three things against the README's own rules. Is the package in one of the twelve supported ecosystems? If not, the contribution will not be accepted, and an issue proposing the ecosystem is the correct move. Is the reference you want to add primary and relevant, or a secondary write-up? The README generally avoids the latter. Does the advisory already carry the fix commit you are about to add? If so, the change is duplicative and will be closed.

What the README does not document is rollback: there is no described process for reverting an advisory once a change is merged, and no stated appeal path beyond opening an issue. Treat a merged edit as durable and get it right the first time.

## Conclusion

Adopt it if you need machine-readable advisory data in OSV format for Composer, Erlang, GitHub Actions, Go, Maven, npm, NuGet, pip, Pub, RubyGems, Rust or Swift, and you can live with the CC-BY-4.0 attribution terms. Do not adopt it if your stack sits outside those ecosystems, since the README states community contributions to advisories outside them are not accepted, or if you need a scanner rather than a dataset. Before relying on it, check the advisories/ directory layout and a sample file's database_specific block, confirm whether the advisories you care about are reviewed or unreviewed, and read the terms linked from the License section in full.

## FAQ

### What is the GitHub Advisory Database?

It is a free, open-source repository of CVEs and GitHub-originated security advisories affecting open source software. Each advisory is stored as an individual file in the Open Source Vulnerability (OSV) format, and pull requests can be submitted to change or update the information.

### How does the GitHub Advisory Database compare with NVD?

The National Vulnerability Database is one of the sources this database imports from, so the relationship is partly downstream. This repository republishes CVE-derived and GitHub-originated advisories as OSV files and also imports from npm, FriendsOfPHP, Go, Python, Ruby and Rust advisory databases, which means it can carry GHSA records that have no matching CVE.

### What is a GHSA vulnerability?

A GHSA ID is the unique identifier assigned to a security advisory when it is created on GitHub or added to the GitHub Advisory Database from a supported source. The format is GHSA-xxxx-xxxx-xxxx, where each character is a letter or digit from the set 23456789cfghjmpqrvwx, randomly assigned and lowercase outside the GHSA prefix.

### What is a GitHub security advisory?

The README lists security advisories reported on GitHub as one of the sources feeding this database, alongside NVD, npm, FriendsOfPHP, Go, Python, Ruby, RustSec and community contributions. When an advisory originated from a GitHub repository, the original publisher is mentioned on the pull request for optional commentary during curation.

## Sources

- [github/advisory-database on GitHub](https://github.com/github/advisory-database)
- [Issues](https://github.com/github/advisory-database/issues)
- [License: CC-BY-4.0](https://github.com/github/advisory-database/blob/main/LICENSE)
- [README](https://github.com/github/advisory-database/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/github-advisory-database
