CLI tool
ossf/malicious-packages avatar
ossf/malicious-packages

ossf/malicious-packages: an OSV-format database of malicious open source packages

A repository of reports of malicious packages identified in Open Source package repositories, consumable via the Open Source Vulnerability (OSV) format.

617 stars146 forksGoApache-2.0

At a glance

What is it?
The OpenSSF repository collects reports of malicious packages from npm, PyPI and other registries and stores each one as an OSV record. It is a data source for scanners and registry maintainers, not a tool you install to clean a machine.
Who is it for?
Adopt it if you maintain a scanner, registry, or dependency policy that can ingest OSV and you want a public, Apache-2.0-licensed feed of malicious package reports. Do not adopt it as an endpoint antivirus product, and do not expect it to cover non-package malware.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What ossf/malicious-packages actually is

The repository is a collection of reports about malicious packages found in open source package registries, stored in the Open Source Vulnerability (OSV) format. Each report is a record about a package, not a detection tool and not a scanner. The README frames the objective plainly: to be a comprehensive, high quality, open source database of reports of malicious packages published on open source package repositories.

The audience is narrow and specific. It is for people who build or operate something that consumes package data: registry maintainers deciding whether to remove a package, dependency scanners that want a public feed, and security researchers who need labelled examples of malicious packages to train or evaluate detection. A developer looking for a way to check their own laptop is not the target reader, and the project does not pretend to be one.

The adjacent project is OpenSSF Package Analysis, which the README names as closely related. The division of labour matters: Package Analysis produces detections, and this repository is where reports are collected and published in a common schema.

The scope rules decide what you get

The most useful part of the README is not the description but the boundary drawing. In scope: any package in an ecosystem supported by the OSV Schema, typosquatting attacks, account takeover, malicious prebuilt binaries shipped with a package, security researcher activity, and dependency or manifest confusion. Out of scope: non-malicious packages, vulnerability reports, compromised infrastructure, and offensive security tools unless they execute a malicious payload on install.

The project then spends most of its length on edge cases, which tells you something about how much argument there has been. Spam and typosquatting packages are admitted even when they are empty or trivial, on the grounds that they are hard to distinguish from dependency confusion. Obfuscated packages are admitted when they belong to an ecosystem whose terms forbid obfuscation, such as PyPI and Crates.IO, even if they are not clearly malicious. Telemetry alone is not malicious, but telemetry used to exfiltrate sensitive data or provide remote access is. Protestware is malicious only when it affects availability, integrity or confidentiality; a message printed to a console is not.

That is a deliberate bias toward recall over precision. If you ingest this dataset as a hard blocklist, you will occasionally block a package that is merely annoying rather than dangerous. The README is honest about this rather than hiding it, which is the right call for a shared dataset, but it shifts the false-positive cost onto the consumer.

How the OSV records move from report to feed

The repository layout shows the shape of the pipeline. Records live under osv/, configuration under config/, the command line tools under cmd/, shared code under internal/, and a site/ directory backs the published statistics. The Makefile exposes the working entry points: validate, preprocess, and preprocess-strict.

Preprocessing runs before IDs are assigned, which is the interesting design choice. Reports arrive from pull requests and from automated sources, and a package may be reported by more than one contributor. Preprocess merges those reports, and preprocess-strict aborts when it encounters reports it cannot merge, so a maintainer can choose between skipping a conflict and failing the run. IDs are assigned after that merge step, so the identifier refers to the merged record rather than to whichever contributor happened to open the first pull request.

Automated ingestion is supported from cloud storage. The README states that if you regularly produce high-quality detections with few false positives and have them accumulating in a database, the project can consume them automatically as OSV from S3 or GCS. The go.mod dependency list is consistent with that: it pulls in the AWS S3 SDK, Google Cloud Storage, and gocloud.dev, plus go-git for repository operations and the OSV schema bindings for Go.

Installing the tooling and running a validation

There is no released binary and no package on a registry. The README does not document an install step, so the way in is to clone the repository and use the Makefile targets, which invoke the Go tools in cmd/ directly. The module targets Go 1.25.8 according to go.mod.

The first useful thing to run is the validator, which checks the OSV files in the repository against the configuration in config/config.yaml. It is read-only, so it is a safe way to confirm your checkout and Go toolchain work before touching anything.

bash
make validate

If the records are well formed, the command exits without complaints. If a record is malformed, the failure names the file, which is how you locate the offending report.

The preprocessing step is the one that mutates state, merging reports and preparing them for ID assignment. It reads the same configuration file.

bash
make preprocess

The strict variant is the same command with an extra flag, and it aborts rather than skipping reports it cannot merge. Use it when you would rather fail loudly than publish a dataset with silently dropped records.

bash
make preprocess-strict

The Makefile also defines a test target that runs the unit tests with the race detector enabled, which is the quickest way to confirm the internal packages behave on your machine.

bash
make test

If you only want the data and not the tooling, you do not need any of this. The records are files in osv/ on the default branch, and you can read or download them without building anything.

Where this dataset is the wrong tool

Several limitations follow directly from what the README says and does not say.

It is not an endpoint security product. Every record concerns a package in a registry. If you want to know whether a phone is infected, or whether a document contains a macro, this repository has nothing for you. Those questions appear in general search traffic around the words malicious packages, and the dataset does not answer them.

Coverage is bounded by the OSV Schema's supported ecosystems. A malicious package in a registry outside that set is out of scope by definition, not by oversight.

The dataset is not a vulnerability feed. The README explicitly puts vulnerability reports out of scope. If your pipeline treats this as a CVE source, you will miss ordinary vulnerabilities entirely and you will misinterpret the records you do get, because a malicious package report is a statement about intent, not about a coding flaw.

Freshness depends on human contribution. Reports arrive through pull requests, and the README notes that bulk imports are accepted but asks contributors to open an issue first. There is no retrieved release history for this repository, and no release artifacts, so there is no versioned snapshot to pin against. You consume the default branch or the automated sources, and you inherit whatever state they are in.

Finally, the borderline cases are admitted on purpose. Empty typosquats and obfuscated-but-not-proven-malicious packages can appear in the data. A consumer that treats every record as a confirmed compromise will produce false alarms, and the project's own scope section is the evidence for that.

How it differs from Package Analysis and from vendor datasets

The closest relative is OpenSSF Package Analysis, which the README describes as closely related. The difference is the output. Package Analysis is the detection system: it watches package registries and produces findings. This repository is the publication layer: a curated, schema-normalised collection of reports, some of which come from Package Analysis and some from outside contributors. If you want to run detection yourself, Package Analysis is the project to look at. If you want a feed to match against your dependency graph, this repository is the one.

The other comparison is with vendor-maintained malicious package datasets, such as those published by commercial security firms. The practical difference is licensing and control. This repository is Apache-2.0 and its records are files in a public git repository, so you can mirror it, diff it across days, and feed it into your own tooling without a contract. Vendor datasets typically arrive through a product, with the coverage and update cadence set by the vendor. The trade-off runs the other way too: a vendor dataset usually comes with a support relationship and a stated service level, and this repository offers neither. What it offers is a schema you can read and a contribution process you can join.

Licence, maintenance and the cost of keeping up

The repository is licensed Apache-2.0, and the Makefile headers carry the same licence notice. That permits commercial use and modification, but it is worth reading the LICENSE file rather than assuming, and this is not legal advice. One thing the licence does not settle is the provenance of individual reports: a report may describe a package whose own licence is unrelated, and the record is a factual statement about that package, not a redistribution of its code.

Maintenance cost for a consumer is mostly ingestion, not code. Because there are no releases, there is no version to pin. You either track the default branch, or you consume the automated OSV sources the README describes, which are delivered through S3 or GCS. The second option is the one the project itself uses for bulk data, and it removes the need to run the Go tooling at all.

If you contribute rather than consume, the cost is the review process. CONTRIBUTING.md and CONTRIBUTION-GUIDELINES.md define what is accepted, and the strict preprocess target exists because merging multiple reports about the same package is a real source of friction. The README points bulk importers at an issue first, which is a signal that large drops need coordination rather than a single pull request.

Editorial conclusion

Adopt it if you maintain a scanner, registry, or dependency policy that can ingest OSV and you want a public, Apache-2.0-licensed feed of malicious package reports. Do not adopt it as an endpoint antivirus product, and do not expect it to cover non-package malware. Before relying on it, read CONTRIBUTION-GUIDELINES.md to see which reports are accepted and which are excluded, and check the daily statistics page at ossf.github.io/malicious-packages/stats for how much of the dataset covers the ecosystems you actually use.

Frequently asked questions

What is ossf/malicious-packages?

It is a repository of reports of malicious packages identified in open source package repositories, stored in the Open Source Vulnerability (OSV) format. It is closely related to the OpenSSF Package Analysis project, which produces detections.

Which ecosystems do the reports cover?

The README puts any package belonging to an ecosystem supported by the OSV Schema in scope, and lists typosquatting, account takeover, malicious prebuilt binaries, security researcher activity, and dependency or manifest confusion as the attack types covered.

How do I validate the OSV files in ossf/malicious-packages?

The Makefile defines a validate target that runs the validator command against config/config.yaml. Running make validate checks the OSV files without modifying them.

Does ossf/malicious-packages include vulnerability reports?

No. The README lists vulnerability reports as out of scope, along with non-malicious packages, compromised infrastructure, and offensive security tools that do not execute malicious payloads on install.

What is the difference between preprocess and preprocess-strict?

The Makefile describes preprocess-strict as behaving like preprocess but aborting if unmergable reports are encountered, which is the same command with the -abort-unmergable flag. Preprocess without that flag continues instead of failing the run.

Official sources

  1. Official README
  2. Project repository