Open-source project
Galeax/CVE2CAPEC avatar
Galeax/CVE2CAPEC

CVE2CAPEC: turning a CVE into a chain of attack techniques

Security tooling that maps CVE identifiers to CAPEC attack patterns and converts repository context into practical threat references for operations.

319 stars52 forksPythonGPL-3.0

At a glance

What is it?
CVE2CAPEC is a Python pipeline that walks from a CVE to CWE, CAPEC, MITRE ATT&CK, D3FEND and ATLAS. It is a data-generation tool, not a scanner, and its daily output lands in results/new_cves.jsonl.
Who is it for?
Adopt CVE2CAPEC if you want a locally regenerable CVE-to-CAPEC-to-technique chain and you are comfortable with GPL-3.0. Do not adopt it if you need a scanner, a severity score or a supported API: it produces JSONL files and nothing else.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What CVE2CAPEC actually produces, and for whom

A CVE identifier tells you a vulnerability exists. It does not tell you how an adversary would use it, and it does not tell you which defensive technique answers it. CVE2CAPEC exists to close that gap by chaining public MITRE datasets together: CVE to CWE, CWE to CAPEC, CAPEC to MITRE ATT&CK Technique, ATT&CK Technique to MITRE D3FEND, and ATT&CK Technique to MITRE ATLAS.

The README frames the audience indirectly. The project stores all CVE data in the database folder and says plainly that it does not need to be run by yourself, because GitHub Actions updates the database every day at 00:05 UTC. That daily artifact is the real product: results/new_cves.jsonl. Someone doing vulnerability triage, building a detection backlog, or feeding a knowledge graph wants that file. Someone who wants a scanner with a severity score does not, and this is not that.

The generated data also feeds an interactive MITRE ATT&CK, D3FEND and ATLAS generator hosted at galeax.github.io/CVE2CAPEC. That page is a consumer of the same output, not a separate service.

The step-by-step data flow from retrieve_cve.py to technique2atlas.py

The architecture is a chain of small scripts, each writing what the next one reads. The README lists the order explicitly, and the repository layout matches it: retrieve_cve.py, cve2cwe.py, cwe2capec.py, capec2technique.py, technique2defend.py, technique2atlas.py, plus five update_*_db.py scripts that refresh the reference datasets.

The direction of travel is one way. A CVE is retrieved, its CWE is looked up, the CWE maps to one or more CAPECs, each CAPEC maps to one or more ATT&CK techniques, and each technique fans out to D3FEND and ATLAS. That means a single CVE can end up associated with several techniques, and a single technique with several defensive mappings. The scripts do not deduplicate for you; the mapping tables are the join.

Two design consequences are worth naming. First, everything depends on the quality of the upstream CWE and CAPEC assignments, so a vague CWE produces a vague chain. Second, the chain is only as fresh as the last update_*_db.py run, which is why those scripts come first in the documented order rather than being optional.

Installing CVE2CAPEC and getting a first CVE chain

The README gives a three-command install. It clones the repository, enters the directory, and installs the pinned dependency list from requirements.txt.

bash
git clone https://github.com/Galeax/CVE2CAPEC.git
cd CVE2CAPEC
pip install -r requirements.txt

Before mapping anything, refresh the reference databases. The README lists these five scripts as the update step, and they pull the CAPEC, CWE, ATT&CK technique, D3FEND and ATLAS datasets respectively.

bash
python update_capec_db.py
python update_cwe_db.py
python update_technique_db.py
python update_defend_db.py
python update_atlas_db.py

Then run the chain in the documented order. Each script feeds the next, so running them out of sequence leaves later stages reading stale or missing intermediate data.

bash
python retrieve_cve.py
python cve2cwe.py
python cwe2capec.py
python capec2technique.py
python technique2defend.py
python technique2atlas.py

The README states that new CVEs with all their data end up in results/new_cves.jsonl. Open that file after the chain completes; you should see one JSON object per line, and the file is the thing you hand to whatever consumes the mapping. The README does not document a CLI flag set, an output schema, or a rollback path if an update script fails partway through.

Where the chain breaks: CWE quality, missing links and no rollback

The weakest joint is the first one. CVE to CWE depends on the CWE assigned in the CVE record, and that assignment is frequently coarse. A CWE that describes a category rather than a specific weakness will map to a broad set of CAPECs, and the resulting ATT&CK techniques will be correspondingly generic. The tool does not flag low-confidence links, because it has no confidence model: it is a join, not an inference engine.

Gaps propagate in the other direction too. If a CWE has no CAPEC entry in the current CAPEC dataset, the chain stops there and that CVE simply has a shorter record. Nothing in the README describes a fallback or a placeholder, so downstream consumers need to handle partial records themselves.

There is also no documented rollback. The update_*_db.py scripts overwrite reference data, and the README does not describe backups, versioned snapshots or a way to pin the datasets to a date. If an upstream dataset changes shape, you find out when a later script fails, not before. For a pipeline that runs unattended on a schedule, that is a real operational gap rather than a theoretical one.

CVE2CAPEC against reading the MITRE sources directly

The obvious alternative is to skip the tool and query the MITRE datasets yourself. ATT&CK, CAPEC, CWE, D3FEND and ATLAS are all published openly, and the README lists their licences: ATT&CK under Apache 2.0, D3FEND under MIT, and ATLAS data from mitre-atlas/atlas-data under Apache 2.0 with the embedded Navigator from mitre-atlas/atlas-navigator also under Apache 2.0.

The difference is in what each side maintains. Going direct means you own the joins: you decide how CWE maps to CAPEC, how a CAPEC maps to a technique, and how you handle one-to-many relationships. That is more work, but it is also the only way to apply your own weighting, since CVE2CAPEC treats every mapping as equally valid. CVE2CAPEC's value is that the joins are already written and run daily, so you get a consistent chain without building it. Its cost is that you inherit its mapping choices and its schedule rather than your own.

A second alternative is simply consuming the hosted generator at galeax.github.io/CVE2CAPEC. That gets you the interactive view without running Python, but it is a UI over the same data, so it does not give you the JSONL file or any control over when the mapping is regenerated.

Licence, dependency weight and what maintenance costs you

CVE2CAPEC is released under GPL-3.0. The README addresses commercial use directly: for commercial use where you need to not be using the GPL, it points to contact [AT] galeax.com for additional options. If you plan to embed the generated data or the scripts in a proprietary product, that sentence is the one to read before anything else, and it is a licensing question rather than a technical one.

requirements.txt is heavier than the script count suggests. It pins pandas 2.2.3, numpy 2.1.2, ijson 3.4.0.post0, openpyxl 3.1.5, PyYAML 6.0.2, requests 2.32.3, tqdm 4.66.5 and a set of transitive pins including certifi, urllib3, idna and charset-normalizer. The pandas and numpy pins matter if you install into an environment that already has different versions, because the file uses exact equality rather than ranges. ijson and openpyxl suggest large JSON and spreadsheet sources are being streamed rather than loaded whole, which is a reasonable choice for MITRE datasets but does mean the pipeline is I/O bound.

The repository is not archived, and the README describes a daily GitHub Actions run at 00:05 UTC. No last push date is given, so the freshness signal to check is lastUpdate.txt in the repository root and the newest entries in the database folder. Those two files tell you whether the schedule is still producing output far more reliably than any claim about the project's activity level.

Deciding whether to depend on the daily new_cves.jsonl feed

The practical question is whether you consume the published artifact or run the pipeline yourself. Consuming results/new_cves.jsonl costs nothing and requires no Python environment, but it ties you to the upstream mapping choices and to the 00:05 UTC schedule. Running it yourself means you control when the reference datasets refresh and you can inspect intermediate output, at the price of a pandas and numpy install and an undocumentedly fragile update step.

If you run it yourself, the order in the README is not a suggestion. The five update scripts must precede the six mapping scripts, and the six mapping scripts must run in sequence. There is no orchestrator in the repository root, so a shell script or a CI job of your own is the only thing keeping that order honest. The README does not ship one.

Editorial conclusion

Adopt CVE2CAPEC if you want a locally regenerable CVE-to-CAPEC-to-technique chain and you are comfortable with GPL-3.0. Do not adopt it if you need a scanner, a severity score or a supported API: it produces JSONL files and nothing else. Before relying on it, check lastUpdate.txt and the newest line in database/ to confirm the GitHub Actions schedule is still producing data, and read requirements.txt before pinning anything in a shared environment.

Frequently asked questions

Do I need to run CVE2CAPEC myself?

No. The README states that GitHub Actions update the database every day at 00:05 UTC, so the new CVEs with all their data are available in results/new_cves.jsonl without running anything locally. Running it yourself is only necessary if you want to control the update timing or inspect intermediate output.

Which MITRE datasets does CVE2CAPEC chain together?

The README lists CVE, CWE, CAPEC, MITRE ATT&CK, MITRE D3FEND and MITRE ATLAS. The mapping runs from CVE to CWE, CWE to CAPEC, CAPEC to ATT&CK Technique, and then ATT&CK Technique to both D3FEND and ATLAS.

What licence is CVE2CAPEC released under?

It is released under the GNU General Public License version 3. The README adds that for commercial use where you need to not be using the GPL, you can contact contact [AT] galeax.com for additional options.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/galeax-cve2capec.svg)](https://hysenlabs.com/projects/galeax-cve2capec)