Malware-Research-Hub: 2,916 specimens, none of them runnable
Self-contained malware research hub: curated catalog of 80 families (1971-2024) + 2,764 real encrypted samples, indexed and searchable. Local Flask app, bilingual.
At a glance
- What is it?
- Malware Research Hub is a local Flask application that indexes a curated catalogue of 84 deduplicated malware families and twenty collections totalling 2,916 encrypted specimens. Every sample ships as a password-protected archive, the server refuses to execute or serve binaries, and the documentation tells you to work inside an isolated virtual machine.
- Who is it for?
- Malware Research Hub is worth a look as a catalogue, because deduplicating variants into one record per family with impact, attribution and chronology is more useful than a directory of sample hashes, and the containment story is designed rather than promised. Three cautions before you clone it.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 23 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Every specimen is a password-protected archive and the server never serves binaries
The containment model is the design, not a warning bolted onto one. Every sample in the arsenal is stored as an encrypted ZIP with the password `infected`, the server does not execute or serve binaries, and the tree is documented as containing no raw executables at all. The rules that follow from that are specific. Never extract a sample on your main machine. Work only inside an isolated virtual machine with host-only networking and snapshots, using a malware analysis distribution and a simulated network so nothing escapes. An extracted specimen is hostile, and the documentation compares handling one to handling digital plutonium. That last framing is doing real work: it tells a reader that the danger survives the archive format. The application also exposes an open-folder action rather than a download, which keeps the sample bytes inside the encrypted archives on disk and keeps the browser out of the delivery path entirely.
The catalogue deduplicates variants into one record per family
The catalogue is the intellectual content here, and its organising principle is deduplication rather than accumulation. Eighty-four families are catalogued with the variants and remixes of the same malware folded into a single record, with Petya, NotPetya and GoldenEye named as the example of three names that are one family, and each record carrying impact, attribution and a chronology. That is a different artefact from a hash list, and it is the part a researcher can actually use when they want to know whether a sample they hold is something they have seen. The timeline is interactive, organised by category from 1971 to 2024, and lives inside the application as its own tab, and it is regenerated when the catalogue is edited rather than being maintained separately. So the timeline cannot drift from the catalogue, which is the usual failure mode of a chronology that lives in a spreadsheet next to a database.
Twenty collections add up to 2,916, and the repository description says 2,764
The collection table is the most checkable part of the documentation, and the numbers do not agree across the repository. Twenty collections are listed with their specimen counts: a thousand and forty-one source files by platform, 364 from a virus exchange collection, 359 from a zoo archive of binaries plus source, 314 historical binaries, 261 documented techniques in C and C++, 144 JavaScript samples grouped by date, 130 modern families, 108 pulled from a bazaar service, 41 campaign samples with metadata, 42 Android samples, then a tail of smaller collections of 33, 19, 15, 12, 10, 10, 8 and 3, and finally one source dropper plus rootkit and one file encryptor. Those rows sum to 2,916, which matches the headline figure in the documentation. The repository description, however, says 80 families and 2,764 samples. So the description is stale by 152 specimens and four families. Not dangerous, but it is the kind of drift that makes a reader doubt the parts they cannot verify.
The collector pulls by signature and encrypts before indexing
Growing the arsenal is a documented two-command flow against a named source, and the pipeline order matters. The collector script is given an API key from a free account on a well-known malware exchange service, then asked for samples by signature rather than by name, with a limit:
cd app
export MB_API_KEY="tu-clave-gratuita-de-auth.abuse.ch"
python3 mb_download.py by-signature LockBit --limit 25
python3 integrate_collections.py # cifra lo nuevo + reindexaThe integration step is the one that enforces the containment invariant, because it encrypts whatever arrived and reindexes the tree in the same run. Nothing lands as a bare executable even when it arrives that way. Fetching by signature rather than by family name is also the better interface for a research workflow, since signatures are what a sandbox or an incident report gives you, and the limit keeps one query from pulling a whole campaign. The downloader is described as expanding the arsenal from reliable sources by API with a verified hash and encrypted delivery, so the verification claim is about the download itself rather than about a later scan.
Startup is an indexer pass, then a local server on one port
The run sequence is four steps and one of them is the part people forget. You move into the application directory, optionally create and activate a virtual environment, install the requirements, then run the indexer to generate the index from the collections directory, and only then start the server:
cd app
python3 -m venv .venv && source .venv/bin/activate # opcional
pip install -r requirements.txt
python3 indexer.py # (re)genera el índice desde collections/
python3 server.py # ▶ http://127.0.0.1:5057The indexer being a separate pass is the design decision worth noticing. The index is derived state, generated by walking the collection tree, which means the same property holds for the specimens as for the catalogue. Nothing about the index is authoritative, so a corrupted index can be rebuilt by running one command. The server then serves metadata only, on the loopback interface, which is what makes the no-binaries rule enforceable rather than aspirational. The documentation claims no external dependencies beyond Flask, though a requirements file is still installed at setup, so the real dependency set is whatever that file contains.
The explorer jumps to the real folder, which only makes sense locally
The specimen explorer is described as searchable and filterable by family, platform, type and collection, with a direct jump to the real folder for each specimen. That last feature is the interesting one and the one with the sharpest edges. A jump to a folder means the browser hands a path to the server and the server acts on the local filesystem, which is a category of operation that does not belong in anything reachable from a network. It is also the reason the rules insist on a host-only machine: the application is a local tool by construction, and the port binding is part of its safety model rather than a convenience. Filtering by four dimensions is where the value is for a researcher, since narrowing 2,916 specimens to the Android ones that match a signature is the actual work. The server side is described as an API for families, samples, resources, statistics and the open-folder action, which is a small surface for what the interface does.
The resource hub lists tooling rather than samples
One module in the project contains no malware at all, and it is the one most likely to be useful to a reader who is not handling samples. The resource hub points at other repositories, sandboxes, machine learning datasets including two named image-classification malware corpora, fuzzy hashing tools with their three named algorithms, and the STIX and MAEC standards for representing malware intelligence. That list is a reading list rather than a download, and it is the one part of the project you can act on without a virtual machine. Naming fuzzy hashing specifically is a good sign, since near-duplicate detection is the problem the catalogue solves at the family level, and pairing the two is coherent. The STIX and MAEC references matter for a different reason: they suggest the catalogue is intended to feed something outside itself eventually, which would also raise the question of what schema the records follow. The documentation does not say whether they do.
There is no licence file for a repository that redistributes other people's samples
The repository tree is short and worth reading literally. It contains the readme in two languages, a security file, the application directory, the collections directory, a scripts directory, two image files and the usual ignore and attribute files. There is no licence file. For a repository whose stated purpose is to aggregate specimens from twenty other people's collections, that is the most consequential gap in the project, because redistribution rights for malware samples are not as simple as they are for source code. The provenance section does name the upstream collections it aggregates from and links several of them, so the sourcing is documented even where the licensing is not, and the collected copy of that section is cut off partway through the final list. The security file being present is a smaller positive signal, and the bilingual readme pair suggests the project is meant to be read by more than one audience. Neither of those changes what is missing.
Editorial conclusion
Malware Research Hub is worth a look as a catalogue, because deduplicating variants into one record per family with impact, attribution and chronology is more useful than a directory of sample hashes, and the containment story is designed rather than promised. Three cautions before you clone it. The counts are not consistent across the repository, since the project description says 80 families and 2,764 samples while the documentation says 84 and 2,916 and the twenty collection rows sum to 2,916. There is no licence file in the tree for a repository that redistributes other people's samples, which is a real problem if you intend to do anything with it beyond reading. And the explorer can open the real folder on disk, so keep the whole application bound to a host-only network as the rules say rather than exposing it.
Frequently asked questions
How does Malware-Research-Hub keep samples from running on your machine?
Every sample is stored as a ZIP encrypted with the password infected, the server neither executes nor serves binaries, and the tree is documented as containing no raw executables. The rules require working inside an isolated host-only virtual machine with snapshots.
How many malware families and specimens does the hub catalogue?
The documentation says 84 deduplicated families and 2,916 specimens across twenty collections, and those collection rows add up to 2,916. The repository description is behind that, saying 80 families and 2,764 samples.
What does Malware-Research-Hub add to the samples it aggregates?
A deduplicated catalogue with impact, attribution and chronology per family, an interactive timeline from 1971 to 2024 inside the app, an explorer filterable by family, platform, type and collection, and a resource hub listing other repositories, sandboxes, datasets and standards.
How do I add new samples to Malware-Research-Hub?
Export an API key for a malware exchange service, run the downloader against a signature with a limit, then run the integration script, which encrypts what arrived and reindexes the tree.
Is Malware-Research-Hub licensed?
The repository tree contains no licence file, and no licence is reported in the repository metadata. The provenance section names the upstream collections the samples are aggregated from, but the licensing question is left open.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/darama22-malware-research-hub)