wordlists: three tags promised, four listed, and no licence file
📜 Yet another collection of wordlists
At a glance
- What is it?
- kkrypt0nn/wordlists is a sorted collection of wordlists for security testing, with custom lists built for a legal practice platform and published as four container images. It is a data repository rather than a tool, and the interesting parts are its provenance and its paperwork.
- Who is it for?
- Treat this as a data repository and judge it as one. The custom lists were built for a platform that exists to let people practise offensive work legally, which is the frame the project itself puts around them, and sorting by content with a line count next to every entry is the right way to publish a corpus.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A legal practice platform is where the custom lists came from
The section that explains the project's own point of view is the one about a commercial training platform. That platform is described as a place to practise penetration testing in a legal and controlled environment, offering challenges and realistic scenarios, and the author says the custom lists in one folder of this repository were made specifically for it after getting into the challenges and enjoying them. Those lists then became useful for enumeration work, and the invitation to anyone else taking on the same challenges is explicit. Read that way, the repository is a practitioner keeping notes from a sanctioned environment and publishing them onward, which is the ordinary provenance for wordlists of this kind. What it does not say is where the other, older lists came from, and those are the bulk of the files.
Three tags promised and four listed
The container section opens by telling you how many image tags there are, and then lists more than that number. The count is three. The list has four: an Alpine-based latest, a Debian-based latest, a scratch-based latest and an Ubuntu-based latest. Whichever number you were expecting, the practical shape is clear enough, since all four follow the same naming convention with the distribution name in front of the tag, and each image carries a folder holding every wordlist in the project. Including one built from nothing at all is a deliberate choice for a data-only image, since there is no runtime to need a base distribution. It is worth knowing that the tag suffix is a moving one on all four, so pinning a specific build rather than the latest variant is the way to keep a test run reproducible.
The index is a line count next to every entry
What makes this repository browsable is that every file in it carries its size in the listing, which is the only metadata offered. The entries are grouped under collapsible headings by category, and the categories visible are discovery, a small group of well-known lists, the custom lists for the practice platform, and a large set of language-specific files. The sizes span four orders of magnitude. In the discovery category the smallest entries are a handful of lines, a handful of sensitive filename patterns for Unix and Windows, and a list of extensions with only forty-three entries, while the largest is nearly a million lines of endpoint paths. Size is not quality, but it is the difference between a list you can run against everything and one you point at a single host.
One entry is a zip and is fourteen million lines
A single entry dominates the whole index. It is the famous password corpus, and it is the only one stored as a compressed archive rather than as a plain text file, which is why its line count can be an order of magnitude above anything else in the repository. That single fact changes how you should think about the rest. A directory list of a couple of hundred thousand lines is cheap to keep on disk and quick to run. A fourteen-million-line password list is a different object: it is large enough that the interesting engineering is in how you feed it to whatever you are using, not in having it. The readme gives no guidance on that, no tool recommendations and no notes on memory, so this repository answers what to download and not what to do with it.
The same file sits in two categories
The category grouping is a navigation aid rather than a partition, and the listing shows where they overlap. A gitignore pattern list appears once under discovery and again under the custom lists, with the same line count in both places, which suggests the file is shared rather than duplicated but indexed twice. Elsewhere the overlap is deliberate and useful: separate lists exist for Unix and for Windows under both a local-file-inclusion heading and a sensitive-filenames heading, because a path traversal probe on one platform is not the same probe as on the other. Nothing in the listing says which files are shared, which are near-copies, and which are unique, and there is no machine-readable manifest in the readme to tell you.
No licence file, but a terms-of-use document
The repository root holds a contributing guide, a notice file and a document of terms of use, and no licence file at all. The recorded licence metadata is empty as well. That combination is the fact to sit with before anything else. Most entries in this collection did not originate with its author: a famous password corpus, a list of phished addresses for a named service, a list of addresses for another, subdomain lists and endpoint wordlists all trace to other people and other collection efforts, and none of that provenance is recorded in the visible text. There is no statement here about which upstream licences apply, which of these corpora may be redistributed inside your own tooling, or what attribution you owe. For a file you are about to bake into an image or a pipeline, that question comes before the line counts.
Sorted by content, versioned by hand, indexed twice
The remaining structure tells you how the collection is maintained. Names carry version numbers, so there is a first and a second user-enumeration list for one server family, and a directory list in three sizes plus lowercase variants of two of them. That is a sensible way to publish the same corpus at several granularities, and it is also a naming scheme that has to be kept consistent by hand. Alongside the human-readable index there is a machine-readable one at the root, a directory of container definitions and a tools directory, which together suggest that the image builds are generated from the metadata rather than maintained separately. The readme's own contribution guidance matches the shape: ask in the issue tracker for something to be added, open a pull request if you already have it.
Editorial conclusion
Treat this as a data repository and judge it as one. The custom lists were built for a platform that exists to let people practise offensive work legally, which is the frame the project itself puts around them, and sorting by content with a line count next to every entry is the right way to publish a corpus. Before you pull any of it into your own tooling, read the terms-of-use document and the notice file at the root, then work out where the upstream corpora came from. There is no licence file in the repository and the recorded licence field is empty, so nothing in it tells you what you may redistribute. That is the gap to close, and the line counts are the other thing to look at: they range from single digits to fourteen million, so choosing a list is mostly a question of what you can afford to run.
Frequently asked questions
How are the kkrypt0nn wordlists organised?
By content, under collapsible category headings, with every file showing its line count. The categories visible are discovery, a small set of well-known lists, custom lists made for a legal pentesting practice platform, and language-specific files. A machine-readable index sits at the repository root alongside the human-readable one.
Which Docker images does the kkrypt0nn wordlists project publish?
Four tags, named for their base: alpine-latest, debian-latest, scratch-latest and ubuntu-latest, each carrying a folder with every wordlist in it. The readme says there are three tags and then lists four. All four use a moving latest suffix.
Where did the custom wordlists in this repository come from?
The author says they were created specifically for a commercial platform that exists to let people practise penetration testing in a legal, controlled environment, and became useful there for enumeration tasks. The older, larger lists in the collection do not have their provenance recorded in the visible text.
What licence are the kkrypt0nn wordlists under?
No licence file is present at the repository root and the recorded licence metadata is empty. The root does carry a notice file and a document of terms of use, but the visible text does not state which upstream licences apply to the individual corpora or what attribution a redistributor owes.
How do I request a new wordlist be added?
Ask in the issue tracker if you have a wordlist you want included, or open a pull request if you already have the file ready. The repository also carries a container directory and a tools directory alongside a machine-readable index at the root.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kkrypt0nn-wordlists)