CeWL: a Ruby spider that scrapes a site into a password cracking wordlist
CeWL is a Custom Word List Generator
At a glance
- What is it?
- digininja's CeWL walks a target site to a configurable depth, harvests unique words plus metadata and email addresses, and hands the result to something like John the Ripper. The whole thing is a single Ruby script with a companion tool called FAB.
- Who is it for?
- CeWL is worth having when your target is a small site with a public blog, because the words a company's own pages contain are exactly the ones no generic wordlist has, and CeWL also gives you author names and email addresses from page metadata that most scraping tools skip. It is the wrong tool against a large site, since the depth default of 2 with offsite crawling will drift across domains, and it is the wrong tool if you are not authorised to test the target.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 105 days ago.
- What is it written in?
- Mainly Ruby, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Installing a single Ruby script with bundler
CeWL is small in a way that is unusual for a security tool. The whole crawler is `cewl.rb`, with `cewl_lib.rb` holding the extraction logic, and `fab.rb` alongside it. There is no gem to publish and no library to require; you run the script.
The install path the README leads with is the repository clone. Three steps:
git clone https://github.com/digininja/CeWL.git
cd CeWL
bundle install
chmod u+x ./cewl.rbThe dependency list is in the README as well, and it explains why Ruby rather than something else: `nokogiri` for parsing, `public_suffix` for domain handling, `spider` for the crawl itself, `mime` and `mime-types`, `rubyzip`, `getoptlong` for the option parser, and `mini_exiftool`, which is the odd one out because it also needs the `exiftool` application installed separately. That last dependency is what makes the metadata extraction work, since image metadata is where author and creator names come from.
If you want it on your PATH, the README suggests a symlink into `/usr/local/bin`. There is also a check step worth copying, since a missing exiftool is the most common install problem:
./cewl.rb --versionOne warning you can ignore. On Ruby 2.7 the mime-types gem emits a message about `_1` being reserved for a numbered parameter, coming from the gem's own logger rather than from CeWL. The README says as much and points at the upstream commit fixing it.
Running the published Docker image instead
There is an official image on GitHub Container Registry, which is the easier route if you would rather not manage Ruby or exiftool on the host:
docker run -it --rm -v "${PWD}:/host" ghcr.io/digininja/cewl [OPTIONS] ... <url>The volume mount is the part that matters. CeWL writes its wordlist to the current directory when you pass the write option, so mapping your working directory to `/host` means the output lands where you can see it rather than inside a discarded container.
The `Dockerfile` in the repository is short and explains the design. It starts from `ruby:3-alpine`, copies just the `Gemfile` first so the dependency layer caches independently of the source, installs `gcompat` plus a build toolchain, runs `gem install bundler` and `bundle install`, deletes the build dependencies again, then copies the repository and sets the working directory to `/host`. The entrypoint is the CeWL script itself, which is why the run command needs no command name after the image.
There is also a `compose.yml` for people who want a service definition, though it does little more than name the image and point at a local build.
What the option list actually gives you
The `--help` output is the best documentation in the project, and it is worth reading as a design statement as well as a reference. The defaults are conservative: depth 2, minimum word length 3, no offsite crawling, output to stdout.
-d <x>,--depth <x>: Depth to spider to, default 2.
-m, --min_word_length: Minimum word length, default 3.
-x, --max_word_length: Maximum word length, default unset.
-o, --offsite: Let the spider visit other sites.The depth and offsite flags are the pair the README warns about together. Set a large depth and allow offsite visits and the crawl drifts across domains, which is both slow and full of words you did not ask for. Two flags narrow it instead: `--exclude` takes a file of paths to skip, and `--allowed` takes a regex the path must match before it is followed. Those two are how you keep a large site crawl honest.
Output control is `--write` for a file rather than stdout, `-c` to show a count for each word, `--lowercase`, `--with-numbers` to accept words containing digits, and `--convert-umlauts` which rewrites Latin-1 characters in the common German way, `ä` to `ae` and `ß` to `ss`. That last flag exists because wordlists are usually run against ASCII-only cracking tools.
There is a group for authentication, `--auth_type` for digest or basic with a user and password, a proxy group with host, port defaulting to 8080 and credentials, and a headers option that takes `name:value` and can be passed more than once. If a target sits behind a login, those are what get you through it.
Metadata, emails and the companion FAB tool
The features beyond word extraction are what separate CeWL from a generic scraping script. The `-a` flag includes metadata, and `--meta_file` writes it to a file. That is where the `mini_exiftool` dependency earns its place: CeWL downloads files the page references and reads their embedded metadata to build a list of authors and creators.
Email harvesting is separate and simpler: `-e` includes addresses found in the pages and `--email_file` writes them out. The 6.2 release fixed a bug here, in mailto hrefs not being detected and the output file not being created when it should.
The URL structure group is the most underrated feature. `--capture-paths` adds path components to the wordlist, `--capture-subdomains` adds subdomain components, `--capture-domain` adds the main domain, and `--capture-url-structure` does all of it together. For an attacker or a penetration tester this matters because company intranets frequently expose hostnames in links that a pure word extractor would throw away.
The associated tool is FAB, Files Already Bagged, which the README says uses the same metadata extraction to create author and creator lists from files you have already downloaded. It ships as `fab.rb` in the same repository and shares the dependencies.
Release history and what has changed
Three releases are visible in the project's history. Version 6.1 from 2023-07-31 added the maximum word length option alongside the existing minimum. Version 6.2 from 2024-06-19 was bug fixes for mailto email detection and the metadata output file. Version 6.2.1 from 2024-07-30, named More Fixes, addressed a single reported issue.
The help output prints the version as `CeWL 6.2.1 (More Fixes)`, which confirms the current release matches. The default branch `master` was last pushed on 2026-06-24 and the repository is not archived, so work is continuing even though the tagged releases are spaced roughly a year apart. `changelog.md` sits in the repository root for anything the release notes miss.
The Ruby 2.7 warning documented in the README is a small example of how this project handles maintenance. Rather than pinning an old Ruby or ignoring the problem, it names the gem, links the upstream commit that fixes it, and states plainly that the warning does not affect CeWL. If you want the warning gone, the README gives you the flag: `ruby -W0 ./cewl.rb`.
Where CeWL sits against other wordlist sources
The usual alternative is a prebuilt wordlist, and CeWL does not compete with those. RockYou and the SecLists collections are general purpose: millions of leaked passwords and common English words, applicable to anything. CeWL produces something different, a list of words specific to one organisation, drawn from the pages they chose to publish. The two are complementary and the interesting workflow is combining them.
The second alternative is writing the crawler yourself, or reaching for a general scraping library and a word-frequency pass. That is more work for the same result, and it is worth remembering that CeWL's value is not the crawling, which is a dozen lines with `spider` and `nokogiri`, but the details: umlaut conversion, metadata extraction, URL structure capture and the email handling.
Where CeWL is genuinely the wrong tool is a large site with a deep URL structure. Depth 2 is fine for a company blog. Pointed at a news site or a documentation portal with thousands of pages, a crawl will take a long time and produce a list with a lot of noise in it. The `--allowed` regex exists for that case, and so does capping depth and word length, but the tuning work is on you.
On the legal side, this is a tool for authorised testing. Crawling a site you do not own to build an attack list is not what it is for, and the README's own framing, a password cracker such as John the Ripper, makes that plain.
Editorial conclusion
CeWL is worth having when your target is a small site with a public blog, because the words a company's own pages contain are exactly the ones no generic wordlist has, and CeWL also gives you author names and email addresses from page metadata that most scraping tools skip. It is the wrong tool against a large site, since the depth default of 2 with offsite crawling will drift across domains, and it is the wrong tool if you are not authorised to test the target. Two practical notes. The current release is 6.2.1 from 2024-07-30 while `master` was pushed on 2026-06-24, so the packaged gem and the source have moved apart. And the licence is Creative Commons Attribution-Share Alike 2.0 UK, or GPL-3+ if you prefer that, which is unusual for a tool and worth reading before you redistribute it. Start with `./cewl.rb --version`, then a depth-2 crawl with no offsite flag.
Frequently asked questions
What does CeWL stand for and what is its purpose?
CeWL is the Custom Word List Generator. It is a Ruby app that spiders a URL to a specified depth, optionally following external links, and returns a list of unique words intended for password crackers such as John the Ripper. By default it stays on the site you gave it, goes two links deep, and outputs words of three characters or more.
How to use CeWL tool?
Clone the repository, run `bundle install`, make the script executable with `chmod u+x ./cewl.rb`, then run `./cewl.rb` with a URL and any options. Useful flags include `-d` for depth, `-w` to write the wordlist to a file, `-e` to include email addresses, `-a` for metadata, and `--capture-url-structure` to keep domain, path and subdomain components. `./cewl.rb --help` lists everything.
Can CeWL run without installing Ruby?
Yes, there is an official image at `ghcr.io/digininja/cewl` and a `Dockerfile` in the repository that builds from `ruby:3-alpine`. Mount your working directory to `/host` so the written wordlist lands where you can see it: `docker run -it --rm -v "${PWD}:/host" ghcr.io/digininja/cewl -w <url>`. The compose file in the repository does the same thing as a service.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/digininja-cewl)