Open-source project
megadose/holehe avatar
megadose/holehe

holehe probes 120 password-reset endpoints, and its rate-limit advice is to change your IP

GitHub describes it as holehe allows you to check if the mail is used on different sites like twitter, instagram and will retrieve information on sites with the forgotten password function.. The repository metadata lists Python as its primary language. The metadata lists the GPL-3.0 license. This article stays within the project description and details documented in the GitHub repository README.

15,089 stars1,944 forksPythonGPL-3.0

At a glance

What is it?
holehe is a Python tool that takes an email address and asks more than 120 sites whether an account exists, by triggering the same forgot-password flow a user would. It has not been pushed to since 2024-09-10, it has no published release despite a version string in its packaging, and its own documentation treats rate limits as the primary obstacle rather than a detail.
Who is it for?
holehe is worth knowing about for what it reveals about how account-existence leaks work, and for the fact that its module list is a catalogue of sites that have not closed that leak. It is a poor tool for anything you need to be reliable, because its coverage is built on endpoints that get patched without notice and its own answer to being blocked is to change IP.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Probably not. The repository last received commits 25 months ago, on September 10, 2024.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Every module asks the same question through the forgot-password form

The mechanism is one idea applied many times, and understanding it tells you what the tool can and cannot tell you.

For each site it has a module for, holehe submits an email address to that site's account-recovery endpoint and reads the difference in the response. A site that returns the same page whether or not the account exists gives nothing away. A site that sends a reset mail only to real accounts, and shows a different confirmation otherwise, gives a yes or no answer.

The module table names the probe used per site, and it is not uniform. Some rows are marked register, some password recovery, some login. That is the interesting part of the design: the same question gets asked through three different doors depending on what the site happens to expose.

What comes back is a yes or no per site. Not a password, not a session, not access to anything. The output is a dictionary with an exists flag and two partially obfuscated fields.

So the ceiling on this tool is account existence. Anyone expecting it to go further than that is looking at the wrong tool.

The packaging carries version 1.61 and the repository has published no release at all

There is a gap between what the packaging claims and what the project has actually shipped.

The setup script declares its version as 1.61, which implies a long sequence of internal iterations. The repository's release list is empty. No tag has been published through the platform's release mechanism, so there is nothing to pin a dependency to and no changelog to read against that version number.

The last push to the default branch was 2024-09-10, so treat that as the point where the code stopped moving rather than assuming anything is being fixed. When a site patches its password-reset behaviour, the corresponding module here has no upstream to pull from unless the maintainer chooses to write one.

This matters more than a normal stale dependency. A library that wraps third-party endpoints rots from the outside in: the wrapped services change whether or not this repository does. A two-year gap since the last push means the module list is, today, a record of what was true when it was last touched.

If you need this in something you maintain, fork it and own the updates. The alternative is a dependency whose behaviour is set by 120 other companies.

The documented answer to being rate-limited is to change IP

Read the rate-limit handling as the project's own statement of its weakest point.

Every module returns a rateLimit flag, and the documentation's guidance for a hit is one sentence: rate limit, change your IP. That is not a workaround, it is an admission. The tool cannot distinguish between a site that will never answer and a site that is refusing to answer you right now, so the flag exists to stop it from reporting a false negative.

The module table makes the scale visible. It carries a column for whether a site rate-limits frequently, and the marks are scattered across the list rather than concentrated. Some large services are marked as limiting often. Several small forums are not. The pattern is not popularity, which tells you the rate limiting comes from the sites' own abuse controls rather than from any load holehe puts on them.

The consequence for a user is direct. A single run can return a partial answer, and nothing in the output distinguishes a site that rate-limited you from a site where the email genuinely is not registered. Both look like exists false to a reader who does not check the rateLimit flag.

So if you act on this tool's output, read both fields per row. Treat exists false as unknown whenever rateLimit is true.

The documented Python API is one module at a time, with no batch entry point

The library surface shown in the documentation is narrower than the command line, and that asymmetry shapes how you would embed it.

The CLI runs a full scan from an address. The Python example does not. It imports a single site's module, opens one HTTP client, awaits that one function, and prints the result.

python
import trio
import httpx

from holehe.modules.social_media.snapchat import snapchat

async def main():
    email = "[email protected]"
    out = []
    client = httpx.AsyncClient()

    await snapchat(email, client, out)

    print(out)
    await client.aclose()

trio.run(main)

Two things follow. The caller supplies the client and owns its lifetime, which is the right shape for a tool that makes a hundred-plus outbound requests and wants connection reuse across them. And the caller owns the output list, which the module appends to rather than returning, so the same list accumulates results across as many modules as you choose to call.

What is not shown is a function that walks the whole module set. If you embed this, you are choosing which modules to call and in what order, and you own the concurrency and the backoff. The documentation does not cover either.

That is a fair amount of glue to write, and it is the only way to control rate limiting rather than just absorb it.

The recovery email and phone fragments are the most sensitive thing in the output

Set the two obfuscated fields beside what the tool is actually for, because they are easy to under-weight.

The documented output shape carries an exists flag, a rateLimit flag, an emailrecovery field for sometimes partially obfuscated recovery emails, a phoneNumber field for sometimes partially obfuscated recovery phone numbers, and an others slot for anything extra.

Recovery addresses and phone numbers are not account metadata. They are the contact details a person uses to get back into accounts they have lost access to. A partially masked value is still enough to confirm which of a person's addresses is on file at which service, and across 120 sites that is a correlation surface.

The consequence is about handling rather than about collection. A scan written to a file or piped into a shared report is a different artefact from a scan run in a terminal and read once, and the artefact outlives the reason for making it.

Collect only what you need. Most questions are answered by the exists flag alone, and the other two fields are worth having only when they are the actual point of the query.

The container build installs with a setup.py call that current tooling deprecates

There are three documented installation routes and they disagree about what runs.

The package index route is a single pip command. The source route clones the repository and runs the setup script directly. The container route builds an image whose three instructions copy the source into a working directory and then run the same setup script install inside it.

code
FROM python:3.11-slim-bullseye
COPY . /opt/holehe
WORKDIR /opt/holehe
RUN python3 setup.py install

So the container and the manual source route are the same mechanism, wrapped differently. Neither goes through pip, and installing through setup.py directly is a path the packaging ecosystem has been moving away from for years.

The container gives you one fixed interpreter, 3.11 on a slim Debian base, which is the main reason to prefer it: it removes the question of what Python version the hundred-plus modules and their four dependencies end up compiled against.

The dependency list is short, termcolor, a HTML parser, the HTTP client, trio for the async layer, a progress bar and a colouring library. Nothing exotic. That is the easy part of this project.

The claim that the target is not alerted is a link to an issue, not a mechanism

The documentation makes one assurance about the people being probed, and it is worth reading precisely.

Under the summary, a bracketed line states that the tool does not alert the target email, and it is a hyperlink to an issue in the repository rather than a description of how that is guaranteed. A read-only password-reset request does not by itself send a notification to the address holder, but sites differ, and some do send one on a recovery request.

Nothing in the repository documents which of the 120 modules fall on which side of that line. So the statement cannot be relied on as a property of the tool, and it is also not something the user controls.

The licence section closes with a single line restricting the tool to educational purposes, and the licence itself is GPL-3.0, which governs the code and not what anyone does with it. That framing puts the judgement on the person running it.

Combined with the rate-limit guidance, the honest summary is that this is a research tool with a real capability and an unresolved compliance question that the project leaves entirely with the user.

Editorial conclusion

holehe is worth knowing about for what it reveals about how account-existence leaks work, and for the fact that its module list is a catalogue of sites that have not closed that leak. It is a poor tool for anything you need to be reliable, because its coverage is built on endpoints that get patched without notice and its own answer to being blocked is to change IP. Before running it on anyone, decide whether you are entitled to, since the tool's own documentation restricts it to educational use and the obligation is yours, not the author's. And if you embed it, use the per-module Python entry point rather than the CLI, because that is the only surface the documentation actually shows you.

Frequently asked questions

What is holehe?

A Python tool that takes an email address and checks whether it is attached to an account on more than 120 sites, by triggering each site's forgotten password function and reading the difference in the response. It runs from the command line or can be embedded per site through a Python module, and it depends on termcolor, a HTML parser, httpx, trio, tqdm and colorama.

Is Holehe legal to use?

The project's own documentation restricts it to educational purposes, and the licence is GPL-3.0, which covers the code rather than its use. The documentation does not make a legal claim and does not document which sites can be probed without breaching their terms. The decision about authorisation is the user's to make, not the project's.

How to install holehe in Kali?

Three routes are documented. From a package index with pip3 install holehe. From source by cloning the repository and running python3 setup.py install. Or as a container built with docker build . -t my-holehe-image and run with docker run my-holehe-image holehe [email protected]. The container is the one that pins the interpreter, at 3.11 on a slim Debian base.

how to use holehe python

Import the module for a single site, for example holehe.modules.social_media.snapchat, create an httpx.AsyncClient, pass the email, the client and an output list to the module function, then close the client. The documented example runs under trio.run. The CLI is the only documented way to sweep every module at once.

Official sources

  1. Official README
  2. Project repository