Self-hosted service
laramies/theHarvester avatar
laramies/theHarvester

theHarvester: Passive Recon with P0/P1/P2 Source Classes

E-mails, subdomains and names Harvester - OSINT

17,721 stars2,635 forksPythonGPL-2.0

At a glance

What is it?
theHarvester collects emails, subdomains and hostnames from public sources, then separates passive lookups from DNS and direct-contact actions. Here is how the install works, where the source taxonomy matters, and what the GPL-2.0 licence means for internal tooling.
Who is it for?
Adopt theHarvester if you run authorized assessments and want a documented boundary between passive provider queries and DNS or direct-contact actions, since the P0/P1/P2 classes and the explicit selection requirement for P1 and P2 give you something concrete to put in a scope document. Do not adopt it as a general-purpose internet scanner: it is built around authorized targets, and the README repeats that constraint.
Can I use it commercially?
Yes, with conditions. GPL-2.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What theHarvester collects and who runs it

theHarvester gathers open-source intelligence about a domain or organization from search engines, certificate transparency logs, DNS datasets, code repositories and threat-intelligence platforms. It normalizes what it finds into hostnames, email addresses, IP addresses, URLs, ASNs, people, breach names and structured evidence from optional actions. That normalization is the actual product: individual providers are replaceable, but a single result shape across all of them is what makes the output usable in a report.

The intended user is a penetration tester or red teamer working the early reconnaissance stage of an authorized assessment. The README is explicit that you run it only against targets you own or have explicit permission to test. Blue team topics are attached to the repository, which suggests defenders use it to see what their own external footprint exposes, but the documentation is written from the operator's side.

It is not a scanner in the vulnerability sense. Nothing in the described feature set probes for a bug. It maps names, addresses and public references, and the active options that do exist (takeover checks, API path scanning, screenshots) are framed as evidence for review.

P0, P1 and P2: the source taxonomy that defines your blast radius

The most useful design decision here is that sources are classified by observable network behavior rather than by confidence. P0 queries an existing provider or dataset without directing traffic toward the target. P1 queries DNS about authorized names or addresses. P2 contacts a target endpoint or causes equivalent direct interaction, covering HTTP, TLS, screenshots, takeover checks, virtual hosts, ports and API paths.

The classes describe what the network sees. They say nothing about how trustworthy a result is, and the README states that plainly. That distinction matters when you write a scope document: you can commit to a P0-only run and mean something specific.

P1 and P2 run only when you select them. A discovered related hostname, network, ASN or URL stays review evidence and does not expand the authorized target on its own. The README also points out that www.example.test and example.test are different targets unless both are authorized. That is a small sentence with large consequences for anyone who treats a domain as a wildcard.

One caveat the README raises rather than buries: passive does not mean invisible. The selected providers still receive the target string. If your engagement forbids disclosing the target to third parties, P0 is not automatically safe.

Installing theHarvester with uv and running a first P0 query

theHarvester requires Python 3.14 and uv. The repository ships a .python-version file so uv selects the interpreter automatically, which removes the usual version-matching step. Clone, sync, and run against a domain you control.

bash
git clone https://github.com/laramies/theHarvester.git
cd theHarvester
uv sync
uv run theHarvester -d example.com -b crtsh,certspotter

That first run uses two P0 passive sources that need no API key. The terminal reports each source outcome and any retained findings. A source can finish with zero findings, stop early with partial evidence, or fail without erasing evidence that other sources retained. Read those per-source lines rather than only the totals; a partial outcome carries a stop reason.

Three discovery sources run at once by default. Use -j or --source-workers to change the worker count, and note that every selected source still runs regardless of the worker count. Capability selectors such as subdomains, emails, ips, asns, urls, people and breaches form a union and choose which sources run; they do not discard other result types those sources return.

bash
uv run theHarvester -d example.com -b emails,urls,certspotter
uv run theHarvester -d example.com -b crtsh,certspotter -f report

The second command writes report.jsonl for automation plus report.json and report.xml compatibility reports. Completed runs are also kept in the local SQLite evidence store. The README notes that -b all runs every cataloged P0 passive source, while P1 DNS and P2 direct sources require explicit selection. The installation guide in docs/wiki covers packaged distributions and platform-specific setup.

HarvestView, the REST API and the shared run engine

The CLI is one of three interfaces over the same finite-run engine and normalized evidence model. HarvestView is a local browser interface for run history, evidence review, hostname changes and schedules, started with uv run harvestview. The REST API serves authenticated local automation and integrations, with interactive documentation at http://127.0.0.1:5000/docs.

Because all three share one engine, an option set on the CLI has a counterpart elsewhere. The README gives --no-hosts as an example: it skips hostname-only sources and omits hostname results while keeping other result types, and HarvestView and the REST API expose the same option as no_hosts. It cannot be combined with actions that depend on hostnames.

bash
uv run theHarvester -d example.com -b emails,ips,urls --no-hosts -f non-host-results

The Docker Compose file shows how the API side is meant to be deployed. The service runs read-only with all capabilities dropped, no-new-privileges set, a 1gb shared memory segment, and a size-limited tmpfs at /tmp. The API key is read from a file through THEHARVESTER_API_KEY_FILE, and the published port binds to 127.0.0.1 only, defaulting to 5000 and mapping to 8000 in the container. API keys and proxy configuration mount read-only from theHarvester/data/api-keys.yaml and theHarvester/data/proxies.yaml.

yaml
environment:
  THEHARVESTER_API_KEY_FILE: /run/secrets/operator-api-key
  THEHARVESTER_HARVESTVIEW_LOCAL_PROXY: enabled

With explicit proxy mode enabled, supported discovery and actions fail closed when no proxy is available rather than sending a direct request. That is a deliberate choice: a misconfigured proxy stops the run instead of leaking traffic.

Where theHarvester is the wrong tool

The limits are structural. P1 and P2 do nothing unless you ask for them, so anyone expecting a single command to produce a full picture of a target will be disappointed by a default run that only touches certificate transparency and similar datasets.

Provider quotas are real and the tool does not hide them. Passing --limit 0 removes the shared per-source result cap and local page ceilings, but the README states that provider quotas and runtime safeguards still apply. If a provider or safety limit stops a source after it has retained results, the run keeps them and records a partial outcome with the stop reason. You get a truncated answer with a label, not a complete one.

The active options are narrower than the names suggest. A takeover indicator is evidence for review, not proof that a provider resource can be claimed. Screenshots, virtual host checks and port activity all create direct interaction with the target, which is exactly the traffic a covert engagement may forbid.

Scope handling is manual. The tool will not stop you from running an action against a hostname that appeared in results but was never authorized. That judgement stays with the operator, and the README's guidance to read the responsible-use page before active work is the only guardrail described.

How it compares with a DNS brute-force and resolution tool

The closest alternative in practice is a dedicated DNS enumeration tool such as a brute-force and resolution utility, and the difference is in where the names come from. A brute-force tool generates candidate labels from a wordlist and asks DNS whether each one exists. Its output is bounded by the wordlist and by the resolver you point it at, and every query is P1 activity under theHarvester's own classification.

theHarvester inverts that. It starts from public datasets and provider APIs, so crtsh and certspotter return names nobody would have guessed from a wordlist because they were published in a certificate. DNS resolution (-r), brute force (-c), reverse DNS (-n) and recursive DNS (--dns-recursive-depth) exist here too, but they are options layered on top of a passive base rather than the primary mechanism.

The practical consequence is different failure modes. A brute-force run fails quietly when the wordlist is thin; you get few names and no signal that more exist. A theHarvester run fails loudly per source, with zero-result and partial outcomes reported separately. The trade-off is dependency on third parties: if a provider changes its response format or rate-limits you, that source degrades, and you are reading stop reasons instead of DNS timeouts.

Maintenance, upgrades and the GPL-2.0 licence

The repository is not archived, and the last push was on 2026-09-19. Releases are frequent enough to plan around: 4.11.1 on 2026-06-03, 4.11.0 on 2026-05-23, and 4.10.1 on 2026-02-22. The changelog file at the repository root is where upgrade notes would live.

The dependency list is pinned exactly, including aiohttp, aiodns, playwright, fastapi, sqlalchemy and uvicorn, with uvloop on non-Windows platforms and winloop on Windows. That pinning makes upgrades a deliberate act rather than a side effect of reinstalling. The Dockerfile builds with uv sync --locked --no-dev, so a container rebuild reproduces the locked set; the runtime image installs a headless Chromium for the screenshot path and runs as a non-root user with a fixed data directory under /var/lib/theharvester.

The licence is GPL-2.0-only per pyproject.toml, with the container image labelled org.opencontainers.image.licenses="GPL-2.0-only". If you embed theHarvester in a distributed product, the copyleft terms of GPL-2.0 are the thing to read in full with your own counsel. Running it internally against your own authorized targets is a different situation from shipping it inside something you hand to customers, and the licence text, not this article, decides which applies.

Editorial conclusion

Adopt theHarvester if you run authorized assessments and want a documented boundary between passive provider queries and DNS or direct-contact actions, since the P0/P1/P2 classes and the explicit selection requirement for P1 and P2 give you something concrete to put in a scope document. Do not adopt it as a general-purpose internet scanner: it is built around authorized targets, and the README repeats that constraint. Before your first run, verify that Python 3.14 and uv are available on the host, confirm which sources need entries in theHarvester/data/api-keys.yaml, and check whether your engagement permits the target string to reach third-party providers at all.

Frequently asked questions

What does theHarvester do?

It gathers open-source intelligence about a domain or organization from search engines, certificate transparency logs, DNS datasets, code repositories and threat-intelligence platforms. It normalizes the results into hostnames, email addresses, IP addresses, URLs, ASNs, people, breach names and structured evidence from optional actions.

Is theHarvester an OSINT tool?

Yes. The repository describes it as an OSINT tool and lists osint among its topics, and the README frames it for the early reconnaissance stage of an authorized security assessment. It is meant for targets you own or have explicit permission to test.

How do I get theHarvester?

Clone the repository from GitHub and install it with uv, which the .python-version file uses to select Python 3.14 automatically. The README's installation guide in docs/wiki covers packaged distributions and platform-specific setup.

How do I install theHarvester?

It requires Python 3.14 and uv. After cloning the repository, run uv sync in the project directory, then invoke it with uv run theHarvester followed by your options.

How do I use theHarvester?

Pass a domain with -d and a source list with -b, for example uv run theHarvester -d example.com -b crtsh,certspotter. Capability selectors such as emails, subdomains or urls can replace explicit source names, and -f writes report.jsonl alongside JSON and XML compatibility reports.

Official sources

  1. laramies/theHarvester on GitHub
  2. License: GPL-2.0
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/laramies-theharvester.svg)](https://hysenlabs.com/projects/laramies-theharvester)
Community notes

Community notes